Most UI/UX design services companies can make an AI feature look finished. Few can design for uncertainty, streaming, and trust. Here are the 20 questions that separate AI-native partners from generalists claiming AI expertise.
20 questions every CTO should ask before hiring an AI product design partner.

Every design agency's website says they "specialize in AI." Their portfolios are gorgeous, their decks are polished, and in a first call they will nod along to every AI buzzword you mention. None of that tells you what you actually need to know: can they design an interface for something that is non-deterministic, that streams, that is sometimes wrong, that acts on its own, and that has to earn a user's trust before anyone will touch it?
Choosing UI/UX design services for an AI product is a fundamentally different exercise than hiring a designer to make your marketing site look good.Most CTOs run the generic evaluation, hire on vibes and visuals, and discover the gap three months into the engagement, usually right after launch, when adoption numbers come in flat despite a beautiful interface.
TL;DR
AI products are non-deterministic, so vetting a design vendor needs AI-specific questions, not a generic portfolio review.
Score every vendor across six areas: AI/ML literacy, trust and transparency, human-in-the-loop design, technical collaboration, portfolio proof-points, and engagement model and pricing.
Watch for shops where every "AI" case study is just a chatbot bubble bolted onto a normal app.
Run all 20 questions before you shortlist, not after you sign a contract.
A vendor with a slightly less polished portfolio but genuine command of non-determinism will outperform a beautiful portfolio that has never had to make an AI feature trustworthy.
This is a 20-question due-diligence checklist built specifically for AI products, organized into six categories a CTO can run through before shortlisting any vendor. Ask these, listen for the reasoning behind the answers rather than the buzzwords, and you will separate the teams who can genuinely design AI products from the ones who just added "AI" to their homepage.
Why vetting UI/UX design services for AI is different

AI products fail on adoption more than they fail on models, and the reasons are almost always design problems that a generalist design partner is not equipped to solve. A traditional product is deterministic: given the same input, it returns the same output, so you design a fixed, predictable interface. An AI product is not. Its output varies, streams in over time, can be wrong, and, in agentic products, acts autonomously.
That difference creates design problems that simply do not exist in conventional software:
How do you communicate confidence and uncertainty to a user?
How do you render a response that is still streaming in?
How do you turn a user's "that's wrong" into a useful correction instead of a dead-end thumbs-down?
How do you keep a user in control of an agent that is taking a dozen autonomous steps?
A design vendor that has only built conventional SaaS dashboards and marketing sites has never had to answer these questions. They can make your AI feature look finished while missing the exact things that determine whether users trust and adopt it. This is the gap our own AI-first UX design work is built around, designing for copilots and agents rather than retrofitting a chat bubble onto a static product. The six categories below map closely to the broader skills a UI/UX agency must have in 2026, and are how you find out whether a vendor can do the same.
1. AI/ML Literacy: Do They Understand What They're Designing For

This is the filter that removes generalists fast. If a vendor cannot talk fluently about how AI behaves, they will design your product like any other SaaS app, and it will show in every downstream decision.
Questions to ask:
How do you design for an AI feature that is sometimes wrong? A strong answer treats error and uncertainty as a first-class design case: confidence signals, easy verification and correction, graceful failure, a human kept in control. A weak answer is "we make the model more accurate," which is not a design answer at all.
How would you handle a response that streams in token by token? Listen for partial-state rendering, handling incomplete output, cursors, and stop or regenerate controls. A blank stare, or "we'd show a spinner," is a red flag.
What is the difference between designing for a deterministic feature and a probabilistic one? A vendor fluent in AI UX will immediately name the design implications: no fixed happy path, multiple valid outputs, and interface states that a conventional design system was never built to handle.
How do you decide between a copilot experience and a fully autonomous agent experience? The two need fundamentally different UX, and a vendor should be able to explain why without you prompting them.
You are not looking for textbook answers. You are looking for a team that clearly thinks about these problems and has opinions formed by doing the work, not opinions borrowed from a conference talk.
2. Trust and Transparency Questions

Adoption of an AI feature depends almost entirely on whether users trust it, and trust is a design decision, not a marketing line. This category tests whether a vendor treats trust as something to be engineered into the interface.
Questions to ask:
How do you communicate a model's confidence level to a non-technical user? Good answers reference visual hierarchy, source citations, and plain-language framing rather than raw probability scores nobody understands.
What happens in your designs when the AI does not know the answer? A vendor should have a real pattern for graceful uncertainty, not a generic error state borrowed from a 404 page.
How do you avoid over-trust, where users blindly accept AI output without checking it? This is the inverse problem most vendors never consider. Strong teams actively design friction into high-stakes decisions.
Can you show me a case where you designed a correction or feedback flow, and what happened to the user's input afterward? If a vendor cannot describe where feedback goes, they have not designed the full loop, only the moment of collection.
If a vendor cannot answer questions in this category with specifics, they are designing decoration, not a trustworthy product.
3. Human-in-the-Loop Design Questions

Most serious AI products, especially agentic ones, need a human somewhere in the loop. How a vendor designs that checkpoint tells you whether they understand control, or whether they think "human in the loop" is a slide in a deck.
Questions to ask:
Where would you place approval gates in an autonomous workflow, and why there specifically? A thoughtful answer ties gate placement to the cost of a mistake, not to arbitrary steps in the flow.
How do you design a kill switch or an interrupt for an agent that is mid-task? This is a question most conventional design teams have never had to answer, and it usually reveals who has genuinely shipped agentic products.
How do you keep a user oriented when an agent has taken several steps on their behalf? Look for answers involving visible step logs, undo paths, and clear summaries, not a silent black box that reports back only at the end.
When should a product ask for permission before acting, versus act and inform afterward? This is a judgment call that separates a vendor with real agent experience from one repeating a framework they read about.
Our UX strategy work leans heavily on this category before any screen gets drawn, because getting the control model wrong early is expensive to unwind later.
4. Technical Collaboration Questions

AI UX lives at the boundary between design and engineering more than almost any other kind of interface work, which is exactly where the design process breaks down if a vendor cannot operate on both sides. A design team that cannot operate at that boundary will produce beautiful specs that your engineers cannot, or will not, build.
Questions to ask:
How closely do your designers work with the engineers building the model or the AI pipeline? Vendors who design in isolation and hand off a static file tend to miss constraints that only surface once real model output hits the interface.
Have you designed for latency, and how did that change the interaction? Streaming, loading states, and perceived performance are UX problems with technical roots, and a vendor should be fluent in both sides.
Did you collaborate well with a client's engineering team on the AI-specific implementation, and can a past client confirm that? This is a reference-check question worth asking directly, because a glowing generic reference tells you nothing about AI delivery specifically.
If you also handle the build, not just the design file, how do you hand off or implement components like streaming renderers and confidence indicators? If a vendor's web development capability sits close to their design team, this handoff tends to be far smoother, and worth asking about directly.
Weak answers here are the ones that stay abstract: "we use Figma and hand off specs," with no mention of how those specs survive contact with a live, non-deterministic system.
5. Portfolio and Proof-Point Questions

By the time you reach this category, you should already be looking at portfolios with a specific lens: evidence of thinking and outcomes, not aesthetics.A portfolio full of beautiful screens is the single biggest trap, since polish is the easiest thing to fake and the least predictive of AI competence.
What to look for, beyond visual polish:
Problem framing and process. Case studies should show how the team got to the solution: research, ideation, prototyping, testing, not just the final UI.
System-level thinking. Evidence they design systems and patterns, not one-off screens.
Business outcomes. Case studies should reference activation, retention, conversion, time-to-value, or velocity. A portfolio that never mentions impact is a red flag; you want metrics, not mockups.
Real AI interface patterns. Conversational interfaces, confidence or source indicators, streaming responses, feedback and correction flows, agent or human-in-the-loop views. If every "AI" case study is a chatbot bubble bolted onto a normal app, that is telling.
Non-deterministic states designed on purpose. Empty, loading, uncertain, wrong, and error states, not just the happy path.
Questions to ask:
Can you walk me through an AI feature you designed where the model got something wrong in production, and how the interface handled it? A vendor who has genuinely shipped AI products will have a real story here, not a hypothetical.
What was the adoption outcome of your last AI feature, and how did you measure it? Activation, repeat use, and time-to-first-value are the numbers that matter, not launch dates.
A proof point worth asking any vendor to match: when we redesigned PathwaysX's AI-powered B2B hiring platform, the brief was to make personality-based assessments and AI-driven candidate matching feel effortless rather than like a black box, so hiring teams could trust the recommendations without having to second-guess every match. That kind of outcome only happens when trust, control, and interface craft are designed together, not bolted on after the model is built. Ask any vendor you are evaluating for an equivalent story, and press for specifics if they cannot give you one.
6. Engagement Model and Pricing Questions

This is the category most CTOs skip, and it is where a lot of AI design engagements quietly go wrong. Pricing structure affects how a vendor behaves once the project is underway, not just what you pay upfront.
Questions to ask:
Is pricing hourly, project-based, retainer, or subscription, and what happens when the scope shifts mid-project, which it will on most AI products? Hourly and fixed-project models can both work, but each has a failure mode: hourly can run over with no ceiling, and fixed-price can lead to corners being cut once a vendor feels squeezed. A retainer or subscription model, where you get a dedicated designer and strategist for a fixed monthly rate with role flexibility built in, tends to handle AI's shifting scope more gracefully than either extreme.
What is actually included at each pricing tier, and what counts as an add-on? Ask for this in writing, in a formal UX design proposal, before you sign anything, especially for AI-specific deliverables like a design system for streaming states or agent control patterns.
How quickly can you start, and what does onboarding look like? AI products move fast, and a vendor who needs six weeks to staff a project is already behind your roadmap.
What is your policy if the assigned designer is not the right fit, or is unavailable partway through? This is a practical question that separates vendors with real operational maturity from ones improvising as they go.
For context, project-based UI/UX pricing across the market commonly ranges from a few thousand dollars for a narrow scope to well over $100,000 for a complex, multi-platform build, with hourly rates generally falling between $50 and $250 depending on region and seniority. Subscription and embedded-designer models are a newer alternative worth asking about directly, since they can offer more predictable monthly costs and faster onboarding than a traditional project quote. Our own pricing page breaks down exactly this kind of flexible, subscription-based engagement if you want a live example to compare a vendor's proposal against.
Turning the checklist into a shortlist
Run all six categories as a funnel rather than deciding on a single call:
AI/ML literacy and trust and transparency filter out generalists early.
Human-in-the-loop and technical collaboration confirm the team can actually ship what they design.
Portfolio proof-points validate the story with evidence.
Engagement model and pricing protect you from a good design partner turning into a bad commercial relationship.
Score every vendor at each stage, and weight the AI-specific evidence most heavily, because for an AI product, that is the competence that determines whether users adopt what you ship. A vendor with slightly less dazzling visuals but a genuine command of non-deterministic design will serve you far better than a beautiful portfolio that has never had to make an AI feature trustworthy. This process takes a few hours of diligence up front, and it saves you a failed engagement and a costly re-do later.
A note on scope: don't separate the "UI" from the "UX" for AI
When you evaluate UI and UX design services for an AI product, resist the instinct to treat visual competence and experience competence as separable. For AI, they are inseparable. Some vendors sell primarily on visual craft, others on research and process, and for a conventional product you might get away with weighting one over the other. For an AI product, you cannot.
The confidence indicator, the streaming renderer, the correction flow, the agent view: each of these is simultaneously a UX decision, how it behaves and earns trust, and a UI decision, how it is rendered and understood at a glance. A team strong on UI but shallow on AI experience design will make your uncertainty states pretty but meaningless. A team strong on UX thinking but weak on UI will not render those states clearly enough to be trusted. This is also why AI products built on top of a broader SaaS platform benefit from a vendor who treats SaaS UX design and AI-specific design as one continuous discipline rather than two separate hires.
Conclusion
Hiring the right design partner for an AI product is not about finding the prettiest portfolio. It is about finding a team that can design for the hard realities of AI: uncertainty, streaming, corrections, autonomy, and the trust that adoption depends on. To recap:
Test AI/ML literacy first; it filters out generalists fastest.
Probe trust and transparency and human-in-the-loop design, since both determine whether users actually adopt the feature.
Confirm technical collaboration and portfolio proof-points with specifics, not general praise.
Get the engagement model and pricing in writing before scope shifts, which it will.
Run all 20 questions before you shortlist, and score vendors on paper rather than deciding on vibes from a single call.
If you are vetting UX/UI design services for an AI product, run this exact checklist on us. We are an AI-native product design team, and we would rather be evaluated on how we handle non-determinism, trust, and adoption than on our prettiest screen. Book a discovery call with Groto and bring the hard questions.
























































































































































































































































