XIV.AI & design · September 2026
AI product design is an evaluation problem. Designers are already good at it.
Building AI products is less about prompting and more about judging. Designers have spent their careers separating 'it functions' from 'it works for a human' — and that is exactly what evals need.
By Gisella Famà · 6 min read · AI & design
Most AI product conversations start with the wrong question. Which model? Which stack? Which prompt technique? Teams treat the work as a build problem and then wonder why the result feels impressive but unusable. The real question is older and harder: how do you know if what you built is any good?
That question is an evaluation problem. And it is where designers, of all people, have a hidden advantage. We have spent years looking at things we did not build — interfaces, flows, visuals, copy — and deciding whether they serve a human. We are trained to separate 'it functions' from 'it works for a person.' That distinction is the entire job of AI evaluation.
"The question isn't 'does it generate an output?' It's 'does it generate the right output, in the right way, for the right person?'"
What evaluation actually means
Evals are not benchmarks. Benchmarks measure models. Evals measure products. A model can top a leaderboard and still produce something your users will hate, because benchmarks do not know your users. They measure an average against an average. Your product lives in the specific.
At Lumina, where I worked on a creative AI agent for brand and marketing work, the gap showed up constantly. A generated asset could look correct and still miss the brief because the brief was implicit: the client's tone, the campaign's history, the thing the user meant but did not type. The model cannot see any of that. The designer can.
The four questions every eval should ask
Intent vs outcome. Does the output match what the user meant, not just what they typed?
Tone and context. Is it appropriate for the moment, the brand, and the relationship?
Failure shape. When it is wrong, is it wrong in a way the user can recover from?
Edge cases. What happens at the boundaries of the user's actual life, not the happy path?
Why designers are naturally good at this
We have a vocabulary for judgement that is underused in engineering-led rooms. We can say 'this feels off' and then unpack it: the hierarchy is wrong, the tone is too familiar, the failure is invisible, the affordance is missing. That translation from gut to actionable detail is what makes an eval useful instead of just an opinion.
Engineers often evaluate whether the system did what was asked. Designers evaluate whether the thing it did was worth asking for. Both matter. But the second one is where AI products live or die, because a system that reliably does the wrong thing is not a good system.
The eval is the new prototype
In traditional product work, the prototype is where you test assumptions. In AI product work, the eval is where you test assumptions — because the prototype changes every time you run it. You cannot click through the same flow twice and expect identical outputs.
A good eval is a small, repeatable scenario that tells you whether the product is getting better or worse. It is not a vibe check. It is not a one-off demo for stakeholders. It is a disciplined way of looking at a non-deterministic thing and saying, with evidence, whether it is fit for the job.
A simple eval framework
Pick ten real tasks your users actually do. Not toy examples. Not the tasks that demo well.
Run each one three times. Non-determinism means you need to see the distribution, not the single good run.
Score outputs against intent, not correctness. Sometimes there are many right answers; the question is whether any of them serve the user.
Track regressions. Fixing one failure mode often creates another. The eval catches that before your users do.
The uncomfortable part
Evaluating AI well requires admitting that some of what the model does is genuinely impressive and still wrong for the product. This is hard for teams because the demo is seductive. Everyone wants to believe the magic.
Designers are used to this. We have all killed a beautiful screen because it did not serve the user. We have all watched a stakeholder fall in love with an idea that the research did not support. The skill transfers directly: separate the craft from the outcome, and judge the outcome by the user's standard.
"The moat in AI product design isn't prompting skill. It's the judgement to know when the output is good enough — and when 'good enough' isn't good enough."
What to do with this
If you are a designer worried about AI replacing you, learn to evaluate it. Build the tests. Define the criteria. Sit in the gap between 'it runs' and 'it's right.' That is where the work is. And it is work only a human — especially a human trained in design — can do well.
The best AI product teams I have seen are not the ones with the most impressive demos. They are the ones with the most honest evals. They know what good looks like, they test against it constantly, and they are not afraid to say no to the clever thing. That discipline looks a lot like design. Because it is.
Disagree? That's the point. Tell me why.
← Back to all thoughts