AI for Designers (UX, Product, Brand)
Proficient · M22 · lesson 22 of 27 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
The Synthetic-User Question: Where AI Personas Fail
📖
now learning

The Synthetic-User Question: Where AI Personas Fail

15 min

Somewhere in 2025 a slide deck went around claiming you no longer needed to talk to users, because you could just ask the AI to pretend to be one. Prompt a model with a persona - "you are Sarah, 34, a busy marketing manager" - and interview it. It answers fluently, in character, with opinions and frustrations and preferences, and it never cancels, never costs $80 an incentive, never makes you schedule around its calendar. By 2026 this pattern has a name, synthetic users, and a real product category, and it is one of the most seductive and most dangerous ideas in AI-augmented research. This lesson does two things. First it demonstrates the pattern and dismantles it, showing precisely why a model role-playing a user is not data and can never be. Then it does the harder, more honest thing: it identifies the three places where the technique is genuinely useful, draws the line cleanly, and gives you a synthetic-user policy doc you can hand your team so the seduction stops being an argument and becomes a settled rule.

The Demo That Feels Like Magic

Run the demo yourself, because the spell only breaks once you have felt it. Open Claude or ChatGPT and write: "You are Sarah, a 34-year-old marketing manager at a mid-size B2B SaaS company. You are evaluating project-management tools. I am going to ask you about your needs and frustrations. Answer as Sarah." Then interview it. Ask what frustrates her about her current tool, what she looks for in a new one, what would make her switch. The answers come back immediate, articulate, and plausible. Sarah hates context-switching. Sarah wants better reporting for her boss. Sarah is nervous about migration cost. It feels like research. It feels, honestly, better than a lot of real interviews, because Sarah is never inarticulate, never goes on a tangent about her dog, never says "I don't know, I guess it's fine."

That last part is the tell, and it is worth sitting with. Real users are inarticulate. They contradict themselves. They tell you the feature they desperately want and then never use it. They say the onboarding was fine and then you watch the recording and they rage-quit at step three. The mess is not noise to be filtered out - the mess is the data. The gap between what users say and what they do, the half-formed frustration they cannot name, the workaround they built without realizing it, the thing they did not mention because it never occurred to them it could be different: this is where real insight lives, and a synthetic user has none of it, because a synthetic user is not a person. It is a model generating the most plausible-sounding answer a marketing manager might give, which is a very different thing from what an actual marketing manager would actually do.

Why a Model Pretending to Be a User Is Not Data

The dismantling has to be precise, because "it's not real" is too easy to wave away. Here is the exact mechanism. When you ask a model to be Sarah, it does not simulate Sarah. It generates text statistically consistent with how the word "Sarah, marketing manager, evaluating PM tools" is described and discussed in its training data. Its answers are a reflection of how marketing managers are written about on the internet - in blog posts, marketing copy, case studies, product reviews, and other companies' persona decks - not how any marketing manager actually thinks, and certainly not how your users behave with your product.

This produces three specific, fatal problems. The first is that it has never used your product. It cannot tell you that your date picker is confusing because it has never touched your date picker; it can only tell you what date pickers are generically said to be like. Every answer about your actual interface is invented. The second is the absence of the say-do gap. Real research is valuable largely because it surfaces the difference between stated preference and revealed behavior, and a synthetic user collapses that gap to zero - it only has "say," generated to sound reasonable, with no "do" underneath. You learn what a plausible user would claim, which is exactly the least reliable part of real research, distilled and served as the whole meal. The third is convergence to the mean and the median. The model produces the average, most-discussed answer, which means it systematically misses the edge cases, the unusual mental models, the specific weird thing your specific users do that is often the most important finding. It cannot surprise you, and the entire point of research is to be surprised.

A synthetic user tells you how users like this are written about on the internet. Real research tells you what your users actually do with your actual product. Confusing the first for the second is how you ship a feature validated by a model that has never met a human.

The Confirmation-Bias Engine

There is a fourth problem severe enough to deserve its own section, because it is the one that makes synthetic users actively harmful rather than merely useless. A model role-playing a user is extraordinarily suggestible. Ask "Sarah, wouldn't this new dashboard make your reporting easier?" and Sarah, generating the agreeable, plausible response, says yes, it would. You have not learned anything about whether the dashboard helps; you have learned that the model will agree with a leading question, which you already knew. The synthetic user is a mirror that reflects your own hypothesis back at you in the voice of a customer, and a mirror that flatters is the most dangerous instrument in research.

This is qualitatively worse than no research, because no research at least leaves you uncertain and cautious. Synthetic-user research leaves you confident and wrong. A team that "validates" a feature against a synthetic user walks into the build with the warm feeling of having checked with customers, when what they actually did was ask a very fluent yes-machine to endorse a decision they had already made. The provenance gets laundered along the way - "we tested it with users" becomes the shorthand, the synthetic part drops off, and three sprints later the real feature ships to real users who behave nothing like Sarah did. The cost of synthetic users is not the time you spend on them; it is the false confidence they manufacture, which is a far more expensive thing.

The Three Legitimate Uses

Now the honest part, because a lesson that only says "never do this" is not useful and is not true. The technique has three legitimate uses, and what unites all three is that the model is never standing in for a user as a source of data. It is being used as a language tool, on text, where its actual capability - generating plausible language - is exactly the right capability and no claim about real human behavior is being made.

Use One: Stress-Testing Copy and Flows for Comprehension

The first legitimate use is checking whether your interface copy, error messages, and flow logic are comprehensible and where they might be misread. Prompt a model to read an error message "as a confused first-time user" and tell you every way it could be misinterpreted, and you get a useful list of ambiguities, because finding the plausible misreadings of a piece of text is a genuine language task the model is good at. You are not asking it to be a user; you are asking it to enumerate how a string of words could be parsed, which is squarely within its competence. The output is hypotheses about your copy, to be confirmed with real users or simply fixed if obvious - not findings.

Use Two: Pre-Flighting a Study Guide Before You Run It

The second is rehearsing a research session before you spend real participants on it. Have the model role-play a participant answering your discussion guide, and you will quickly surface leading questions, confusing instructions, dead-end branches, and questions that produce un-synthesizable answers. This is the same logic as the previous lessons: you are testing the instrument, not collecting data. If the synthetic participant cannot answer a question without confusion, a real one will struggle too, and you have caught a study-design flaw for free before it cost you a single recruited respondent. The value is entirely in the dry run; the synthetic answers themselves are discarded.

Use Three: Hypothesis Generation to Widen Your Thinking

The third is generating hypotheses to investigate, explicitly labeled as hypotheses. Ask the model to brainstorm "twenty reasons a user might abandon this checkout flow" and it will produce a wide list, some obvious, some you had not considered, drawn from the broad pattern of how abandonment is discussed everywhere. This widens your thinking and seeds your real research with things to look for. The critical discipline is the label: these are hypotheses to test against real users, never conclusions. The model is good at divergent generation and bad at telling you which of the twenty actually applies to your users - that second part is what the real research is for, and nothing about the synthetic list shortcuts it.

The Line, Stated Once So It Holds

The three legitimate uses share a single property, and naming it makes the boundary memorable enough to enforce. In every legitimate use, the model operates on text and produces hypotheses or finds flaws in your own artifacts; in every illegitimate use, the model stands in for a human and produces something you treat as a finding about real behavior. The legitimate uses ask "what could this text mean, what might be wrong with my question, what should I look into?" The illegitimate use asks "what do my users want, does my feature work, should I build this?" - questions only real humans interacting with your real product can answer.

The test you apply, every time someone proposes a synthetic-user technique, is one question: are we using the model's language ability on our own artifacts, or are we using its role-play as a substitute for evidence about real human behavior? The first is fine and often valuable. The second is theater, and the more convincing the theater, the more dangerous it is. Everything else - all the nuance about which prompts and which products and which models - reduces to this one line, and a team that internalizes it stops needing to relitigate the question per project.

The Synthetic-User Policy Doc

The named artifact is a one-page policy your team adopts, and its job is to convert a recurring, seductive argument into a settled decision so no one has to win it again every quarter. A workable policy has five parts.

The principle, stated plainly. "A model role-playing a user is not a source of data about real human behavior. It does not use our product, it has no say-do gap, it converges to the mean, and it agrees with leading questions. We do not treat synthetic-user output as research findings, and we do not describe it as 'testing with users.'" Stating it this baldly is the point; ambiguity is where the laundering happens.

The three permitted uses, named and bounded. Stress-testing copy and flows for comprehension; pre-flighting a study guide before recruiting; generating labeled hypotheses to investigate. Each with a one-line scope and the reminder that the output is hypotheses or flaws, never conclusions.

The disclosure rule. Any artifact informed by synthetic-user work labels it as such, and synthetic output is never aggregated into a findings doc alongside real-user data without an explicit marker. This prevents the provenance laundering where "we asked a synthetic Sarah" silently becomes "users told us."

The decision the policy makes for you. The one-question test - language-tool-on-our-artifacts versus substitute-for-real-evidence - written down, so any new proposal gets sorted in seconds rather than debated.

The escalation note. What to do when a stakeholder or a vendor pushes synthetic users as a replacement for real research: a short, non-defensive script that cites the four failure modes and offers the three legitimate uses as the constructive alternative, so the policy gives you a way to say no that also says yes to something real.

What the Vendors Are Actually Selling, and How to Read Their Claims

By 2026 there are funded products selling synthetic users as a research category, with confident marketing about validating designs against AI personas at scale. Read these claims through the line above and they sort themselves immediately. Where the product helps you pre-flight a study, enumerate copy ambiguities, or brainstorm hypotheses, it is doing a real and useful thing and the line endorses it. Where the product claims to replace user interviews, validate a design, or tell you what your users want, it is selling the most dangerous version of the theater, dressed up with enough scale and polish to make the false confidence feel like rigor. The scale makes it worse, not better - a thousand synthetic interviews converge to the same mean a single one does, with a thousand times the spurious authority.

This is not an argument against the vendors wholesale; it is an argument for reading their claims against the line and buying only the part that survives it. A senior IC's job here is to be the person who can sit in the room when the synthetic-user vendor demos, watch the CPO get excited about cutting research cost to zero, and say the calm, specific thing: "This is genuinely useful for testing our study guides and stress-testing copy, and we should use it for that. It cannot tell us what our users want, because it has never met our users, and if we treat it as if it can, we will ship features validated by a model that has never used our product." That sentence, delivered without anti-AI defensiveness and with a clear yes attached to a clear no, is the entire value of having understood this lesson.

Why This Matters More as the Models Get Better

The instinct is to assume this problem shrinks as models improve, and it is exactly backwards. A better model makes a more convincing Sarah - more fluent, more consistent, more emotionally plausible, more able to maintain character across a long interview. None of that touches the underlying problem, because the problem is not that the model is a bad actor playing a user. The problem is that it is an actor at all - it is generating plausible text about a user, not being one, and no amount of acting skill converts performance into the lived experience and actual behavior that research exists to capture. The model can get arbitrarily good at sounding like a user and remain exactly as useless as a source of data about your real users, because it has still never used your product and still has no behavior underneath the words.

So as the demos get more seductive, the discipline matters more, not less. The thing that protects you is not skepticism about whether the model is good - it is increasingly, genuinely good - but clarity about what kind of thing it is producing. It is producing language, and language about users is not the same as users, no matter how good the language gets. Hold that line and you can use the technique for everything it is genuinely good for while never once mistaking a fluent performance for the messy, contradictory, surprising truth that only real humans interacting with your real product can give you. That clarity, written into a policy your team actually follows, is what this lesson exists to leave you with.

Key Takeaways

  • The synthetic-user demo feels like magic because the model is articulate, available, and free - but that fluency is the tell. Real users are inarticulate, contradictory, and surprising, and that mess is the data; a synthetic user has none of it because it is generating plausible text, not being a person.
  • A model role-playing a user is not data, for four specific reasons: it has never used your product (every answer about your interface is invented), it has no say-do gap (only "say," generated to sound reasonable), it converges to the mean (missing the edge cases that are often the real finding), and it agrees with leading questions (a confirmation-bias mirror).
  • The confirmation-bias problem makes synthetic users worse than no research, not merely useless: no research leaves you cautious, while synthetic research leaves you confident and wrong, and the provenance launders into "we tested with users" along the way.
  • Three uses are genuinely legitimate, united by the model operating on text as a language tool rather than standing in for a human: stress-testing copy and flows for comprehension, pre-flighting a study guide before recruiting, and generating labeled hypotheses to investigate. In all three the output is flaws or hypotheses, never findings.
  • The line, stated once: are we using the model's language ability on our own artifacts (legitimate), or using its role-play as a substitute for evidence about real human behavior (theater)? The whole nuance reduces to this one question.
  • The synthetic-user policy doc converts a recurring seductive argument into a settled rule: the principle stated baldly, the three permitted uses bounded, a disclosure rule against provenance laundering, the one-question test written down, and a non-defensive escalation script for when someone pushes it as a replacement for real research.
  • This matters more as models improve, not less: a better model makes a more convincing Sarah, but it remains an actor generating language about a user, not a user with behavior underneath. The protection is clarity about what kind of thing the model produces, not skepticism about how good it is.