AI for Designers (UX, Product, Brand)
Proficient · M5 · lesson 5 of 27 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Designing an Unmoderated Test for AI Synthesis
📖
now learning

Designing an Unmoderated Test for AI Synthesis

15 min

Most unmoderated tests fail before the first respondent ever clicks anything, and they fail at the moment you write the question. You can run a flawless thirty-person Maze study, hand the recordings to the best synthesis model in 2026, and get back a confident, well-formatted summary that is quietly useless - not because the AI is bad at synthesis, but because you asked questions the AI had no way to synthesize. The fix is not a better prompt at the end. It is a better study design at the start. This lesson teaches you to write a Maze or UserTesting study whose question structure an AI can actually turn into findings: specific behaviors instead of vague opinions, observable moments instead of remembered impressions, and quantitative scaffolds the model can count rather than invent. You walk away with a real test plan and a synthesis-readiness checklist you run before you ever hit launch.

The Study That Synthesized Into Nothing

Start with the failure, because you have probably shipped it. A product designer needs to validate a new checkout flow before the Thursday review. They spin up a Maze study, recruit thirty respondents, and write the questions the way they have always written them: "What did you think of the checkout experience?" "How easy or hard was it to complete your purchase?" "Is there anything you would change?" Reasonable-sounding questions. Thirty people answer them. The designer feeds the whole thing - task completion data, click paths, and a pile of open-text responses - to Claude or Maze AI and asks for the top findings.

What comes back reads beautifully and says almost nothing. "Users generally found the checkout flow intuitive, though some experienced friction during payment. Several respondents suggested improvements to clarity and trust." Every word of that is defensible and none of it is actionable. The designer cannot tell which step caused the friction, how many people hit it, or what specifically broke. They go into the review with a summary that the CPO could have written from a stock photo of a checkout page. The study cost a week and produced a vibe.

Here is the part that matters: the AI did not fail. It synthesized exactly what it was given, and what it was given was thirty people's remembered impressions of a whole experience. You cannot count an impression. You cannot trace "intuitive" to a screen. The raw material was un-synthesizable, so the synthesis was vapor. The lesson the designer should take is not "AI synthesis is unreliable." It is "I designed a study that produced opinions when I needed behaviors, and the model can only ever be as specific as my questions let it be."

Why AI Synthesis Needs Different Questions Than a Human Moderator

A skilled human moderator can salvage a badly designed study live. When a participant says "the checkout felt confusing," a good moderator hears it, leans in, and asks "confusing how - what were you looking at when that happened?" The vague answer becomes a specific one through a follow-up the moderator improvises in real time. The whole craft of moderation is turning impressions into evidence on the fly.

An unmoderated test has no moderator. Nobody leans in. Whatever the respondent says or does is the entire record, and an AI synthesizing that record cannot improvise a follow-up - it can only work with what was captured. This changes the design constraint completely. In a moderated study, the question can be loose because the human tightens it live. In an unmoderated study synthesized by AI, the question structure has to do the tightening, because there is no second chance. The structure you build into the task is the only follow-up the study will ever get.

This is why "write a better prompt at synthesis time" is the wrong instinct. By the time you are prompting, the data is frozen. If the data is thirty remembered impressions, no prompt on earth extracts a behavioral finding from it, because the behavior was never observed - only the opinion about it was recorded. The leverage is entirely upstream. You are not designing a test and then synthesizing it. You are designing a test for synthesis, and the two are inseparable.

An AI can only synthesize what your study observed. Vague questions record opinions; opinions cannot be counted, traced, or fixed. The synthesis is decided at the moment you write the task, not the moment you run the prompt.

The Three Properties an AI-Synthesizable Question Has

Across the studies that synthesize cleanly and the ones that turn to vapor, the difference comes down to three properties. A question an AI can actually synthesize is specific about behavior, anchored to an observable moment, and scaffolded with something quantitative. Miss any one and the synthesis degrades toward the confident-but-empty summary above.

Property One: Specific Behaviors, Not Opinions

"What did you think of the navigation?" asks for an opinion. "Find where you would update your billing address, and tell us what you clicked first" asks for a behavior. The difference is that the second one produces a fact - the respondent either clicked the right thing or clicked the Account icon expecting billing to live there - and a fact is synthesizable. Twenty-eight of thirty people clicking the wrong thing first is a finding. "Most people thought the nav was fine" is not.

The reframe is mechanical once you see it. Every opinion question has a behavioral twin. "Was the error message clear?" becomes "After you saw the error, what did you do next?" - which reveals whether the message actually recovered them or sent them in a circle. "Did you trust the payment screen?" becomes "Before you entered your card, did you look for anything? What were you looking for and did you find it?" - which surfaces the missing trust signal by its absence. You are trading a rating you cannot act on for a behavior you can. The AI can count behaviors. It cannot count feelings, and when you ask it to, it averages adjectives into mush.

Property Two: Observable Moments, Not Remembered Impressions

The second property is timing. A question that asks someone to summarize a whole experience after the fact - "How was the overall flow?" - collects a remembered impression, and memory is lossy, generous, and recency-biased. People forget the moment they got stuck and remember that they eventually succeeded, so they rate a painful flow as "pretty good." The AI synthesizing those ratings inherits all the distortion and presents it as data.

The fix is to attach questions to specific, observable moments in the task rather than to the experience as a whole. Maze and UserTesting both let you place a follow-up immediately after a particular step or screen. A question fired right after the payment step - "What just happened on that screen?" - catches the moment while it is still real, before memory smooths it over. Even better, the behavioral data itself is the observable moment: the click path, the time-on-task spike, the misclick on step three, the rage-quit before completion. These are observed, not remembered, and the AI can synthesize them directly because they are events with timestamps, not stories told after the fact. Design the study so the most important moments are captured as events, and the synthesis has hard ground to stand on.

Property Three: Quantitative Scaffolds the Model Can Count

The third property is giving the AI something to count. A pile of open-text answers forces the model to invent its own categories ("several respondents mentioned trust"), and invented categories are where hallucinated precision creeps in - "several" might be three or might be eleven, and the model will pick whichever reads well. A quantitative scaffold removes the invention. If every respondent rates task difficulty on a one-to-five scale at the same observable moment, the model is no longer estimating; it is reporting that nineteen of thirty rated the payment step a four or five for difficulty, and that is a number you can verify and defend.

The scaffolds that synthesize well are the ones tied to behaviors and moments: a single-difficulty rating after each task (the SEQ, or Single Ease Question, is the workhorse here), a success/fail flag per task that Maze captures automatically, a misclick count, a time-on-task figure. These give the open-text answers something to attach to. Now "users found payment confusing" becomes "the payment step had a 63 percent success rate and a median SEQ of 2.1, and the open-text answers cluster on the CVV field" - a finding with a number, a behavior, and a quote, which is the shape of every finding that survives a design review. The scaffold does not replace the qualitative; it anchors it so the AI cannot drift.

The Test Plan, Structured for Synthesis

Now assemble these into a real test plan. The named artifact for this lesson is a test plan plus a synthesis-readiness checklist, and the test plan has a specific shape when it is built for AI synthesis. Take a real example: validating that a redesigned checkout flow reduces payment-step abandonment.

Objective, stated as a behavior you can observe. Not "validate the new checkout." Rather: "Determine whether users can complete payment without hesitation, and identify any step where success rate drops below 85 percent or median difficulty exceeds 2.5." The objective itself is now measurable, which means the synthesis has a target to report against.

Tasks, each a single observable behavior. Break the flow into discrete tasks: "Add the highlighted item to your cart." "Proceed to checkout and enter the shipping address." "Complete the payment using the test card provided." Each task is one behavior with a clear success condition Maze can flag automatically, so the AI gets a per-task success rate without you doing anything.

A scaffold after every task. Immediately after each task, a single difficulty rating (SEQ, one to five) and one targeted open-text question tied to that exact moment - "What, if anything, slowed you down on that step?" The rating gives the model a number to count; the targeted open-text gives it quotes that attach to a specific step rather than floating free.

One behavioral probe at the highest-risk moment. At the payment step, the riskiest moment, add a probe that captures behavior rather than opinion: "Before submitting, did you check anything on this screen? If so, what?" This surfaces trust-seeking behavior (or its absence) at the exact place abandonment happens, as an observed action rather than a remembered feeling.

Notice what this plan refuses to do. It does not ask "what did you think of the experience?" anywhere. It does not collect a single overall-impression rating. Every question is bolted to a task, a moment, and a number. When thirty people run it, the AI receives a structured record - success flags, SEQ scores, time-on-task, misclick counts, and step-anchored quotes - and the synthesis it produces is specific because the data was specific. You engineered the finding's specificity into the study before launch.

The Synthesis-Readiness Checklist

Here is the artifact you run before you ever hit launch. It is six questions, and any "no" sends you back to the test plan, because fixing it now costs five minutes and fixing it after thirty people have responded costs a re-run.

  1. Behavior, not opinion? Does every task ask the respondent to do something with a clear success condition, rather than to rate or opine on something? If a question can be answered without taking an action, rewrite it as the action.
  2. Observable moment, not remembered impression? Is each follow-up fired at the specific step it asks about, rather than at the end of the whole flow? Move any "overall" question to the moment it actually concerns, or cut it.
  3. A number to count? Does each task carry a quantitative scaffold - a difficulty rating, a success flag, a time or misclick capture - so the AI reports counts instead of inventing words like "several"?
  4. Quotes that attach? Is every open-text question tied to a specific step, so the answers cluster on something rather than floating as general commentary the model has to categorize blindly?
  5. A target to report against? Does the objective state a threshold (success rate, difficulty score) the synthesis can pass or fail each step against, rather than a vague "validate"?
  6. Would two people synthesize it the same way? If you handed this study's raw output to two different analysts, would they reach the same findings? If the data is ambiguous enough that they would diverge, the AI will diverge too, and you have not designed for synthesis yet.

The sixth question is the whole checklist in one move. Synthesizability is just inter-rater reliability with a model as one of the raters. If the data forces a single obvious reading, the AI gives you that reading; if the data is loose, the AI fills the gaps with plausible-sounding invention. The checklist exists to force the data into the obvious-reading shape before a single respondent touches it.

What This Does and Does Not Fix

Be honest about the boundary. Designing for synthesis makes the AI's job possible; it does not make verification optional. Even a perfectly structured study still requires you to sample-verify the synthesis against the raw recordings, because the model can still mis-cluster a quote or over-weight a vocal minority - that is the next lesson in this chapter. What synthesis-ready design buys you is that the verification becomes cheap and fast, because you are checking specific claims ("the payment step had a 63 percent success rate") against specific data, rather than trying to confirm whether a vibe like "generally intuitive" is true, which you cannot do at all.

It also does not turn every research question into an unmoderated test. Some questions are genuinely about feeling, motivation, and the messy why behind behavior, and those still want a human moderator who can follow the thread live. The synthesis-ready unmoderated test is the right tool when the question is about observable behavior at volume - can people do this, where do they get stuck, how often - which is a large and important slice of the work, and the slice where AI synthesis genuinely multiplies your throughput. Knowing which questions belong in which bucket is itself a craft decision, and it is yours, not the tool's.

The Deeper Move: You Are Designing the Data, Not Just the Test

Step back and the principle generalizes past usability testing. Whenever an AI is going to synthesize something you collect, the quality of the synthesis is capped by the structure of the collection, and that structure is a design decision you own. A research study, a survey, a feedback form, a set of interview prompts - each one is a data-generating instrument, and you can build it so the output is synthesizable or build it so the output is mush, and the AI cannot tell the difference until it is too late to fix.

The senior move is to think one step ahead of the synthesis: before you write a single question, picture the finding you need to be able to defend in the review, and reverse-engineer the question structure that would produce the evidence for it. If you need to say "the payment step is the abandonment point," you need per-step success rates, which means per-step tasks with success conditions, which means you cannot write "how was checkout?" anywhere. The finding dictates the data, the data dictates the questions, and you design backward from the claim you will have to stand behind. That is what it means to design an unmoderated test for AI synthesis: not to add AI to the end of an old process, but to build the whole instrument so the machine can do the part it is good at and you can defend the part that ships.

Key Takeaways

  • An AI can only synthesize what your study observed. Vague, opinion-based questions record impressions that cannot be counted, traced to a screen, or fixed - so the synthesis comes back confident and empty no matter how good the model is. The failure is in the study design, not the AI.
  • Unmoderated tests synthesized by AI need tighter questions than moderated ones, because there is no moderator to improvise the follow-up. The structure you build into the task is the only follow-up the study will ever get.
  • A synthesizable question has three properties: it asks for a specific behavior (with a success condition), it is anchored to an observable moment (fired at the step, not after the whole flow), and it carries a quantitative scaffold (a difficulty rating, success flag, or time/misclick capture the model can count).
  • Build the test plan backward from the finding: state the objective as a measurable threshold, break the flow into single-behavior tasks with auto-flagged success conditions, add a difficulty rating and a step-anchored open-text question after each, and place one behavioral probe at the highest-risk moment.
  • Run the six-question synthesis-readiness checklist before launch. The deciding question is "would two analysts synthesize this the same way?" - synthesizability is just inter-rater reliability with the model as one rater, and loose data makes the model invent.
  • This makes verification cheap, not optional, and it does not turn every research question into an unmoderated test. It is the right tool for observable behavior at volume, which is exactly where AI synthesis multiplies your throughput.