AI for Designers (UX, Product, Brand)
Capable · M14 · lesson 14 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Prompt-to-Prototype A/B Testing Without Anchoring
📖
now learning

Prompt-to-Prototype A/B Testing Without Anchoring

15 min

Type one prompt into Galileo, watch it generate a polished prototype in nine seconds, feel the small relief of having something to react to, and commit to that direction before lunch. That is the anchor trap, and almost every designer falls into it because the first generated direction arrives with a kind of authority the blank canvas never had. This lesson is about refusing that authority on purpose. You will run one prompt through three different tools, Galileo, UX Pilot, and Visily, to force a comparison the first output would have skipped, then run a structured A/B test with five colleagues between the two strongest directions. The artifact is an A/B test report with a clear go or no-go and a debrief line that designers find oddly hard to write: "we picked the model we would have rejected on Tuesday."

The First Direction That Quietly Became the Only Direction

A product designer needs a concept for a new dashboard empty state. They open Galileo, type "empty state for an analytics dashboard, friendly, encourages the user to connect their first data source," and nine seconds later there it is: a centered illustration, a warm headline, a single primary button. It is good. Not amazing, but good, and good arrived instantly with zero effort, which makes it feel like a gift. The designer tweaks the copy, adjusts the spacing, picks an icon. By the time they look up, two hours have passed and they are no longer evaluating a direction. They are polishing it. The decision to use this direction was never made; it was inherited from the first thing the tool happened to produce.

This is the anchor trap, and it is more insidious than the beauty bias from sketch-to-variant work because there is no second option in the room to lose to. The first generated direction does not win a comparison; it skips the comparison entirely. The designer never asked "compared to what?" because there was no "what" on screen. The model produced one plausible answer, the answer looked finished, and finished-looking is psychologically almost indistinguishable from decided. Two weeks later, in a review, someone asks "did we consider a more guided, step-by-step empty state instead of a single button?" and the honest answer is no, not because it was rejected, but because it never existed to be considered.

The cost of the anchor trap is not that the first direction is bad. Often it is fine. The cost is that "fine, and first" beats "better, but never generated," every single time, and you have no way of knowing what you gave up because you never put it on the wall. The discipline this lesson teaches is to manufacture the comparison the single-prompt workflow erases, so that whatever you ship, you shipped it because it won, not because it arrived first.

The first generated direction does not win the argument. It avoids the argument. Anchoring is not choosing the wrong option; it is never noticing there was a choice.

Why Anchoring Is Worse With AI Than It Ever Was By Hand

Anchoring is an old cognitive bias; the first number, the first sketch, the first estimate pulls everything after it toward itself. What changed in 2026 is that AI made the first option arrive faster, look more finished, and cost less effort than any first option in the history of design work. Each of those three properties deepens the anchor independently.

Speed deepens it because there is no struggle to produce the first direction, and struggle used to be the thing that kept you honest. When generating a concept took an hour, you naturally held it loosely, because you knew it was one expensive attempt among possible others. When it takes nine seconds, you skip the part of your brain that asks "is this the right direction" and go straight to "let me refine this direction," because refining is the path of least resistance and the tool made the direction free.

Finish deepens it because a polished output triggers a completion instinct. A rough sketch invites alternatives; it visibly says "I am one idea, there could be others." A finished-looking prototype says "I am the answer," and the more production-ready Galileo, UX Pilot, and Visily make their output, the more strongly it broadcasts that false finality. You are not anchoring on an idea; you are anchoring on something that looks like a decision that has already been made.

Low effort deepens it through a sunk-cost inversion: because the direction cost you nothing to generate, it feels wasteful to throw it away and start over, which is backwards. The thing you got for free should be the easiest to discard, but the immediacy of having something makes the blank alternative feel like a step backward. Understanding these three forces is half the defense, because once you can name why the first output feels authoritative, you can deliberately discount that authority.

The Three-From-Three-Models Pattern

The structural fix is simple and slightly annoying, which is why it works: never let one tool produce the only direction. Take your single prompt and run it, as close to identically as the interfaces allow, through three different tools. Galileo, UX Pilot, and Visily are a good trio precisely because they have different defaults, different training biases, and different opinions about what a screen should be, so the same prompt produces three genuinely different directions rather than three versions of one. The friction of opening three tools is the point. It forces the comparison the single-prompt workflow erased, and it does it before you have anchored on any one output.

The reason to use three different models rather than three runs of the same model is that one model has one center of gravity. Ask Galileo for the same thing three times and you get three variations clustered around Galileo's idea of an empty state. Ask three different tools and you get three different centers of gravity, which is what divergence actually requires. Each tool's bias becomes a feature: Galileo tends toward polished hi-fi, UX Pilot tends toward flow and structure, Visily tends toward clean wireframe-level clarity, and the spread between them is the design space you would otherwise never have seen.

Keeping the Prompt Constant Is the Whole Experiment

For the three-from-three pattern to teach you anything, the prompt has to stay constant across the three tools, because the variable you are studying is the tool's interpretation, not your prompt engineering. If you tune the prompt differently for each tool, you have confounded the experiment and you no longer know whether the differences come from the tools or from your edits. Write one prompt, paste it into all three, and accept that each tool will interpret it through its own lens. Those three interpretations are your three directions. Resist the urge to "help" the weaker output by improving its prompt; the weak output is data, and sometimes the tool you helped least produces the direction that wins, which is exactly the surprise the pattern exists to surface.

The Structured A/B Test With Five Colleagues

Three directions is one too many to decide between cleanly, so the next move is to narrow to the two strongest on your own judgment and then take those two to a structured A/B test with five colleagues. Five is a deliberate number: enough that you are not deciding alone, few enough that you can run it in twenty minutes and the signal does not drown in scheduling. The colleagues do not need to be designers; in fact a mix of a PM, an engineer, and a couple of designers gives you a more honest read than five people who all share your aesthetic training.

The structure matters because an unstructured "which do you like better" just relocates the beauty bias from your eyes to theirs. Instead, give the two directions neutral labels (A and B, never "the Galileo one" and "the Visily one," because naming the tool reintroduces bias), and ask each colleague a fixed set of task-oriented questions. Show them the user goal first, then both directions, then ask: which one gets you to the goal faster, and why; where would you get stuck on each; which one's primary action did you notice first; and which would you trust more if it were a real product. Record the answers verbatim, the same verbatim discipline from the research lessons, because "B felt cluttered" is a finding and "I like A" is noise.

Why Blind Labels Protect the Result

The single most important methodological choice in this test is to strip the tool names before showing the directions. The moment a colleague knows direction A came from "the expensive hi-fi tool" and B came from "the wireframe tool," their judgment is contaminated by their priors about the tools, by office politics about which tool the team prefers, and by the same finish-equals-quality bias you are trying to defeat. Blind labels force them to evaluate the direction on its merits for the user task. This is the same reason a wine tasting hides the labels: not because the label is irrelevant to the final decision, but because it corrupts the judgment of the thing in front of you. You can reveal the tools afterward, in the debrief, where the reveal becomes the most instructive part.

The A/B Test Report With Go or No-Go and the Tuesday Debrief

The deliverable is an A/B test report a stakeholder can read in two minutes and trust. It has four parts. First, the setup: the one prompt, the three tools it ran through, and the two directions that advanced, with the third documented and a one-line reason it did not. Second, the results: the five colleagues' task-oriented answers, summarized into which direction won on each question, with the sharpest verbatim quotes kept intact. Third, the verdict: a clear go or no-go naming which direction advances and, critically, which tool produced it. Fourth, the debrief, which is where this report earns its keep.

The debrief line is the one that designers find genuinely uncomfortable to write, and that discomfort is the signal that the process worked: "we picked the model we would have rejected on Tuesday." It means the direction that won the structured, blind A/B test was not the one you anchored on, not the polished Galileo output you would have refined into the ground if you had run the single-prompt workflow. The wireframe-level Visily direction, or the structurally-clearer UX Pilot flow, won on the user task once five people evaluated it blind. Writing that sentence is an admission that your instinct anchored you toward the wrong direction and the process corrected you. A report that contains that sentence is proof the anchor trap was sprung and survived; a report whose winner is exactly the direction you started with should make you suspicious that the test merely ratified your anchor.

When the First Direction Genuinely Wins

Sometimes the direction you anchored on does win the blind test fairly, and the discipline does not require you to pretend otherwise. The point of the three-from-three pattern and the A/B test is not to guarantee that the first direction loses; it is to make sure that whichever direction ships, it shipped because it won a fair comparison rather than because it arrived first. If the Galileo output you generated first wins the blind five-colleague test on task speed and clarity, advance it with confidence, and write the honest debrief: "the first direction won, and now we know it won rather than assuming it." That report is just as valuable, because it converts an anchored guess into a verified choice. The trap is not choosing the first direction; the trap is never finding out whether you should have.

The Three Anchoring Failures to Catch

Across many prompt-to-prototype workflows, three failures recur. Catch them by name.

Failure One: The Single-Tool Monopoly

You run the prompt through one tool, get one direction, and refine it without ever generating an alternative. The comparison never happens because there is nothing to compare to. Catch it with the three-from-three rule: no direction advances until it has at least two siblings from different tools sitting beside it. If there is only one option on the wall, you have not designed, you have accepted.

Failure Two: The Loaded Label

You take two directions to colleagues but tell them which tool made each, and their feedback reflects their tool priors rather than the design's merit. Catch it by stripping tool names to blind A and B labels before any colleague sees the directions, and only revealing the tools in the debrief after the verdict is recorded.

Failure Three: The Ratification Test

You run an A/B test but design it so loosely (vague questions, your own framing, a friendly audience) that it simply confirms the direction you already preferred. A test that can only ever agree with you is theater. Catch it by using task-oriented questions tied to the user goal, recording verbatim answers, and treating a result that matches your anchor as a prompt for extra scrutiny rather than relief.

The Meta-Anchor on the Tool That Keeps Winning

There is one more anchor lurking a level up, and experienced practitioners of this pattern eventually trip over it. If Visily wins your blind tests three sprints in a row, you will start to think of Visily as "the good tool," and you will begin favoring its output before the test even runs, which quietly recreates the single-tool monopoly you worked so hard to break. The guard is to re-earn the verdict every time: keep the prompt constant, keep the labels blind, and run the comparison regardless of which tool you expect to win. The lesson's own example is the antidote in miniature, the wireframe-level tool wins for a fast settings task, not universally, so a tool that wins one brief carries no presumption on the next. Anchoring does not disappear when you defeat it once; it relocates, and the discipline has to be reapplied at every level it reappears, including the choice of which three tools you put in the ring.

Putting It to Work This Week

Take a real screen you are about to design and refuse the single-prompt workflow. Write one prompt, run it through Galileo, UX Pilot, and Visily without tuning it per tool. Narrow to the two strongest directions on your own judgment, strip the tool names, and run a twenty-minute blind A/B test with five colleagues using task-oriented questions tied to the user goal. Record the verbatim answers. Write the report with a clear go or no-go that names the winning tool, and write the debrief line honestly, even when it reads "we picked the model we would have rejected on Tuesday."

You will know the practice has landed the first time the blind test sends you to a direction your gut had quietly dismissed, and you can point at five colleagues' task-based answers to explain why the gut was wrong. The tool will always hand you a fast, finished, free first direction, and that direction will always feel like a decision that has already been made. Your job is to treat it as one option among three, force the comparison the workflow tried to skip, and ship the direction that won rather than the one that arrived first.

Key Takeaways

  • The anchor trap is accepting the first generated direction not because it beat alternatives but because it skipped the comparison entirely; the cost is that "fine, and first" beats "better, but never generated," and you never learn what you gave up.
  • AI deepens anchoring three ways: speed removes the struggle that kept you honest, finish triggers a false completion instinct, and low effort creates a sunk-cost inversion where the free thing feels wasteful to discard.
  • Use the three-from-three-models pattern: run one constant prompt through Galileo, UX Pilot, and Visily so three different centers of gravity produce three genuinely different directions, rather than three variations of one tool's idea.
  • Keep the prompt constant across tools so the variable you study is the tool's interpretation, not your prompt engineering; the output you helped least sometimes wins, which is the surprise the pattern exists to surface.
  • Run a structured, blind A/B test between the two strongest directions with five colleagues, using task-oriented questions tied to the user goal and stripping tool names to A and B so tool priors and finish-equals-quality bias cannot contaminate the read.
  • Ship an A/B test report with a go or no-go that names the winning tool and a debrief line: "we picked the model we would have rejected on Tuesday." That uncomfortable sentence is proof the anchor was sprung; if the winner is exactly your starting direction, scrutinize whether the test merely ratified your anchor.