The Forty-Concept Sprint: Ideating Onboarding Flows in 30 Minutes
Forty onboarding concepts in thirty minutes sounds like a party trick, and if you stop at forty it is one. The actual skill, the one that makes this an AI-augmented design practice instead of a slot machine, is the discarding. Generating volume is the easy half; AI does it for free. The hard half, the half that separates a designer from a prompt-typer, is keeping six and killing thirty-four with a stated reason for each kill. This lesson uses Whimsical AI, Sticky-AI, and Tldraw with Makereal to generate forty onboarding-concept thumbnails fast, then teaches the kill-criteria discipline that turns a pile of variations into a concept board you can defend, with an explicit reason next to every rejection.
The Board That Was All Volume and No Decision
A designer runs the sprint, generates forty onboarding thumbnails, and proudly drops the whole board into a team channel. Forty concepts. The reaction is not admiration; it is paralysis. A PM scrolls, glazes over around concept twelve, and asks the question that deflates the whole exercise: "okay, but which one are we doing?" The designer does not have an answer, because they did the generation and skipped the design. Forty undifferentiated concepts is not forty times the value of one. It is roughly the same value as zero, because nobody can act on a wall of options with no point of view.
This is the central trap of AI-augmented ideation. Volume feels like productivity. A full board looks like a hard day's work. But volume without curation is noise, and noise is worse than silence because it costs other people time to wade through. The deliverable of an ideation sprint is never the forty. It is the six you kept and, just as importantly, the thirty-four you killed and why. The kills are not waste; they are the visible evidence that a designer made decisions, and decisions are the thing the team actually needed.
Anyone can generate forty concepts now. The job was never the generating. The job is the killing, and the reason written next to each kill is the artifact that proves you designed instead of just prompted.
Why Volume Is Genuinely the Easy Half Now
It is worth being honest that the generation really is trivial in 2026, and pretending otherwise is nostalgia. Whimsical AI will spin up flow variations from a one-line prompt. Sticky-AI fills a board with concept stickies faster than you can read them. Tldraw with Makereal turns rough boxes into clickable approximations in seconds. The marginal cost of the fortieth concept is essentially zero, which is exactly why generating forty is not an achievement. When something is free, having a lot of it is not a signal of effort or quality.
This inverts where your time and credibility should go. In the old world, generating ten concepts was the work, and you had little time left to evaluate. In the new world, generating forty takes minutes, and all of your time and judgment goes into evaluation. The sprint is not a generation exercise with a little evaluation at the end. It is an evaluation exercise with a little generation at the start. If you find yourself proud of the forty, you have misplaced the work.
The Thirty-Minute Sprint Structure
Here is the time-boxed structure that produces a defensible board, not just a full one.
Minutes 0 to 3, write the brief and the kill criteria first. Before generating anything, write one sentence stating the user need this onboarding must serve and three to five kill criteria, the conditions under which a concept is automatically out. Writing kill criteria before you see the concepts is essential, because it stops you from rationalizing whatever the model happens to produce. Criteria written after the fact bend to fit the output.
Minutes 3 to 13, generate forty thumbnails. Use Whimsical AI, Sticky-AI, and Tldraw with Makereal. Deliberately vary the prompt across the three tools so you get range, not forty variations of the same idea. The goal of generation is coverage of the possibility space, not polish.
Minutes 13 to 25, kill thirty-four. Run every concept against the kill criteria. Most will fail one immediately. Drag each killed concept to a "killed" zone and write the one-line reason. This is the bulk of the real work, and it should feel like the hardest part, because it is.
Minutes 25 to 30, sharpen the six survivors. For the six you kept, write one line on what each is betting on, the distinct hypothesis it represents. Six concepts that each make a different bet are worth infinitely more than forty that blur together.
What Makes a Good Kill Criterion
A kill criterion is specific, tied to the user need, and testable on sight. "Requires account creation before the user sees any value" is a good kill criterion: you can check it against any thumbnail in two seconds, and it encodes a real conviction about onboarding. "Feels off" is a bad one: it is vague, unfalsifiable, and unteachable. Good kill criteria read like design principles with teeth. They might include "more than three fields before first value," "no clear single next action," "relies on a tour the user must sit through," or "assumes desktop-only interaction." Each one lets you kill fast and explain why.
The Concept Board: Six Survivors, Thirty-Four Reasons
The deliverable is a concept board built around the kills, not the keeps. It has three zones. The kept zone holds the six survivors, each with a one-line bet ("this one bets that users will trust a product faster if they can use it before signing up"). The killed zone holds the thirty-four, each with its one-line kill reason and the criterion it violated. And a header holds the brief and the kill criteria themselves, so anyone reading the board can see the rules you judged against.
The killed zone is the part that earns the board its credibility. When a stakeholder asks "did you consider a video walkthrough," you point to the killed zone where three video-walkthrough concepts sit, each tagged "killed: relies on a tour the user must sit through, violates the no-passive-onboarding criterion." You did consider it. You killed it on purpose, for a stated reason, and here is the receipt. That transforms the conversation from "why didn't you think of X" to "here is why X loses," which is the conversation a designer wants to be having.
The Kill Reason Is the Real Output
It bears repeating because it is the entire point: the kill reasons are the real output of this sprint. They are the proof of design judgment. A board with six keeps and no kill reasons is indistinguishable from luck. A board with thirty-four stated kill reasons is a documented argument about what this onboarding should and should not be. The reasons are also reusable: next sprint, your kill criteria are sharper because you have seen what they catch, and your team starts internalizing the principles, which is how taste spreads through an organization.
The Three Sprint Failures to Avoid
Across many AI ideation sprints, three failures recur. Avoid them deliberately.
Failure One: The Forty Clones
The model, prompted lazily, produces forty variations of essentially one idea, differing only in color and copy. The board looks full but the possibility space is barely covered. Avoid it by varying the prompt hard across the three tools and across framings ("what if onboarding were a single screen," "what if there were no onboarding," "what if the team set it up, not the user"). Range, not repetition, is the point of volume.
Failure Two: The Retroactive Criteria
You skip writing kill criteria up front, generate the forty, and then invent reasons that conveniently justify keeping the concepts you already liked. This is rationalization wearing the costume of judgment. Avoid it by writing the criteria first, before a single concept exists, so the criteria judge the concepts and not the other way around.
Failure Three: The No-Kill Board
You keep fifteen because killing feels wasteful and every concept has "something." A board of fifteen lukewarm keeps is a failure to decide. The discomfort of killing a concept you mildly like is exactly the discomfort of designing. Avoid it by committing to a hard keep number (six) before you start, so the constraint forces real prioritization instead of polite hedging.
When to Generate More, and When to Stop
Forty is a default, not a law. The real stopping rule is coverage of the possibility space, not a count. If your forty concepts cluster into only two genuinely different approaches, you have not generated enough range and you should prompt for more divergence, not more volume. If your forty already span six distinct strategic bets, you have plenty, and generating more would just add clones to kill. Generate until the space is covered, then stop and switch entirely into evaluation mode. The signal to stop generating is when new concepts stop surprising you.
This is the L2 discipline applied to divergence. The model supplies the volume and the surprise; you supply the judgment that turns volume into a decision. The kill criteria are your verification mechanism, the same role the verification log played in research synthesis. They are how you prove the decision was made on principle, not on whim, and they are what makes a forty-concept sprint a design act rather than a generation stunt.
Putting It to Work This Week
Take a real onboarding problem and run the sprint exactly. Write the brief and kill criteria first. Generate forty thumbnails across Whimsical AI, Sticky-AI, and Tldraw with Makereal, varying the framing for range. Kill thirty-four against the criteria, writing a one-line reason for each. Sharpen the six survivors into six distinct bets. Present the killed zone as proudly as the kept zone.
You will know it is working when a stakeholder stops asking "which one" and starts engaging with the bets, and when "did you consider X" gets answered by pointing at the killed zone instead of going back to the drawing board. The forty was never the achievement. The thirty-four reasons are. Generate freely, kill ruthlessly, and let the kill reasons be the artifact that proves a human designed this, not a model that happened to fill a board.
Prompting for Range, Not Repetition
The quality of your forty concepts is set almost entirely in the first three minutes, by how you prompt, and most designers prompt for repetition without realizing it. If you ask any of these tools for "onboarding flow concepts for our app," you will get forty polite variations on the single most statistically common onboarding pattern, because that is what the model averages toward. The board fills, but the possibility space barely moves. The fix is to prompt along deliberately orthogonal axes, forcing the tools to explore directions they would never reach on their own. Instead of asking forty times for "an onboarding flow," ask for an onboarding flow that assumes the user is in a hurry, then one that assumes the user is skeptical, then one that assumes the user was invited by a teammate, then one that assumes the user has used three competitors already. Each framing pulls the generation into a different region of the space.
A useful technique borrowed from traditional ideation is to seed the prompts with provocations, the more extreme the better, because extreme provocations produce range even when most of the results get killed. "Design an onboarding with no onboarding screen at all." "Design an onboarding the user's manager completes for them." "Design an onboarding that is a single irreversible decision." "Design an onboarding that teaches by letting the user break something safely." Most of these will violate a kill criterion and land in the killed zone, and that is fine, because the job of a provocation is not to win; it is to push the surviving concepts further from the obvious than they would otherwise be. When you vary the provocation across Whimsical AI, Sticky-AI, and Tldraw with Makereal, you also exploit the fact that each tool has a different default aesthetic and structure, which adds range for free. The discipline is to spend your prompt budget on coverage, never on polish, because polish is what the evaluation phase is for and coverage is the only thing generation can give you that you cannot add later.
The Anti-Anchoring Benefit of Volume
There is a deeper reason the forty matters beyond coverage, and it connects directly to the next lesson on the anchor trap. When you generate one concept and refine it, you anchor hard to that first idea, and every subsequent decision is a small adjustment to a starting point you never chose deliberately. Volume breaks the anchor. When forty concepts sit on the board at once, no single one has the privileged status of "the idea," and you are forced to choose among genuine alternatives rather than polishing the first thing you saw. This is the real psychological gift of cheap generation: not that you get more options, but that you are protected from committing to the first option before you have seen the field. The designer who generates forty and kills thirty-four has, almost as a side effect, inoculated themselves against the most common failure in AI-assisted design, which is mistaking the first plausible output for the best available direction.
From Thumbnails to the Next Fidelity Without Losing the Discipline
A concept board is a divergence artifact, and the danger comes at the handoff to convergence, when the six survivors have to become two or three things worth prototyping. The kill discipline does not stop at the board; it carries forward, and forgetting that is how a clean sprint turns into mush a week later. The mistake is to treat the six survivors as six things to build, or worse, to let stakeholders cherry-pick favorites and quietly resurrect a killed concept because someone senior liked the look of it. Protect the board's decisions into the next phase. When a killed concept comes back, it should have to clear the criterion it failed, in writing, not just win on enthusiasm. The killed zone is not a graveyard; it is a record of arguments you already had and resolved, and reopening one should cost something.
The convergence move is to take the six survivors, each carrying its distinct bet, and ask which bets are actually testable against real users and which are testable only against opinion. A bet like "users will trust the product faster if they can use it before signing up" is testable; you can prototype both and measure activation. A bet like "this feels more premium" is not testable without a lot of careful work, and should be flagged as such rather than smuggled in as fact. Rank the survivors by how cheaply you can learn whether their bet is right, and prototype the two or three where the learning is cheapest and the stakes are highest. This keeps the sprint honest all the way through: you generated for range, you killed on principle, and you converge on the bets you can actually verify rather than the ones that merely sound good in a review. That straight line from divergence through documented kills to a testable hypothesis is what makes the whole exercise a design act rather than a brainstorm that produced a pretty wall.
Key Takeaways
- Generating forty onboarding concepts with Whimsical AI, Sticky-AI, and Tldraw with Makereal is the easy half and essentially free; the design work is keeping six and killing thirty-four with a stated reason for each.
- Volume without curation is noise, and noise is worse than silence because it costs the team time; a wall of forty undifferentiated concepts has roughly the value of zero because nobody can act on it.
- Write the kill criteria before generating anything, so the criteria judge the concepts rather than bending to justify whatever the model produced; good criteria are specific, tied to the user need, and testable on sight.
- The sprint is an evaluation exercise with a little generation at the start, not the reverse; if you are proud of the forty rather than the six and their kill reasons, you have misplaced the work.
- The three sprint failures to avoid: the forty clones (volume without range), the retroactive criteria (rationalization disguised as judgment), and the no-kill board (fifteen lukewarm keeps that fail to decide).
- Ship a concept board with three zones: six survivors each tagged with their distinct bet, thirty-four kills each tagged with a reason and the criterion it violated, and a header showing the brief and criteria. The kill reasons are the real output and the receipt that proves you designed.
Skill.re