AI for Designers (UX, Product, Brand)
Proficient · M11 · lesson 11 of 27 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Mapping the Real Sprint: Where AI Helps, Where It Hurts
📖
now learning

Mapping the Real Sprint: Where AI Helps, Where It Hurts

15 min

By L3 the question is no longer "should I use AI in my design work?" - that argument is over and you lost it on both sides, because the answer is "yes, but not everywhere." The real question, the one that separates an AI-integrated designer from a prompt-monkey, is surgical: which exact steps of your actual two-week sprint does AI accelerate, which does it quietly degrade, and where is the line a generated artifact must not cross without a human between it and a user? This lesson hands you a method to answer that for a real sprint, not a hypothetical one. You map every step of a sprint you actually ran onto a five-column board - definitely manual, manual-with-AI-check, AI-with-manual-check, AI-led, definitely AI - and you justify every single placement out loud, as if your design lead is sitting across the table asking "why is research synthesis in column three and not column four?" The artifact is a sprint-flow audit you can defend. The skill is the judgment that produces it.

Why a Map, and Not a Policy

Most teams that try to "figure out AI" reach for a policy. A document that says "we use AI for ideation and research, we do not use it for final visual design," pinned in Notion, read once, obeyed never. Policies fail at this because they operate at the wrong altitude. The interesting decisions in design work are not at the level of disciplines ("research," "visual design"), they are at the level of steps ("pulling verbatim quotes from eight transcripts" versus "deciding which three pain points the readout leads with"). Within a single afternoon of research synthesis, one step is something AI does better than you and another step is something that, if you let AI do it, ships a worse readout faster. A discipline-level policy cannot see that. A step-level map can.

The map is also honest in a way a policy is not. A policy is aspirational - it says what you wish were true. A map of a sprint you actually ran is forensic - it says what was true, step by step, including the steps where you used AI and should not have, and the steps where you did everything by hand out of habit when the model would have saved you forty minutes with no risk. The exercise is uncomfortable for exactly the right reason: it makes you confront the gap between the workflow you tell people you run and the one you actually run. That gap is where your next quarter of leverage lives.

And the map is defensible. When your design lead asks why your team's velocity changed, or when a skeptical engineering manager asks whether the synthesis that drove a roadmap decision was "just AI," you do not wave your hands. You point at a board where every step has a column and every column has a reason. That is the difference between operating with AI and being operated on by it.

The Five Columns, Defined So They Cannot Be Fudged

The board has five columns, and the entire value of the exercise collapses if the columns are vague. Two designers should be able to take the same sprint step, apply the column definitions, and land on the same placement. So define them concretely, by who does the work and who checks it, not by vibe.

Column One: Definitely Manual

The human does the work, start to finish, and AI is not in the loop at all - not as a drafter, not as a checker. This column is not nostalgia. A step belongs here when the work is the thinking, and outsourcing any part of it would outsource the judgment that is the point. Deciding which three of nine pain points the readout leads with is definitely manual, because that decision is a bet on what the business should care about, and it requires holding the user, the roadmap, the politics, and your own taste in your head at once. The model can list nine pain points. It cannot decide which three are load-bearing for your specific company this quarter, and pretending it can is how strategy gets laundered into "the AI suggested it."

Column Two: Manual-With-AI-Check

The human does the primary work; AI runs a check over it afterward. The authorship is yours, the verification is partly the model's. You hand-design a hi-fi screen, then run a Claude or Stark pass to catch contrast failures and a Claude prompt to flag inconsistent copy tone. The human is the author and the model is a second pair of eyes that never gets tired. This column is where AI is genuinely safe and genuinely useful, because the model's failure mode in a checking role is a false positive (it flags something fine) which costs you ten seconds, not a false authorship (it invents something wrong) which costs you a shipped defect.

Column Three: AI-With-Manual-Check

AI does the primary work; the human verifies every consequential piece of it before it advances. This is the most important column and the one teams get wrong most often, because it is where the verification tax lives (the next lesson is entirely about that). Research-transcript synthesis belongs here: Claude drafts the synthesis in forty minutes, and then you spend real time checking every quote against the raw video, because the model paraphrases the verbatim that matters into a bland summary. The model produced the artifact; you are responsible for it being true. The defining test for this column is: if the AI output shipped unchecked, would it cause harm? If yes, and the harm is recoverable through your review, it is AI-with-manual-check.

Column Four: AI-Led

AI does the work and the human reviews lightly or by sampling rather than exhaustively, because the stakes are low enough and the model reliable enough that full verification is not worth the tax. Renaming four hundred Figma layers to a consistent convention is AI-led: Magician or Figma AI does it, you spot-check a sample, and if it got three layers wrong the cost is that you rename three layers. Generating forty divergent onboarding-concept thumbnails to react to is AI-led, because you are going to discard thirty-four of them anyway and a hallucinated concept is just a concept you reject faster. The column is defined by reversibility and low blast radius, not by the model being "good at" the task.

Column Five: Definitely AI

The work is fully delegated and the human is not meaningfully in the loop at all, because verifying it would cost more than the worst-case error. Auto-tagging assets for searchability, generating alt-text drafts that a separate audit step will catch later, transcribing interview audio - mechanical, high-volume, where an error is cheap and self-correcting downstream. Be honest that this column is small in serious design work. If you find yourself putting visual design or final copy here, you have confused "the model can produce this" with "the model can be trusted to ship this," which is exactly the confusion this whole level exists to cure.

The columns are not ranked by how good AI is at the task. They are ranked by how much a human has to stand between the AI and the user. That distance is set by stakes and reversibility, not by model capability.

Mapping a Real Two-Week Sprint, Step by Step

Take a concrete sprint: a checkout-flow redesign, two weeks, the kind you actually run. The instruction is to list every step at a useful grain - not "do research" but the eight or twelve sub-steps research actually decomposes into - and place each. Here is the spine of one such map, with the reasoning that makes each placement defensible.

Transcribing eight user-interview recordings: definitely AI. Mechanical, high-volume, errors are cheap and visible. The model wins outright and nobody should spend a designer-hour on it.

Drafting a first-pass synthesis from those transcripts: AI-with-manual-check. Claude produces a structured draft in forty minutes; you then verify every quote and every claimed pattern against source. The draft is the model's; the truth is your responsibility.

Deciding the three lead pain points: definitely manual. This is a strategic bet, not a summarization. The model can surface candidates; it cannot make the call.

Building the JTBD map structure: manual-with-AI-check. You author the jobs and their priority; Notion AI checks for gaps and proposes jobs you may have missed, which you accept or reject by hand.

Generating forty wireframe variants to react to: AI-led. UX Pilot floods you with structure; you keep six, discard the rest, and a hallucinated layout is just a faster discard.

Designing the hi-fi checkout screen: definitely manual, with one carve-out. The hierarchy, the destructive-action treatment on "remove item," the error states - all hand-designed, because this is the low-frequency, high-stakes work the model fumbles. The carve-out: an empty-state illustration generated by Firefly drops into column three, AI-with-manual-check, because it is a bounded asset you can verify on sight.

Writing the microcopy for the payment-error state: manual-with-AI-check. You write it; Claude checks tone consistency against your existing strings and flags where "Something went wrong" should be "Your card was declined - no charge was made."

Generating an interactive prototype from the hi-fi: AI-with-manual-check. Figma Make builds it; you verify the focus order, the responsive breakpoints, and that the destructive action still behaves before anyone clicks it.

Running a contrast and target-size audit on the final flow: manual-with-AI-check moving toward AI-led. Stark and a Claude pass do the math; you verify the handful of edge cases the tools flag. Over time, as you trust the tooling on your token set, this drifts toward AI-led.

Notice the texture: a single sprint touches all five columns, and the placements are not by discipline. Research synthesis spans column five (transcription), column three (draft synthesis), and column one (the lead-pain-point call). That spread is the whole insight a policy could never capture.

The Justification Is the Artifact, Not the Board

The board is the visible deliverable, but the board is worthless without the justification. A placement with no reason is an opinion, and opinions do not survive a design lead asking "why?" For every step, write one sentence that answers two questions: what is the worst-case error if AI does this step, and is that error caught before it reaches a user? Those two questions determine the column.

"Drafting the synthesis is AI-with-manual-check because the worst case is a paraphrased quote that distorts a finding, and that error is caught only if I verify quotes against source, which I do." That sentence is defensible. "Drafting the synthesis is AI-led because Claude is good at synthesis" is not - it names the model's strength and ignores the stakes, which is the exact error the five-column frame exists to prevent. The justification forces you to reason about consequence and reversibility rather than capability, and that reasoning is the transferable skill. A designer who can write the justification can map a sprint they have never seen.

Hunting for the Honest Mismatches

The most valuable output of the exercise is the steps where your actual behavior and the correct column disagree. There are two flavors, and you should hunt both. The first is over-trust: a step you currently run AI-led that the stakes say belongs in AI-with-manual-check. These are your latent shipped defects - the places where you are one bad generation away from a misplaced destructive action or a distorted finding reaching a user. The second is over-labor: a step you currently do definitely-manual out of habit that the stakes say could safely move to manual-with-AI-check or AI-led. These are your reclaimable hours. A good sprint-flow audit names at least one of each, because a map that only confirms what you already do is a map you did not need to draw.

Defending It to Your Design Lead

The artifact is not done until you can defend it in conversation, because the defense is where the judgment gets tested. Your design lead will not accept the board at face value, and good - their pushback is the eval. Expect three challenges and have an answer for each.

"Why isn't this faster? If AI does half the steps, where are the savings?" The answer is the verification tax (the next lesson), and your honest reply is that columns two and three add review time that partly offsets the generation savings, and the net win comes from moving the right steps to columns four and five, not from cramming everything into AI-led. A map that promises across-the-board speedup is lying; a map that shows where speed is real and where it is taxed is true.

"You put visual design in definitely-manual - aren't you just protecting your job?" The answer is stakes and frequency, not self-interest. Hi-fi hierarchy and destructive-action treatment are low-frequency, high-stakes decisions where the model ships convincing wrong answers, and you can point at a specific past example - the red "Save view" in the delete position - where the average was dangerous. You are not protecting the discipline; you are placing the steps within it by consequence, and you proved it by moving the empty-state illustration into column three.

"What changes if you're wrong about a placement?" The answer is that the map is versioned and revisited. A placement is a hypothesis, and the verification tax ledger and the reversibility test (the next two lessons) give you the data to move steps between columns as the tooling and your trust change. The audit is dated; you run it again next quarter and the placements migrate. A designer who treats the map as permanent has missed that the floor is still moving.

Why This Makes You Both Faster and Safer

The reflexive read of a five-column map is that it slows you down by adding ceremony. The opposite is true, for the same reason the Generated-Mock Audit makes a reviewer faster: it tells you precisely where to spend attention and, more importantly, where not to. Without the map, you either verify everything (and drown in tax) or verify nothing (and ship defects). The map lets you verify exhaustively in column three, lightly in column four, not at all in column five, and reinvest the saved attention into the column-one work where your judgment is the entire value.

It also makes you legible to your team and your org. When a junior asks "should I use AI for this?" you do not give them a vibe, you give them two questions - worst-case error, is it caught? - and a board. When leadership asks whether AI changed your team's craft posture, you hand them a forensic map instead of a defensive paragraph. The sprint-flow audit is the foundational L3 artifact precisely because everything downstream - the verification tax ledger, the human-in-the-loop matrix, the reversibility test - is a refinement of the judgment you exercise here. Map the sprint first. The rest is engineering.

Putting It to Work This Week

Pick a sprint you finished in the last month, not one you are imagining. Recent and real beats clean and hypothetical, because the value is in the honest mismatches and you cannot lie to yourself about a sprint you actually ran. List its steps at the grain of "the smallest unit of work I could hand to someone else," which is usually fifteen to thirty steps for a two-week sprint. Place each in one of the five columns. Then, for every placement, write the one-sentence justification answering worst-case-error and is-it-caught. Flag every step where your placement and your actual past behavior disagree.

Bring the board to your next 1:1 with your design lead and defend it out loud. You will find that two or three placements that felt obvious on paper collapse under one question, and that is the point - the conversation is the verification step for the map itself. Version it, date it, and put a reminder to redraw it next quarter, because the tools will have moved and so will the right answers. You will know the exercise worked when the board changes how you start your next sprint: you will catch yourself about to AI-lead a step that belongs in column three, and you will stop, because now you can see the column.

Key Takeaways

  • Map the work at the level of steps, not disciplines. A single afternoon of research synthesis spans multiple columns; a discipline-level policy cannot see that and a step-level map can.
  • The five columns - definitely manual, manual-with-AI-check, AI-with-manual-check, AI-led, definitely AI - are ranked by how much a human must stand between the AI and the user, set by stakes and reversibility, not by how capable the model is.
  • Place each step by answering two questions: what is the worst-case error if AI does it, and is that error caught before it reaches a user? Capability-based reasoning ("the model is good at this") is the exact error the frame prevents.
  • The justification is the real artifact. A placement without a one-sentence reason is an opinion that will not survive your design lead asking "why?"
  • Hunt the honest mismatches: steps you currently run AI-led that the stakes say need a manual check (latent shipped defects), and steps you do by hand out of habit that could safely move to AI (reclaimable hours).
  • The map makes you both faster and safer by telling you where to spend verification attention and where not to, and it is the foundation the verification-tax ledger, HITL matrix, and reversibility test all refine.