The Use-Case Inventory and Scoring Matrix
You have diagnosed where the team is, written the bet, and drawn the line. Now you have to decide what the team actually does, in what order, for the next twelve months - and the failure mode here is not laziness, it is enthusiasm. A design team in 2026 has thirty plausible AI use cases it could pursue, every one of them defensible in isolation, and a strategist who chases all thirty achieves none. This lesson teaches you to build the artifact that converts thirty good ideas into one executable sequence: a use-case inventory that captures every candidate, a scoring matrix that rates each on impact, risk, and reversibility, and a twelve-month roadmap your design leads can actually execute. The output is not a wish list. It is a sequenced, scored, defensible plan where every initiative can be traced to why it is happening now rather than later, which is the difference between a roadmap and a list of things someone read about on LinkedIn.
Why You Inventory Before You Roadmap
The instinct, once you have a strategy, is to jump straight to "here is what we will do this quarter." Resist it, because a roadmap built without a complete inventory is a roadmap built from whatever happened to be salient - the tool a competitor mentioned, the workflow that frustrated you last week, the thing the loudest IC is excited about. Salience is not the same as importance, and the gap between them is exactly where strategic mistakes live. The inventory exists to make the full opportunity space visible before you choose, so that what you choose is chosen against everything you could have chosen, not against the three things you happened to remember.
The inventory is also a completeness check on your own strategy. When you list every candidate use case across the team - research, UI, prototyping, brand, design systems, accessibility, design ops, hiring, documentation - patterns emerge that no single-quarter view reveals. You discover that five of your candidates all depend on the same prerequisite (a machine-readable design system), which tells you that prerequisite is not one initiative among many but the unlock for a whole cluster. You discover that your highest-enthusiasm candidate is also your highest-risk one, which tells you to slow down. The inventory's value is not the list itself; it is the relationships the list makes visible, which a roadmap-first approach hides.
There is a discipline to capturing candidates well. A use case is not "use AI for research" - that is a category. A use case is a specific, scoreable unit of work: "AI-augmented synthesis of unmoderated usability studies, with a human verification pass on every flagged finding." It names what gets done, where AI touches it, and what the human still owns. Capture them at that grain or the scoring is meaningless, because you cannot rate the impact and risk of a category, only of a specific application. Cast the net wide first - every team member's candidates, every workflow that recurs - and refine the grain second.
The Three Scoring Axes and Why These Three
Every candidate gets scored on three axes, and the choice of these three is deliberate. Many frameworks score only on impact, or on impact and effort, and both produce dangerous roadmaps because they ignore the dimensions that determine whether an AI initiative blows up. Impact tells you whether a use case is worth doing. Risk tells you what it costs you if it goes wrong. Reversibility tells you how cheaply you can undo it if it does. Together they answer the only question that matters for sequencing: not just "is this valuable" but "is this a safe early bet or a thing we should do later once we have learned more."
Impact: How Much Does This Matter
Impact is the value the use case creates if it works: time saved at volume, quality raised, a capability unlocked, a gap from the maturity assessment closed. Score it on a simple scale (say one to five) and force yourself to denominate it the way the thesis taught - in outcomes the business tracks, not in "it would be cool." The discipline is to distinguish broad impact (a workflow many people run often) from narrow impact (a tool one person uses occasionally), because a high-quality improvement to a rare task scores lower than a modest improvement to a constant one. Volume times quality-gain, roughly, is what impact measures.
Risk: What Does It Cost If It Goes Wrong
Risk is the magnitude of harm if the use case fails or is done badly: a user harmed, a brand damaged, an accessibility liability created, a legal exposure opened, the team's trust in AI burned by a visible failure. This is where the What Stays Human statement feeds in directly - a use case that touches one of your protected decisions carries elevated risk by definition, and a candidate that proposes delegating something on your human-only list should score so high on risk that it sequences last or off the roadmap entirely. Risk is not the same as difficulty; a hard-to-build use case can be low-risk, and an easy-to-build one can be catastrophic if it touches customer-facing destructive flows.
Reversibility: How Cheaply Can We Undo It
Reversibility is the axis most frameworks omit and the one that most determines safe sequencing. It asks: if this use case turns out to be a mistake, how expensive is it to back out? An internal workflow change is highly reversible - you stop doing it and lose nothing permanent. Adopting a tool that becomes load-bearing across the team, training everyone on a vendor's proprietary workflow, or letting AI generate customer-facing output at scale are low-reversibility, because backing out means retraining, rebuilding, or recalling. The strategic principle, borrowed straight from the L3 reversibility test, is that high-reversibility use cases are safe to try early even when uncertain, because the cost of being wrong is small, while low-reversibility use cases should wait until you have reduced their uncertainty, because being wrong is expensive to undo.
Impact tells you whether a use case is worth doing. Risk tells you what it costs if it goes wrong. Reversibility tells you how cheaply you can undo it. Sequencing is the art of doing the high-impact, high-reversibility bets first and earning the right to the low-reversibility ones.
From Scores to Sequence: The Logic That Turns a Matrix Into a Roadmap
Three scores per use case do not automatically produce an order; you need a sequencing logic, and the logic is where strategic judgment lives. The naive move is to rank by impact alone and do the highest-impact things first. This is a trap, because your highest-impact use cases are frequently your lowest-reversibility ones (letting AI generate customer-facing work has huge impact and is hard to undo), and leading with them means betting big before you have learned anything. The sequencing logic instead reads all three axes together.
The first wave is high-impact, high-reversibility, low-risk: the use cases that are worth a lot, safe to try, and cheap to undo if wrong. These build momentum and team confidence while teaching you how your team actually adopts AI, at almost no downside. The second wave is the high-impact, low-reversibility use cases whose uncertainty the first wave reduced - now that you have learned how adoption goes, you can commit to the things that are expensive to back out of. The third wave, or the never-wave, is the high-risk use cases, especially any touching the What Stays Human lines, which either wait for capability to improve or stay off the roadmap. And the prerequisite logic overrides all of it: if five high-impact use cases all depend on a machine-readable design system, that prerequisite sequences first regardless of its own impact score, because it is the gate the cluster waits behind.
The Prerequisite Insight
The single most valuable thing the inventory surfaces is the prerequisite cluster. When the matrix shows that a third of your high-impact use cases are blocked behind one enabling capability, that capability stops being a use case and becomes the foundation of the roadmap. This is the same leverage logic from the maturity assessment, now made rigorous: the AI-readable design system is rarely the highest-impact line item in isolation, but it is the unlock for the cluster, and a roadmap that does not sequence the prerequisite first will see its whole second half stall waiting for a foundation that was never built. Reading the inventory for prerequisites is how you avoid building a roadmap whose later quarters quietly depend on work no quarter was assigned.
Building the Matrix So Leads Can Execute It
The artifact has to be more than a scored spreadsheet; it has to be a roadmap your design leads can pick up and run without you in the room. That means each sequenced initiative carries the things an executor needs: an owner, a success metric drawn from the impact score's denomination, a rough quarter, and the prerequisite it waits behind if any. The scoring matrix is the reasoning; the roadmap is the output, and the roadmap is what you hand to the leads. Keep both - the matrix so the sequencing is defensible when challenged ("why is this in Q4 not Q1?" answers from the scores), the roadmap so the work is executable.
A practical structure: the inventory is a table with every use case and its three scores plus a one-line note on each score's reasoning. The roadmap is the subset you committed to, ordered into quarters, each initiative with owner, metric, and dependency. The use cases you scored but did not sequence are not deleted; they go into a visible backlog with their scores, so that when a quarter frees up or a tool's reliability improves (raising a use case's reversibility or lowering its risk), you re-sequence from a scored backlog rather than starting the thinking over. The matrix is a living instrument, not a one-time exercise, and the backlog is what makes re-planning cheap.
Use coarse scales, not false precision. Score each axis low, medium, or high rather than on a one-to-ten scale, because the value of the matrix is the decomposition and the visible reasoning, not the decimal. A one-to-ten scale invites arguments about whether something is a six or a seven that the underlying judgment cannot actually support, and it dresses fuzzy estimates in a precision they do not have. Three coarse levels per axis force a clear call, keep the matrix honest about its own uncertainty, and still produce all the sequencing information you need, because what drives the order is the pattern across axes - high-impact-high-reversibility versus high-impact-low-reversibility - not the difference between a 6.5 and a 7. The one-line reasoning note attached to each score is what carries the real defense; the level is just the headline. When someone challenges a placement, you answer from the note, not from the number, which is exactly why the matrix is a decision aid rather than a decision oracle - it exposes the reasoning so it can be argued with rather than hiding it behind a computed total.
A Worked Example: Scoring and Sequencing a Real Inventory
Take the ten-person product team from the earlier lessons. The inventory captures, among others: AI-augmented usability synthesis with verification; a machine-readable token layer; AI-drafted design-review writeups; Storybook MCP component docs so agents read the system; AI-generated first-pass icons; letting AI generate customer-facing empty-state copy at scale; and an AI brand-asset pipeline. Score them.
AI-drafted review writeups: moderate impact (saves time on a frequent chore), low risk (internal, edited before send), high reversibility (stop anytime). A clean first-wave candidate. The token layer: moderate direct impact but it is the prerequisite for the Storybook docs and three other use cases, so its effective importance is the cluster it unlocks - it sequences first as the foundation regardless of its standalone score. AI-augmented usability synthesis with verification: high impact (research is a bottleneck), moderate risk (mitigated by the mandatory verification pass), high reversibility - first wave. AI-generated first-pass icons: moderate impact, low risk, high reversibility - an easy early win.
Now the dangerous ones. Letting AI generate customer-facing empty-state copy at scale: high impact but high risk (customer-facing, brand voice, the cheery-bot failure mode) and low reversibility (it ships to users and recalling it is costly), and it brushes against the brand-voice line in the What Stays Human statement - this sequences late and only behind a verification workflow, or it waits. The AI brand-asset pipeline: high impact but it touches brand distinctiveness, a protected decision, so it scores high on risk and is deferred until the brand-anchor system exists to constrain it. The roadmap that falls out: Q3 builds the token layer (prerequisite) plus the safe first-wave wins (review writeups, synthesis, icons); Q4 builds the Storybook MCP docs (now unblocked) and a constrained brand-anchor pipeline; the at-scale customer-facing copy sits in the scored backlog awaiting a verification workflow. Every placement traces to the scores, so when a lead or an executive asks "why this order," the matrix answers.
Defending the Sequence Against the Enthusiasm Pull
The hardest part of owning a scored roadmap is defending it against the constant pull to do the exciting thing now. An IC will want to start the brand-asset pipeline because it is the most fun; an executive will want to do the at-scale copy generation because the time savings look enormous on a slide; you yourself will feel the pull of the high-impact, low-reversibility bets because impact is seductive. The matrix is your defense, and it works because it externalizes the reasoning. "I'm not saying no to the brand pipeline; I'm saying it scores high-risk because it touches a protected decision and low-reversibility because it ships to customers, so it sequences after we build the anchor system that de-risks it" is a defensible position. "I don't feel ready" is not.
The deeper point is that the inventory-and-matrix converts strategy from taste into a method anyone can audit. A roadmap defended by the strategist's intuition is only as durable as the strategist's standing in the room; a roadmap defended by impact, risk, and reversibility scores is durable because the reasoning is visible and re-runnable by anyone who doubts it. That is what makes it a roadmap your design leads can execute - not just that it has owners and dates, but that when they hit a fork they can re-derive the priority from the same scores you used, which means the strategy survives you being on vacation, in a different meeting, or, eventually, in a different job.
Key Takeaways
- The failure mode in roadmapping is enthusiasm, not laziness: a team has thirty defensible AI use cases and a strategist who chases all thirty achieves none. The inventory-and-matrix converts thirty good ideas into one executable sequence where every initiative traces to why it is happening now rather than later.
- Inventory before you roadmap, because a roadmap built from salience (the tool a competitor mentioned, the loudest IC's excitement) is built from whatever you happened to remember, not from everything you could have chosen. Capture use cases at a scoreable grain - a specific application, not a category - and cast the net across the whole team.
- Score every candidate on three axes: impact (value if it works, denominated in business outcomes, volume times quality-gain), risk (cost if it goes wrong - and anything touching a What Stays Human line scores high by definition), and reversibility (how cheaply you can undo it). Most frameworks omit reversibility, and it is the axis that most determines safe sequencing.
- The sequencing logic reads all three together: first wave is high-impact, high-reversibility, low-risk (safe bets that build momentum and teach you how the team adopts); second wave is the high-impact, low-reversibility bets whose uncertainty the first wave reduced; third wave or never-wave is the high-risk use cases, especially those touching protected decisions.
- The most valuable thing the inventory surfaces is the prerequisite cluster: when a third of your high-impact use cases are blocked behind one enabling capability (often the machine-readable design system), that capability sequences first regardless of its standalone impact, because it is the gate the cluster waits behind.
- Ship both the matrix and the roadmap: the matrix is the defensible reasoning (it answers "why this order"), the roadmap is the executable output (owners, metrics, quarters, dependencies). Unsequenced use cases go into a scored backlog, so re-planning is cheap when a quarter frees up or a tool's risk drops.
- The matrix is your defense against the enthusiasm pull. It converts strategy from the strategist's taste into a method anyone can audit and re-run, which is what lets design leads execute the roadmap without you in the room and what makes the strategy survive your absence.
Skill.re