โ†
AI for Designers (UX, Product, Brand)
Visionary ยท M9 ยท lesson 9 of 16 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
The Design-AI Pilot Portfolio
๐Ÿ“–
now learning

The Design-AI Pilot Portfolio

15 min

The fastest way to waste a design org's AI budget is to run one big pilot, bet the team's credibility on it, and have it fail in a way that lets the skeptics say "see, AI does not work for design." The second fastest way is to run twelve scattered experiments with no kill criteria, so that none of them ever conclude and the org cannot tell what it learned. The discipline that beats both is a portfolio: three pilots over six months, deliberately chosen at different risk levels - one safe, one stretching, one bet - each with kill criteria written before it starts, so that the portfolio as a whole returns learning whether any single pilot succeeds or fails. This lesson teaches you to build that portfolio, calibrate the three risk levels so they are genuinely different, and write the kill criteria that turn "we tried some AI stuff" into a defensible, fundable program. The artifact is a pilot portfolio doc your CPO will recognize as the work of someone who has run experiments before.

Why a Portfolio, Not a Pilot

The single-pilot instinct is natural and wrong. You pick the most promising AI use case, pour the team's energy into proving it works, and stake your credibility on the outcome. The problem is statistical: any individual pilot has a real chance of failing for reasons that have nothing to do with whether AI is useful for design - bad tool fit, a team that was already underwater, a use case that was wrong for reasons you could not see in advance. When the single pilot fails, the organizational lesson is not "this particular bet did not pay off" but "AI does not work for us," and that lesson is both wrong and very expensive to unlearn. A single pilot puts the entire credibility of design's AI program on one roll of the dice.

A portfolio fixes this the way a portfolio fixes any risky-asset problem: by diversifying across risk levels so the expected return is positive even though individual outcomes vary. With one safe pilot likely to succeed, one stretching pilot that might, and one genuine bet that probably will not, the portfolio is designed so that you learn something valuable from every outcome and you almost certainly book at least one clear win. The win funds the program's continuation; the failures, conducted cheaply and concluded cleanly, teach the org where the boundaries are. The portfolio converts "is AI useful for design" from a binary the whole program rides on into a set of calibrated experiments that collectively answer "where, specifically, is AI useful for design at our org."

There is a leadership dimension too. A portfolio signals to your executives that you are running a real program, not chasing hype. Anyone can say "we are experimenting with AI." A leader who shows up with three pilots at deliberately chosen risk levels, each with pre-written kill criteria and a measurement plan, is demonstrably doing the thing executives trust: managing risk, allocating effort proportionally, and committing in advance to how they will judge success. The portfolio doc is as much a credibility artifact as an experimental one, and in a year when design's AI budget is contested, the credibility matters as much as the experiments.

The Three Risk Levels, Calibrated

The portfolio's power comes from the three pilots being genuinely different in risk, and the most common mistake is running three "stretching" pilots and calling it diversified. Calibrate deliberately. The three levels differ along three axes: probability of success, magnitude of payoff if it works, and the cost of failure to the team and the org.

The Safe Pilot

The safe pilot is a use case where AI is already known to work, applied to a real but low-stakes part of the team's work, with a high probability of success and a modest but reliable payoff. Think: using a structured Claude prompt to draft alt-text across the image library, or auto-naming and re-indexing a Figma component library, or running AI-assisted contrast audits on a flow. These are not exciting. That is the point. The safe pilot exists to bank a clear, early win that builds the team's confidence and gives you a concrete success to point to when the bet pilot inevitably wobbles. A safe pilot that "succeeds" is not a surprise; its value is momentum and proof-of-concept for the program, not learning at the frontier.

The discipline in the safe pilot is to keep it genuinely safe - resist the temptation to make it more ambitious because it feels too easy. The safe pilot's job is to succeed, visibly and early. If you stretch it, you have lost the anchor the portfolio needs, and a portfolio of three uncertain pilots has no floor.

The Stretching Pilot

The stretching pilot is a use case where AI plausibly works but you do not yet know if it works for your team, your stack, and your standards - a real uncertainty with a meaningful payoff. Think: an end-to-end research-to-prototype loop compressing two weeks into three days, or a design-to-code handoff pattern through Anima or Locofy on a real feature, or a brand-system-as-code migration that lets AI agents generate against your tokens. The stretching pilot might succeed or might surface a fundamental obstacle, and either outcome is genuinely informative. This is the heart of the portfolio - the pilot most likely to change how the team works if it lands, and most likely to teach you something real if it does not.

The Bet

The bet is a use case at the frontier of what is possible in 2026 - low probability of success, but a payoff large enough that even a small chance justifies the effort. Think: a fully AI-agentic design-system workflow where agents read the system and propose compliant components, or a generative-brand-system where the rules rather than the marks are the asset, or AI-driven cross-platform token round-tripping across five surface families. The bet probably will not work this cycle. Its value is twofold: occasionally a bet pays off and reshapes the whole program, and even when it fails, it maps the frontier - it tells you precisely how far the technology is from your need, which is information you cannot get any other way and which positions you to move first when the frontier moves. A portfolio with no bet is a portfolio that will be surprised by where the technology goes next.

Kill Criteria: Writing the Ending Before the Beginning

The single most important discipline in the entire portfolio is writing kill criteria before each pilot starts. A kill criterion is a pre-committed, specific, observable condition under which you will stop the pilot - written down, agreed to, and dated, before anyone has fallen in love with the work. Without kill criteria, pilots do not end; they fade, consuming team energy and budget while never producing a clean verdict, and the org learns nothing because nothing concluded. With them, every pilot has a built-in decision point and a clear answer to "should this continue."

The reason kill criteria must be written in advance is the same reason they are hard to write in advance: once a team has invested in a pilot, sunk cost and identity make it nearly impossible to judge objectively whether it is working. The designer who built the workflow will always see promise around the next corner; the leader who championed the bet will reframe any failure as "early." A kill criterion written before the emotional investment accrued is the only objective judge available, because it was authored by a version of you who had no stake in the outcome. Writing the ending before the beginning is how you protect the program from your own future bias.

What Makes a Kill Criterion Good

A good kill criterion is specific, observable, and tied to the pilot's actual hypothesis rather than to vibes. "Kill if the team does not like it" is useless. "Kill if, after four weeks, the verification overhead on AI-generated handoffs exceeds the time saved by generation - measured by the ledger we are keeping - for three consecutive features" is a kill criterion, because it names what to measure, the threshold, and the window. Good kill criteria also distinguish the type of failure: a pilot can fail because the technology is not ready (kill and revisit later), because the use case was wrong (kill and do not revisit), or because the team's adoption broke down (kill the pilot but fix the adoption problem before trying again). Naming the failure type in advance turns a kill into a lesson rather than a defeat.

Each risk level gets calibrated kill criteria. The safe pilot's kill criterion should almost never fire - if it does, something is badly wrong and you want to know immediately. The stretching pilot's kill criterion is the real working tool, the threshold that tells you whether the meaningful uncertainty resolved for or against you. The bet's kill criterion should be generous on timeline but strict on the core question - you give a bet room to be slow, but you hold it hard to "did the frontier-defining thing actually happen," because a bet that drifts into a comfortable mediocre outcome is worse than a bet that fails cleanly and teaches you the frontier's location.

Write the ending before the beginning. A kill criterion authored before the team falls in love with the pilot is the only objective judge available, because it was written by a version of you with no stake in the outcome. Pilots without kill criteria do not fail - they fade, and fading teaches the org nothing.

Measurement, Effort, and the Budget Across the Portfolio

The portfolio needs a single measurement frame so the three pilots are comparable and the whole thing rolls up into a story your CPO can read. For each pilot, define the hypothesis (what we believe AI will do here), the metric (how we will know), the baseline (what the metric is today, without AI), the kill criterion (when we stop), and the resourcing (who, how much of their time, for how long). The discipline of a shared frame is what turns three experiments into a portfolio - without it you have three differently-shaped efforts that cannot be compared or summarized.

Effort should be allocated inversely to risk in the early weeks: most of the team's pilot capacity on the safe and stretching pilots, a deliberately small slice on the bet. This is counterintuitive to people who want to chase the exciting frontier, but it is correct portfolio management - you fund the bet enough to learn, not enough to bet the program on it. The safe pilot is cheap and quick; the stretching pilot gets the real investment because it is where the program's working future probably lies; the bet gets a small, time-boxed allocation with a generous timeline and a strict question. The measurement frame and the effort allocation together are what make the portfolio defensible to finance: you can show that you are spending proportionally to expected value and that every dollar has a pre-committed exit.

One measurement note specific to design-AI pilots: the verification tax is real and must be in the ledger. A pilot that "saves time" by generating output but adds more verification overhead than it removes is a failure dressed as a success, and only a measurement frame that counts the verification cost will catch it. Many design-AI pilots fail exactly here - the generation is fast, the verification is slow, and the net is negative - so make the verification tax an explicit line in every pilot's metric, not an afterthought.

The Pilot Portfolio Doc

The artifact is a pilot portfolio doc, and it has a deliberately simple structure because its job is to be read and acted on, not admired. It opens with the portfolio thesis - one paragraph on why a portfolio rather than a single pilot, and what the six-month program is meant to learn. Then it presents the three pilots in a consistent template: name, risk level, hypothesis, metric and baseline, kill criterion (with the failure-type distinction), resourcing, and timeline. Then a single roll-up table showing all three side by side, so a CPO can see the whole bet-allocation at a glance. Finally, a decision cadence: when the portfolio is reviewed, who decides on kills and continuations, and how the learning gets captured regardless of outcome.

The doc's most important property is that it commits in advance. By the time it is signed off, the org has agreed not just to run three pilots but to the specific conditions under which each will be stopped and the specific metrics by which each will be judged. This pre-commitment is the entire value - it is what prevents the safe pilot from being quietly over-resourced because it is comfortable, the bet from being killed early because it is scary, and the stretching pilot from fading into a permanent zombie experiment. A portfolio doc that names the kills in advance is a governance instrument as much as an experimental plan, and the governance is what makes it survive contact with the organization.

A Worked Example: A Six-Month Portfolio for a Product Design Team

Make it concrete for a forty-person product design org heading into the next six months. The safe pilot: AI-assisted accessibility audits across the team's top five flows, using Stark plus a structured Claude pass, with the hypothesis that it cuts audit time by half without missing findings a manual audit would catch. Metric: audit hours and findings parity against a manual control. Kill criterion - which should never fire - is if findings parity drops below the manual baseline, indicating the AI is missing real issues; failure type would be "tool not ready." Resourcing: one designer, two weeks. This banks an early, visible win and ships a real accessibility improvement regardless.

The stretching pilot: a research-to-prototype loop on one squad, compressing the standard two-week cycle to three days using Granola plus Claude for synthesis, UX Pilot for wireframes, and Figma Make for the interactive prototype, with the hypothesis that it holds research rigor while tripling speed. Metric: cycle time and a research-rigor check (verbatim traceability, contradicting-evidence callouts) against the team's normal standard. Kill criterion: if the verification tax on synthesis exceeds the time saved for two consecutive cycles, or if rigor measurably drops, kill it; failure type would be "use case wrong for our standard." Resourcing: two designers and the research lead, eight weeks. This is the pilot most likely to reshape how the team works.

The bet: an agentic design-system workflow where an MCP-enabled agent reads the team's tokens and Storybook docs and proposes compliant components, with the hypothesis that agents can generate genuinely system-compliant work rather than plausible drift. Metric: percentage of agent-proposed components that pass system-compliance review with no edits. Kill criterion: generous timeline (twelve weeks) but strict question - if zero agent proposals pass clean review by week eight, the frontier is not here yet; failure type "technology not ready, revisit in two quarters." Resourcing: one senior systems designer, a small time-boxed slice. The bet probably fails, and when it does, the team knows precisely how far agentic systems are from production - which is exactly the frontier map that lets them move first when it closes. Three pilots, three risk levels, three pre-written endings, one defensible program.

Putting It to Work This Quarter

Resist the single-pilot instinct even though it feels focused and brave. Pick three use cases at genuinely different risk levels - and check honestly that your "safe" one is boring enough to nearly guarantee a win and your "bet" is ambitious enough to probably fail, because a portfolio of three stretching pilots is not diversified and will not protect the program. Then write each pilot's kill criterion before you write anything else about it, because the kill criterion authored before the team's emotional investment is the only honest one you will ever get.

Put the three on one page with a shared measurement frame, allocate effort inversely to risk, make the verification tax an explicit metric line, and circulate the doc for sign-off so the kills are pre-committed organizationally and not just in your head. The deliverable is not three experiments; it is a six-month program designed so the org learns something valuable from every outcome and books at least one clear win - a program that survives a failed bet without anyone concluding that AI does not work for design, because the portfolio was built to make exactly that conclusion impossible.

Key Takeaways

  • A single big pilot stakes the whole program's credibility on one roll of the dice; when it fails for incidental reasons, the org wrongly concludes "AI does not work for design." A portfolio of three pilots at different risk levels returns learning from every outcome and almost certainly books at least one win.
  • Calibrate three genuinely different risk levels. Safe: AI already known to work, low stakes, high probability, modest reliable payoff - keep it boring; its job is an early visible win. Stretching: plausible but unproven for your team, meaningful payoff - the heart of the portfolio. Bet: frontier use case, low probability, large payoff - maps the frontier even when it fails.
  • Write kill criteria before each pilot starts. A criterion authored before the team's emotional investment is the only objective judge available. Pilots without kill criteria do not fail - they fade, consuming budget while teaching the org nothing.
  • Make kill criteria specific, observable, and tied to the hypothesis, and name the failure type in advance (technology not ready, use case wrong, adoption broke) so a kill becomes a lesson rather than a defeat. Calibrate them: the safe pilot's should never fire, the bet's should be generous on time but strict on the core question.
  • Use one shared measurement frame (hypothesis, metric, baseline, kill criterion, resourcing) so the three are comparable and roll up into a CPO-readable story. Allocate effort inversely to risk - fund the bet to learn, not to bet the program. Make the verification tax an explicit metric line, because many design-AI pilots fail exactly there.
  • The pilot portfolio doc is a governance instrument as much as an experimental plan: by sign-off, the org has pre-committed to the kills and the metrics, which prevents the safe pilot from being over-resourced, the bet from being killed early, and the stretching pilot from becoming a zombie.