AI for Designers (UX, Product, Brand)
Strategic · M16 · lesson 16 of 26 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
The Four-Tool Bake-Off Framework: Figma Make vs. Lovable vs. v0 vs. Bolt
📖
now learning

The Four-Tool Bake-Off Framework: Figma Make vs. Lovable vs. v0 vs. Bolt

15 min

Sooner or later someone above you will say "just pick a tool," and the wrong way to answer is to pick the one with the best demo, the loudest user numbers, or the most LinkedIn buzz. Lovable crossed roughly 8 million users and a reported $206M to $400M ARR at a valuation around $6.6 billion; Bolt hit about 5 million users and $40M ARR in five months; v0 grew to around 4 million users on Vercel. Those numbers are real and they tell you almost nothing about which tool is right for your team, because a tool's market traction measures its appeal to its whole market, not its fit to your specific use case against your specific stack. This lesson teaches you to run a real bake-off - a structured, time-boxed comparison on a real use case, run by two ICs across two weeks - that produces a vendor decision doc with named verdicts per tool, using the market data as context rather than as a leaderboard. The output is not "Lovable won." It is "for this use case, against our stack, here is what each tool is for and what it is not, and here is the decision."

Why a Bake-Off, Not a Leaderboard

The instinct when choosing between four hyped tools is to find the ranking - the article or the benchmark that says which is best - and adopt the winner. This fails for a structural reason: there is no context-free best, because these tools have genuinely different shapes. Figma Make sits closest to the designer's existing Figma workflow and is strongest for the design-review prototype. v0 is component-first and strongest at producing React output that engineering can actually take. Lovable is app-first and strongest at producing a working full-stack thing fast, which is also its risk. Bolt is the high-fidelity throwaway specialist. A leaderboard collapses these different shapes into one axis and hands you a winner that may be wrong for the thing you actually need to do. The bake-off refuses the collapse and asks the only question that matters: which tool is best for this use case, against our stack, judged by our people.

The market data belongs in the doc, but as context, not verdict. Lovable's 8 million users tell you it is real, well-funded, unlikely to vanish, and that its app-first framing resonates with a huge non-designer audience - which is useful to know and is also exactly the source of its framing risk on a design team, because the same thing that makes non-designers love it (it produces a working app from a sentence) is what makes them mistake a prototype for a build commitment. v0's 4 million users on Vercel tell you it has deep platform integration and a component-first philosophy that fits an engineering-handoff use case. The numbers contextualize each tool's character and staying power; they do not rank fitness for your job. Using them as a leaderboard is the exact mistake the L1 stack-map lesson warned against, now with a budget attached.

Designing the Bake-Off: One Real Use Case, Specified Once

A bake-off is only as good as its use case, and the cardinal rule is that you pick one real use case from your actual backlog and specify it once, identically, for all four tools. Not a toy. Not four different briefs. One real feature - a customer-portal settings page, an onboarding flow, a dashboard - specified to the same level of detail and handed to all four tools the same way. The single-specification discipline is what makes the comparison fair: any difference in output is a difference in the tool, not a difference in what you asked. A bake-off where each tool got a slightly different or more forgiving brief produces a verdict that is really a verdict about the briefs, which is worthless.

The use case also has to be representative of work you will actually do repeatedly, because you are choosing a tool for a category of work, not for one feature. Picking an unusually simple or unusually hard feature skews the result toward whichever tool happens to suit that extreme. The right use case is a median example of the work the tool would do in production, complex enough to exercise the things that matter (real data, real states, design-system tokens, a real interaction) and ordinary enough that the verdict generalizes. Spend real care on the spec, because the spec is the experiment's controlled variable and everything downstream depends on it being held constant.

Two ICs, Two Weeks: The Resourcing That Makes It Real

The brief specifies the bake-off is run by two ICs across two weeks, and both numbers are deliberate. Two ICs rather than one, because a single person's facility with a tool confounds the result - someone who already knows Lovable will make Lovable look best regardless of its actual fit, so you want at least two people so individual familiarity averages out and you can see whether a tool is genuinely easy or just easy for the one person who already knew it. Pairing the ICs across the tools, or having each run all four, surfaces the difference between "this tool is good" and "this tool is good in this person's hands," which is a distinction a solo evaluation cannot make.

Two weeks rather than an afternoon, because the things that matter about these tools do not show up in a demo. In an afternoon every tool looks magical, because the demo path is the happy path. The failures that determine real fitness - where the tool drifts from your design-system tokens, where it invents components, where the output is hard to hand to engineering, where the second and third iterations get worse instead of better, where the framing confuses a stakeholder - these emerge over days of real use on a real feature. A two-week box is long enough to hit the failures and short enough to stay bounded and fundable. An unbounded evaluation never ends and a one-day evaluation only sees the magic; two weeks is the window where the truth shows up.

In an afternoon every tool looks magical, because the demo path is the happy path. The failures that determine real fitness - token drift, invented components, handoff friction, the framing that confuses a stakeholder - only emerge over days of real use. Two weeks is the window where the truth shows up.

What You Score, and Against Your Stack

The scoring dimensions are not generic; they are the things that determine whether a tool actually serves your team, and several of them are about fit to your existing stack rather than about the tool in the abstract. Score each tool on: fidelity to the spec (did it build what you asked), design-system token preservation (did it use your tokens or invent its own values - the single most important dimension for a team with a real system), time-to-clickable (how fast to a usable artifact), handoff quality (can engineering take the output, or does it have to be rebuilt), iteration behavior (do the second and third passes improve or degrade), and stakeholder-confusion risk (will a non-designer mistake this artifact for something it is not). That last dimension is where Lovable's strength becomes a risk, and scoring it explicitly is how you catch a problem a feature-comparison would miss.

The phrase "against your stack" is the heart of the framework. A tool that preserves design-system tokens beautifully is worth more to a team with a mature token system than to a team without one; a tool that produces clean React is worth more to a React shop than to one on a different framework; a tool whose output integrates with your existing Figma library matters more if that library is your source of truth. The bake-off scores fitness to your specific context, which is exactly what a market leaderboard cannot do, because the leaderboard does not know your stack. The verdict you produce is therefore not transferable to another team with a different stack, and that non-transferability is a feature: it means the verdict is actually about your decision and not a borrowed opinion.

Named Verdicts: The Output That Makes the Doc Useful

The deliverable is a vendor decision doc, and its core is a named verdict per tool - not a score, a verdict, stated in plain language a CPO can act on. A named verdict says what the tool is for and what it is not, against your stack, in a sentence someone can execute. The L2 build-vs-buy lesson modeled the shape: "Use Figma Make for the design-review prototype; use v0 for the engineering-handoff component; do not use Lovable on this feature because the framing risk exceeds the speed benefit." Notice that the verdict is not a single winner - it can assign different tools to different jobs, which is usually the honest answer, because these tools have different shapes and the right decision is often "this one for this, that one for that, none of them for this third thing."

The named verdict's discipline is that it must be actionable and defensible. Actionable means a reader can do something with it without re-running the bake-off - "use v0 for handoff components" is actionable, "v0 scored 7.5" is not. Defensible means the verdict traces to the scores and the two-week experience - when someone asks "why not Lovable for this," the answer is "it scored well on time-to-clickable but high on stakeholder-confusion risk and poorly on token preservation, and on this feature those outweigh the speed," which is a defense, not an opinion. A doc full of scores with no named verdicts forces the reader to do the synthesis the strategist was supposed to do; a doc of named verdicts backed by scores has done the job.

The "Do Not Use" Verdict Is the Valuable One

The verdict that earns the most trust is the honest negative - naming a tool you will not use for this, and why. It is tempting to find a use for every tool you evaluated, because it makes the bake-off feel complete, but a strategist who concludes "Lovable is not the right tool for this feature on our stack, despite its market traction, because the framing risk exceeds the benefit" demonstrates exactly the judgment the bake-off exists to produce. The negative verdict is also where you most clearly use the market data as context rather than leaderboard: Lovable's enormous user base is acknowledged and then explicitly set aside as irrelevant to this decision, which is the move that proves you are choosing for your team rather than following the crowd.

A Worked Example: A Bake-Off for a Customer-Portal Settings Page

Run it concretely. The use case is a customer-portal settings page from the real backlog - tabs, toggles, a form with validation, a destructive "delete account" action, real design-system tokens. Specified once, handed to all four tools, two ICs over two weeks, scored on the six dimensions.

Figma Make: high fidelity to the spec and strong token preservation because it lives near the Figma library, fast to a clickable prototype, but the output is a prototype not production code, so handoff quality is moderate. Verdict: "Use Figma Make for the design-review version of this page; it is the fastest path to a stakeholder-ready clickable artifact on our stack." v0: moderate token preservation (needed token mapping), excellent handoff quality because it produces real React the engineers can take, slower to first clickable. Verdict: "Use v0 to produce the handoff component once the design is settled; its React output is what engineering can actually build on." Bolt: very fast to a high-fidelity working demo, but the output is a throwaway and token preservation was weak. Verdict: "Use Bolt only for a throwaway sales-style demo of this page, retired after the meeting; do not use it for handoff."

Lovable: produced a working full settings page fastest of all, genuinely impressive, but token preservation was weak, the output framing made a stakeholder in the test ask "so this is built?", and the destructive delete action it generated sat in the wrong place with no confirmation - the exact generation-versus-understanding failure. Verdict: "Do not use Lovable for this feature. Its speed is real and its market traction is real, but on our stack the token drift, the handoff distance, and especially the stakeholder-confusion risk - a non-designer mistaking the demo for a build - exceed the speed benefit. If we use Lovable at all, it is for early exploratory conversations with an explicit framing memo, never as a build artifact." Four named verdicts, three of them assigning a tool to a job, one an honest no, every one traceable to the two-week scores and none of them derived from the user-count leaderboard.

The Doc as a Reusable, Re-Runnable Instrument

The vendor decision doc is not a one-time purchase justification; it is a dated decision that names its own assumptions so it can be re-run when they change. These tools ship major updates constantly - a verdict from a bake-off six months ago may be stale because v0 improved its token handling or Lovable added a framing control - so the doc should state the date, the tool versions tested, the use case, and the scores, so that re-running it next year is cheap and the staleness is visible. A vendor decision doc that does not name its expiry conditions becomes the stale recommendation a team clings to long after the tool landscape moved, which is the design-tooling version of the lock-in the next lesson addresses.

The doc also generalizes as a method even though its verdicts do not. The specific verdicts are non-transferable - they are about your stack and your use case - but the framework (one real use case specified once, two ICs over two weeks, scored on fit-to-stack dimensions, ending in named actionable verdicts with an honest negative) is reusable for any tool decision the team faces. A strategist who runs one rigorous bake-off has not just chosen a prototyping tool; they have established the method by which the team makes every subsequent tool decision, which is the deeper deliverable, because the next four-hyped-tools decision is always coming and the method is what stops it from being made on demo-dazzle and user counts again.

Key Takeaways

  • Market traction is context, not verdict. Lovable's roughly 8 million users and $206M-$400M ARR at a $6.6B valuation, Bolt's 5 million users and $40M ARR in five months, and v0's 4 million users on Vercel tell you each tool is real and staying - and nothing about which fits your use case against your stack. Using the numbers as a leaderboard is the budgeted version of the stack-map mistake.
  • Run a bake-off, not a leaderboard, because these tools have genuinely different shapes - Figma Make for the design-review prototype, v0 component-first for engineering handoff, Lovable app-first and fast (which is also its framing risk), Bolt the high-fidelity throwaway. A leaderboard collapses different shapes into one axis and hands you a winner that may be wrong for your job.
  • Specify one real, representative use case once, identically, for all four tools. The single specification is the experiment's controlled variable; different briefs produce a verdict about the briefs, not the tools. Pick a median example of real work, not an unusually easy or hard one, so the verdict generalizes.
  • Resource it as two ICs over two weeks. Two ICs so individual familiarity averages out and you can tell "good tool" from "good in this person's hands." Two weeks because in an afternoon every tool looks magical on the happy path; token drift, invented components, handoff friction, and stakeholder confusion only emerge over days of real use.
  • Score on fit-to-stack dimensions: fidelity to spec, design-system token preservation (the most important for a team with a real system), time-to-clickable, handoff quality, iteration behavior, and stakeholder-confusion risk. "Against your stack" is the heart of the framework and exactly what a leaderboard cannot judge - which makes your verdict non-transferable, and that is a feature.
  • Produce named verdicts, not scores: a plain-language statement of what each tool is for and is not, actionable without re-running the bake-off and defensible by tracing to the scores. The verdict is often not a single winner but different tools for different jobs, and the most trust-earning verdict is the honest "do not use" that sets a tool's market traction explicitly aside as irrelevant to this decision.
  • The doc is a dated, re-runnable instrument: it names its tool versions, use case, scores, and expiry conditions so re-running it when the tools update is cheap and staleness is visible. And the framework generalizes as a method even though the verdicts do not, establishing how the team makes every future tool decision rather than just this one.