AI for Designers (UX, Product, Brand)
Capable · M8 · lesson 8 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Figma Make vs. Lovable vs. v0: A Build-vs-Buy Memo on a Real Feature
📖
now learning

Figma Make vs. Lovable vs. v0: A Build-vs-Buy Memo on a Real Feature

15 min

By 2026 you can build a clickable prototype of a real feature three different ways in an afternoon, which means the question is no longer "can AI prototype this?" but "which tool, for which job, and how do I prove it?" Lovable has reportedly crossed roughly 8 million users and hundreds of millions in ARR; Bolt hit $40M ARR in six months; v0 anchors the app-generation wave on Vercel. The hype is loud and the marketing is louder. This lesson teaches you to cut through it the only way that survives a CPO's scrutiny: pick one real backlog feature, spec it once, build the same prototype in Figma Make, Lovable, and v0, score each on four named dimensions, hand-fix the three highest-impact errors per tool, and write a one-page build-vs-buy memo with a named verdict. The artifact is the memo, three prototype URLs, a scoring table, and a per-tool fix-log.

The Memo That Ends the Argument

Every design team in 2026 is having the same unproductive argument. Someone read that Lovable is worth billions and wants to standardize on it. Someone else shipped a v0 component last sprint and swears by it. The design lead saw Figma Make at Config and wonders if it makes the others redundant. The argument goes in circles because nobody has run the same feature through all three with a fixed rubric. It is opinion versus opinion, and opinion loses to whoever is most senior or most recently impressed.

The memo ends that argument. It replaces "I like Lovable" with "on our customer-portal settings feature, Lovable scored highest on time-to-clickable and lowest on token preservation, here is the table, here are the three errors I had to hand-fix, and here is why I would use it for the day-one PM conversation but not the engineering handoff." A named verdict backed by a scoring table and a fix-log is unarguable in the way an opinion never is. The point of this lesson is not to crown a winner for all time; the tools change every quarter and any ranking is stale in ninety days. The point is to own a repeatable method that produces a defensible verdict for your feature, on your stack, this quarter. The method outlives the tools.

This is the L2 discipline scaled up to a decision. The three tools supply the speed: three working prototypes in an afternoon, which would have been a week of work in 2023. You supply the verification: the spec they are all measured against, the scoring, the fixes, and the judgment that turns three demos into a recommendation. The tools build; you decide.

Spec Once: The Fairness Rule That Makes the Comparison Real

The single most important methodological move is to spec the feature exactly once, in writing, before you open any tool, and then give all three tools the identical spec. This is the fairness rule, and skipping it is how most tool comparisons become worthless. If you prompt Lovable richly because you are excited about it and prompt Figma Make lazily because you are tired, you are not comparing the tools, you are comparing your two moods. A real comparison holds the input constant and varies only the tool.

Pick one real feature from your actual backlog, not a toy. A customer-portal account settings page is a good choice because it has real structure: a profile section, a notifications section with toggles, a billing section with a plan display and a destructive "cancel subscription" action, plus the states that matter (loading, saved, error). Write the spec to include the layout, the components, the design-system tokens that should be used, the copy, and the states. The spec is also your scoring key: every dimension you will grade is something the spec specified, so "fidelity to spec" has an objective referent rather than a vibe. Write it once, save it, and paste the same thing into all three.

The Four Scoring Dimensions, Defined So They Cannot Be Fudged

You score each tool on four dimensions, and the discipline is in defining each one concretely enough that two designers would score the same prototype the same way. Vague dimensions produce vague verdicts.

Fidelity to Spec

Does the prototype contain what the spec specified, structurally? Count it: of the sections, components, states, and copy the spec called for, how many are present and correct? A tool that omits the error state and invents a section the spec did not ask for scores lower than one that built exactly what was specified. This is the closest thing to an objective dimension because the spec is the answer key.

Design-System Token Preservation

Did the prototype honor your tokens, or did it invent its own spacing, type, and color? This is the dimension where the marketing is most misleading, because a prototype can look polished while being built entirely off-system. Check the actual values: are the spacings on your scale, is the type your ramp, are the colors your tokens, or has the tool generated a generic, plausible, off-system approximation? Token preservation is where the difference between "looks done" and "is ours" lives.

Time to Clickable

How long from pasting the spec to a working, clickable prototype you could show someone? Measure it honestly, including the prompt iterations it took to get something coherent, not just the first generation. This is the dimension the tools optimize for and market on, so it is usually their strongest, but measuring it yourself keeps the number real rather than the vendor's.

Stakeholder-Confusion Risk

This is the dimension nobody else scores and the one that matters most politically. If you put this prototype in front of a stakeholder, how likely are they to mistake it for a finished build, a commitment, or a decision already made? A tool that produces something so polished and app-like that a PM assumes it is two weeks from shipping carries high confusion risk, even if it scored well elsewhere. This dimension is why the verdict is rarely "use the highest total score"; a high-fidelity, high-confusion tool is the right choice for some conversations and exactly the wrong choice for others.

The marketing measures time-to-clickable. The job measures token preservation and stakeholder-confusion risk. A verdict that only looks at speed is a verdict that ships drift and misunderstanding faster.

The Two-Day Build and the Per-Tool Fix-Log

Run the build across two days, not two hours, because fatigue and novelty both distort a same-day comparison. Day one: build the feature in two of the tools from the identical spec, logging as you go. Day two: build it in the third and revisit the first two with fresh eyes. For each tool, you do two things. First, you log every meaningful thing it got wrong against the spec: the omitted error state, the off-token spacing, the invented component, the destructive "cancel subscription" action rendered as a cheerful primary button. Second, you hand-fix the three highest-impact errors, because a prototype you have not corrected at all is not a fair representation of what the tool plus a competent designer can produce, and because the fixes themselves are data, they tell you how hard each tool is to wrangle toward your system.

The fix-log is the artifact that makes the memo credible. It is a short table per tool: the error, its impact, and what the fix took. "Lovable: invented its own button component instead of using ours; high impact; required manually replacing and rebinding tokens, fifteen minutes." "v0: focus ring missing on primary CTA; high impact for accessibility; added focus-visible styles, five minutes." "Figma Make: destructive cancel action styled as primary; high impact; restyled to destructive treatment, three minutes." The fix-log does two jobs at once: it shows the work you did so the scoring is honest, and it surfaces the pattern of each tool's characteristic failures, which is often more decision-relevant than the raw scores.

The Named Verdict: Why It Is Rarely a Single Winner

The memo ends in a named verdict, and the discipline is that the verdict names tools to jobs, not a single overall champion. A real verdict reads like this: "Use Figma Make for the design-review prototype, because it preserves our tokens best and stays inside Figma where the team already works. Use v0 for the engineering-handoff component, because its React output is closest to what engineering will actually build and the token drift is correctable. Do not use Lovable on this feature, because its time-to-clickable is fastest but its stakeholder-confusion risk is highest, and on a billing feature with a destructive action, a stakeholder mistaking the demo for a commitment is the most expensive failure available."

That verdict is more useful than "Figma Make won" because it matches each tool to the job its profile fits. Figma Make's token preservation suits a design review. v0's React fidelity suits a handoff. Lovable's speed-plus-polish suits a fast exploration but endangers a high-stakes stakeholder conversation. The four dimensions are not meant to be summed into one number; they are meant to be read as a profile, and the verdict assigns each tool to the context where its profile is an asset rather than a liability. A memo that crowns one winner for all jobs has misunderstood the question.

Reading the Market Numbers Without Being Sold

The market data is real and worth knowing, but it is context, not a verdict. Lovable reaching roughly 8 million users and a reported valuation in the billions tells you the category is enormous and the tool is capable and well-funded; it does not tell you whether Lovable preserves your tokens on your billing feature. Bolt hitting $40M ARR in six months tells you the speed-to-prototype value proposition is resonating with a huge market; it does not tell you whether Bolt's output is safe to put in front of your CPO. v0's position on Vercel tells you it is anchored to a serious engineering platform with real React credibility; it does not tell you how much its output drifts from your design intent.

The trap is letting a valuation substitute for a verdict, because a valuation measures market traction, not fit-for-your-feature. A tool can be worth billions because it is extraordinary at making non-designers feel productive and still be the wrong tool for your specific high-stakes handoff. The market numbers belong in the memo as the one-paragraph context that explains why this comparison is worth running at all, and then the scoring table and fix-log do the actual deciding. Cite the numbers to show you are informed; rely on the rubric to show you are rigorous.

Keeping the Rubric Honest: Inter-Rater Reliability and Anti-Gaming

A scoring rubric is only as defensible as its resistance to two failure modes: two designers scoring the same prototype differently, and one designer quietly steering the result toward a favored tool. A memo that falls to either is no better than the opinion argument it was supposed to replace, so the discipline includes making the dimensions concrete enough to be inter-rater reliable and the inputs neutral enough to be ungameable.

Inter-rater reliability comes from anchoring each dimension to something countable or evidence-based rather than a feeling. Fidelity to spec becomes a checklist of specified sections, components, states, and copy, scored as the fraction present and correct. Token preservation becomes a sampled audit of actual spacing, type, and color values against the token set, scored as the fraction on-token. Time to clickable becomes a logged elapsed time from spec paste to a coherent clickable build, including iterations. Stakeholder-confusion risk, the softest dimension, becomes a rubric of defined levels tied to observable cues: low if the prototype obviously reads as a sketch, high if it is indistinguishable from a shipped app, with the level pinned to concrete signals like realistic versus placeholder data, app-like routing, working auth, and production-grade deploy URLs. When two designers disagree, they resolve it by re-checking the artifact against the cues, not by debating impressions, which is what makes the verdict survive scrutiny.

Anti-gaming comes from neutralizing the three inputs a motivated author could bend. The first is feature choice: pick the real backlog feature before any tool is favored, and require it to include the high-stakes elements the team actually faces, so the comparison cannot be rigged by choosing a feature that flatters one tool's strengths. The second is the input itself: the identical written spec given to all three tools removes the chance to prompt a favored tool more richly. The third is fix selection: choose the three hand-fixes per tool by a fixed impact rule rather than discretion, so a favored tool is not advantaged by picking flattering fixes to apply. Have a second designer review the spec for tool-neutral coverage, and require evidence for every score. Neutralize feature choice, input, and fix selection, and the verdict becomes a function of the tools rather than the author's preference, which is the entire point of running a controlled comparison instead of having an argument.

One last guard keeps the memo honest over time: timestamp it and state a validity window. A verdict is true for a feature, a stack, and a quarter, and any of those three can shift under you. Figma Make may gain better React export, v0 may improve token handling, Lovable may add a design-system import, and the moment any of them does, last quarter's verdict is stale even though it still reads as rigorous, which is the dangerous part. So treat the verdict as perishable and the method as durable: store the reusable spec and rubric as the lasting assets, mark the verdict with the date and the conditions it assumed, and re-run the same controlled comparison whenever a tool ships a major change or the evaluation cycle comes around. A team that institutionalizes the method instead of the verdict always holds a fresh, defensible recommendation rather than an aging conclusion mistaken for a standing decision.

Putting It to Work This Week

Run the real thing. Pick one feature from your actual backlog with genuine structure and at least one high-stakes element like a destructive action. Write the spec once, including layout, components, tokens, copy, and states, and save it as your scoring key. Over two days, build it in Figma Make, Lovable, and v0 from that identical spec, logging every meaningful error and hand-fixing the three highest-impact ones per tool. Score each on fidelity-to-spec, token preservation, time-to-clickable, and stakeholder-confusion risk. Then write the one-page memo with a named verdict that assigns tools to jobs, attach the three prototype URLs, the scoring table, and the per-tool fix-log, and put the market numbers in a single context paragraph.

You will know the method has landed when the circular team argument stops, because you have replaced opinion with a defensible artifact, and when a CPO reads your memo and makes a tooling decision in five minutes instead of five meetings. That is what this skill buys: not a permanent answer about which AI prototyping tool is best, because there is no permanent answer, but a repeatable method that produces a defensible verdict every time the tools shift under you. The tools build the prototypes. The memo is the part only you can build, and it is the part that decides.

Key Takeaways

  • The question in 2026 is not "can AI prototype this?" but "which tool, for which job, and how do I prove it?" The memo replaces a circular opinion argument with a defensible, named verdict for your feature on your stack this quarter.
  • Spec the feature exactly once before opening any tool, and give all three the identical spec. Holding the input constant is the fairness rule that makes the comparison real rather than a comparison of your moods.
  • Score on four dimensions defined concretely enough that two designers would agree: fidelity to spec (the spec is the answer key), design-system token preservation (looks-done versus is-ours), time to clickable (measured honestly, including prompt iterations), and stakeholder-confusion risk (the one nobody else scores and the one that matters most politically).
  • Build across two days from the identical spec, log every meaningful error per tool, and hand-fix the three highest-impact ones. The fix-log makes the scoring honest and surfaces each tool's characteristic failure pattern, which is often more decision-relevant than raw scores.
  • End in a named verdict that assigns tools to jobs, not a single champion: Figma Make for the token-faithful design review, v0 for the React-fidelity handoff, Lovable not on this feature because its confusion risk outweighs its speed on a destructive-action billing flow. Read the four dimensions as a profile, not a sum.
  • Treat the market numbers (Lovable ~8M users and a multi-billion valuation, Bolt $40M ARR in six months, v0 on Vercel) as context, not verdict. A valuation measures market traction, not fit-for-your-feature. The method outlives the tools.