โ†
AI for Translation & Localization
Strategic ยท M6 ยท lesson 6 of 20 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Building the Business Case
๐Ÿ“–
now learning

Building the Business Case

15 min

The email lands at 4:52 on a Friday, and it is three sentences long. "Finance is pushing back on the localization-AI budget line. They want the ROI model before Monday's committee, not the deck. Can you build something that survives Elena?" Elena is the CFO, and Elena has spent nineteen years learning that every operational team that walks into her office with a "productivity transformation" is quietly asking her to fund a hope. She does not care that machine translation is fluent. She does not care that adoption "jumped from 26% to 46%." She cares about three questions, and she will ask them in the same order every time: what does this cost me all-in, what does it return, and what happens when it goes wrong. If your business case cannot answer those three from a spreadsheet she can pull apart cell by cell, it is not a business case, it is a vibe with a chart on it, and she will send it back. This lesson is how you build the one that survives her. We are going to construct the cost model, the value model, and the risk model for a localization-AI program from the ground up, with real numbers, honest payback, and the two assumptions a CFO will attack first, so that when you sit across the table on Monday you are not selling speed. You are defending an investment.

The Three Questions a CFO Actually Asks

Before a single number goes on a page, understand what a business case is for, because most localization business cases fail not on their arithmetic but on their frame. A business case is a structured argument that an investment returns more than it costs at an acceptable level of risk, expressed in the language of the person who controls the money. That last clause is the one localization leaders forget. You live in words-per-day, fuzzy matches, and MQM severities. The CFO lives in cost-per-unit, payback period, and downside exposure. A business case is a translation job, and the source language is your operation while the target language is finance. If you deliver it in your own dialect, it fails the same way a fluent-but-wrong translation fails: it reads confident and means nothing to the reader who has to act on it.

The three questions decompose into three models, and a serious business case builds all three because leaving one out is how the whole thing collapses under scrutiny. Return on investment (ROI) is the ratio of net benefit to cost, usually stated as a percentage or a payback period, and it is the headline the committee remembers, but it is the least interesting part of the work. The interesting parts are the two models underneath it. First, the cost model: the fully-loaded, honest total of what the program consumes, not the seductive per-word savings alone but the tooling, the training, the governance, and the ramp. Second, the value model: what the program returns, and here is where amateurs stop at "we save money on translation" and professionals add the two harder-to-quantify axes a CFO will actually pay for, more languages shipped faster and lower catastrophic-error liability. Third, wrapped around both, the risk model: what the downside is, how likely it is, and what it costs, because a CFO trusts the person who names the risk before she has to.

A business case is a translation job. The source language is your localization operation; the target language is finance. Deliver it in cost-per-unit, payback period, and named downside, or it fails the way a fluent-but-wrong rendering fails: confident, and meaningless to the reader who must act on it.

Why the Speed Story Alone Loses the Room

The trap that kills most localization-AI business cases is that they are built entirely on the speed axis. "A hybrid workflow lifts a linguist from 2,000 words a day to 5,000 or more; therefore we save money; therefore fund us." Every word of that is true, and it will still lose, for a reason that is worth internalizing before we build anything. A CFO has seen a hundred speed stories, and she has learned that speed alone is a commodity claim: if the only thing your program does is make translation cheaper per word, then it is a cost-reduction play, and cost-reduction plays get squeezed, benchmarked against the cheapest vendor, and eventually offshored or automated to the floor. Worse, a pure speed story invites the exact question that guts it: "if the machine is so good, why do I still need the linguists at all, and why is your headcount not going down?" You have handed her the argument against your own team.

The business case that survives is the dual-axis story, and the whole strategic move of this lesson is to build it. The program does two things at once that a naive automation play does not: it captures speed and volume, and it reduces the probability and cost of a catastrophic quality failure. Those are two different axes of value, and the second one is the one a CFO cannot get from the cheapest vendor and cannot get from raw machine output. A shipped Critical error in a drug label or an indemnity clause is not a rework line item; it is a recall, a lawsuit, a regulatory finding, a client lost. When you can put a probability and a dollar figure on that downside and show that your program measurably lowers it, you have stopped selling a cost reduction and started selling risk-adjusted value, which is the only kind of value a CFO funds without a fight. Speed gets you in the door. Quality-risk reduction is what keeps the linguists employed and the budget approved.

Building the Cost Model, Honestly

Start with cost, because it is where credibility is won or lost, and because the fastest way to lose a CFO is to show her a cost model that is obviously incomplete. The seductive version of the localization-AI cost story is the per-word one: full human translation runs a certain rate, machine-translation post-editing (MTPE), the workflow where a human edits machine output rather than translating from a blank target, runs at roughly 50 to 75% of that human rate, and therefore every word gets cheaper. That is real and it is the engine of the whole case, but a CFO who has been burned before knows that the per-word line is never the whole cost, and if you present it as though it is, she will assume you either do not understand your own operation or you are hiding something. So we build the cost model in full, in four layers, and we are honest about every one.

Layer One: The Production Cost (the Per-Word Story, Told Correctly)

The first layer is the direct production cost, and here the per-word economics do most of the work, but they must be stated with their real terms. Cost-per-word is the fully-burdened cost to move one source word from source language to delivered, verified target, and it is the atomic unit of localization finance. Full human translation, on our worked example, runs at $0.20 per word all-in (a round number chosen for clean arithmetic; use your own rates). MTPE at the program's anchor of 50 to 75% of the human rate lands between $0.10 and $0.15 per word for full post-editing, and light post-editing on low-consequence content can drop as low as $0.02 per word. The savings are real, but notice what drives them: not the machine doing the work for free, but the linguist's throughput, the number of source words a linguist processes to delivery quality per day, rising from roughly 2,000 words a day on raw human translation to 5,000 or more on a hybrid MT-first workflow. Throughput is the lever. Cost-per-word is throughput's shadow.

Here is the layer-one arithmetic on a concrete annual volume. Suppose the operation localizes 10 million source words a year across its languages. At full human translation of $0.20 per word, that is $2,000,000 in production. Move the eligible content to full MTPE at $0.12 per word (the midpoint of the 50-to-75% band) and that same volume costs $1,200,000, a gross production saving of $800,000. But not all 10 million words are eligible, and this is the first place an honest model diverges from a fantasy one. High-liability content (medical labeling, legal clauses, financial disclosures) stays on full human translation or full post-editing by rule, not by preference, because a fluent Critical error there costs a life or a lawsuit. If 30% of the volume is high-liability and stays at or near human cost, then only 7 million words move to the MTPE rate, and the gross saving is closer to $560,000, not $800,000. A CFO respects the model that carves out the ineligible volume before she has to ask why the whole 10 million was assumed to be automatable.

Layer Two: The Tooling and Engine Cost

The second layer is the one the pure per-word story ignores, and its omission is the single most common reason a localization business case gets sent back. Running an MT-first pipeline is not free. There is the MT or LLM engine cost (per-character or per-token API pricing, or a licensed or hosted engine, or a fine-tuned custom engine with its training and hosting bill). There is the CAT tool and translation-management system (TMS) cost, the platform the linguist works inside and the system that routes the jobs, defined here so a finance reader is not lost: a CAT tool (computer-assisted translation) is the linguist's editing environment, and a TMS is the orchestration layer that manages projects, files, and workflows. There is the quality-estimation (QE) and evaluation tooling, the automatic scoring that routes effort. There is integration and localization-engineering time to connect the engine to the pipeline and keep it connected. On a mid-size operation these are not rounding errors; a realistic annual figure for engine plus TMS plus QE tooling and the engineering to run them can land in the low-to-mid six figures, and the model must carry it as a real line, not a footnote.

Layer Three: The Training, Ramp, and the J-Curve

The third layer is the one that separates an honest business case from a dishonest one, because it is the layer that hurts in year one and disappears from optimistic models entirely. A linguist does not hit 5,000 words a day the moment the engine is switched on. There is a learning curve: post-editing well, catching the silent Critical error, working against source and termbase rather than the smooth surface, is a distinct skill from translating, and the revised ISO 18587 exists precisely because a post-editor must hold full professional-translator competence and then some. There is a ramp cost: training the linguists, training the PMs and engineers, building the risk-tiering and quality-gate discipline into the workflow, and absorbing the lower throughput while people learn. There is a real J-curve: an investment where results dip before they rise, because you pay the tooling and training costs up front and the throughput gains arrive over quarters, not on day one. A business case that shows a clean line straight up and to the right is a business case a CFO does not believe, because she knows the J-curve is always there. The one that shows the dip, names it, and dates the crossover is the one she trusts.

The J-curve is not a weakness to hide; it is the tell that you understand your own program. Costs land first (tooling, training, absorbed ramp), gains arrive over quarters. Show the dip, name it, and date the crossover, because the CFO knows it is there and is testing whether you do.

Layer Four: The Governance and Quality Cost (the Cost of Being Defensible)

The fourth layer is the one the naive model treats as free and the serious model treats as the price of the value model's second axis. Reducing catastrophic-error liability is not automatic; it costs money to run the controls that produce it. There is the cost of the severity-scored quality gate, the human evaluation time to score output against the MQM/ISO 5060 typology (the analytic error framework, formalized by ISO 5060:2024, that classifies errors by category and by Critical/Major/Minor severity). There is the cost of terminology governance, keeping the termbase current and enforced. There is the cost of the quality record, the per-segment provenance that makes the delivery auditable and conformant. There is the cost of a governance group and incident-response readiness. Here is the pivotal insight for the whole business case: this fourth layer is not overhead to be minimized. It is the mechanism that produces the risk-reduction value in the value model, and cutting it does not save money, it destroys the second axis of the value story and turns your defensible program back into a raw-MT gamble. The governance cost is the premium you pay for the risk-reduction benefit, and a CFO understands premiums.

Building the Value Model on Two Axes

Now the return, and this is where the business case becomes strategic rather than clerical. The value model has two axes, and a business case that includes only the first is a cost-reduction play that will be squeezed forever, while a business case that credibly includes the second is a risk-adjusted value play that a CFO will defend to the board.

Axis One: Throughput, Cost, and Time-to-Market

The first axis is the one that got you in the door, and it has three components, not one. The obvious component is the cost saving, the layer-one production saving net of the layer-two-through-four costs, which we will total in the worked case below. The less obvious and often larger component is capacity: the same linguist team, at 5,000 words a day instead of 2,000, can ship two and a half times the volume without hiring, which means the operation can localize into more languages and more content types with the headcount it already has. A CFO reads this not as "cheaper" but as "more output per dollar of fixed cost," which is a different and stronger claim. The third component is time-to-market: faster localization means a product ships into a new market weeks or months sooner, and the revenue from those weeks or months is real money the CFO can model. If entering a market a quarter earlier is worth a quantifiable amount of revenue, that number belongs in the value model, and it is often larger than the per-word saving that got the meeting.

Axis Two: Quality-Risk Reduction (the Axis That Wins)

The second axis is the one that separates a business case that survives from one that gets squeezed, and it is the harder one to quantify, which is exactly why doing it well is the strategic edge. The claim is this: a governed MT-first program with a severity-scored quality gate measurably lowers the probability and the cost of a shipped Critical error, and a shipped Critical error in high-liability content is a catastrophic-cost event, not a rework line. Ground it in the program's own numbers. LLM output on medical content shows error rates around 59% on drug names, 60% on dates and times, and 66% on adverse events, every one delivered in fluent, confident prose. Raw machine output, ungated, ships those errors. A governed program with risk-tiered intake, source-grounded post-editing, and a one-Critical-fails gate is the control that stops them. The value of that control is the expected cost of the failures it prevents.

You quantify it the way an insurer quantifies a policy: expected value equals probability times cost. A single shipped Critical error in a drug label can trigger a recall, a regulatory finding, and a product-liability exposure running into the millions, plus the client relationship, which for an LSP is the entire account. Even if you assign a conservatively low annual probability to such an event under an ungoverned workflow, the cost is so large that the expected annual liability is a serious number. The governed program does not drive that probability to zero, and you must not claim it does (a CFO distrusts zero), but it drives it down by a defensible factor, and the difference between the ungoverned expected liability and the governed one is the risk-reduction value. That number, probability-times-cost avoided, sits in the value model as a real benefit, and it is the benefit the cheapest vendor and the raw engine can never claim, because they have no gate.

Quantify the risk-reduction axis the way an insurer quantifies a policy: expected liability equals probability times cost. The governed program does not zero the probability (never claim zero to a CFO); it lowers it by a defensible factor. The reduction, in dollars, is a benefit the cheapest vendor structurally cannot offer.

The Honesty That Buys Credibility on the Second Axis

The quality-risk axis is powerful and it is also where a business case can lose all its credibility by overreaching, so it must be built with visible restraint. Do not put a single dramatic number on the risk axis and lean the whole case on it, because a CFO will discount a benefit that depends on one scary scenario. Instead, present it as a range with an explicit, conservative probability you are willing to defend, and state plainly that it is the least certain part of the model. Paradoxically, naming the uncertainty makes the number more persuasive, not less, because it signals you are modeling rather than selling. The strongest version of the second axis is not "this saves us from a $10 million lawsuit," which sounds like fear. It is "under our current ungoverned workflow the expected annual liability from a shipped Critical in regulated content is X, using a deliberately conservative probability; the governed program lowers it to Y; the difference of X minus Y is the risk-reduction benefit, and here is the assumption behind the probability so you can challenge it." That last clause is the whole game.

The Worked Business Case, End to End

Now we assemble it, on one concrete mid-size operation, so you see the whole machine turning. The numbers are illustrative and round for clarity; the structure is what transfers. The operation: 10 million source words a year, currently all full human translation at $0.20 per word, so $2,000,000 in annual production cost as the baseline. The proposal: stand up a governed MT-first program.

The Cost Side, Totaled

Production, layer one. 30% of volume (3 million words) is high-liability and stays at the human rate of $0.20, costing $600,000. The remaining 7 million words move to full MTPE at $0.12 per word, costing $840,000. New annual production cost: $1,440,000, versus the $2,000,000 baseline, a gross production saving of $560,000.

Now subtract the layers the naive model forgets. Layer two, tooling: engine, TMS, CAT, and QE tooling plus the localization-engineering time to run them, call it $180,000 a year. Layer three, training and ramp: a one-time-heavy first-year cost of $150,000 to train linguists, PMs, and engineers and to absorb the throughput dip while the J-curve plays out, tapering in later years. Layer four, governance: the severity-scored gate, terminology governance, the quality record, and incident readiness, call it $120,000 a year of human evaluation and governance time. Total non-production cost in year one: $450,000.

So the year-one net production picture is a $560,000 gross saving against $450,000 of new cost, a net cash benefit of $110,000 on the cost axis alone, thin and honest. And here is where the J-curve becomes visible and where you must not flinch: in year one the throughput gains are still ramping, the training cost is at its heaviest, and the net benefit is small or, on a more conservative ramp, possibly negative. That is not a flaw in the case. That is the case being told truthfully.

The Value Side, Added

Axis one, the rest of it. The capacity gain: the same team now has the throughput headroom to localize into three additional target markets without new headcount, and the business has quantified the incremental revenue from those markets at, say, $400,000 a year once ramped. The time-to-market gain: products now reach existing markets a quarter earlier, and finance values that acceleration at another $200,000 a year in pulled-forward revenue. These are axis-one benefits beyond the raw cost saving, and they are the ones that turn a thin cost case into a real one.

Axis two, the risk reduction. Under the current ungoverned trajectory (the proposal is partly a defense against a cheaper competitor pushing raw MT), the expected annual liability from a shipped Critical error in regulated content, using a deliberately conservative probability and a mid-range cost estimate, is modeled at $300,000 (a low annual probability multiplied by a multi-million-dollar per-event cost). The governed program with its gate lowers that expected liability by a defensible factor to roughly $75,000. The risk-reduction benefit is the difference, $225,000 a year, presented explicitly as the least certain line in the model with its probability assumption exposed for challenge.

The ROI, the Payback, and the Honest Crossover

Stack it. Steady-state annual benefit once the J-curve crosses: $560,000 gross production saving, plus $400,000 capacity revenue, plus $200,000 time-to-market revenue, plus $225,000 risk reduction, totaling roughly $1,385,000 in annual benefit. Steady-state annual cost: $180,000 tooling plus $120,000 governance plus a reduced ongoing training cost of, say, $50,000, totaling $350,000. Steady-state net annual benefit: roughly $1,035,000. Against a first-year investment weighted by the heavy ramp, the program pays back inside the second year, and you state the payback period as a range, not a point, because the crossover date depends on how fast the throughput ramps and how conservatively you set the risk probability. A CFO does not want a single heroic ROI percentage. She wants a payback range with the sensitivities exposed, and giving her that is what makes the number believable.

Do not hand a CFO one heroic ROI percentage. Hand her a payback range with the two or three sensitivities that move it exposed on the page. A single confident number invites a single confident attack; a range with named assumptions invites a negotiation you can win.

The Assumptions a CFO Will Attack, and How to Survive Them

A business case is not the spreadsheet; it is the spreadsheet after it has survived interrogation. Elena will not read your model politely. She will find the two or three assumptions the whole thing rests on and push on them until they break or hold. Anticipate them, and pre-answer them on the page, because the assumption you raised and defended yourself is the assumption that does not sink you.

Attack One: "Where Does 5,000 Words a Day Come From?"

The single most load-bearing assumption in the cost model is the throughput jump from 2,000 to 5,000 words a day, because it drives the entire per-word saving. A CFO will ask where that number comes from and whether your specific content and languages will actually hit it. The honest answer is that it is an industry benchmark for a hybrid MT-first workflow, not a guarantee for your operation, and the difference matters enormously to the model: if your real steady-state throughput is 3,500 rather than 5,000, the per-word saving shrinks substantially. Survive this by never presenting the throughput figure as a fact. Present it as a benchmark you will validate in a pilot before scaling, model the case at a conservative throughput (say 3,500) as the base and 5,000 as the upside, and show that the case works even at the conservative number. A business case that only works at the optimistic throughput is a business case that fails the first honest sensitivity test.

Attack Two: "How Much of the Volume Actually Automates?"

The second attack lands on the eligible-volume assumption. You carved out 30% as high-liability and non-automatable, but a CFO will ask both directions: is it really only 30%, and could you automate more? The trap is to over-promise automatability to inflate the saving, and it is a trap because the content you wrongly classify as automatable is exactly the content where a shipped Critical error detonates the risk model. Survive this by tying the eligible-volume figure directly to the risk-tiering discipline: the split is not a guess, it is the output of a documented risk-tiered intake process, and the high-liability carve-out is a control that protects the very risk-reduction value the case depends on. If anything, err toward classifying more content as high-liability, because a slightly smaller cost saving with an intact risk model beats a larger saving that the first shipped Critical error vaporizes.

Attack Three: "You Made Up the Risk-Reduction Number"

The third attack lands on the second value axis, and it is the sharpest, because the risk-reduction benefit rests on a probability you estimated. A CFO is right to be skeptical of a benefit that depends on a scary hypothetical. Survive this not by defending the number but by defending the method. State the probability conservatively, show the per-event cost is grounded in real recall and liability precedent for the industry, present the benefit as a range, and, most importantly, frame the risk axis as the reason the governance cost exists rather than as a standalone windfall. The governance layer costs $120,000 a year; the risk-reduction benefit is what that $120,000 buys. Presented that way, the risk number is not a fantasy revenue line, it is the return on a specific, visible cost, and a CFO can evaluate a return on a cost even when the return is probabilistic. That reframing, from "trust my scary number" to "here is the cost and here is what it defends against," is what makes the second axis survive.

Attack Four: "If the Machine Is So Good, Why Isn't Headcount Falling?"

The fourth attack is the one that endangers your team, and it comes from the pure speed story you must have avoided from the start. If the case was built only on cheaper words, the natural CFO conclusion is that the linguists are a cost to be reduced. The dual-axis story is the answer, and this is where it earns its place: the linguists are not the cost the program reduces, they are the mechanism that produces the risk-reduction value the program sells. The engine produces fluency for free and guarantees accuracy on nothing; the linguist is the control that catches the silent Critical error, and that control is the entire second axis of the value model. Cutting the linguists does not save money, it deletes the governance layer and turns the defensible program back into a raw-MT gamble that the risk model just showed is worth hundreds of thousands in expected liability. The correct headcount story is not "fewer linguists," it is "the same linguists moved up the value chain from words-per-hour to quality ownership, producing more output and defending against a catastrophic downside." That is a story a CFO can take to the board.

Presenting the Case So It Survives the Room

The model can be right and still lose if it is presented wrong, so a few disciplines about the delivery, because at L4 you are not just building the case, you are defending it in the room. Lead with the dual-axis frame, not the per-word saving, so the first thing the committee hears is "this captures speed and lowers catastrophic-error liability," not "this makes translation cheaper," because the frame you open with is the frame the whole discussion happens inside. Present the cost model in full, all four layers, before the value model, because volunteering the complete cost first is what earns you the credibility to be believed on the value. Show the J-curve honestly, with the year-one dip visible and the crossover dated, because the person who hides the dip loses trust the moment the CFO finds it herself, and she will. Give payback as a range with sensitivities, never a single percentage. And name the risks before she does: the throughput may underperform, the eligible volume may be smaller, the risk probability is an estimate, and the ramp may take longer than planned. The leader who names the four risks herself is the leader the CFO funds, because the unnamed risk is the one that ambushes the committee, and a CFO's entire job is avoiding ambushes.

One final discipline, and it is the one that turns a funded pilot into a funded program. Build the case so its own success is measurable against the exact numbers you promised. If you claimed 3,500 words a day at the conservative base, instrument the pilot to report actual throughput. If you claimed a risk-reduction benefit resting on the gate catching Criticals, instrument the gate to count the Criticals it caught, because every caught Critical is a shipped-error-prevented and a live data point that validates the second axis. A business case that specifies how it will be proven is a business case a CFO can say yes to twice: once to fund the pilot, and once, when the pilot reports back against its own promised numbers, to fund the scale-up. The numbers you promise on Monday are the numbers you will be measured against next quarter, so promise conservatively and instrument thoroughly, and let the pilot's real results make the second, larger argument for you.

A business case that specifies how it will be proven is one a CFO can approve twice: once to fund the pilot, and again when the pilot reports back against its own promised numbers. Promise conservatively, instrument thoroughly, and let real results make the larger argument for you.

Key Takeaways

  • A business case is a translation job into the CFO's language: it must answer what the program costs all-in, what it returns, and what happens when it goes wrong, expressed in cost-per-word, payback period, and named downside, not in words-per-day and MQM severities.
  • The speed story alone loses the room, because a pure cost-reduction play gets squeezed, benchmarked to the cheapest vendor, and invites the question "why keep the linguists at all." The winning frame is the dual-axis story: the program captures speed and volume AND lowers the probability and cost of a catastrophic quality failure, and the second axis is the one the cheapest vendor cannot match.
  • Build the cost model in four honest layers, not one: layer one is production (MTPE at 50 to 75% of the $0.20 human rate, driven by throughput rising from ~2,000 to 5,000+ words a day, with high-liability volume carved out and kept at human cost); layer two is tooling (engine, CAT, TMS, QE); layer three is training and ramp with a visible J-curve; layer four is governance, which is not overhead but the mechanism that produces the risk-reduction value.
  • Build the value model on two axes: axis one is cost saving plus capacity (more languages and content on the same headcount) plus time-to-market revenue; axis two is quality-risk reduction, quantified like an insurance policy as probability times cost avoided, grounded in the LLM medical error rates (~59% drug names, ~60% dates/times, ~66% adverse events) that a governed gate stops and raw MT ships.
  • On the worked mid-size case (10M words, $2M baseline), a 30% high-liability carve-out yields a $560K gross production saving; subtract ~$450K of first-year tooling, training, and governance for a thin, honest year-one cost-axis net; add capacity revenue, time-to-market revenue, and a conservatively modeled risk-reduction benefit for a steady-state net around $1M and a payback inside the second year, stated as a range with sensitivities exposed.
  • Pre-answer the four assumptions a CFO attacks: the 5,000-words-a-day throughput (present as a benchmark to validate in a pilot; model the case at a conservative 3,500 and prove it still works); the eligible-volume split (tie it to documented risk-tiering, err toward more high-liability, never inflate automatability); the risk-reduction number (defend the method not the figure, and frame it as the return on the governance cost); and the headcount question (linguists are the control that produces the risk-reduction value, not the cost to cut).
  • Never hide the J-curve or hand over a single heroic ROI percentage: costs land first and gains arrive over quarters, so show the year-one dip, date the crossover, give payback as a range, and name the risks yourself, because the leader who names the downside before the CFO does is the leader the CFO funds.
  • Instrument the case to prove itself: specify the exact throughput, saving, and Criticals-caught numbers the pilot will report, so the CFO can approve twice, once to fund the pilot and once to fund the scale-up when the pilot reports back against its own conservative promises, letting real results make the larger argument.