Success Metrics for Localization AI
The quarterly business review is in six days, and the deck has one slide that terrifies you. It is the localization-AI scorecard, and right now it says one thing: "Throughput up 137%." Your VP loves that number. She has already put it in the email to her boss. And you know, sitting alone with the actual data at 9pm, that the number is a trap. Because in the same quarter that throughput jumped, two files went out with a flipped negation in a safety warning, a medical device client opened a formal complaint about a mistranslated contraindication, and your termbase conformance quietly slid because the engine kept preferring a synonym the client explicitly banned. The 137% is true. It is also, on its own, a lie, because it tells the half of the story that flatters the machine and buries the half that gets you sued. This lesson is about the other half. It is about choosing the small set of numbers that tell the whole story, the four that capture speed and quality-risk in the same breath, so that when your VP asks "how is the AI program doing," you can hand her a scorecard that is honest enough to survive an audit and clear enough to survive a board meeting. We are going to define those four numbers, prove why these four and not the seductive vanity metrics the vendors push, set targets and baselines you can defend, show exactly how they interlock so a gain in one that damages another is caught rather than celebrated, and then build a complete metrics scorecard for a real operation, cell by cell.
Why Four Numbers, and Not Forty
Every localization-AI dashboard you have ever seen has too many numbers on it, and that is not an accident, it is a symptom. When you do not know which numbers actually tell the story, you hedge by showing all of them: edit-distance, BLEU, segments-per-hour, fuzzy-match leverage, cost-per-word, MQM score, engine confidence, reviewer pass rate, and eleven more, each with a sparkline, none of them load-bearing. The result is a dashboard nobody can read and nobody trusts, because a metric that means everything means nothing, and a scorecard that measures forty things measures your inability to decide which four matter. The discipline of this lesson is subtraction. A great scorecard is defined as much by what it refuses to show as by what it shows, and the strategist's job is to find the smallest set of numbers that, taken together, cannot be gamed without the gaming showing up somewhere on the same page.
The four numbers are throughput, cost-per-word, critical-error rate, and terminology conformance. They are not chosen because they are the easiest to measure (two of them are genuinely hard) and not because a vendor recommended them (the vendors recommend the opposite). They are chosen because of a single structural property: together they span both axes of the localization-AI value story at once, the speed axis and the quality-risk axis, and they span them in a way where you cannot move one axis without the other axis's numbers reacting on the same scorecard. Throughput and cost-per-word are the speed and economics axis, the part that flatters the machine. Critical-error rate and terminology conformance are the quality-risk axis, the part that keeps you out of court. A scorecard that shows only the first pair is the 137% slide that gets someone hurt. A scorecard that shows only the second pair is a quality report that finance will not fund. The four together are the whole story, and the whole story is the only story worth reporting.
A great scorecard is defined by what it refuses to show. Four numbers that span both the speed axis and the quality-risk axis, measured on the same page, beat forty numbers that let a speed gain hide a quality collapse.
The Non-Negotiable: They Must Be Reported Together
The most important rule about these four numbers is not how you define any one of them. It is that you never report one without the other three. This is the rule that survives contact with the real world, because the real world is full of people who want to lift one number out of the set and wave it around. Your VP wants to quote throughput. A cost-focused finance partner wants to quote cost-per-word. A nervous quality manager wants to quote the critical-error rate. Each of them, quoting one number alone, is telling a fractional truth that becomes a whole lie the moment it leaves the room without its companions. The scorecard is a unit. Its power is entirely in the interlock, which we will build in detail later, where a change in one number is only meaningful against the simultaneous change in the others. A throughput jump is a triumph if critical-error rate held flat and a catastrophe if it did not, and the only way to know which is to see both numbers on the same line, in the same period, for the same content.
Throughput: The Speed Number That Does Not Lie
Start with the number your VP already loves, because it is the one that got the program funded and it is the easiest to define, and getting it exactly right is how you earn the credibility to be believed on the three that follow. Throughput is the volume of source content a linguist or a team moves from raw machine-translation output to delivered, verified target quality in a unit of time, most usefully expressed as source words per linguist per day. Note every word of that definition, because each one is doing work. It is source words, not target words, so a language pair that expands text is not credited with phantom productivity. It is per linguist per day, a rate normalized to the person and the time, not a raw batch total that hides how many people worked how long. And it is to delivered, verified target quality, which is the clause that stops throughput from being a lie: words that shipped but failed the quality gate do not count, because a word you had to redo was not throughput, it was rework wearing throughput's clothes.
The reason throughput is the honest speed number, and edit-distance is not, comes down to what each one measures. Throughput measures the thing the operation actually sells: finished, deliverable words. It is denominated in the unit the client pays for and the CFO models. When throughput rises from a baseline of roughly 2,000 words per linguist per day on raw human translation to 5,000 or more on a hybrid machine-translation post-editing (MTPE) workflow, the workflow where a human edits machine output rather than translating from a blank target, that is a real capacity gain the business can spend: more languages, more content, the same headcount. The number is legible to everyone, from the linguist to the board, and it cannot be quietly inflated without the inflation showing up as failed deliveries somewhere else on the scorecard. That legibility and that coupling to real delivered value are exactly why it earns a seat and why the vendor's favorite speed proxies do not.
The Throughput Trap: Speed Without the Quality Denominator
Throughput has one failure mode, and it is the whole reason the other three numbers exist. Throughput measured alone rewards the wrong behavior, because the fastest way to raise words-per-day is to edit less, and editing less is exactly how a silent critical error slips through. A linguist under pressure to hit 6,000 words a day learns, consciously or not, to trust the smooth surface of the machine output and stop cross-checking every number, negation, and term against the source segment. Their throughput soars. Their critical-error rate soars with it, invisibly, until a shipped mistranslation surfaces as a client complaint or a recall. This is why throughput can never be a solo metric and why the phrase "to delivered, verified target quality" is load-bearing in its definition. Throughput is a real number and a good one, but it is only trustworthy when it is chained to a quality number that moves in the opposite direction when the linguist starts cutting corners. Speed you cannot verify is not speed. It is a debt.
Cost-Per-Word: The Economics Number, Fully Burdened
Cost-per-word is the fully-burdened cost to move one source word from source language to delivered, verified target, and it is the atomic unit of localization finance, the number every other financial claim ultimately reduces to. It sounds simple and it is where most scorecards quietly cheat, because the tempting version of cost-per-word counts only the linguist's per-word rate and calls the drop from full human translation to MTPE the whole savings. That version is a lie of omission, and a CFO who has seen a localization business case before will spot it instantly. The honest cost-per-word carries the full stack: the linguist's post-editing labor, the machine-translation or large-language-model engine cost (the per-token or per-character API bill, or the licensed or hosted or fine-tuned engine's amortized cost), the computer-assisted translation (CAT) tool and translation-management system (TMS) platform costs (the linguist's editing environment and the orchestration layer that routes jobs), the quality-estimation and evaluation tooling, and, critically, the governance overhead that produces the quality-risk reduction. Cost-per-word that excludes governance is not a cheaper number, it is a dishonest one, because it books the savings while hiding the cost that makes the savings safe.
Told correctly, cost-per-word is the number that translates the throughput story into money. If full human translation runs $0.20 per word all-in and full MTPE lands at the program anchor of 50 to 75% of that, roughly $0.10 to $0.15 per word, the per-word saving is real and large. But the honest scorecard also shows that light post-editing on low-consequence content can drop as low as $0.02 per word while high-liability content stays at or near the full human rate by rule, so the blended cost-per-word depends entirely on the content mix. A scorecard that reports a single flattering cost-per-word without the content mix behind it is reporting an average that could hide anything. The strategic use of cost-per-word is not to celebrate a low number, it is to show the blended number alongside the mix that produced it, so that a suspiciously low cost-per-word triggers the right question: what content did we move to the cheap workflow to get there, and was any of it content that should never have gone?
Cost-per-word that excludes engine, tooling, and governance is not a cheaper number, it is a dishonest one. It books the savings while hiding the cost that makes the savings safe.
Why Cost-Per-Word and Throughput, Not Edit-Distance or BLEU
This is the moment to say plainly why two metrics the whole industry reaches for are absent from the four, because their absence is a strategic choice, not an oversight. Edit-distance, usually measured as a post-edit-distance percentage, counts how much the linguist changed the machine's output, character by character or word by word. BLEU (bilingual evaluation understudy) scores machine output by how closely it matches a reference translation using n-gram overlap. Both are seductive because they are automatic, cheap, and produce a tidy number without a human in the loop. Both are dangerous as scorecard metrics for the same reason: they measure similarity, not correctness, and in localization those two things come apart at exactly the point where the money and the liability live.
Consider what edit-distance actually rewards. A low post-edit distance means the linguist changed little, which the naive reader interprets as "the machine was good." But a linguist who is rushing, or who trusts the fluent surface, also produces a low edit-distance, and so does an engine that produced a fluent, confident, completely wrong sentence that the linguist skimmed past. Edit-distance cannot tell the difference between "the machine was right so I changed nothing" and "the machine was wrong and I missed it." It rewards under-editing, which is the precise behavior that ships the silent critical error. Optimize for low edit-distance and you are optimizing for the failure mode the entire discipline exists to prevent. BLEU fails the same way from a different angle: a rendering can score high on n-gram overlap while flipping a dosage or dropping a negation, because a single swapped word barely dents the overlap score but completely inverts the meaning. Neither number can see the one error that matters most, the fluent mistranslation, and a metric that is blind to your worst failure mode is not a quality metric, it is a comfort blanket. The four we chose are harder to measure precisely because they refuse to look away from the thing that hurts.
Critical-Error Rate: The Number That Keeps You Out of Court
Here is the number the vendors will never put on the default dashboard, and the one that matters more than the other three combined on high-liability content. Critical-error rate is the frequency of Critical-severity errors in evaluated output, expressed per unit of volume, most defensibly as Criticals per 1,000 words or per 10,000 words evaluated. A Critical error, under the MQM/ISO 5060 error typology (the analytic evaluation framework formalized by ISO 5060:2024 that classifies every error by category and by Critical/Major/Minor severity), is an error severe enough that it renders the content unfit for use and can cause real-world harm: a flipped dosage, a dropped negation in a safety instruction, an inverted indemnity clause, an invented number in a financial disclosure. The defining property of a Critical error is that it fails the file regardless of how clean everything else looks. One Critical is not a deduction from a score, it is a gate: the file does not ship. So the critical-error rate is not measuring average quality, it is measuring how often your process produces the specific failure that ends relationships, triggers recalls, and produces lawsuits.
Why this number and not an overall quality score? Because averages hide catastrophes. A file can score 98% on an aggregate quality metric and still contain the one flipped dosage that kills someone, and the 98% will actively work to reassure everyone that the file is fine. The critical-error rate refuses that comfort. It isolates the tail-risk event and counts it directly, which is the only honest way to measure a failure mode whose cost is not proportional to its frequency. This connects to the hard ground truth the whole program is built on: studies of large-language-model output on medical content found error rates around 59% on drug names, 60% on dates and times, and 66% on adverse events, every one of them delivered in grammatically perfect, confident prose. Raw, ungated machine output ships those errors at those rates. The entire economic and safety case for a governed MT-first program is that the process drives the shipped critical-error rate down toward zero, and the only way to prove that claim, to a client, an auditor, or a court, is to measure it and report it as a first-class number, not to bury it under a flattering average.
Measuring Critical-Error Rate Honestly, Including the Sampling Problem
Critical-error rate is the hardest of the four to measure well, and pretending otherwise is how scorecards lie. The measurement requires a human MQM evaluation, because the whole point is to catch the fluent error an automatic metric cannot see, which means it costs real evaluator time and cannot be run on 100% of volume in most operations. So you sample, and the sampling is where the number can be quietly corrupted. If you sample only the easy content, or only the segments the engine was confident about, or only after the post-editor has already cleaned the file, you will measure a critical-error rate near zero that means nothing, because you sampled away the risk. An honest critical-error rate is measured on a representative, risk-weighted sample: oversample high-liability content, sample the raw-plus-post-edited output that actually ships, and sample enough volume that a rate of "one Critical per 20,000 words" is statistically meaningful rather than an artifact of a tiny denominator. A single evaluated file proves nothing. The scorecard must state the sample size and the sampling method alongside the rate, because a critical-error rate without its denominator and its sampling frame is not a measurement, it is a guess dressed as one.
Averages hide catastrophes. A file can score 98% and still carry the one flipped dosage that kills someone. Critical-error rate isolates the tail-risk event and counts it directly, because its cost is not proportional to its frequency.
Terminology Conformance: The Consistency Number the Engine Fights
The fourth number is the quiet one, the one that never causes a single dramatic incident but bleeds quality and client trust in a thousand small cuts. Terminology conformance is the percentage of in-scope terminology occurrences in delivered output that match the client's approved term in the termbase, measured against the total occurrences of terms that have an approved entry. If the client's termbase mandates "infusion pump" and forbids "IV pump," and across the delivery there are 400 occurrences of that concept, and 388 use the approved term, terminology conformance for that term is 97%. Aggregate across all governed terms and you have the operation's terminology-conformance rate, a direct measure of whether the approved language actually held through the machine and the post-editing, or drifted segment by segment into whatever synonym the engine preferred.
Terminology conformance earns its seat for two reasons the other three cannot cover. First, it measures a failure mode that is neither catastrophic like a Critical nor economic like cost-per-word, but corrosive: an engine that cheerfully ignores the client's approved device name because it "prefers" a more common synonym produces output that is fluent, technically not wrong, and completely off-brand or off-spec, and it does this consistently, across every segment, because the engine's preference is stable. This is a systematic error, not a random one, which makes it both more damaging and more measurable than a one-off slip. Second, terminology conformance is the leading indicator on the scorecard, the number that moves first when engine grounding degrades, when a termbase goes stale, or when a new content type is routed to an engine that was never taught the client's terms. A dip in terminology conformance is often the early warning that predicts a later rise in critical-error rate, because the same loss of control that lets an approved term drift is the loss of control that eventually lets a negation flip. It is the canary, and a scorecard without it is a scorecard that finds out about a grounding failure only after it has already caused a Critical.
A Subtlety: Conformance Is Not the Same as Terminological Correctness
Terminology conformance measures adherence to the approved termbase, which is a different and narrower thing than terminological correctness, and conflating the two is a mistake a strategist must not make. Conformance is high when the output matches the termbase. But if the termbase itself is wrong, out of date, or missing an entry, high conformance can mean the operation is consistently, measurably delivering the wrong term with perfect fidelity. This is why terminology conformance is a process-control metric, not a truth metric: it tells you the approved language held, which is exactly what you want to measure at the scorecard level, while the correctness of the termbase itself is a separate governance responsibility upstream. The scorecard measures whether the operation did what the termbase said. Keeping the termbase right is the terminologist's job, and the two must not be collapsed, or a stale termbase will produce a beautiful conformance number that certifies the wrong thing.
Baselines, Targets, and How the Four Interlock
A number without a baseline is a rumor, and a number without a target is a decoration. Before any of the four means anything on a scorecard, it needs two reference points: where you started (the baseline, measured before the AI program or at the start of the period) and where you are trying to get to (the target, set deliberately and defended). And here is the strategic core of the entire lesson: the targets for the four numbers are not independent. You do not set a throughput target in one room and a critical-error target in another. You set them as a system, with explicit rules about how a movement in one constrains the acceptable movement in the others, because the entire reason to report the four together is that they interlock.
Setting Defensible Baselines and Targets
For throughput, the baseline is your pre-AI rate, realistically around 2,000 source words per linguist per day for full human translation, and a defensible target is a range, not a point: 3,500 as a conservative base and 5,000+ as the upside, validated in a pilot before it is promised at scale. For cost-per-word, the baseline is your current fully-burdened rate (say $0.20/word all-in on human translation), and the target is a blended rate that reflects the real content mix, not the best-case MTPE rate applied to all volume. For critical-error rate, the baseline is the honest measured rate under your current workflow, and here the target is unusual: it is not "lower is better without limit," it is a hard ceiling that the file-level gate enforces at zero shipped Criticals, with the measured rate on the pre-gate sample trending toward zero as the process matures. For terminology conformance, the baseline is your current measured adherence (often a humbling number the first time you measure it honestly, frequently in the low 90s or worse), and the target is high and specific: 98%+ on governed terms for most content, with high-liability content held higher still.
The Interlock Rules: When a Gain Is Not a Win
Now the rules that make the scorecard more than four unrelated gauges. These are the constraints that turn four numbers into one honest verdict, and they are the reason the whole lesson exists.
- A throughput gain that raises critical-error rate is not a win, it is a regression. This is the master rule. If words-per-day rose 40% but Criticals per 10,000 words also rose, the program got faster at producing the exact failure it exists to prevent. The correct reading of the scorecard is not "throughput up," it is "throughput up, quality-risk up, net negative," and the corrective action is to slow down, not to celebrate. The two numbers on the same line make this reading unavoidable, which is precisely why they must never be separated.
- A cost-per-word drop that came from moving high-liability content to a cheap workflow is not a saving, it is a deferred liability. If cost-per-word fell but the content mix shifted regulated content into light post-editing, the scorecard should show the blended cost improving while the critical-error rate on that content segment worsens or its sample thins suspiciously. Cheap words that carry catastrophic risk are the most expensive words you will ever ship.
- A terminology-conformance dip is an early warning even if the other three still look good. Because conformance is the leading indicator, a drop there should trigger investigation before it propagates into critical-error rate. A scorecard reader who sees throughput up, cost down, Criticals flat, and conformance sliding should not relax on the strength of three green numbers; the amber one is telling them the grounding is degrading and the green will not last.
- The critical-error rate has veto power over the other three. No throughput number, no cost number, and no conformance number can redeem a rising shipped-Critical rate on high-liability content. The critical-error rate is the metric with a gate behind it, and the gate does not negotiate with efficiency. This is the scorecard expression of the program's cardinal rule: one Critical fails the file, and no amount of speed or savings buys it back.
A speed gain that raises the critical-error rate is not a win, it is a regression: the operation got faster at producing the exact failure it exists to prevent. The interlock is the whole point, which is why the four numbers must share one page.
The Worked Scorecard for an Operation
Now assemble it, on one concrete mid-size operation, so the whole machine turns in front of you. The operation localizes 10 million source words a year across a language set, with a content mix of roughly 60% marketing and UI (low-to-moderate liability), 10% product documentation (moderate), and 30% regulated content (medical labeling, legal clauses, financial disclosures, high liability). The scorecard is reported quarterly, per content tier and blended, with baseline and target columns beside the current-period actuals, because a number with nothing to compare it to is not a metric.
Reading the Scorecard Across, Not Down
The discipline of reading this scorecard is to read across each metric with its baseline and target, and then to read the four metrics against each other for the period, applying the interlock rules. Here is the quarter's story in numbers.
- Throughput. Baseline 2,000 words/linguist/day (pre-AI). Target range 3,500 base to 5,000 upside. This quarter's actual: 4,100, verified against delivered-and-passed volume only, with reworked words excluded. Reading: solidly above the conservative base, below the upside, healthy and believable.
- Cost-per-word (blended). Baseline $0.20/word all-in. Target blended $0.14 reflecting the 30% high-liability carve-out held at human cost. This quarter's actual: $0.145 blended, with the mix shown beneath it (regulated content held at $0.19, documentation at $0.12 full MTPE, marketing/UI at $0.09 including some light PE). Reading: on target, and the mix beneath the blend proves the saving did not come from cheating the high-liability tier.
- Critical-error rate. Baseline (measured, pre-gate, on a risk-weighted sample) 1 Critical per 12,000 words. Target: zero shipped Criticals (gate-enforced), pre-gate sample trending down. This quarter's actual: zero shipped Criticals across all tiers; pre-gate sample rate improved to 1 per 20,000 words on a 40,000-word risk-weighted sample (sample size and method stated). Reading: the gate held, and the underlying process is genuinely improving, not just being caught at the gate.
- Terminology conformance. Baseline 93% on governed terms. Target 98%+, high-liability 99%+. This quarter's actual: 97.2% blended, but 99.1% on regulated content and 95.8% on marketing/UI. Reading: the blended number is below target, and the amber lives in the marketing/UI tier, where a newly onboarded engine has not been fully grounded on the client termbase. This is the leading indicator doing its job.
Now read the four against each other, which is the part a single-metric dashboard makes impossible. Throughput is up and the critical-error rate held at zero shipped with an improving underlying rate: by the master interlock rule, this is a genuine win, speed and quality-risk moving the right way together. Cost-per-word is on target and the mix beneath it confirms no high-liability content was quietly moved to a cheap workflow to hit the number: the saving is real, not deferred liability. But terminology conformance is sliding in the marketing/UI tier, and the interlock rules say a conformance dip is an early warning even when the other three look good. So the honest verdict on this quarter is not "three green, one amber, mostly great." It is: "the program is delivering speed and cost on target with quality-risk controlled, and there is one specific, addressable degradation, engine grounding on the new marketing content, that must be fixed this quarter before it propagates into the critical-error rate." That verdict, actionable and honest, is what four interlocked numbers produce and forty disconnected ones never can.
What This Scorecard Lets You Say in the Room
Return to the quarterly business review that opened this lesson, and see what the four-number scorecard lets you say instead of "throughput up 137%." You can say: "Throughput is up to 4,100 words a day, comfortably above our conservative target, and it is real throughput, delivered and passed, not rework. Blended cost-per-word is on target at 14.5 cents, and here is the mix that proves we did not get there by cutting corners on regulated content. We shipped zero Critical errors this quarter across every tier, and our underlying pre-gate error rate is improving, which is the number that keeps us out of court and keeps the medical client. And we have one thing to fix: terminology conformance dipped on the new marketing engine, we caught it early because we measure it, and it is scoped for correction this quarter before it can affect anything downstream." That is a scorecard that survives an audit, survives a board meeting, and survives the CFO who pulls it apart cell by cell, because every number is honest, every number has its baseline and target, and the four together tell the whole story: fast, economical, safe, and controlled, with the one soft spot named by you before anyone had to find it. The 137% slide could never say any of that. It could only flatter the machine until the machine got someone hurt.
Key Takeaways
- Four numbers tell the whole localization-AI story, and the discipline is subtraction: throughput and cost-per-word (the speed and economics axis) plus critical-error rate and terminology conformance (the quality-risk axis). Together they span both axes on one page; any two alone tell a fractional truth that becomes a whole lie the moment they leave the room without their companions.
- Throughput is source words per linguist per day to delivered, verified target quality, not raw batch totals and not words that failed the gate; the "verified quality" clause is load-bearing because it stops rework from masquerading as throughput. Baseline ~2,000 words/day human, target 3,500 base to 5,000+ upside, validated in a pilot.
- Cost-per-word is the fully-burdened cost per source word including engine, CAT/TMS, QE tooling, and governance, reported as a blend with the content mix shown beneath it; a cost-per-word that excludes governance or hides the mix is dishonest, because a suspiciously low number usually means high-liability content was moved to a cheap workflow it should never have entered.
- Edit-distance and BLEU are deliberately excluded: both measure similarity, not correctness, and both reward under-editing, the exact behavior that ships the silent critical error. A single swapped word barely dents either score while completely inverting the meaning, so a metric blind to the fluent mistranslation is a comfort blanket, not a quality gate.
- Critical-error rate counts Critical-severity (MQM/ISO 5060) errors per unit of volume because averages hide catastrophes: a file can score 98% and still carry one flipped dosage. It must be measured on a representative, risk-weighted human-evaluated sample with the sample size and method stated, and its target is a hard zero shipped Criticals (gate-enforced) with the pre-gate rate trending down.
- Terminology conformance is the percentage of governed-term occurrences matching the approved termbase; it is the corrosive-failure metric and the leading indicator, moving first when engine grounding degrades or a termbase goes stale, often predicting a later rise in critical-error rate. It measures conformance to the termbase, not the correctness of the termbase itself, which is a separate upstream responsibility.
- The four interlock, and the interlock is the whole point: a throughput gain that raises critical-error rate is a regression, not a win; a cost-per-word drop from moving high-liability content to a cheap workflow is deferred liability, not saving; a conformance dip is an early warning even when the other three look good; and critical-error rate has veto power, because one shipped Critical fails the file no matter how good the speed and cost look.
- Report the four together, always, each with a baseline and a target, per content tier and blended, and read across the metric before reading down the dashboard. On the worked operation this turns "throughput up 137%" into an honest verdict: speed and cost on target, zero Criticals shipped, and one scoped, self-caught grounding issue named before anyone had to find it, a scorecard that survives an audit, a board, and a CFO pulling it apart cell by cell.
Skill.re