โ†
AI for Translation & Localization
Strategic ยท M3 ยท lesson 3 of 20 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Avoiding Metrics That Hide Risk
๐Ÿ“–
now learning

Avoiding Metrics That Hide Risk

15 min

The quarterly business review deck was, by every visible measure, a triumph. On the third slide, the head of localization at a medical-device company projected the localization-AI scorecard she had spent a quarter building. Throughput was up 168%: her linguists had gone from roughly 2,000 words a day to more than 5,000 once the machine-translation-first workflow landed. Cost per word had dropped by half. The post-edit distance chart, the one that showed how much editing each machine-translated segment needed, trended down and to the right, a clean descending line that read, to every executive in the room, as "the machine is getting better and our people are barely touching it." The aggregate pass rate across all languages sat at 98.7%. The CFO nodded. The COO asked when they could double the volume. Nobody in that room knew that eleven weeks earlier a Turkish instructions-for-use file had shipped with an inverted warning, a sentence that told a clinician the device was safe to use during a procedure the English original explicitly warned against, and that the file had passed every metric on the slide. It scored a low post-edit distance because the post-editor barely touched the fluent, confident, wrong sentence. It landed inside the 98.7% because a single flipped clause does not move an average built from thousands of clean segments. The dashboard did not hide the risk by accident. It hid the risk because of what it chose to measure. This lesson is about how a metrics program buries the one error that matters, why the most reassuring numbers in localization are often the most dangerous, and how a strategist rebuilds the scorecard so the critical error cannot hide inside it.

The Seduction of the Clean Number

Before we take apart any specific metric, you have to understand why a decision-maker reaches for the wrong ones in the first place, because the pull is structural and it will act on you exactly when you are most exposed. A metric is a proxy: a number that stands in for something you actually care about but cannot observe directly. You care about quality, safety, and defensibility. You cannot put those on a slide. So you reach for a number that correlates with them, or seems to, and you put that on the slide instead. The entire risk of a metrics program lives in the gap between the thing you care about and the proxy you chose to represent it.

A vanity metric is a number that reliably goes up, is easy to measure, makes the operation and its owner look good, and does not actually track the outcome that would hurt you if it went wrong. Vanity metrics are seductive precisely because they are honest about the easy things and silent about the hard ones. Throughput genuinely rose. Cost per word genuinely fell. Those numbers are not lies. They are true answers to questions that were never the ones keeping you up at night. The danger is not that a vanity metric is false. The danger is that it is true, prominent, and irrelevant to the failure that ends a contract or a life, and its very truth lends the whole dashboard a credibility that the dashboard has not earned on the dimension that matters.

A vanity metric is not a lie. It is a true answer to the wrong question, displayed where it will be mistaken for an answer to the right one.

There is a second force at work, and it is worse than laziness. The metrics you choose do not just describe the operation. They steer it. This is Goodhart's law in its localization form: the moment a number becomes a target that people are rewarded against, the number stops measuring what it used to measure, because everyone starts optimizing the number instead of the thing the number was a proxy for. If you pay a post-editor by throughput, you will get throughput, and you will get it by having them stop reading the source segment closely enough to catch the flipped warning, because reading for meaning is slow and the incentive rewards speed. If you celebrate a falling post-edit distance, you teach your linguists that touching the machine output less is good, which is precisely the behavior that lets a fluent critical error sail through untouched. A metric is never a passive mirror. It is an instruction to your organization about what to do more of, and if you point it at the wrong target, it will faithfully manufacture the wrong behavior. Choosing metrics is therefore not a reporting decision. It is an operational-design decision with a body count.

Why the Strategist Owns This, and Not the Analyst

An analyst can compute any number you ask for. What an analyst cannot do, and what only the person accountable for the operation can do, is decide which numbers are allowed on the scorecard that leadership sees and clients are quoted against. That is a governance act, not an analytics act. The strategist owns the metrics program because the strategist owns the consequence of the metric steering the organization wrong. If a critical error ships and the post-incident review finds that the dashboard was built to make it invisible, "the numbers looked good" is not a defense; it is the indictment. The revised ISO 18587 (the standard for post-editing of what it now calls "non-human translation output," expanded to cover AI and large language models, in DIS ballot with publication targeted for late 2025 into 2026) makes the human accountable for quality and insists the post-editor hold full professional-translator competence. That accountability does not evaporate because a metric said the file was fine. It attaches to whoever chose to let a metric speak for quality it could not see.

Edit Distance: The Effort Number Wearing a Quality Costume

Start with the metric that fooled the medical-device dashboard, because it is the most common and the most misunderstood trap in the entire localization stack. Edit distance, reported in localization as post-edit distance or "edits per segment," measures how much a post-editor changed the raw machine output on the way to the final delivered text: the count of insertions, deletions, and substitutions it took to turn the machine's draft into what shipped. It is easy to compute automatically, it produces a clean trend line, and it looks, to anyone who has not thought hard about it, like a quality number. It is not a quality number. It is an effort number, and effort is not correctness.

Sit with what edit distance actually records. It records the distance between two strings of text: the machine's version and the human's version. It says nothing about whether the human's version is right, and nothing about whether the machine's version was wrong in a way the human failed to fix. It measures how much the text moved, and it is structurally silent about whether the text is correct. That silence is the whole problem, and it produces two failure modes that point in opposite directions and both bury risk.

The Under-Edited Catastrophe

The first failure mode is the one that shipped the Turkish warning. A segment arrives from the engine fluent, confident, and wrong: a dropped negation, an inverted warning, a flipped dosage, all in grammatically perfect prose. The post-editor, moving fast because throughput is what pays and the sentence reads perfectly, corrects a comma and moves on. The edit distance for that segment is tiny. On the dashboard, it looks pristine: barely touched, therefore, the reasoning goes, barely flawed. In reality a Critical error passed straight through, and the low edit distance is not evidence of quality but evidence of the exact failure that produced the catastrophe. Low edits on a fluent critical error is the signature of the disaster, and a metrics program that reads low edit distance as good quality has inverted the reading of its most important signal. It is not merely failing to catch the error. It is congratulating the operation for the behavior that let the error through.

The Over-Edited Illusion

The second failure mode runs the other way and burns money instead of shipping risk. A segment arrives already correct, and a fussy post-editor rewrites it to personal preference: a synonym here, a reordered clause there, none of it adding any accuracy, none of it fixing a real defect. The edit distance is high. On a naive dashboard this reads as "the machine was bad, the human worked hard," when in fact the machine was fine and the human burned budget producing a preferential rewrite that added no quality and, on a bad day, introduced a new error into a segment that did not have one. High edit distance is just as ambiguous as low edit distance. It can mean the machine was genuinely poor, or it can mean your post-editors are over-editing and destroying the MTPE economics that justified the workflow in the first place.

Edit distance measures how much the text moved, never whether it is right. Low edits can be a barely-touched Critical error; high edits can be a fussy rewrite that added nothing. Effort is not correctness.

The strategic error is not that edit distance is a useless number. It has a legitimate use: as an input to productivity and pricing analysis, where it genuinely helps you understand how much human effort a given engine and content type demand, which feeds your cost model and your MTPE rate. The error is promoting an effort-and-cost metric into a quality slot on the scorecard, where its downward trend gets read as rising quality by an audience that does not know it is looking at an effort number in disguise. Keep edit distance. Move it to the cost section of the dashboard, label it as effort, and never let it appear anywhere near the word "quality."

BLEU and the Benchmark That Cannot See the Flip

The next trap lives one layer up, in how the operation selects and reports on its engines. BLEU (Bilingual Evaluation Understudy) is a reference-based automatic metric from the early 2000s that scores machine output by counting how many short word sequences, called n-grams, it shares with a human "reference" translation. More overlap, higher score. For two decades BLEU was how the field reported whether one engine "beat" another, and it still shows up in vendor pitch decks and internal engine-comparison slides as if it were a verdict on quality. Understanding exactly what BLEU can and cannot see is essential for any strategist evaluating engines, because a BLEU-driven engine decision can look rigorous and be dangerously blind.

Consider what n-gram overlap rewards. A translation that renders the source correctly but with different wording than the reference scores low, because its n-grams do not match the reference, even though it is a perfectly good translation. Meanwhile, a translation that matches the reference word for word except for a single flipped "not" scores extremely high, because almost every n-gram still matches, even though that one flip is a Critical error that reverses the meaning of the sentence. BLEU literally cannot distinguish a creative-but-correct rendering from a near-identical-but-catastrophic one, because it counts surface overlap and not meaning. A high BLEU score and a shipped Critical error are entirely compatible. This is not a tuning problem to be solved in the next release; it is the definitional boundary of a metric that counts string overlap without reading for meaning.

Why BLEU Misleads the Engine Decision

There are two additional strategic hazards specific to using BLEU to choose an engine. First, BLEU scores are only comparable within a single test set, tokenization, and language pair. A vendor quoting a BLEU of 62 is quoting a number that is meaningless without the exact reference set and scoring configuration behind it, and vendors rarely volunteer that the configuration was chosen to flatter their engine. Second, and more damaging, BLEU aggregates over an entire test set, which means it tells you about average surface similarity and is completely silent about the tail, and the tail is where the critical errors live. An engine can post an excellent corpus-level BLEU and still be the engine that flips negations on exactly the high-liability segments you care about most, because those segments are a rounding error in the aggregate. The metric is strong at exactly the resolution that does not matter to you (the average) and blind at exactly the resolution that does (the individual catastrophic segment).

The strategic replacement is not "a better automatic metric." It is a purpose-built evaluation. When you evaluate an engine for your domain, you score it against an MQM-aligned error typology, on a representative sample of your actual high-risk content, and you count Critical errors as a first-class result, not as noise inside an aggregate. An engine's critical-error rate on your content, the number of Critical errors per thousand segments or per file, tells you the thing BLEU structurally cannot: how often this engine will hand your post-editors a fluent landmine. That is the number that belongs in a build-versus-buy decision, and it is precisely the number BLEU averages into invisibility.

Throughput Without a Gate Is a Speedometer With No Brakes

Now the most seductive vanity metric of all, because it is the one leadership asks for by name. Throughput, words per day per linguist, is the headline of every localization-AI business case, and for good reason: the whole promise of the machine-translation-first workflow is that it lifts a linguist from roughly 2,000 words a day to 5,000 or more. That lift is real, it is worth celebrating, and reporting it is not the mistake. The mistake is reporting throughput without an error gate attached to it, because throughput measured alone does not just fail to capture risk; it actively rewards the behavior that creates risk.

Trace the incentive. If throughput is the number your linguists are measured and paid against, and quality is not gated with equal force, then the rational move for a post-editor under deadline is to go faster, and the fastest way to go faster is to stop reading the source segment closely, which is the one activity that catches the silent critical error. You have built a metric that pays people to skip the step that protects you. Speed and the discipline that catches the killer error are in direct tension for a post-editor's attention, and a naked throughput number resolves that tension in favor of speed every single time. A dashboard that shows rising throughput and no critical-error rate beside it is not a neutral omission. It is a policy: go faster, and we will not look at what you dropped.

Throughput without an error gate is not a productivity metric. It is a standing instruction to trade the source read for speed, and the source read is what catches the error that ends the contract.

The Only Honest Way to Report Speed

Throughput is only defensible when it is reported as a paired number, chained to a quality gate that has veto power over it. The honest statement is never "we shipped 5,000 words a day." It is "we shipped 5,000 words a day through a gate that failed any file containing a Critical error, and here is the critical-error rate that gate caught and the count that got through." Speed that survived the gate is an achievement. Speed reported without the gate is a liability wearing an achievement's clothing. The error gate, the rule that a file cannot ship until it has been evaluated against a severity typology and cleared, is what converts throughput from a vanity metric into a real one, because it makes the speed conditional on the safety rather than traded against it. Throughput past the gate is the only throughput number that belongs on a strategist's scorecard.

The Average That Drowns the Critical

We arrive at the deepest and most mathematical of the traps, the one that explains why the medical-device dashboard's 98.7% pass rate and its clean QE averages were not just unhelpful but structurally incapable of showing the risk they were meant to summarize. The problem is averaging itself. Averaging is the wrong operation for the risk that defines localization, and understanding why is the single most important quantitative insight a localization strategist can hold.

Recall the cardinal rule of the entire quality system, the one the whole program builds toward: a Critical error, an error that can cause physical, legal, financial, or safety harm, fails the file regardless of how clean every other segment is. One Critical is dispositive. Now look at what an average does to a Critical. An average blends every segment into a single central value, which means it treats the file as a population where each segment contributes proportionally and outliers wash out. A file of 4,000 clean segments and one inverted contraindication has an average quality that is, to three decimal places, indistinguishable from a flawless file, because one catastrophic segment against thousands of clean ones moves the mean by essentially nothing. The average is not slightly wrong here. It is answering a completely different question from the one that matters. It answers "what is this file like on average," when the only question that decides whether the file ships is "does this file contain even one Critical error," and averaging is the exact operation that dissolves the answer to the second question into the first.

Aggregate Pass Rates Hide the Same Way

The aggregate pass rate, the percentage of segments or files that passed across the whole operation, is the same failure at a higher altitude. A 98.7% pass rate sounds like a near-flawless operation. But that 1.3% is not distributed like harmless static; it can contain every Critical error the operation shipped, and a single Critical in a drug label or an indemnity clause is not 1.3% of a problem. It is a recall, a lawsuit, or a death. Reporting quality as an aggregate pass rate takes the one class of event that must be counted individually, the Critical, and dilutes it into a percentage where it becomes invisible. The higher the volume, the more effective the concealment, which produces the perverse result that the operation buries its critical errors more thoroughly precisely as it scales, exactly when the stakes are highest and the volume of high-liability content is largest.

What to Count Instead of What to Average

The fix is a change of mathematical operation, from averaging to counting, and it is not subtle. For the dimension that can kill you, you do not report a central tendency; you report a count, and the count you report is the count of Criticals. The headline quality number on a defensible scorecard is not a pass rate or an average. It is the critical-error rate and, alongside it, the absolute count of Critical errors, which for high-liability content should be reported as zero and treated as a hard gate rather than a trend to be minimized over time. A Major-error rate and a Minor-error rate can be reported as distributions, because Majors and Minors are the kind of thing where accumulation and central tendency actually carry meaning. But the Critical is never averaged and never expressed as a rate you are content to see hover near zero. It is counted, and any nonzero count for high-liability content is an incident, not a data point.

Averaging is the correct operation for a risk that accumulates and the catastrophically wrong operation for a risk where one instance is fatal. Count Criticals. Never average them into a pass rate.

The Before-and-After Dashboard

Abstract principles do not change an operation; a rebuilt scorecard does. So let us take the medical-device company's dashboard and rebuild it in front of you, panel by panel, so the difference between a dashboard that hides risk and one that surfaces it is concrete and copyable.

The Before: The Dashboard That Hid the Warning

The original scorecard had four panels, and every one of them was a trap:

  • Throughput: 5,200 words/day, up 168%. A real number, reported naked, with no gate beside it. It rewarded speed and was silent about what speed cost.
  • Cost per word: down 51%. True, and irrelevant to safety. A pure efficiency number sitting where an audience would read the whole board as a quality story.
  • Post-edit distance: trending down. An effort metric promoted to a quality slot, whose downward trend the room read as "the machine is getting better," when it partly measured post-editors touching fluent errors less.
  • Aggregate pass rate: 98.7%. An average that mathematically could not show a single inverted warning inside 40,000 clean segments. The one panel labeled "quality," and the one least able to represent it.

Notice the shape of the deception. Nothing on the board was false. Every number was accurately computed. The board hid the risk not through error but through selection: it measured effort, speed, and cost, all of which improved, and it represented quality with an average that structurally cannot surface a Critical. The inverted Turkish warning was invisible on this dashboard not because someone lied but because no panel on the board was built to see it. That is the essence of a metrics program that hides risk: it is usually not corrupt, it is miscomposed, and miscomposition is a design failure the strategist owns.

The After: The Dashboard That Surfaces the Critical

The rebuilt scorecard keeps the speed and cost story, because leadership legitimately needs it, but it demotes those numbers to an efficiency section and puts a dual-axis structure at the top: every efficiency number is paired with a risk number, and the risk numbers have veto power. The new board reads:

  • Critical-error rate: 0 per file, gated. The headline. Not an average, not a pass rate, a count, reported as an absolute and enforced as a gate. Any nonzero value on high-liability content is flagged as an incident with a root-cause entry, not smoothed into a trend.
  • Major and Minor error rates, by content risk tier. Reported as distributions, broken out by risk tier so a Major in a drug label is never blended with a Major in internal documentation. This is where central tendency is legitimately informative.
  • Terminology conformance: % of approved terms honored, by tier. A direct measurement of whether the engine and post-editors held the client's termbase, one of the dimensions an automatic score is blind to and a real evaluation checks by design.
  • Throughput past the gate: 5,200 words/day, of files that cleared the severity gate. Speed, but conditional on safety. The gate is named in the metric so no one can read the speed as unconditional.
  • Post-edit distance, in the efficiency section, labeled "effort." Retained for cost modeling, explicitly not a quality metric, physically separated on the board from the quality panels so its downward trend can never again be misread as rising quality.
  • Coverage: % of high-liability content that received a full source-verified evaluation. The metric that catches the gap the old board never showed, whether the scrutiny the risk tier demanded actually happened, so "we didn't check" can never masquerade as "it was clean."

The difference between the two boards is not more numbers. The after-board has roughly the same number of panels. The difference is that the after-board pairs every efficiency claim with a risk claim, counts the thing that must be counted instead of averaging it, moves the effort metrics out of the quality section and labels them honestly, and adds a coverage metric so that silence about a high-risk file is itself visible. A dashboard built this way cannot hide the inverted warning, because the warning would either be caught by the coverage-mandated evaluation and show up as a nonzero Critical count, or its absence from evaluation would show up as a coverage gap. Either way, the risk becomes visible on the board rather than dissolving into an average. That is the entire objective: not a prettier dashboard, but one where the failure that ends a contract has nowhere to hide.

Designing Metrics That Cannot Be Gamed Into Hiding Risk

A scorecard is a static artifact; the harder strategic work is designing a metrics system that resists the drift back toward vanity, because the pull toward clean, rising, reassuring numbers is constant and it will reassert itself the moment your attention moves elsewhere. A few design principles hold the line.

Pair Every Speed Number With a Risk Number

The foundational rule is that no efficiency metric appears alone, ever. Throughput is chained to critical-error rate. Cost per word is chained to terminology conformance and coverage. The pairing is not cosmetic; it is what prevents the organization from optimizing the efficiency number in isolation, because the moment speed rises while the paired risk number degrades, the board shows it in the same glance. This is the dual-axis discipline the program teaches for reporting to leadership, and it is the structural antidote to Goodhart's law: you cannot cheat a speed target by dropping the source read if the source read's absence lights up a paired risk metric on the same slide.

Count the Catastrophic, Average the Accumulative

Match the mathematical operation to the shape of the risk. For any error class where a single instance is disqualifying, the Critical, you count, you report an absolute, and you gate. For error classes where quality is a matter of accumulation, Majors and Minors, distributions and averages are legitimate and informative. The strategic discipline is refusing to let the accumulative operation, averaging, ever touch the catastrophic class. The instant someone proposes a "quality score" that blends Criticals into a weighted average with everything else, you have a metric that can show green while a Critical ships, and you reject it on those grounds alone.

Measure Coverage So Silence Is Visible

The subtlest way a metrics program hides risk is by only measuring the files it looked at, so a high-liability file that never received a proper evaluation simply does not appear as a problem; it appears as nothing. A coverage metric, the percentage of high-liability content that actually received the full source-verified evaluation its risk tier demands, turns that silence into a visible number. Without coverage, an operation can post a perfect critical-error rate simply by not evaluating the files where Criticals would be found. Coverage is the metric that makes "we did not check" impossible to hide behind "it was clean," and it is the one most operations forget, because it measures the absence of work rather than its output.

Report the Record, Not the Proxy

When a client or an auditor asks whether the operation's quality is good, the answer is never a proxy metric. It is not "our post-edit distance is low" or "our BLEU is 62" or "our pass rate is 98.7%." The answer is the severity-scored quality record: here are the errors found, categorized and severity-rated against the MQM and ISO 5060:2024 typology, here is the Critical count, which is zero, here is terminology conformance, and here is the coverage that proves we looked. Proxies allocate your effort and inform your cost model; they are inputs. The defensible record is an output of human judgment against an error typology, and it is the only thing a strategist should let stand as the answer to "is it good." A metrics program that cannot produce that record on demand is a metrics program that will, sooner or later, hide the error that matters.

Key Takeaways

  • A vanity metric is a true answer to the wrong question: throughput and cost per word genuinely improve, which is exactly why their prominence lends false credibility to a dashboard that says nothing about the critical error. The danger is not that the number is false but that it is true, prominent, and irrelevant to the failure that ends a contract or a life.
  • Metrics steer the organization, they do not just describe it (Goodhart's law): reward throughput and you get people who stop reading the source segment; celebrate a falling post-edit distance and you teach linguists that touching fluent errors less is good. Choosing metrics is an operational-design decision with a body count, and the strategist owns it.
  • Edit distance and post-edit distance measure effort, not correctness. Low edits can be a barely-touched fluent Critical error (the disaster signature), and high edits can be a fussy rewrite that added no quality and burned budget. Keep edit distance in the cost section, labeled "effort," never in a quality slot.
  • BLEU counts n-gram overlap with a reference and cannot distinguish a creative-but-correct rendering from a near-identical-but-catastrophic one; it aggregates over a test set and is blind to the tail where Critical errors live. Evaluate engines on their critical-error rate against your own high-risk content, not on a BLEU score.
  • Throughput reported without an error gate is a standing instruction to trade the source read for speed. Report only "throughput past the gate," speed made conditional on a severity evaluation that can fail the file, never speed traded against the safety that catches the killer error.
  • Averaging is the correct operation for accumulative risk and the catastrophically wrong one for a risk where a single instance is fatal. An average quality score or aggregate pass rate cannot surface one inverted contraindication inside thousands of clean segments; the concealment gets more effective as the operation scales. Count Criticals, do not average them.
  • The before-and-after dashboard shows the fix: pair every efficiency number with a risk number, count the catastrophic and average only the accumulative, demote effort metrics out of the quality section and label them honestly, and add a coverage metric so a high-liability file that was never evaluated shows up as a gap instead of as silence.
  • When a client or auditor asks whether quality is good, the answer is never a proxy metric. It is the severity-scored record against MQM and ISO 5060:2024: the errors found, the Critical count of zero, terminology conformance, and the coverage that proves you looked. Proxies are inputs; the defensible record is the output you stand behind.