Measuring Transformation at the Operation Level
It is the third Tuesday of the quarter, and you are sitting in a conference room with the Chief Operating Officer, the Chief Financial Officer, and the General Counsel, presenting the health of the localization transformation you have spent two years building. On the wall behind you is a single slide. It says the enterprise moved 94 million source words this quarter across 38 languages, up 61% year over year, at a blended cost that fell from 19 cents to 12 cents per word. The COO is delighted. The CFO is doing the arithmetic on the savings and smiling. And you, the person who built this operation, are the only one in the room who knows that this slide is a beautifully lit photograph of an iceberg taken from above the waterline. Because underneath those two magnificent aggregate numbers sit fourteen delivery teams, six machine-translation engines, and a content portfolio that ranges from throwaway marketing banners to drug labels that can kill someone if a decimal moves. And somewhere in that portfolio, hidden inside a rolled-up average that looks fantastic, a regional team in one language pair shipped three Critical errors on regulated medical content last month, because their newest engine was never grounded on the client termbase and their throughput target was set as if all words carried the same risk. The enterprise number is up. A specific, dangerous liability is also up. And the entire discipline of this lesson is refusing to let the first fact hide the second. At the single-file and single-team level you learned four numbers that must be reported together. Now you must learn something harder: how to roll those numbers up across an entire multilingual operation, across teams and tiers and languages, without the act of aggregation itself averaging away the one Critical that ends a relationship or a life. We are going to build the operation-level dashboard, define what belongs on it and what must never be blended into it, separate the indicators that warn you early from the ones that only confirm the damage, and then assemble a complete enterprise scorecard for a real operation, tier by tier and language by language, so that when leadership asks "how is the transformation going," you can answer with a number that is both magnificent and honest.
The Aggregation Problem: Why Enterprise Metrics Lie Differently
At the single-team level, the danger was that a speed number would flatter the machine while a quality number quietly collapsed. At the enterprise level, the danger mutates into something more insidious, because the enterprise reports in aggregates, and aggregation is a mathematical operation that is very good at hiding exactly the thing you most need to see. When you roll fourteen teams into one number, you are not summarizing them, you are averaging them, and an average is a machine for making a catastrophe in one component invisible behind excellence in the others. This is the central problem of measuring a transformation at the operation level, and it is not a reporting inconvenience, it is a structural trap that has ended real localization programs and real careers.
Start with the definitions, because at enterprise scale each of the four core numbers acquires a scope it did not have at the team level. Throughput, at the operation level, is aggregate delivered-and-verified volume: the total source words the entire operation moved from raw machine-translation output to delivered, verified target quality across all teams, languages, and content tiers in a period, most usefully expressed both as a total (94 million words a quarter) and as a normalized rate (source words per linguist per day, portfolio-wide). Note the same load-bearing clause from the team level survives the roll-up: it counts only delivered, verified words, so a team that shipped fast and failed the gate does not inflate the enterprise total with rework wearing throughput's clothes. Machine-translation post-editing (MTPE), the workflow where a human edits machine output rather than translating from a blank target, is the baseline production mode across the whole portfolio, so aggregate throughput is fundamentally a measure of how much finished multilingual content the human-plus-machine system delivered.
Cost-per-word, at the operation level, becomes the blended cost-per-word: the fully-burdened cost to move one source word to delivered, verified target, averaged across the entire portfolio's content mix. The word "blended" is doing enormous work here and it is the most dangerous word on the enterprise dashboard, because a blend is an average, and an average of a portfolio that contains both nine-cent marketing words and nineteen-cent regulated words can land on any number you like depending on the mix. A blended cost-per-word reported without the mix beneath it is not a metric, it is a mood.
Critical-error rate, at the operation level, becomes the portfolio critical-error rate: the frequency of Critical-severity errors across all evaluated output in the operation, per unit of volume, typically Criticals per 10,000 or per 100,000 words. A Critical error, under the MQM/ISO 5060 error typology (the analytic evaluation framework formalized by ISO 5060:2024 that classifies every error by category and by Critical, Major, or Minor severity), is an error severe enough to render content unfit for use and capable of causing real-world harm: a flipped dosage, a dropped negation in a safety warning, an inverted indemnity clause, an invented figure in a financial disclosure. And here is where aggregation becomes lethal, because a Critical error is not a quantity that averages meaningfully. One Critical in a drug label is not "half as bad" as two; it is a gate event that fails a file and can trigger a recall. Rolling Criticals into a portfolio rate can make three Criticals in a high-liability language pair vanish into a denominator of 94 million words, producing a portfolio rate so low it reads as success. This is the averaging-away-a-Critical problem, and it is the single most important thing this lesson exists to prevent.
Terminology conformance, at the operation level, is the percentage of governed-term occurrences across the portfolio that match the client's approved term in the termbase, aggregated across all teams and languages. And a fifth number appears that had no meaning at the single-team level: coverage, the percentage of the portfolio that is actually under measurement at all, broken out by tier and by language. Coverage is the metric that measures your metrics, and at enterprise scale it is indispensable, because the most dangerous number on any operation-level dashboard is the volume you are not evaluating and therefore reporting nothing about.
An average is a machine for making a catastrophe in one component invisible behind excellence in the others. At enterprise scale the reporting is aggregate, and aggregate reporting hides exactly the local failure you most need to see.
The Simpson's Paradox of Localization Quality
There is a formal name for the way aggregation can invert the truth, and it is worth naming because it will happen to you. In statistics it is called Simpson's paradox: a trend that appears in every subgroup can reverse or disappear when the subgroups are combined. In a localization operation it shows up like this. Suppose your regulated-content critical-error rate got worse this quarter in every single high-liability language pair, but the overall portfolio also grew, and most of that growth was low-liability marketing content with near-zero Criticals. Blend it all together and the portfolio critical-error rate can actually improve, quarter over quarter, while the specific thing that can get someone killed got worse in every place it matters. The aggregate improved because the mix shifted, not because the risk fell. A leader reading only the portfolio number would conclude the program is getting safer at the exact moment it is getting more dangerous where danger counts. This is not a hypothetical statistical curiosity; it is the default failure mode of any enterprise dashboard that reports quality as a single blended rate. The defense is not cleverer statistics. The defense is a dashboard architecture that refuses to let the high-liability tier be diluted by the low-liability volume, which is what the rest of this lesson builds.
The Operation-Level Dashboard: What Belongs and What Is Forbidden
The operation-level dashboard is not the team dashboard with bigger numbers. It is a different instrument, designed around one architectural principle: nothing high-consequence may ever be reported only as a portfolio-wide blend. Every metric that can hide a catastrophe when averaged must appear both blended, for the leadership story, and disaggregated, by tier and by language, so the blend can never become the whole truth. The dashboard has five components, and the discipline is in how each is required to be shown.
Aggregate Throughput, Shown as Total and as Rate
Aggregate throughput is the friendliest number on the dashboard and the one leadership will quote, so define it precisely and show it two ways. Show the total delivered-and-verified volume for the period, the headline the COO wants, and show it normalized as source words per linguist per day across the portfolio, the rate that tells you whether the operation is actually more productive or just larger. A common enterprise self-deception is celebrating a rising total volume that came entirely from adding headcount or languages while the per-linguist rate stayed flat or fell, which means the transformation delivered scale but not efficiency, and the CFO who funded an efficiency thesis is being told a scale story. Throughput must also carry its verified-quality clause into the aggregate: the enterprise total counts only words that passed the gate and shipped, because an operation that books 94 million words including a few million that bounced back for rework is reporting a fiction that will unwind the moment someone reconciles delivered volume against invoiced volume.
Blended Cost-Per-Word, Never Without Its Mix
Blended cost-per-word is the number finance loves and the number most likely to be quietly dishonest, so the dashboard rule is absolute: the blend is never shown without the content mix and the per-tier rates that produced it. The honest blended cost-per-word carries the full burdened stack across the whole operation: post-editing labor, the machine-translation or large-language-model engine cost (the per-token or per-character bill, or the amortized cost of a licensed, hosted, or fine-tuned engine), the computer-assisted translation (CAT) tool and translation-management system (TMS) platform costs (the linguists' editing environment and the orchestration layer that routes jobs across teams), the quality-estimation and evaluation tooling, and the governance overhead that produces the quality-risk reduction. At the enterprise level a sixth cost appears that the single team never carried: the cost of running the measurement operation itself, the evaluators, the sampling, the dashboard. That is not overhead to be hidden; it is the cost of knowing whether the transformation is safe, and a cost-per-word that excludes it is booking the savings while hiding the price of the safety. The blend must always sit above a breakdown that shows what each tier and each engine actually cost, so a suspiciously attractive blended number immediately provokes the right question: which content moved to the cheap workflow to produce this, and should any of it have.
Portfolio Critical-Error Rate, Disaggregated by Tier and Language
This is the component where the dashboard architecture earns its existence, and the rule is the strictest on the board: the portfolio critical-error rate is always shown alongside its disaggregation by risk tier and by language pair, and it is never permitted to stand alone. The blended portfolio rate exists only to answer "how is the operation doing overall," and it is the least trustworthy number on the dashboard precisely because it is the one most capable of averaging a catastrophe into invisibility. What actually governs is the tier-level and language-level rate. High-liability content (medical, legal, financial, life-safety) gets its own critical-error rate, evaluated on its own risk-weighted sample, reported on its own line, held to its own hard target, and this number is never blended into the marketing volume that would dilute it. The same discipline applies by language: a single regional team producing Criticals on regulated content must appear as its own red cell, not as a rounding error in a portfolio denominator. If the dashboard shows only a blended portfolio critical-error rate, it is not a quality instrument, it is a liability-concealment device, and it will eventually conceal the liability that ends the program.
Terminology Conformance, the Leading Indicator
Terminology conformance is reported blended and disaggregated like the others, but it earns a special role on the operation-level dashboard: it is the earliest warning the operation has. A leading indicator is a metric that moves before the outcome you actually care about, giving you time to act; a lagging indicator only confirms the outcome after it has already happened. Terminology conformance is the operation's premier leading indicator because it moves first when engine grounding degrades, when a termbase goes stale, when a new content type is routed to an engine that was never taught the client's approved terms, or when a newly onboarded language team has not yet wired its termbase into the pipeline. A dip in terminology conformance in a specific language or tier is frequently the tremor that precedes the earthquake of a rising critical-error rate, because the same loss of control that lets an approved term drift into a synonym is the loss of control that eventually lets a negation flip or a dosage invert. On the enterprise dashboard, conformance is watched per language and per tier precisely so a localized grounding failure lights up as an amber cell weeks before it can become a red one somewhere downstream.
Coverage: The Metric That Measures Your Measurement
The fifth component exists only at operation scale, and neglecting it is how enterprise dashboards lie by omission. Coverage is the percentage of delivered volume that is actually under quality measurement, broken out by tier and by language. It answers the question a good auditor asks first: not "what is your critical-error rate," but "on how much of what you shipped do you even have a critical-error rate, and where are the blind spots." An operation can post a spotless portfolio critical-error rate while measuring only 40% of its volume and none of its newest language pair, in which case the spotless number describes the measured 40% and says exactly nothing about the unmeasured 60% where the real risk is accumulating. Coverage must be reported by tier, because the non-negotiable is that high-liability content approaches 100% coverage while low-liability content can be sampled far more sparingly, and it must be reported by language, because a brand-new language team with 5% coverage is a stated, visible risk rather than a silent one. Coverage is the humility metric. It converts "we found no Criticals" into the honest "we found no Criticals in the 82% of high-liability volume we evaluated, and here is the 18% we did not, and here is the plan to close it." That sentence is the difference between a dashboard that defends you in an audit and one that indicts you.
Coverage is the metric that measures your metrics. "We found no Criticals" is a boast; "we found no Criticals in the 82% of high-liability volume we evaluated, and here is the plan for the other 18%" is a defense that survives an auditor.
Rolling Up Without Averaging Away a Critical
The heart of operation-level measurement is a roll-up rule that most dashboards get exactly backwards. The instinct, when you have fourteen teams, is to compute each metric per team and then average the teams into an enterprise number. For throughput and cost, sensible weighted averaging is fine, because volume and money genuinely add up and average meaningfully. For critical-error rate, weighted averaging across the whole portfolio is precisely the operation that hides the catastrophe, so the roll-up must be built on a different principle entirely.
The Worst-Cell-Governs Rule
Here is the rule that keeps aggregation honest for quality-risk: for any high-consequence metric, the enterprise status is governed by the worst-performing cell in the high-liability tier, not by the blended average. If thirteen language teams shipped zero Criticals on regulated content and one team shipped three, the enterprise quality status on regulated content is not "excellent, 3 Criticals across 30 million regulated words, a negligible rate." The enterprise quality status is "one language pair is in breach on regulated content, three shipped Criticals, remediation in progress," and the operation-level dashboard must surface that one cell as red at the top level, not let it be averaged into green. The blended portfolio rate is still computed and shown, because leadership needs the overall picture, but it is subordinate to the worst-cell status for anything in the high-liability tier. This is the enterprise expression of the program's cardinal rule that one Critical fails the file: at the operation level, one team in breach on high-liability content fails the operation's quality status on that tier, no matter how flawless the other thirteen were. The gate does not average, and neither does the dashboard that governs the gate.
The mechanics of this are concrete. You maintain a matrix, tier on one axis and language on the other, and every cell holds that segment's metrics. Throughput and cost roll up by volume-weighted sum, top to bottom, because scale and money add. Critical-error status rolls up by a different function: the high-liability row's enterprise status is the maximum severity present in any cell, not the mean, so a single red cell makes the row red. Terminology conformance rolls up as a volume-weighted average for the overall story but is simultaneously scanned cell by cell for any value below the tier's threshold, and any breaching cell is flagged regardless of how good the blend looks. This dual roll-up, additive where addition is honest and worst-cell-governs where averaging lies, is the entire technical trick of measuring a transformation without letting the measurement betray you.
Weighting by Consequence, Not by Volume
There is a subtler weighting error the enterprise dashboard must avoid, which is weighting purely by volume when consequence is what matters. Marketing and UI content is usually the largest volume in the portfolio and the lowest consequence per word; regulated content is usually a minority of volume and carries nearly all of the catastrophic risk. If you weight every quality metric by word volume alone, the low-risk majority dominates every blended number and the high-risk minority, the content that can actually sink the enterprise, contributes almost nothing to the picture. The corrective is to think in two currencies at once: words, for throughput and cost, and consequence, for quality-risk. On the quality axis the dashboard weights toward consequence, which in practice means the high-liability tier gets its own always-visible rows and its own targets and is never allowed to be diluted by the marketing volume, even though marketing is where most of the words are. An enterprise that measures quality-risk in proportion to word count is measuring the wrong thing, because the risk was never proportional to the word count. Thirty regulated words can carry more liability than three million marketing words, and the dashboard has to be built so that arithmetic is impossible to forget.
Words add up; consequence does not. Weight throughput and cost by volume, but weight quality-risk by consequence, because thirty regulated words can carry more liability than three million marketing words, and a volume-weighted quality metric buries exactly that.
Leading Versus Lagging Indicators of Transformation Health
A transformation is a moving system, and the difference between running it and merely reporting on it is whether your dashboard warns you early or only confirms the damage. This is the distinction between leading and lagging indicators, and at operation scale it is the difference between steering and doing forensics. Both belong on the dashboard, but they answer different questions and must never be confused.
The lagging indicators are the outcome numbers, the ones that tell you what already happened. Shipped critical-error rate is the ultimate lagging indicator: by the time a Critical has shipped, the harm is either done or one client-side review away from being done, and the number is confirming a failure, not preventing one. Blended cost-per-word is lagging; it settles after the period's work is invoiced. Aggregate delivered throughput is lagging; it counts what was finished. These numbers are essential for accountability and for the leadership story, but a dashboard made only of lagging indicators is a rear-view mirror, and you cannot steer an enterprise transformation by looking exclusively at where it has already been.
The leading indicators are the numbers that move before the outcome, giving you the window to act. Terminology conformance, as established, is the premier leading indicator of quality-risk: it degrades first, weeks before the critical-error rate reacts. But at operation scale there are others the strategist must watch, and they are the difference between a transformation you control and one that controls you. Coverage trend is leading: coverage falling in a growing language pair means measurement is not keeping pace with delivery, and unmeasured volume is where the next Critical is quietly incubating. The gap between raw engine output quality and post-edited output quality, sampled continuously, is leading: if the raw output is degrading (a new engine version, a new content type the engine was not grounded on), the post-editors are absorbing more risk per file and the critical-error rate will follow unless the trend is caught. Evaluator agreement, the consistency between two evaluators scoring the same sample, is leading on the health of the measurement system itself: when evaluators start disagreeing, the scores are drifting and every quality number downstream is losing its meaning. Post-editor edit-effort trend by tier is leading: if edit effort on a tier is falling while volume rises, linguists may be under-editing under throughput pressure, the exact behavior that ships the silent critical error.
Assembling the Canary Set
The operation-level dashboard should carry an explicit set of leading indicators, the canaries, watched per language and per tier, whose job is to turn amber before any lagging number turns red. The discipline is to treat any canary turning amber as an actionable event even when every lagging number is still green, because the entire value of a leading indicator is that it fires while you still have time. A strategist who waits for the shipped critical-error rate to rise before acting has chosen to steer a transformation using only the rear-view mirror, and on high-liability content that choice is measured in recalls and lawsuits. The canaries are cheap to watch and expensive to ignore.
- Terminology conformance by language and tier. The first mover; a dip signals grounding degradation before it becomes a Critical.
- Coverage trend. Falling coverage in a growing segment means the blind spot is expanding faster than the measurement.
- Raw-versus-post-edited quality gap. A widening gap means the engine got worse and the humans are silently absorbing the difference until they cannot.
- Evaluator agreement. Drift here means the measurement instrument itself is losing calibration and every downstream number is suspect.
- Edit-effort trend by tier. Falling effort under rising volume is the fingerprint of under-editing, the behavior that ships the fluent mistranslation.
The Worked Operation-Level Scorecard
Now assemble the whole instrument on one concrete enterprise, so the machine turns in front of you exactly as it would in the quarterly review. The operation localizes 94 million source words a year across 38 languages, run by fourteen delivery teams, on six machine-translation engines, with a content portfolio of roughly 55% marketing and UI (low-to-moderate liability), 15% product documentation (moderate), and 30% regulated content (medical labeling, legal clauses, financial disclosures, high liability). The dashboard is reported quarterly, and every high-consequence metric is shown blended for the leadership story and disaggregated by tier and by language for the truth. Here is the quarter, read the way the dashboard demands: across each metric against baseline and target, then across the metrics against each other, then down into the tier and language cells where the aggregates try to hide things.
The Aggregate Row: The Story Leadership Sees First
- Aggregate throughput. Total delivered-and-verified: 94 million source words, up 61% year over year. Normalized rate: 4,050 source words per linguist per day, portfolio-wide, against a pre-AI baseline of roughly 2,000 and a target range of 3,500 base to 5,000 upside. Reading: the total is up substantially and the per-linguist rate confirms the growth is genuine efficiency, not just added headcount. Honest and healthy.
- Blended cost-per-word. $0.12 all-in, against a baseline of $0.19 and a blended target of $0.13 that reflects the 30% regulated carve-out held near human cost. Beneath the blend: regulated content at $0.185 (full human or full MTPE by rule), documentation at $0.11, marketing/UI at $0.075 including some light post-editing, and the measurement operation itself carried explicitly at roughly $0.006 per word across the portfolio. Reading: on target, and the mix beneath it proves the saving did not come from cheating the high-liability tier.
- Portfolio critical-error rate. Blended: 1 Critical per 61,000 words on the evaluated sample, which as a lone number looks superb. This is the number the dashboard forbids you to trust on its own, and the reason is in the disaggregation below.
- Terminology conformance. Blended: 97.6% on governed terms, against a target of 98%+ overall and 99%+ on high-liability content. Reading: blended just under target, and the dashboard rule says find where the amber lives before relaxing.
- Coverage. Blended: 71% of delivered volume under quality measurement. By tier: high-liability at 96%, documentation at 78%, marketing/UI at 61%. Reading: the high-liability coverage is strong but not yet at the 100% target, and that gap is itself a stated risk, not a hidden one.
Reading Down Into the Cells: Where the Aggregate Was Hiding a Fire
Now disaggregate, which is the entire point of an operation-level dashboard, and the quarter's real story appears. The blended portfolio critical-error rate of 1 per 61,000 words is a Simpson's paradox artifact. Thirteen of the fourteen teams shipped zero Criticals on regulated content, and the enormous marketing volume with near-zero Criticals dominates the denominator and drags the blended rate to a beautiful number. But one team, the newly onboarded language pair added two quarters ago, shipped three Critical errors on regulated medical content: two flipped negations in dosage instructions and one inverted contraindication, every one of them fluent and confident and missed under throughput pressure, on an engine that was never grounded on the client termbase. In the blended number those three Criticals vanish into 94 million words. On the disaggregated dashboard, applying the worst-cell-governs rule, they light up as a red cell that makes the entire high-liability row red, and the enterprise quality status on regulated content is not "excellent," it is "one language pair in breach, remediation in progress."
And the leading indicators had already called it. Look back at the canary set for that team and the story was written weeks earlier: terminology conformance on that language pair's regulated content had dipped to 94.1% while the portfolio blend sat comfortably at 97.6%, coverage on that pair was only 63% against the 96% high-liability average, and the raw-versus-post-edited quality gap on its engine was the widest in the operation. Three canaries were amber on that single cell while every aggregate number on the leadership slide was green. That is precisely the scenario the operation-level dashboard is built to expose and the blended slide is built to hide. The transformation was genuinely succeeding across thirteen teams and simultaneously incubating a serious liability in the fourteenth, and only a dashboard that refuses to average away the worst cell, and watches leading indicators per language and per tier, could hold both truths at once.
What the Enterprise Scorecard Lets You Say to the Board
Return to the conference room and the two-number slide, and see what the operation-level dashboard lets you say instead. You can say: "The transformation delivered 94 million words this quarter, up 61%, at a genuine per-linguist efficiency gain, not just added scale, and blended cost fell to 12 cents with the mix beneath it proving we did not get there by cutting corners on regulated content. Across thirteen of fourteen teams, quality-risk is controlled and improving. And there is one specific, serious issue I am surfacing to you before an auditor or a client does: one newly onboarded language pair shipped three Critical errors on regulated medical content, because its engine was never grounded on the client termbase and its coverage was too low. Our leading indicators flagged it weeks ago, we caught it in evaluation before it reached the client on the affected files, delivery on that pair's regulated content is paused pending re-grounding and a coverage lift to 100%, and here is the remediation plan and timeline. Every other number on this dashboard is honest because this one is on it too." That is a statement that survives the General Counsel, the auditor, and the client, because it does not ask anyone to trust an average. It shows the magnificent aggregate and the dangerous cell on the same board, governs the enterprise status by the worst cell on high-liability content, and proves the operation can see its own worst problem before anyone outside it can. The two-number slide could never say any of that. It could only photograph the iceberg from above the waterline until the enterprise hit the part it refused to measure.
From Scorecard to Operating Rhythm
A dashboard is an artifact; a transformation is a habit, and the operation-level scorecard only protects the enterprise if it is wired into a rhythm that forces action on what it reveals. The quarterly board view is the visible tip, but the instrument governs only if it drives a cadence beneath it. Leading indicators are reviewed weekly, per language and per tier, so an amber canary triggers investigation while the window to act is still open. The tier-and-language matrix is reviewed monthly with the delivery leads, worst cell first, so remediation starts on the reddest cell before it is escalated rather than after. The full blended-and-disaggregated scorecard is assembled quarterly for leadership, always with coverage stated, always with the worst high-liability cell surfaced at the top rather than buried. And an out-of-cycle trigger overrides all of it: any shipped Critical on high-liability content convenes an incident review immediately, regardless of where the quarter sits, because the cardinal rule does not wait for a reporting date.
The strategic point is that measurement at the operation level is not a scoreboard you check, it is a control system you run. The four core numbers plus coverage, disaggregated by tier and language, rolled up by the worst-cell-governs rule on anything high-consequence, watched through leading indicators that fire before the lagging ones, and wired into a weekly-monthly-quarterly cadence with an incident override, is what separates a transformation you are steering from one you are merely watching happen. Leadership will always want the two magnificent aggregate numbers, and you should give them, because the speed and the savings are real and hard-won. But the discipline of this level, the thing that makes you the person who can run an enterprise multilingual operation rather than just report on one, is the refusal to let those two numbers stand without the disaggregated truth beneath them, and the architecture that makes averaging away a Critical structurally impossible rather than merely discouraged.
Key Takeaways
- At enterprise scale the metrics lie differently than at the team level: the reporting is aggregate, and aggregation is a machine for hiding a local catastrophe behind blended excellence. The operation-level dashboard's founding principle is that nothing high-consequence may ever be reported only as a portfolio-wide blend; it must appear both blended for the leadership story and disaggregated by tier and by language for the truth.
- The operation-level dashboard has five components: aggregate throughput (total and per-linguist rate, verified-quality only), blended cost-per-word (never without the content mix and the cost of measurement itself beneath it), portfolio critical-error rate (always disaggregated, never standing alone), terminology conformance (the premier leading indicator), and coverage (the percentage of volume actually under measurement, by tier and language, the metric that measures your measurement).
- Beware the Simpson's paradox of localization quality: high-liability critical-error rate can worsen in every language pair while the blended portfolio rate improves, simply because low-risk marketing volume grew and shifted the mix. The aggregate got safer-looking while the content that can kill got more dangerous, which is why quality-risk must be weighted by consequence, not by word volume.
- Roll up honestly with a dual rule: throughput and cost roll up by volume-weighted sum because scale and money add, but critical-error status on high-liability content rolls up by the worst-cell-governs rule, where the enterprise status is the worst cell, not the mean. One team in breach on high-liability content fails the operation's status on that tier no matter how flawless the other thirteen were, the enterprise expression of "one Critical fails the file."
- Coverage converts a boast into a defense: "we found no Criticals" becomes the honest "we found no Criticals in the 82% of high-liability volume we evaluated, and here is the plan for the other 18%." High-liability coverage must approach 100%; a new language team at 5% coverage is a stated, visible risk rather than a silent one, and unmeasured volume is where the next Critical incubates.
- Separate leading from lagging indicators, because one lets you steer and the other only confirms damage. Shipped critical-error rate, blended cost, and delivered throughput are lagging (a rear-view mirror). Terminology conformance, coverage trend, the raw-versus-post-edited quality gap, evaluator agreement, and edit-effort trend by tier are leading canaries that fire weeks before a lagging number turns red, and any amber canary is an actionable event even when every lagging number is still green.
- On the worked 94-million-word operation, the blended portfolio critical-error rate of 1 per 61,000 words was a beautiful lie: thirteen teams shipped zero regulated Criticals while one newly onboarded language pair shipped three on medical content, its engine never grounded on the client termbase. Three leading indicators (conformance, coverage, raw-versus-post-edited gap) were already amber on that single cell while every aggregate on the leadership slide was green.
- Measurement at the operation level is a control system, not a scoreboard: leading indicators reviewed weekly, the tier-and-language matrix monthly worst-cell-first, the blended-and-disaggregated scorecard quarterly with coverage always stated, and an incident override that convenes a review on any shipped high-liability Critical regardless of the calendar. Give leadership the two magnificent aggregate numbers, but never let them stand without the disaggregated truth that makes averaging away a Critical structurally impossible.
Skill.re