Scaling Without Breaking Quality
The number that ended the celebration was 0.4 percent. For eleven months the pilot had been the pride of the localization operation: one language pair, English into German, one content type, product knowledge-base articles, run through an MT-first pipeline with a severity-scored gate that had produced a clean quality record on every batch. The critical-error rate had held at zero for eleven straight months. So leadership did what leadership does with a success: they scaled it. Nineteen languages, six content types, four business units, all switched to the same "proven" pipeline over a single quarter, because the pipeline had proven itself. Twelve weeks later the head of localization was sitting in a conference room explaining to the general counsel why a machine-translated safety warning had shipped in Japanese, Korean, and Brazilian Portuguese with a negation dropped, why the German termbase that had been immaculate at pilot scale was now fighting three regional teams who had each quietly forked their own version, and why the critical-error rate across the whole operation had climbed to 0.4 percent. That sounds small. On a run rate of two million words a quarter, 0.4 percent is eight thousand words carrying a critical defect, spread across nineteen languages and four business units, most of it fluent, grammatical, and confident. The pipeline had not broken. The pipeline was fine. What broke was the assumption that a control which holds at pilot volume holds at enterprise volume, and this lesson is about why that assumption is false, why the critical-error rate is the one number that must not climb as you scale, and how to add languages and content types without eroding the gate that keeps that number at zero.
The Scaling Lie: Quality Holds Because It Held
Begin with the vocabulary, because scaling is a conversation about specific quantities and a strategist who blurs them will make the exact mistake this lesson exists to prevent. Scale, in an enterprise localization operation, is the total multilingual-content workload measured across three dimensions at once: volume (words per period), breadth (number of languages and locales), and variety (number of distinct content types and business units). A pilot is small on all three; the enterprise is large on all three; and the failure this lesson dissects is what happens when a control validated on the small version is assumed to hold on the large one. Critical-error rate is the proportion of delivered content that carries at least one Critical error, an error that inverts meaning, drops a negation, flips a dosage, swaps an approved term in a way that misleads, or breaks an obligation in a contract, the class of defect that, under the analytic error typology this program is built on, fails a file regardless of how clean the rest reads. It is expressed as a rate because at enterprise scale you cannot inspect everything, so you sample, score, and reason about proportions. And the quality gate is the go/no-go control that scores output against that typology and blocks delivery when a Critical error is present: the single most important control in the operation, and the one most silently degraded by volume.
Fix a few more terms while we are here, because the enterprise conversation uses them constantly. Machine translation (MT) renders text from source to target with no human writing the words; machine-translation post-editing (MTPE), often shortened to post-editing (PE), is a qualified human editing that machine output rather than translating from a blank segment. A large language model (LLM) is a general text predictor that translates as a side effect of broad competence, more fluent and more confidently wrong than classic MT. A translation memory (TM) is the database of previously approved source-target segment pairs an operation leverages across jobs; a termbase is the controlled glossary of approved terms. MQM (Multidimensional Quality Metrics) is the analytic error typology, formalized for translation output by ISO 5060:2024, that scores errors by category (accuracy, terminology, locale, fluency) and severity (Critical, Major, Minor). ISO 18587 is the post-editing standard whose revision (in DIS ballot, publication targeted late 2025 into 2026) expands scope to AI and LLM "non-human translation output," retires the rigid light-versus-full split for an effort spectrum, and requires the post-editor to hold full professional-translator competence. A locale is the language-and-region convention bundle: en-US is not en-GB, de-DE is not de-CH. An evaluator is the qualified human who scores output against the typology at the gate. A language-service provider (LSP) is a translation vendor or agency. Hold those, and the failure modes below become legible.
Now the lie itself, because it is seductive and almost everyone believes some version of it. The pilot succeeded, therefore the pipeline is proven, therefore scaling the pipeline scales the success. Every clause in that sentence is true except the last one, and the last one is false for a structural reason that has nothing to do with the technology. The pipeline is a set of steps. The success of the pilot was not produced by the steps alone; it was produced by the steps plus a set of invisible conditions that held at pilot scale and quietly stop holding as you grow. At pilot scale, one evaluator could look at every batch. At pilot scale, one termbase served one team and drifted nowhere because there was nowhere to drift. At pilot scale, the TM was fed by a handful of trusted linguists whose work you had personally reviewed. At pilot scale, the gate was never bypassed because the volume never exceeded the evaluator's capacity to score it. None of those conditions is part of the pipeline diagram. All of them were load-bearing. And every one of them fails as a function of scale, not as a function of any flaw in the pipeline.
A pilot's quality is produced by the pipeline plus a set of invisible conditions that hold at small scale and silently stop holding at large scale. Scaling the pipeline does not scale those conditions. It exposes their absence.
This is why the critical-error rate must anchor the entire scaling conversation, and why leadership's instinct to lead with throughput and cost is the instinct that ships the dosage error. Throughput scales with the pipeline: more languages and more content types genuinely produce more words per period, and the cost-per-word economics genuinely improve. Those numbers go the right way almost automatically. The critical-error rate is different. It does not scale with the pipeline. It scales with the weakest of the invisible conditions, and at enterprise volume the weakest condition is always something the pilot never stressed. So the correct mental model for scaling is not "the pipeline that worked, bigger." It is "the pipeline that worked, plus a deliberate rebuild of every invisible condition the pilot got for free, at a scale where none of them come for free anymore." A strategist who scales the pipeline and forgets the conditions is not scaling success. They are scaling until the weakest condition breaks, and then discovering, at maximum blast radius, which condition it was.
The Four Scaling Failure Modes
There are four ways an operation breaks quality as it scales, and they are not random. Each is the failure of a specific invisible condition, each is invisible at pilot scale precisely because the condition still held there, and each is preventable if you name it before you grow rather than after. Walk through all four slowly, because the entire back half of this lesson is a set of controls, and a control only makes sense once you understand exactly which failure it is built to stop.
Failure One: The Gate Bypassed Under Volume
The first and most dangerous failure mode is the quality gate getting bypassed under volume, and it almost never happens by decision. It happens by arithmetic. At pilot scale, the gate scored every batch because the volume fit inside the evaluation capacity. As volume grows, a moment arrives when the words arriving at the gate exceed the number of words the available evaluators can actually score in the time the deadline allows. At that moment the operation faces a choice it usually does not make consciously: score everything and miss the deadline, or ship the unscored remainder and hit it. Under deadline pressure, with a client waiting and a pipeline that has "always been fine," the unscored remainder ships. The gate was not removed. It was silently made optional for the content that overflowed capacity, and that overflow is, by definition, the content nobody scored.
The insidious part is that this failure is invisible in exactly the way a critical error is invisible: it produces no signal until it produces a disaster. The batches that got scored come back clean, so the dashboard looks healthy. The batches that shipped unscored carry whatever critical-error rate the raw MT actually had, which on high-liability content in an immature language can be substantial, but nobody scored them, so nobody knows. The rate you report is the rate of the scored sample, and the scored sample is systematically the safest content, so your reported rate is not just wrong, it is optimistically wrong, biased low by the very mechanism that let the dangerous content escape. A strategist who does not understand this will find a healthy dashboard and a spike in shipped critical errors impossible to reconcile, because the dashboard measures the content the gate caught and the disaster came from the content the gate never saw.
Failure Two: The Evaluator Shortage
Failure one has a root cause, and it is the second failure mode: the evaluator shortage. The gate is only as real as the human capacity to run it, because the whole program rests on the principle that accountability stays human and a qualified evaluator scores output against the typology. At pilot scale, one evaluator was enough. Scaling multiplies the demand for evaluation along every dimension at once: more languages means evaluators qualified in more language pairs, more content types means evaluators competent in more domains, more volume means more evaluator-hours per period. The demand for qualified evaluation grows at least as fast as the volume and often faster, because a new language or a new high-risk content type does not just add words, it adds a requirement for a specific, scarce competence. And qualified evaluators do not scale like compute. You cannot provision another evaluator qualified in Korean pharmaceutical content the way you provision another server. The revised ISO 18587 sharpens this further by requiring the post-editor, and by extension the evaluator, to hold full professional-translator competence, which means the pool you can draw from is exactly the pool of scarce senior linguists, not a pool of cheaper juniors you can hire in a hurry.
The evaluator shortage is the true bottleneck of enterprise localization quality, and most operations discover it in the worst possible way: not as a plan but as a crisis, when a scaling deadline collides with a capacity ceiling nobody had measured. The shortage does not announce itself. It shows up as evaluators working longer hours (fatigue raises the miss rate on exactly the fluent errors that are hardest to catch), as evaluations getting shallower under time pressure, and finally as failure one: the gate quietly going optional because there is no human left to run it on the overflow. Evaluator capacity is not a support function to the scaling plan. It is the scaling plan's binding constraint, and a roadmap that scales volume without scaling qualified evaluation capacity in lockstep has scheduled its own gate to fail.
Failure Three: Terminology Drift Across Teams
The third failure mode is terminology drift across teams, and it is the failure of a condition so quiet at pilot scale that most operations do not even know it was a condition. At pilot scale, there was one team, one termbase, one source of truth for what the approved term was. Consistency was not a control; it was a side effect of there being nowhere for the term to diverge. Scale breaks that by adding the one thing the pilot did not have: other teams. A second business unit stands up its own localization workflow, inherits a copy of the termbase, and starts making local decisions, an approved term that does not quite fit their product, a regional preference, a new term the central termbase has not caught up to. A third team does the same. Within a quarter you do not have one termbase, you have four forks that agree on most terms and disagree on the ones that matter, and the engine, which honors whichever termbase it is grounded on, faithfully propagates each team's divergence across every segment it touches.
Terminology drift is uniquely corrosive at scale for three reasons. First, it is a Major or Critical error under the typology (the wrong approved term for a regulated device or a legal concept is an accuracy and terminology defect that can mislead), so drift directly raises the critical-error rate. Second, it compounds silently through the TM: every drifted segment that passes an under-resourced gate gets written back into the memory, so the fork does not just produce errors, it teaches them to every future job that leverages that memory. Third, it is nearly invisible to a per-batch gate, because each team's output is internally consistent with its own forked termbase, so a batch scored in isolation looks clean. The drift only appears when you compare across teams, which a batch-level gate never does. An operation can have four teams each shipping internally consistent work and a serious enterprise-wide terminology problem that no single evaluation would ever surface. Governance, not evaluation, is the control here.
Failure Four: TM Contamination at Scale
The fourth failure mode is translation-memory contamination at scale, and it is the one that turns a temporary quality dip into a permanent quality debt. The translation memory is supposed to be the operation's compounding asset: every approved segment written back into it makes future jobs faster and more consistent, because the leverage of a clean TM is real speed and real quality at once. That virtuous cycle runs in reverse the moment the gate weakens. When failure one lets unscored content ship, and failure three lets drifted terminology through, those defective segments do not just ship once. They get written back into the TM as if they were approved, and from that moment every future job that pulls a fuzzy or exact match from the contaminated memory inherits the defect, pre-populated into the segment before a linguist even opens the file, wearing the authority of a "previously approved" translation.
Contamination violates the intuition that quality problems are transient. A bad batch, caught, is a bad batch. A bad batch written into the TM is a slow infection: it re-emerges in later jobs, in other languages that share the memory, months after the batch that introduced it. At pilot scale this barely mattered because the TM was small, fed by a few trusted linguists, and easy to inspect. At enterprise scale the TM is large, fed by many teams and vendors, leveraged across everything, and effectively impossible to fully re-inspect, so a contaminating segment can hide indefinitely and surface at the worst possible time, in a high-liability job in a different business unit, as a "high-confidence" match that a rushed post-editor trusts precisely because the TM said it was approved. Provenance, knowing where every segment came from and whether it passed a real gate, is the control, and it is why the controls section treats the TM as an asset governed like a financial ledger, not a bucket that fills itself.
The four scaling failures are the gate bypassed under volume, the evaluator shortage beneath it, terminology drift across teams, and TM contamination that turns a transient error into permanent debt. Each is the failure of an invisible condition the pilot got for free.
The Controls That Actually Scale
The four failures share a structure, so the controls do too. Each failure is a condition that held at small scale by luck and breaks at large scale by arithmetic; each control rebuilds that condition deliberately so it holds at any scale. The organizing principle for every control is the same: at scale, quality cannot depend on any capacity that grows linearly with volume while the volume grows faster. Where a human capacity is the constraint, automate the mechanical part so the scarce human capacity is spent only where judgment is required. Where consistency depended on there being one team, replace the accidental consistency with governed consistency. Where the asset was trusted because it was small, replace trust with provenance. Take the controls in the order the failures demand.
Control One: Automate the Checks a Machine Can Run
The first control attacks failure one at its root by changing what the human evaluator has to spend their scarce capacity on. A large fraction of the checks a gate performs are mechanical: they do not require professional-translator judgment, they require a rule reliably applied. Whether a segment preserved its number of placeholders, whether a date or a currency or a unit was rendered in the correct locale format, whether an approved term from the termbase actually appears where the source term appeared, whether a numeric value in the source survived into the target unchanged, whether a negation word is present when the source has one, whether the segment blew past a length budget. These are the checks that catch a meaningful share of Critical and Major errors, and every one of them can be run by an automated quality-assurance check at machine speed across the entire volume, not a sample. Automating them does not replace the evaluator. It does something more valuable: it lets the evaluator stop spending human hours on the mechanical checks a machine does better and never tires of, and spend their entire scarce capacity on the one thing no automated check can do, which is judge whether a fluent, grammatical, plausible sentence actually means what the source meant.
This is the pivot that makes enterprise-scale quality arithmetically possible. The silent critical error, the fluent mistranslation that reads perfectly, is precisely the error an automated check cannot catch, because there is no mechanical rule for "this grammatical sentence inverts the source's meaning." That error requires a human who understands both languages and both meanings. But the automated checks clear away the mechanical errors that would otherwise consume the human's time, and they run across the full volume rather than a sample, so human capacity is concentrated exactly where it is irreplaceable. An operation that automates its mechanical checks can put its evaluators on semantic judgment across far more content than one whose evaluators are hand-checking placeholder counts. This is how you widen the gate's throughput without widening the hole in it: not by asking humans to work faster on everything, but by removing from the human everything a rule can check, so the human's speed on the irreplaceable part is what limits you, not their speed on the mechanical part a machine should have owned all along.
Control Two: Plan Evaluator Capacity as the Binding Constraint
The second control treats the evaluator shortage as the binding constraint it is and plans it explicitly, the way a factory plans against its bottleneck machine. This starts with a number most operations have never calculated: the qualified-evaluation capacity, in words scored to a real MQM standard per period, that you actually have, broken down by language pair and by domain competence, because a global evaluator-hours number is a fiction when a Korean pharmaceutical file cannot be scored by a French marketing evaluator. Once you have that number, the scaling rule becomes a hard constraint rather than an aspiration: you may not schedule more high-risk volume in a language-domain cell than you have qualified evaluation capacity to gate in that cell within the deadline. Volume that exceeds the capacity of the cell does not ship unscored (that is failure one); it either waits, or triggers the deliberate expansion of capacity in that cell before the volume is accepted, or is declined. The capacity plan makes the trade-off visible and decided in advance instead of invisible and decided under deadline pressure by whoever is holding the file at midnight.
Capacity planning at scale is a portfolio problem with three levers, each with its own lead time. The first lever is the automation from control one, which raises the effective capacity of every existing evaluator by removing mechanical work; it is the fastest lever and the one to pull first, but it has a ceiling because it cannot touch semantic judgment. The second lever is qualifying more evaluators, which is slow because the revised ISO 18587 competence bar means you are recruiting or developing scarce senior linguists, not filling seats; a strategist plans this in quarters, not weeks, and starts before the volume arrives, because an evaluator qualified reactively is qualified too late. The third lever is risk-based allocation: not all content needs the same depth of evaluation, so you spend your scarcest capacity, deep human MQM scoring, on the high-risk cells, and let automated checks plus lighter human review carry the low-risk cells, which is legitimate precisely because the consequence of a fluent error on a knowledge-base article is recoverable and the consequence on a drug label is not. The capacity plan is the document that says, for every language-domain cell, how much qualified evaluation it has, how much the scheduled volume demands, and which lever closes the gap, and it is the single most important artifact in a scaling program because it is the one that keeps failure one from ever happening.
Control Three: Govern Terminology and Standards Centrally
The third control replaces the accidental consistency of the single-team pilot with the deliberate, governed consistency an enterprise requires, because drift is not a linguistic problem, it is an organizational one. The mechanism is a single authoritative termbase with a real ownership and change-control process: one central function owns the approved terms, teams request additions and changes through a defined workflow rather than forking a local copy, and the authoritative termbase is the one every engine in the operation is grounded on. This is not bureaucracy for its own sake; it is the recognition that at scale, the question "what is the approved term?" must have exactly one answer, enforced by governance, because the moment it has four answers you have four forks and a rising critical-error rate. The governance function does not slow teams down if it is designed well: it gives them a fast path to propose a term, a clear owner who decides, and a single place the decision propagates from, which is faster than the alternative of every team maintaining a divergent copy and discovering the divergence in a client escalation.
Governance extends past terminology to every standard that must hold across teams: the risk-tiering rules that decide what content is MT-forbidden, the locale conventions each language must follow, the definition of the quality gate itself and its severity thresholds. The failure this prevents is the enterprise version of the pilot's silent success: at pilot scale, the standards held because one person carried them in their head; at enterprise scale, standards that live in someone's head drift the moment a second team cannot read that person's mind. Governance writes them down, assigns them an owner, and makes them enforceable, so that a new business unit inherits the same risk tiers, termbase, gate definition, and locale rules rather than inventing its own and discovering the incompatibility in production. The distinction to hold is that evaluation catches errors in a batch and governance prevents a class of error across the whole operation; the batch gate would never catch terminology drift because each fork is internally consistent, so governance is the only control that can.
Control Four: Track Provenance on Every Segment
The fourth control is provenance, and it is the control that keeps the translation memory an asset instead of letting it become a liability. Provenance is a durable record, attached to every segment that enters the TM, of where it came from and what it passed through: which engine drafted it, which qualified human post-edited it, whether it passed a real severity-scored gate with zero Criticals, which termbase version it was checked against, and when. The rule that provenance enforces is simple and non-negotiable at scale: a segment may be written back into the authoritative translation memory only if it passed the real gate. Content that shipped unscored under deadline pressure, content from a team using a forked termbase, content that was never evaluated, does not earn a place in the compounding asset, because the TM's entire value is that a match from it can be trusted, and a match that cannot be trusted is worse than no match at all, since it arrives wearing the authority of prior approval and a rushed post-editor will trust it precisely because the TM vouched for it.
Provenance is what makes contamination recoverable rather than permanent. When you know the source and gate status of every segment, a contaminating batch can be found and quarantined: you query the TM for every segment that entered from the compromised source, in the compromised window, and remediate exactly those, rather than facing the impossible task of re-inspecting an enterprise-scale memory you can no longer trust as a whole. Without provenance, a discovered contamination poisons your confidence in the entire TM, because you cannot distinguish the clean segments from the infected ones, and an asset you cannot trust in part you cannot trust at all. Provenance also makes the operation defensible under audit, which the revised ISO 18587 and ISO 5060 increasingly expect: when a client or a certifier asks how you know a delivered segment was properly post-edited and gated, provenance is the answer, and "the TM said it was approved" is not. Treat the translation memory the way a finance function treats a ledger: every entry has a source, every entry is auditable, and nothing enters without passing the control, because a ledger that accepts unverified entries is not an asset, it is a growing exposure.
Adding Languages and Content Types Without Eroding the Gate
With the four controls in place, the question becomes operational: how do you actually add a language or a content type to the enterprise without the act of adding it degrading the gate? The answer is a discipline, not an event. The instinct that shipped the dosage error was to add everything at once because the pipeline was proven; the discipline is to treat every addition as a controlled expansion along exactly one axis, gated by a readiness check that confirms the four controls are in place for the new cell before any volume flows through it.
One Axis at a Time Still Holds at Enterprise Scale
The single-axis expansion discipline from the roadmap does not stop applying because you are at enterprise scale; it becomes more important, because the blast radius of a mistake is now enterprise-wide. Adding a language is one axis. Adding a content type is another. Adding a business unit is a third. Raising a content type's risk tier into MT-first is a fourth. You change one at a time, hold the others fixed, and watch the critical-error rate for that specific cell, because a single-variable change gives you a clean signal about whether the new cell is behaving and a bounded problem if it is not. An operation that adds a language and a content type and a business unit in the same quarter has thrown away its ability to attribute a quality regression to a cause, exactly the ability it most needs when something goes wrong at scale. The temptation to parallelize is strong because leadership wants coverage fast, and the strategist's job is to explain that parallel expansion does not get you to safe coverage faster; it gets you to a tangled failure you cannot diagnose, which is slower, because you have to stop everything to find the cause.
The New-Cell Readiness Check
Before volume flows into a new language-content cell, it passes a readiness check that verifies the four controls are actually in place for that specific cell, not in general. A control that exists for German does not exist for Korean until you have built it for Korean. The readiness check asks, concretely:
- Terminology: Does the authoritative termbase cover this content type in this language, approved and governed, or are we about to run an engine grounded on gaps? A new cell with a thin termbase will drift on day one.
- Evaluation capacity: Do we have qualified evaluation capacity for this language-domain cell, sufficient to gate the scheduled volume within the deadline, or are we scheduling failure one? If the capacity is not there, the cell waits until it is, or the volume is reduced to fit.
- Automated checks: Are the mechanical checks configured for this locale, its placeholder formats, its date and unit and currency conventions, its length budgets, so the human capacity is spent on judgment, not mechanics?
- Provenance: Is the TM for this cell tracking provenance, so that whatever we write back is auditable and a contamination is recoverable?
- Risk tier: Has this content type been risk-tiered for this cell, and if it is high-liability, is it correctly routed to full post-editing or human-only rather than sliding into the cheap workflow because the cell is new and nobody set the tier?
The readiness check is the enterprise-scale equivalent of the pilot's first-safe-win discipline: it refuses to let a new cell start until the conditions that made the pilot succeed have been deliberately rebuilt for that cell. It is fast to run once the controls exist, because you are checking that four known things are present, not inventing them each time. And it is the specific mechanism that prevents the scaling lie, because it forces the operation to confront, for every single addition, the question the big-bang rollout never asked: are the invisible conditions actually present here, or are we about to assume them and find out the hard way that they are not?
The Highest-Risk Content Never Enters
One rule survives all of the above unchanged, and it is the rule that scale pressures hardest: the highest-liability content, the regulated life-safety, medical, and legal material the program forbids the machine to touch, does not enter the MT-first operation at any scale, in any language, for any volume incentive. Scale creates enormous pressure to sweep this content in, because it is often high-volume and the savings look large, and because a new business unit joining the operation may not know the exclusion exists. Governance is what holds the line: the exclusion is written down, owned, and enforced by the readiness check, so that no cell can quietly route MT-forbidden content into MT-first because a local team decided the volume was too tempting to leave on the table. The risk axis keeps its veto power at enterprise scale exactly as it held it at pilot scale, and the strategist's job is to make that veto structural, embedded in governance and the readiness check, rather than dependent on someone remembering the rule under deadline pressure.
Every addition is a single-axis expansion gated by a readiness check that verifies terminology, evaluation capacity, automated checks, provenance, and risk tier are actually present for the new cell. The MT-forbidden exclusion holds at every scale, enforced by governance, never by memory.
A Worked Scale-Up: One Language to the Enterprise
Return to the operation from the opening, but this time scale it the way the controls demand, and watch the critical-error rate stay flat while the volume, breadth, and variety all climb. The starting point is the successful pilot: English into German, knowledge-base articles, one business unit, one evaluator, a critical-error rate held at zero across eleven months, a run rate of roughly 200,000 words a quarter. The goal is the same enterprise leadership wanted: nineteen languages, six content types, four business units, roughly two million words a quarter. The difference is entirely in the sequencing and the controls, and the difference is the whole game.
Phase One: Instrument Before You Widen
The first move is not to add anything. It is to build the controls the pilot succeeded without, because the pilot succeeded on invisible conditions and you cannot scale invisible conditions. In this phase you automate the mechanical checks for the German cell (placeholder preservation, locale formats, number and negation preservation, term-presence against the termbase, length budgets), which immediately frees the single evaluator from mechanical work and roughly doubles their effective semantic-judgment capacity on the same content. You establish the authoritative termbase with a governance owner and a change-control workflow, even though there is still only one team, because you are about to have more and the governance must exist before the second team arrives, not after it has forked. You turn on provenance tracking so every segment written back to the TM carries its source and gate status. And you calculate, for the first time, the qualified-evaluation capacity you have and will need, in words per quarter per language-domain cell. You have added zero languages and zero content types, and gained not a single word of throughput, and this is the most important phase in the entire scale-up, because everything after it depends on these controls existing. Leadership will find this phase frustrating because it produces no coverage. The strategist's job is to hold it, because scaling before it is scaling the lie.
Phase Two: Widen on One Axis and Watch the Cell
Now expand, one axis at a time, each expansion gated by the readiness check. Add the second content type in German first (say, UI strings), not a second language, because German is your mature language and holding the language fixed while you vary content type keeps the signal clean. The readiness check confirms the termbase covers UI strings, the automated checks handle the string formats, the evaluator has capacity for the added volume (verified against the capacity number from phase one, not assumed), and the risk tier is set. You watch the critical-error rate for the new German-UI cell specifically. If it holds at zero, that cell is proven and you have learned that the controls transfer to a new content type in the same language. Only then do you add the second language (say, French) on the proven content type, again through the readiness check, which this time flags that French evaluation capacity is thinner than German, so you pull the capacity-planning levers, automate the French locale checks and qualify a French evaluator, before the volume flows, not after. Each cell you light up is a single-variable change with a clean signal and a bounded blast radius, and the critical-error rate is watched per cell, so a regression in French-UI is a French-UI problem you can contain, not an enterprise mystery.
Phase Three: Scale the Binding Constraint Deliberately
As cells multiply, the evaluator shortage becomes the pacing constraint, exactly as the theory predicts, and phase three is the deliberate scaling of qualified evaluation capacity ahead of the volume that will demand it. This is where the capacity plan earns its keep: it tells you, quarters in advance, that reaching nineteen languages at two million words will require a specific number of qualified evaluator-hours in specific cells, and that the ISO 18587 competence bar means those evaluators must be recruited or developed on a quarters-long lead time. So you start qualifying evaluators for the languages the roadmap will reach in two quarters, now, before the volume arrives, and use risk-based allocation to spend the scarce deep-MQM capacity on the high-risk cells while automated checks and lighter review carry the low-risk content. The business units join one at a time through the readiness check, each inheriting the authoritative termbase and governed standards rather than forking their own, so terminology drift never starts, because governance was built in phase one, before there was a second team to drift. The TM grows across all the cells, but every segment carries provenance, so it compounds as an asset rather than an infection, and a contamination in any cell is findable and recoverable rather than a poisoning of the whole.
The Outcome: The Rate That Did Not Climb
At the end of the scale-up the operation has reached the enterprise target: nineteen languages, six content types, four business units, two million words a quarter, roughly a tenfold increase in volume and a sixfold increase in variety over the pilot. And the critical-error rate is still at or near zero, because it was never allowed to become a function of volume. It was held flat by four controls that scaled the invisible conditions the pilot got for free: automated checks that concentrated scarce human judgment where only judgment works, an evaluator-capacity plan that made the binding constraint visible and scaled it ahead of demand, central governance that gave every team one authoritative termbase and one set of standards so consistency was designed rather than accidental, and provenance that kept the TM an auditable asset rather than a contaminating liability. The contrast with the opening is the entire lesson. Same technology, same pipeline, same engines. The difference between an operation that scaled to 0.4 percent and one that scaled to zero is whether the strategist understood that scaling a pipeline means rebuilding, deliberately, every condition that pipeline quietly depended on, at a scale where none come for free.
Key Takeaways
- Quality does not scale with the pipeline; it scales with the weakest invisible condition. A pilot's clean critical-error rate is produced by the pipeline plus a set of conditions (one evaluator can see everything, one termbase drifts nowhere, the TM is small and trusted, the gate is never over capacity) that hold for free at small scale and break by arithmetic at large scale. Scaling the pipeline exposes their absence.
- The critical-error rate is the one number that must not climb with volume. Throughput and cost improve almost automatically as you scale; the critical-error rate is different, and 0.4 percent on two million words a quarter is eight thousand defective words, mostly fluent and confident. Anchor the entire scaling conversation on that rate, not on coverage or savings.
- Know the four scaling failure modes by name. The gate bypassed under volume (the overflow ships unscored), the evaluator shortage beneath it (qualified human capacity is the true bottleneck and cannot be provisioned like compute), terminology drift across teams (forked termbases the engine faithfully propagates), and TM contamination at scale (defective segments written back become permanent debt). Each is a specific invisible condition failing, and each is preventable if named before you grow.
- Automate the mechanical checks so scarce human judgment goes where only judgment works. Placeholder, locale, number, negation, term-presence, and length checks run by machine across the full volume free the evaluator to spend all their capacity on the silent fluent mistranslation no automated rule can catch. This is what makes enterprise-scale quality arithmetically possible.
- Plan evaluator capacity as the binding constraint, per language-domain cell. A global evaluator-hours number is a fiction; a Korean pharmaceutical file cannot be scored by a French marketing evaluator. Calculate qualified capacity per cell, never schedule more high-risk volume than a cell can gate within the deadline, and scale capacity (automation, qualification on a quarters-long lead time, risk-based allocation) ahead of the volume that will demand it.
- Govern terminology and standards centrally; provenance-track the TM like a ledger. One authoritative termbase with change control kills drift that a per-batch gate can never catch, because each fork is internally consistent. Provenance on every segment (source, post-editor, gate status, termbase version) keeps the TM an auditable asset, makes contamination recoverable rather than permanent, and answers the audit question that "the TM said it was approved" cannot.
- Add languages and content types one axis at a time, gated by a readiness check. Every addition changes exactly one variable (language, content type, business unit, or risk tier) so the critical-error rate signal stays clean and the blast radius stays bounded. Before volume flows into a new cell, verify terminology, evaluation capacity, automated checks, provenance, and risk tier are actually present for that specific cell, not in general.
- Instrument before you widen, and never let scale sweep in the MT-forbidden content. The most important phase of a scale-up adds zero coverage: it builds the controls the pilot succeeded without. And the regulated life-safety, medical, and legal exclusion holds at every scale, enforced structurally by governance and the readiness check, never by someone remembering the rule under deadline pressure.
Skill.re