The AI Transformation Playbook for Localization
A Head of Localization at a global software company was handed a slide deck and a number. The slide said "AI-native localization" in a font two sizes too large, and the number was a forty percent cost reduction the CFO had already promised the board for the next fiscal year. She had eighteen months, forty-one target locales, a translation-management system with three years of clean translation memory in it, a vendor pool that had heard the phrase "the machine will replace you" one too many times, and a legal team that would end her career if a mistranslated data-processing clause shipped in German. The temptation, the one nearly every operation gives in to, was to move fast: switch the whole pipeline to machine-translation post-editing overnight, cut the per-word rates to hit the number, and let quality find its own level. She had watched a competitor do exactly that. For one quarter the dashboards were beautiful, throughput doubled, cost per word collapsed, and the executives were delighted. Then a fluent, grammatically perfect, completely wrong rendering of a medical device's contraindication shipped in three languages, a regulator got involved, and the entire program was frozen while lawyers counted the exposure. The competitor did not have a speed problem. It had a sequencing problem. It scaled before it could prove quality, and the collapse was not bad luck, it was the predictable result of the order in which they did things. This lesson is about the order. It is the staged path an enterprise localization function actually walks to get from a pilot to an operating model without the speed-only collapse, and the central claim is that in localization the sequence is the strategy: quality-first is not a slogan on a values page, it is the specific ordering of moves that lets you capture the speed without inheriting the liability.
Why Transformation Is a Sequence, Not a Switch
There is a seductive mental model of AI transformation that treats it as a switch: the operation is "before AI" on Monday and "AI-native" on Tuesday, and the work in between is procurement and a training webinar. That model is wrong in every industry, and it is dangerous specifically in localization, because of the one asymmetry this whole program turns on. Machine-translation and large-language-model output is fluent first and accurate second. The machine produces a rendering that reads perfectly whether or not it means what the source meant, which means the failures do not announce themselves. A speed-only transformation therefore does not fail loudly and early, the way a mistranslation into garbled grammar would. It fails silently and late: the throughput climbs, the costs fall, everyone celebrates, and the flipped dosage or the inverted indemnity clause sits in the shipped output for weeks until a regulator, a customer, or a lawsuit finds it. By the time the failure surfaces, you have shipped it at scale, into many locales, and the cleanup is not a re-edit, it is an incident.
Before we go further, let us fix the working vocabulary this lesson leans on, because an enterprise transformation touches every role and the terms have to be shared. MT is machine translation; NMT is neural machine translation, the classic sentence-to-sentence engine; LLM is a large language model, more fluent and more confidently wrong than NMT. PE is post-editing and MTPE is machine-translation post-editing, the workflow where the machine drafts and a human revises. QE is automatic quality estimation, a machine's own confidence signal. MQM is the Multidimensional Quality Metrics error typology, and ISO 5060:2024 is the standard that formalizes the MQM-aligned analytic scoring of translation output by dimension and by severity (Critical, Major, Minor). ISO 18587 is the post-editing standard, whose revision (in Draft International Standard ballot, publication targeted for late 2025 into 2026) expands scope to AI and LLM "non-human translation output," retires the rigid light-versus-full split for an effort spectrum, and requires the post-editor to hold the full linguistic competence of a professional translator. ISO 17100 is the baseline standard for professional human translation services. TM is translation memory, TMS is the translation-management system, CAT is the computer-assisted translation tool, a segment is the sentence-sized unit the CAT tool works in, locale is the language-plus-region target such as de-DE, i18n is internationalization and l10n is localization, and LSP is a language-service provider. The operating model is the eventual steady-state design: the roles, processes, governance, and metrics by which the transformed function runs every day, not a project but the new normal.
In localization, transformation failures are silent and late, not loud and early. That is precisely why the sequence you follow is the strategy, and why quality has to be proven before speed is scaled.
The Speed-Only Collapse as a Predictable Pattern
It helps to name the failure precisely, because it is not random and you will see it coming if you know its shape. The speed-only collapse has four stages, and they always run in the same order. First, an operation under cost pressure switches to MTPE broadly and cuts per-word rates to match, without first building a quality gate that can catch the silent critical error. Second, throughput and cost metrics improve immediately and visibly, so the transformation looks like a triumph and gets held up as a model. Third, because the quality gate does not exist or is a token "looks fine to me" check, one or more silent critical errors ship into high-consequence content and propagate across the locales the operation scaled into. Fourth, the error is discovered externally, by a regulator, a client, or a lawsuit, and the response is a freeze that costs more than the entire "saving" the speed move produced, plus a reputational hit that follows the operation into every future deal. The lesson of the pattern is not "AI is dangerous." It is "scaling before you can prove quality guarantees the collapse," and the entire staged playbook exists to invert that order so the proof comes first.
The Five Stages of a Localization AI Transformation
An enterprise localization function does not walk from "before" to "after" in one step. It walks through five stages, each of which produces something the next stage depends on, and the discipline is to refuse to start a stage until the previous one has delivered its output. The stages are: assess, pilot with a quality gate, scale by tier and language, operating model, and continuous improvement. Think of them not as a timeline you march through on a fixed date, but as a set of gates you earn the right to pass. The single most common cause of transformation failure is not any of the stages done badly, it is stages done out of order: piloting before assessing, scaling before piloting, declaring an operating model before scaling has taught you where it breaks. We will take each stage in turn, slowly, because the content of each one and the reason it precedes the next is the substance of the playbook.
Stage One: Assess, or Knowing What You Are Standing On
The assessment stage answers one question before a single engine touches a live file: what is the true current state of the operation, across assets, content, quality maturity, and people. Skipping it feels efficient and is the most expensive shortcut in the sequence, because every later decision, which engine, which content first, what the quality gate must catch, is only as good as the map you are working from, and an operation that pilots on a wrong map pilots the wrong thing. The assessment has four axes.
The first axis is linguistic assets: the translation memory, the termbases, and the style guides. These are the enterprise's accumulated, approved language, and in an AI transformation they are not archives, they are the ground truth you will hold the engine to and the material you will ground it on. An operation with three years of clean, deduplicated, well-attributed TM is standing on a very different foundation from one whose TM is a landfill of unverified leverage and propagated errors, and the transformation strategies available to them differ accordingly. Part of assessment is honestly grading the assets, because grounding a fluent engine on a poisoned TM industrializes the poison.
The second axis is the content portfolio, classified by two variables at once: volume and consequence. Volume tells you where the economic prize is (the high-volume, repetitive content where speed pays off most). Consequence tells you where the liability is (the regulated, life-safety, legal, and financial content where a silent critical error costs a life or a lawsuit). The assessment produces a portfolio map that places every content type on both axes, because the transformation sequence is going to move through that map in a deliberate order, and you cannot sequence a map you have not drawn. The third axis is quality maturity: does the operation already score against an MQM/ISO 5060 typology, or does it "vibe" quality with a senior reviewer's gut. An operation with no analytic scoring capability has to build one before it can pilot honestly, because you cannot prove a pilot's quality with a tool you do not have. The fourth axis is people and change readiness: the vendor pool's trust level, the internal team's AI literacy, and the political weather, because a transformation is a change program wearing a technology costume, and a linguist pool that believes the transformation is a plan to replace them will quietly ensure it fails.
Assessment is not paperwork before the real work. It is the map every later stage navigates by, and a pilot run on a wrong map is a confident answer to the wrong question.
Stage Two: Pilot With a Quality Gate, or Proving Before Scaling
The pilot stage is where quality-first sequencing does its most important work, and where most transformations quietly go wrong by doing the pilot for the wrong reason. The failure mode is to pilot in order to prove that AI is fast. That proof is trivial and worthless: everyone already knows the machine is fast, and a pilot that only measures throughput measures the thing that was never in doubt while ignoring the thing that kills you. The correct purpose of the pilot is to prove that the operation can capture the speed and simultaneously catch the silent critical error, on real content, with a real quality gate, before any of it is scaled. The deliverable of a good pilot is not "AI works." It is a defensible quality record on a bounded slice of content, showing throughput up, cost down, and a severity-scored evaluation with zero Criticals, all provable to a skeptic.
That means the quality gate has to exist before the pilot starts, not after. The gate is the MQM/ISO 5060 scoring model with its dimensions (accuracy, terminology, locale convention, style and fluency), a numeric pass/fail threshold expressed as a normalized error score per thousand words, and the non-negotiable absolute rule that a single Critical error fails the file regardless of how clean the average looks. Piloting without this gate is not a pilot, it is an uncontrolled experiment on live content, and it produces exactly the false confidence that precedes the collapse. Choose the pilot content deliberately: a slice that is representative enough that success predicts scaled success, but bounded enough that a failure is contained and cheap. The classic right choice is a high-volume, lower-consequence content type in a well-resourced language pair, which is why subtitling, UI strings, and high-volume support content are the industry's standard entry points. You are learning to run the gate on content where a caught error is a lesson, not a lawsuit.
A pilot also has to be honest about its own limits, and the discipline here is to define the exit criteria before you begin, in writing, so that a motivated executive cannot redefine "success" after the fact to match the throughput number they wanted. Good pilot exit criteria state the minimum throughput lift, the maximum acceptable cost per word, the required MQM/5060 score against threshold, the absolute zero-Critical requirement, and, crucially, the terminology-conformance bar, because an engine that drifts off approved terms at pilot scale will drift catastrophically at enterprise scale. A pilot that clears written criteria has earned the right to scale. A pilot that clears only the throughput metric has proven nothing that matters, and scaling it is the collapse in slow motion.
Stage Three: Scale by Tier and Language, or Expanding Along the Safe Axes
Scaling is the stage where the speed-only operations blow themselves up, because they scale along the wrong axis: they scale by volume, chasing the biggest content first because that is where the cost saving looks largest, and volume is uncorrelated with safety. The disciplined operation scales along two deliberately chosen axes instead, risk tier and language, and it moves outward from proven ground rather than jumping to the biggest prize.
Scaling by risk tier means the transformation moves up the consequence axis only as the quality gate proves it can hold. You do not begin the MTPE-first workflow on drug labels and indemnity clauses because they have high volume. You begin on the lower-consequence tier the pilot proved, then extend to the next tier up only when the gate has demonstrated, on that lower tier at scale, that it catches what it must. The regulated, life-safety, legal, and financial content is scaled last and most cautiously, and some of it never enters the MTPE-first flow at all: it stays full human translation or full post-editing, or it is flagged MT-forbidden, because matching effort to consequence is the first control and the one that prevents the catastrophic mismatch. Scaling by language means expanding from the well-resourced language pairs (where the engine is strong, the TM is deep, and the reviewer pool is mature) outward to the long-tail locales (where engine quality is lower, assets are thinner, and a silent error is both more likely and harder to catch). The order protects you: you learn the workflow where the conditions are forgiving before you apply it where they are not.
The governing principle of the scale stage is that the quality gate is the throttle. Scale advances only as fast as the gate proves it can hold at the new tier or the new language, and the gate's evidence, not the throughput dashboard, is what authorizes the next expansion. An operation that lets the throughput dashboard drive the scaling will always outrun its gate, because throughput is easy and gate-proving is work, and the gap between them is exactly the space where the silent critical error ships. The strategist's job during scale is to keep the gate ahead of the volume, which sometimes means telling an impatient executive that the next tier is not authorized yet, not because AI cannot do it, but because the operation has not yet proven it can catch the failure if it does.
Scale by risk tier and by language, not by volume. Volume is where the money looks biggest and where safety is least correlated with the prize. The quality gate is the throttle, and its evidence, not the throughput dashboard, authorizes each expansion.
Stage Four: The Operating Model, or Making It the New Normal
A transformation that stops at "we scaled AI across the portfolio" has not transformed, it has run a large project that will decay the moment the project team disbands. The fourth stage converts the scaled workflow into an operating model: the permanent design of roles, processes, governance, and metrics by which the function runs every day, so that the new way of working is not a heroic push but the ordinary shape of the operation. The distinction matters because a project has an end and an operating model does not, and the silent critical error does not respect the end of a project. If the discipline that caught it during the transformation lives in a project team's vigilance rather than in the permanent structure, it evaporates when the team moves on, and the operation drifts back toward speed-only by entropy.
The operating model makes four things permanent. It makes the roles permanent: the qualified post-editor held to full professional-translator competence, the evaluator who is structurally independent of the post-editor for the same file, and the quality owner who owns the program itself, all as standing positions with named holders, not project assignments. It makes the process permanent: one documented workflow, risk-tiered, that any qualified linguist can run to the same outcome, applied across the whole portfolio rather than reinvented per account. It makes the governance permanent: the standing gate, the calibration cadence that keeps evaluators from drifting apart, and the internal audit that behaves like an external one. And it makes the metrics permanent as a dual-axis instrument: throughput and cost on one axis, quality-risk (critical-error rate, terminology conformance, MQM/5060 score) on the other, reported together always, because reporting speed without the paired quality-risk number is how an operation lies to itself and its board. The operating model is the answer to "what does this function look like on an ordinary Tuesday two years from now," and if you cannot describe that Tuesday, you have a project, not a transformation.
Stage Five: Continuous Improvement, or Why the Work Is Never Finished
The fifth stage is the one operations most often skip, treating the operating model as a destination rather than a living thing, and the skip is costly because the ground under a localization operation moves constantly. Engines are retrained and their behavior shifts, sometimes improving and sometimes regressing on your specific content in ways your last evaluation did not predict. New content types enter the portfolio and have to be tiered and routed. Languages get added. Standards evolve: the very ISO 18587 revision this program is built around is itself a change the operating model must absorb. Evaluators drift and must be recalibrated. Terminology grows and the termbases must be maintained or the engine's conformance decays. A continuous-improvement stage is the standing discipline of re-assessing the engine, the content, the quality data, and the standards on a cadence, and feeding what it learns back into the tiering, the gate, the grounding, and the training. Without it, the operating model is a photograph of a moment that has already passed, and the gap between the photograph and reality is where the next silent critical error lives.
Notice that continuous improvement closes the loop back to assessment: it is, in effect, assessment run perpetually rather than once, which is why the five stages are better drawn as a cycle than a line. The transformation is never "done" in the sense a project is done. It reaches a steady state in which the operating model runs the daily work and the continuous-improvement discipline keeps that model honest against a moving world. An enterprise localization leader who understands this stops asking "when will the transformation be finished" and starts asking "is our operating model still true this quarter," which is the right question and the one that keeps the collapse permanently at bay.
The Quality-First Sequencing Rule
Everything above rests on a single sequencing rule, and it is worth stating in isolation because it is the load-bearing idea of the entire playbook: you must be able to prove quality before you are allowed to scale speed. The proof is the gate, and the gate must precede the volume at every stage, not follow it. This is not risk-aversion for its own sake and it is not a brake on ambition. It is the recognition that in localization the speed move and the quality move are the same move only when they are sequenced correctly, and they become opposite moves when they are not. Sequenced right, the gate lets you scale confidently, because every expansion is authorized by evidence that the operation can catch the failure at the new frontier. Sequenced wrong, speed without proof is not fast, it is fast toward a cliff, and the throughput you gained is the exact measure of how much wrong content you will have shipped when the cliff arrives.
The rule has a concrete operational form: at every stage boundary, the artifact that authorizes crossing is a quality record, not a throughput number. To move from pilot to scale, you present the pilot's severity-scored evaluation with zero Criticals and its terminology-conformance data, not its words-per-day. To move from one risk tier to the next during scale, you present the gate's evidence that it held at the current tier. To declare the operating model, you present the governance that makes the gate permanent. To claim continuous improvement is working, you present the re-assessment cadence and what it changed. In every case the currency of authorization is provable quality, and an operation that tries to pay for a stage crossing with throughput alone is counterfeiting. The strategist's spine, the thing they are paid to hold when an executive wants to skip ahead, is the refusal to accept throughput as payment for a gate that quality was supposed to buy.
The rule that prevents the collapse: prove quality, then scale speed, at every stage boundary. The currency of authorization is a quality record, never a throughput number. Speed without proof is not fast, it is fast toward a cliff.
The Leadership and Change Dimensions
A localization AI transformation is often described as a technology and process program, and that description is half the truth and the less important half. The harder half is that it is a change program that succeeds or fails on two human relationships: the one with the executives who fund it and the one with the linguists and vendors who execute it. A transformation that is technically perfect and politically naive will die, and the leader has to manage both relationships as deliberately as the pipeline.
Managing Upward: The Executive and Client Story
The executive who sponsors the transformation almost always arrives with a single-axis mental model: AI equals cost reduction, and the measure of success is the cost line. The leader's most important upward job is to install a second axis before the first one drives the operation off the cliff, and to do it in the executive's own language. The frame that works is not "quality is important too," which sounds like a linguist asking to slow down. The frame that works is risk-adjusted cost: the true cost of the transformation includes the expected cost of the failures it produces, and a speed-only approach that ships a silent critical error into a regulated market has a catastrophic tail cost that dwarfs the per-word saving. Told this way, the quality gate is not a tax on the cost saving, it is the insurance that makes the cost saving real rather than a loan against a future incident. The competitor's frozen program is the case study that makes this concrete to a board: they hit the cost number for one quarter and then lost far more than they saved. The leader who can tell the throughput story and the quality-risk story as one story, always paired, never separated, is the leader who keeps the sponsor's expectations survivable.
Managing Across: The Linguist and Vendor Trust Contract
The people who will actually catch the silent critical error are the linguists, and a transformation that frightens or insults them removes the exact human judgment the whole quality-first strategy depends on. The vendor pool that has heard "the machine will replace you" is not being irrational, it is responding to a real economic shift, and pretending otherwise poisons the relationship the transformation needs most. The honest contract, and the one that both retains people and matches the standards, is the one this program's spine rests on: the AI drafts, the human owns the quality. The revised ISO 18587 makes this contract real rather than rhetorical, because it insists the post-editor hold full professional-translator competence, which means the transformation is not deskilling the role, it is moving it up the value chain from words-per-hour production to quality ownership, terminology stewardship, and risk judgment. The leader who communicates the transformation as "we are making you the owner of the thing the machine cannot do" and backs it with real role redesign, real re-qualification, and career paths that reward quality ownership, keeps the pool. The leader who communicates it as "we are cutting your rate because the machine does most of it now" loses the pool, and with it loses the human gate that was the entire safeguard against the collapse. The change dimension is not soft. It is the availability of the judgment your quality-first strategy is built on.
The Failure Modes to Name Before They Happen
A playbook is as much a list of the ways it goes wrong as a list of the stages, because naming the failure modes in advance is how a leader recognizes them early enough to correct. Five recur across enterprise localization transformations, and each maps to a violation of the sequencing rule.
- Scaling before proving. The archetypal collapse: switching broadly to MTPE and cutting rates before a quality gate exists, so the silent critical error ships at scale. The tell is a transformation whose earliest and proudest metrics are all throughput and cost, with no paired quality-risk number. The correction is to refuse the scale-stage crossing until a pilot has cleared written quality criteria.
- Piloting for the wrong proof. Running a pilot that measures only speed, declaring victory on the metric that was never in doubt, and scaling a pilot that proved nothing about the failure that matters. The tell is a pilot report with impressive words-per-day and no severity-scored evaluation. The correction is written exit criteria dominated by the MQM/5060 score, the zero-Critical rule, and terminology conformance.
- Metrics that hide risk. Reporting edit-distance, speed, and cost as if they were quality, so the dashboard looks healthy while the critical-error rate is unmeasured. The tell is a leadership dashboard with no critical-error rate and no terminology-conformance line. The correction is the permanent dual-axis metric that never reports speed without its paired quality-risk number.
- The hero-dependent operation. Quality that lives in one exceptional reviewer's head and personal records rather than in a documented, staffed, calibrated program, so the operation's conformance evaporates the day the hero leaves. The tell is an inability to produce the full quality record for a randomly chosen file worked by anyone other than the star. The correction is the operating-model stage: roles, documented process, logged evidence, calibration, and audit made permanent.
- Engine concentration risk. Building the entire transformed operation on a single engine or vendor, so a price change, a quality regression after a retrain, or a service outage becomes an operational crisis with no alternative. The tell is an operating model with no engine-evaluation cadence and no fallback. The correction lives in continuous improvement: re-assess engines on a cadence and treat single-engine dependence as a named risk, not a convenience.
Every failure mode is the same disease with different symptoms: the volume got ahead of the proof. Name them in advance and you catch them while they are still a course-correction, not an incident.
A Worked Multi-Stage Transformation
Return to the Head of Localization from the opening, with her forty-one locales, her eighteen months, her three years of clean TM, her wary vendor pool, and her CFO's promised forty percent. Watch her walk the five stages in order, and watch the sequencing rule do its work at each boundary.
She Assesses Before She Touches a File
She spends her first six weeks not switching anything, which the CFO finds alarming and she defends as the cheapest insurance in the plan. She grades the four axes. The linguistic assets are strong in her top eight languages (deep, clean TM; maintained termbases) and thin in the long-tail thirty-three. The content portfolio maps into four tiers: high-volume support content and UI strings (high volume, low consequence), marketing and web (medium both), product documentation (medium-high consequence), and regulated medical-device and legal content (highest consequence, and some of it she marks MT-forbidden on the spot). Quality maturity is low: the operation has never scored against MQM/5060, so she notes that a scoring capability must be built before any honest pilot. People readiness is fragile: the vendor pool is skilled but frightened. Her assessment output is a map, and the map already tells her the sequence: pilot on support content in a top-eight language, scale up the tiers and out the languages, never let the regulated content near the MTPE-first flow until the very end, and start the change work with the linguists immediately, in parallel, because trust takes the full eighteen months to build.
She Builds the Gate, Then Pilots Behind It
Before the pilot, she stands up the quality gate that did not exist: an MQM/ISO 5060 scoring model, a threshold per thousand words, the absolute zero-Critical rule, and a terminology-conformance bar. She qualifies two evaluators and separates them from the post-editors by rule. Only then does she pilot, on high-volume support content in German, one of her strong-asset languages. She writes the exit criteria first: a defined throughput lift, a maximum cost per word, an MQM/5060 score against threshold, zero Criticals, and a terminology-conformance floor. The pilot runs for six weeks. Throughput lifts hard, cost per word drops toward the CFO's number, and, critically, the severity-scored evaluation comes back with zero Criticals and terminology conformance above the floor, on real content, provable to the legal team and the board. She now holds the artifact that authorizes the next stage: not a throughput slide, a quality record. When the CFO asks to skip straight to scaling all forty-one languages, she declines, and the pilot's quality evidence, not her opinion, is what lets her decline defensibly.
She Scales by Tier and Language, Gate as Throttle
She scales along the two safe axes, not by volume. First she extends the proven support-content workflow across her top-eight strong-asset languages, watching the gate hold at each new locale before adding the next. Then she moves up one tier, to marketing and web, only after the gate has proven itself on support content at scale, and up again to product documentation only after marketing holds. The long-tail thirty-three languages come after the top eight, and more cautiously, because the assets are thinner and the engine weaker, so the gate is set tighter and more human effort is budgeted. Through all of it, the gate is the throttle: when the critical-error rate ticks up on a long-tail language, she pauses that expansion and does not resume it until the cause is found and the gate proves it catches it. The regulated medical-device and legal content never enters the MTPE-first flow; it stays full human or full post-editing, and the MT-forbidden slice stays forbidden. Eleven months in, the throughput and cost numbers are close to the CFO's target across most of the portfolio, and the critical-error rate is measured and near zero, because the volume never got ahead of the proof.
She Converts the Project Into an Operating Model
With most of the portfolio scaled, she refuses to let the discipline live in her transformation team's vigilance, because she knows the day that team disbands is the day the operation drifts back toward speed-only. She makes the roles permanent (standing post-editor, evaluator, and quality-owner positions with named holders), the process permanent (one documented, risk-tiered workflow across the whole portfolio), the governance permanent (the standing gate, a quarterly evaluator calibration, and an internal audit that behaves like an external one), and the metrics permanent (a dual-axis dashboard that never shows the CFO a cost number without the paired critical-error rate and terminology-conformance number beside it). She can now describe the ordinary Tuesday two years out, which is how she knows she has an operating model and not just a finished project.
She Keeps It Honest Against a Moving World
Finally, she institutes continuous improvement as a standing cadence, not a closing ceremony. Engines are re-evaluated quarterly on her specific content, because a vendor's retrain that improves general quality might regress on her medical terminology, and she will not learn that from the vendor's press release. New content types get tiered on intake. The ISO 18587 revision, when it publishes, is absorbed into the process and the post-editor qualification bar rather than ignored. Evaluators are recalibrated on the quarterly cadence so the gate keeps meaning the same thing. And because she treats single-engine dependence as a named risk, she keeps a second engine qualified and ready, so a price change or an outage from her primary vendor is a switch, not a crisis. Eighteen months after the alarming six weeks of doing nothing, she has hit the CFO's cost number, shipped zero regulated-content incidents, kept a vendor pool that now sees itself as quality owners rather than the machine's cleanup crew, and, unlike the frozen competitor, built something that keeps working the quarter after the celebration. The sequence was the strategy, and the proof came before the speed at every gate.
The worked transformation succeeds for one structural reason: at every stage boundary the volume was authorized by a quality record, never by a throughput number. The competitor failed for the mirror reason. Same technology, opposite sequence, opposite outcome.
Key Takeaways
- In localization, the sequence is the strategy. Because MT and LLM output is fluent first and accurate second, a speed-only transformation fails silently and late, not loudly and early. The quality-first ordering of moves is what lets an operation capture the speed without inheriting the liability, and it is not a slogan, it is the specific order the stages run in.
- The transformation walks five stages, and doing them out of order is the primary failure cause. Assess, pilot with a quality gate, scale by tier and language, operating model, continuous improvement. Each stage produces something the next depends on, and each is a gate you earn the right to pass, not a date on a timeline.
- Assessment is the map every later stage navigates by: linguistic assets (TM, termbases, style guides graded honestly), the content portfolio mapped on volume and consequence, quality maturity (do you score against MQM/ISO 5060 or vibe it), and people and change readiness. A pilot run on a wrong map answers the wrong question confidently.
- The pilot exists to prove you can catch the silent critical error while capturing the speed, not to prove AI is fast. The quality gate (MQM/ISO 5060 model, a per-thousand-word threshold, the absolute one-Critical-fails rule, a terminology-conformance bar) must exist before the pilot, and written exit criteria dominated by quality, not throughput, are what authorize scaling.
- Scale by risk tier and by language, never by volume, and let the quality gate be the throttle. Move up the consequence axis and out to the long-tail languages only as the gate proves it holds; keep the regulated, life-safety, and legal content in full-human or MT-forbidden lanes; and authorize each expansion with the gate's evidence, not the throughput dashboard.
- The operating model makes the discipline permanent so it does not evaporate when the project team disbands. Standing roles (qualified post-editor, independent evaluator, quality owner), one documented risk-tiered process, permanent governance (gate, calibration, internal audit), and a dual-axis metric that never reports speed without its paired quality-risk number. If you cannot describe the ordinary Tuesday two years out, you have a project, not a transformation.
- Continuous improvement closes the loop back to assessment and keeps the model honest against a moving world: engines retrain and regress, content and languages are added, standards evolve (including the ISO 18587 revision itself), evaluators drift, and single-engine dependence is a named risk to manage, not a convenience to enjoy.
- The transformation is a change program in a technology costume, and it lives or dies on two relationships. Upward, reframe cost as risk-adjusted cost so the quality gate reads as insurance, not tax. Across, honor the contract that the AI drafts and the human owns the quality, backed by real role redesign under the revised ISO 18587, because the linguists are the human gate the whole quality-first strategy depends on, and a frightened pool is a removed safeguard.
Skill.re