โ†
AI for Translation & Localization
Strategic ยท M5 ยท lesson 5 of 20 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Building a Localization-AI Roadmap
๐Ÿ“–
now learning

Building a Localization-AI Roadmap

15 min

The board deck had one slide that mattered, and it was the slide that nearly ended the program. "Machine translation, all 19 languages, all content, live by end of quarter." A clean bar chart, a confident cost saving, a single green arrow pointing up and to the right. The VP of Localization who inherited that slide called you in week two, after the first regulated client's quarterly review flagged a mistranslated dosage warning that had shipped in three languages, after two senior linguists resigned rather than clean up after an engine for half their old rate, and after the legal team discovered an indemnity clause had been machine-translated in a contract nobody had tiered as high-liability. The technology was not the problem. The engines worked. The post-editing workflow you had spent a year building at L3 worked. What failed was the sequencing: somebody switched everything on at once, in every language, across every content type, with no order, no gates, and no way to stop. This lesson is about the document that should have existed before that slide was ever drawn. It is a localization-AI roadmap: a phased, sequenced plan that rolls the program out by language maturity, content type, and risk, deliberately, with milestones you can defend and gates you can stop at, instead of a big-bang launch that bets the program's credibility on everything working perfectly on day one. We are going to build that roadmap the way a strategist actually builds it, slowly, with a worked multi-quarter example you could adapt to your own operation, because the difference between a localization-AI program that compounds and one that collapses is almost never the engine. It is the order in which you turned things on.

Why Sequencing Is the Strategy, Not a Detail

Start with the failure mode, because it is the thing the roadmap exists to prevent and most people do not see it until it has already cost them. The instinct when a localization-AI program gets funded is to maximize coverage immediately: more languages, more content, more savings, faster. That instinct is the same one that shipped the dosage error. A big-bang rollout fails not because the technology is immature but because it removes your ability to learn, contain, and stop. When you switch everything on at once, every problem arrives at once, in every language, across every content type, and you have no clean signal about what caused it, no contained blast radius when it goes wrong, and no defensible place to pause. You inherit a system you cannot debug and cannot defend.

Before we go further, fix the vocabulary, because a strategist's roadmap is only as good as the precision of the terms it sequences. Machine translation (MT) is any system that renders text from source to target with no human writing the words; machine-translation post-editing (MTPE), often shortened to post-editing (PE), is a human editing that machine output rather than translating from blank. A large language model (LLM) is a general text predictor that translates as a side effect of broad competence, more fluent and more confidently wrong than classic MT. A translation memory (TM) is the client's database of previously approved source-target segment pairs you can leverage; a termbase is the client's controlled glossary of approved terms. MQM (Multidimensional Quality Metrics) is the analytic error typology, formalized for translation output by ISO 5060:2024, that scores errors by category (accuracy, terminology, locale, fluency) and severity (Critical, Major, Minor); one Critical error fails a file. ISO 18587 is the post-editing standard, whose revision (in DIS ballot, publication targeted late 2025 into 2026) expands scope to AI and LLM "non-human translation output," retires the rigid light-versus-full split for an effort spectrum, and requires the post-editor to hold full professional-translator competence. A locale is the specific language-and-region convention bundle: en-US is not en-GB is not de-DE. A language-service provider (LSP) is a translation vendor or agency. Hold those, and the roadmap becomes legible.

A roadmap is the opposite of a big-bang launch in four specific ways, and each one is a survival mechanism rather than a nicety. It contains the blast radius: when MT-first goes wrong on one content type in one language, the damage is bounded to that cell rather than spread across the whole operation. It isolates the signal: when you turn on one thing at a time, a quality regression has one obvious cause, so you can actually learn from it instead of guessing. It builds proof you can spend: each phase produces a defensible result, a quality record with a clean error score, that earns the political capital and the budget for the next phase. And it preserves the right to stop: a phased plan has gates between phases, so a regulated client's bad audit or a spike in Critical errors triggers a contained pause instead of a program-wide crisis. Sequencing is not project-management hygiene bolted onto a strategy. Sequencing is the strategy.

A big-bang rollout removes the three things a strategist needs most: a bounded blast radius when something breaks, an isolated signal about what broke, and a defensible place to stop. A roadmap exists to preserve all three.

The Crawl, Walk, Run Spine

The shape of every credible localization-AI roadmap is the same three-phase spine: crawl, then walk, then run. The names are familiar, but their meaning here is precise and load-bearing, so define them in the operation's own terms rather than as motivational slogans.

Crawl is the phase where you prove the workflow works at all, on the safest possible content, in your most mature language, with a human gate on every segment and the cheapest economics turned off. The goal of crawl is not throughput. The goal of crawl is a clean quality record on a small, contained, low-consequence pilot that nobody can get hurt by, so that you learn the operational reality of MT-first, the grounding, the post-editing, the severity-scored gate, the terminology and locale enforcement, on content where a mistake is recoverable. You are buying knowledge and proof, not savings. If the crawl phase shows a Critical error rate you cannot drive to zero on the easy content, you have learned something priceless before it cost you a regulated client.

Walk is the phase where you widen on one axis at a time, deliberately, having proven the workflow. You add a content type, or a language, or a higher risk tier, but never all three at once, because the discipline of walk is that every expansion changes exactly one variable so the signal stays clean. Walk is where the economics start to appear: you begin to capture the throughput lift, the move from roughly 2,000 words a day to 5,000 or more, on content that has earned MT-first treatment, while the gates hold. Walk is also where you discover the dependencies you underestimated in crawl: the termbase that was good enough for marketing strings is not good enough for technical documentation, the language that was mature for a European market is immature for a new one, the linguist pool that trusted the workflow on easy content needs retraining for higher tiers.

Run is the phase where MT-first is the standard operating posture across the qualified content and languages, the gates are routine, the quality record is automatic, and the human effort has migrated up the value chain to the judgment the engine cannot supply. Run does not mean everything is machine-translated. Run means the program has matured to the point where MT-first is the default for content that has earned it, human-only is the rule for content that requires it, and the organization can tell the difference defensibly, by policy, on every job. Crucially, some content never leaves crawl, and some content never enters the roadmap at all: the high-liability legal, medical, and life-safety material that the program's own rules keep human-only is not a phase to graduate, it is a permanent exclusion the run phase enforces.

The Three Axes You Sequence Along

A roadmap sequences a rollout along three axes at once: language maturity, content type, and content risk. The art of the roadmap is ordering these three so that each phase advances the program while keeping the blast radius small and the signal clean. Take each axis in turn, because each one carries a different kind of risk and a different sequencing logic, and then we will combine them into the worked plan.

Axis One: Language Maturity

Languages are not equally ready for MT-first, and treating them as if they were is one of the most common and most expensive roadmap mistakes. Language maturity is the readiness of a given target language for an MT-first workflow, and it is a function of four concrete things you can actually assess: the volume and quality of your translation memory in that language, the depth and approval status of your termbase, the raw MT or LLM engine quality for that language pair, and the depth of your qualified linguist pool able to post-edit and evaluate it. A mature language is one where you have a large clean TM, a deep approved termbase, strong engine output, and several linguists you trust to run the gate. An immature language is one where any of those four is thin.

The sequencing logic is to lead with your most mature languages, not your highest-volume or highest-revenue ones, because the mature language is where the workflow has the best chance of producing a clean result, which is exactly what the crawl phase needs. A common error is to start MT-first in the language with the most volume, on the theory that it has the biggest savings, when that language happens to be a newer market with a thin TM and a shallow linguist pool. You have then chosen to learn the workflow on the hardest possible language, maximizing the chance of an early failure that kills the program's credibility. Lead with maturity. The high-volume immature language is a later phase, after the workflow is proven and the assets for that language have been deliberately built up, which is itself a dependency the roadmap must schedule.

Axis Two: Content Type

Content types differ in volume, in repetitiveness, in how much the engine helps, and in consequence, and the roadmap sequences them so that the early phases land on the content where MT-first helps most and hurts least. Think of the spread your operation actually handles: high-volume, repetitive, low-consequence content like UI strings, knowledge-base articles, and product descriptions at one end; brand and marketing content where transcreation matters and the engine is weaker in the middle; and high-liability technical, legal, medical, and financial content at the other end. The early phases of the roadmap belong to the first category, because that content has the best MT leverage (high TM match rates, strong engine quality, repetitive structure) and the lowest cost of error, so it is where the economics appear fastest and the risk is smallest.

The sequencing logic on the content axis is leverage-high, consequence-low first. A knowledge-base article is an almost ideal first content type: there is a lot of it, it repeats, the engine handles it well, the TM leverage is high, and a fluent error in it is recoverable. A drug label is the opposite: lower volume, less repetition, catastrophic consequence, and a fluent error that injures a patient. The roadmap never lets the cheap workflow reach the drug label, and it reaches the knowledge-base article first. Brand and marketing content sits in a deliberate middle: it is not high-liability in the safety sense, but the engine is genuinely weaker at it because transcreation requires human judgment about culture and intent, so it enters the roadmap later than UI strings not because of risk but because the MT leverage is lower and the human effort higher.

Axis Three: Content Risk

The third axis is the one the whole program is built to respect: content risk is the consequence of a fluent error, what a confident, grammatical mistranslation in this content would actually do to someone. This is the risk-tiering discipline from L3, elevated to a sequencing principle. Low-risk content, where a fluent error is recoverable, leads the roadmap. High-risk content, where a fluent error can injure a patient or lose a lawsuit, enters last, under full post-editing or human-only rules, after the workflow and the gates have a long clean track record. And the highest-risk content, the regulated life-safety and legal material the program forbids the machine to touch, never enters the MT-first roadmap at all; it stays a permanent human-only exclusion that the roadmap names explicitly so that no future pressure to expand coverage quietly sweeps it in.

The crucial insight about the risk axis is that it is the one you must never sacrifice to the others. There will be pressure to move a high-volume, high-risk content type into MT-first early because the volume promises savings. The roadmap's job is to refuse that, by rule, and to make the refusal defensible: this content is tiered high-liability, it stays human-only or full-PE regardless of its volume, and no phase advances it past the gate its consequence demands. The risk axis has veto power over the other two. A content type can be high-volume and mature in language and still be excluded from MT-first, permanently, because its consequence forbids it. Sequencing by volume alone is how the dosage error shipped. Sequencing by risk first is how it does not.

You sequence along three axes at once: language maturity, content type, and content risk. Risk has veto power over the other two. A content type can be high-volume and mature in language and still be permanently excluded from MT-first because its consequence forbids it.

Picking the First Safe Wins

The single most consequential decision in the entire roadmap is the choice of the first pilot, the crawl-phase cell, because it sets the program's credibility before anyone trusts it. Get the first win right and you earn the political capital and the proof to fund everything after it. Get it wrong, by choosing content too risky or a language too immature, and a single early Critical error can end the program before it has shown what it can do. So the first safe win deserves a deliberate selection method rather than a guess.

The intersection you are looking for is the cell where all three axes point to "safest": your most mature language, your highest-leverage content type, and your lowest risk tier, all at once. That is the corner of the cube where MT-first has the best chance of producing a clean quality record on content nobody can get hurt by. Concretely, the first safe win usually looks like this:

  • Language: your single most mature language pair, the one with the deepest clean TM, the most approved termbase, the strongest engine output, and the most trusted linguists. Not the highest-volume one unless it also happens to be the most mature.
  • Content type: high-volume, repetitive, high-TM-leverage content where the engine is strong, knowledge-base articles, UI strings, product descriptions, or similar. Content the engine handles well and that repeats, so leverage is high and the human gap to close is small.
  • Risk tier: the lowest, where a fluent error is recoverable and costs polish rather than a life or a lawsuit. Nothing regulated, nothing life-safety, nothing legal.
  • Scale: small and contained. A bounded batch with a real deadline and a real client expectation, large enough to be a genuine test and small enough that a problem is contained and a pause costs nothing catastrophic.

Notice what the first safe win is optimized for, because it is counterintuitive to a leadership team focused on savings. It is not optimized for maximum throughput or maximum cost reduction. It is optimized for the cleanest possible learning and the most defensible possible proof. The savings come later, in the walk phase, on the back of the trust the crawl phase earned. If you choose the first pilot for its savings, you choose it for its volume, and high-volume content is rarely the lowest-risk, highest-leverage, most-mature-language cell. The first win buys credibility, not money. The money is what credibility unlocks.

The Anti-Patterns That Kill First Wins

It is worth naming the specific bad first-win choices, because they are tempting and they are common. The first anti-pattern is leading with the highest-volume content regardless of risk, because that is where the savings spreadsheet points, even when that content is a regulated technical manual. The second is leading with the strategically important new-market language, the one leadership cares most about, even though it has the thinnest TM and the shallowest linguist pool, so you are learning the workflow on your worst-equipped language under the most executive scrutiny. The third is making the pilot too big to stop, scoping the first phase so large that pausing it would itself be a visible failure, which destroys the right to stop that the whole roadmap is built to preserve. The fourth is choosing content with no clean quality baseline, so that even if the pilot goes well you cannot prove it went well, because you have nothing to compare the error score against. Each anti-pattern trades the program's credibility for an illusion of early ambition, and each one is how a roadmap becomes a big-bang launch wearing a roadmap's name.

The Dependencies No Phase Survives Without

A roadmap that sequences only the content and languages and forgets the dependencies will stall in its second phase, because every phase rests on four foundations that must be in place before that phase begins, and building them takes time the roadmap has to schedule. The dependencies are assets, tooling, training, and governance, and the strategist's discipline is to treat each one as a deliverable with a deadline rather than an assumption.

Assets: TM, Termbase, and Style Guide

The linguistic assets are the foundation everything else stands on, because a grounded MT-first workflow is only as good as the TM, termbase, and style guide it is grounded on. A phase that introduces a new language depends on that language having a clean TM and an approved termbase, and if those do not exist, building them is a prerequisite task that must appear in the roadmap with its own timeline, not a thing you discover is missing when the phase starts. The same is true for content types: a phase that introduces technical documentation depends on a technical termbase, which may not exist if your only termbase was built for marketing. Asset readiness is the dependency that most often surprises a program, because the crawl phase succeeded on a mature language with good assets and nobody noticed the assets were the reason, so the walk phase into a new language fails and looks like an engine problem when it is actually an asset gap. The roadmap names asset-building as a scheduled dependency for every phase that touches a new language or content type.

Tooling: Grounding, Gates, and Records

The tooling dependency is the workflow infrastructure: the grounding pipeline that feeds the engine your assets, the severity-scored gate that scores output against the MQM/ISO 5060 typology, the terminology and locale enforcement controls, and the quality-record system that captures provenance per segment. A phase cannot run MT-first on content it cannot ground, cannot score, and cannot record, so the tooling for a given content type or language must be in place before that content enters a phase. Some tooling is built once and reused, the gate, the record system; some is per-language or per-content-type, the grounding sources, the locale rules. The roadmap distinguishes the two and schedules the per-phase tooling as a dependency of the phase it serves.

Training: The Linguists and PMs Who Run It

The human dependency is the most underestimated and the most damaging when it is missed, because a workflow nobody is trained to run produces exactly the failure the two resigned linguists in the opening represented. Every phase depends on the linguists, evaluators, and project managers who run it being trained for the content and the tier that phase introduces. A linguist who post-edited knowledge-base articles cleanly in the crawl phase is not automatically ready to run the full-PE gate on technical documentation in the walk phase; that is a different tier with different verification discipline, and it requires training the roadmap must schedule before the phase begins. Training is also where you earn the trust that retention depends on: the "AI drafts, you own the quality" contract is not a slogan, it is a thing you teach, and a phase that expands the workflow without training the people who run it is a phase that expands the resignation risk. The roadmap treats training as a gated prerequisite, not a thing that happens informally once the work is already flowing.

Governance: The Gates and the Stop Authority

The governance dependency is the decision structure that owns the gates between phases and holds the authority to stop. A roadmap with phases but no governance is a roadmap with no brakes, and a localization-AI program without brakes is the big-bang launch the roadmap was supposed to replace, just slower. Governance is the cross-functional group, quality, terminology, engineering, and PM at the table, that reviews each phase against its gate criteria and decides go, hold, or roll back. It is also the body that owns the permanent human-only exclusions and refuses the pressure to sweep high-liability content into MT-first for its volume. Governance must exist before the first phase ships, because the first phase needs a gate to pass through, and a gate with nobody empowered to fail a file or pause a phase is not a gate. The roadmap stands up governance as a phase-zero dependency, before crawl.

Every phase rests on four dependencies the roadmap must schedule as deliverables, not assume: assets (TM, termbase, style guide), tooling (grounding, gates, records), training (the people who run it), and governance (the gates and the authority to stop).

Milestones and Gates: How a Phase Earns the Next One

The mechanism that turns a sequence of phases into a defensible program is the gate between them: the explicit, measurable criteria a phase must meet before the next phase begins. A milestone is the achievement; a gate is the decision point where governance checks the milestone against criteria and rules go, hold, or roll back. Without gates, a roadmap is just a list of things you intend to do in order, and the order means nothing because nothing stops you from starting the next phase before the current one has proven itself. The gate is what gives the sequencing its teeth.

A phase gate in a localization-AI roadmap should be defined by quality, dependency, and human criteria together, never by a date alone. A date-only gate ("we move to phase two on the first of the quarter") is how programs advance into a phase they are not ready for, which is how the dosage error shipped. The criteria that should gate a phase advance include:

  • A clean quality record over a meaningful sample. The phase has produced MT-first output scored against ISO 5060 with zero unresolved Critical errors and Major and Minor counts within the agreed tier threshold, across enough volume to be a real signal rather than a lucky batch. One clean file is not a gate; a clean record across the phase's content is.
  • Terminology and locale conformance at the agreed level. The approved terms held across the phase's segments and the locale conventions were correct, demonstrated by the enforcement controls, not asserted.
  • The next phase's dependencies are in place. The assets, tooling, training, and governance for the phase you are about to enter are ready and verified, not in progress. You do not gate into a phase whose foundations are still being poured.
  • Linguist and reviewer trust holding. The people running the workflow are not leaving, the training landed, and the quality owners trust the gate, because a phase that advanced on a clean error score while the linguist pool quietly eroded has advanced into a hidden failure.
  • A defensible record the next phase can build on. The quality record from this phase is the proof you spend to fund the next, and it must be assembled and reviewed before the gate, not reconstructed later.

The gate has three possible outcomes, and a roadmap that only allows one of them is not a real gate. Go means the criteria are met and the next phase begins. Hold means a criterion is not yet met, the assets are not ready or the error score is not clean, so the current phase continues until it is, without advancing. Roll back means something went wrong enough that the current phase itself should contract or pause, an unexpected Critical error pattern, a regulated client's adverse audit, a spike in linguist attrition, and the program steps back to a safer posture while the cause is found. The existence of the roll-back outcome is what makes the right to stop real. A roadmap whose gates can only say "go" has no brakes, and a localization-AI program without brakes is exactly the thing that shipped a dosage error in three languages because nobody was empowered to stop it.

Leading and Lagging Signals at a Gate

A sophisticated gate watches both lagging and leading signals, because by the time a lagging signal goes bad the damage may already have shipped. The lagging signals are the obvious ones: the Critical error rate, the terminology conformance, the locale conformance, the rework rate, all measured on delivered output. The leading signals are the ones that predict trouble before it ships: a rising trend in low-QE segments reaching delivery, an increase in the time linguists spend on the high-consequence blocks (which can mean the engine is degrading or the content is getting harder), a drift in termbase coverage, early signs of linguist disengagement in the trust survey. A gate that watches only lagging signals is a gate that catches problems after they cost something. A gate that watches leading signals too can hold a phase before the lagging signal goes red. The roadmap defines both for every gate.

A Worked Multi-Quarter Roadmap

Now assemble everything into a concrete, multi-quarter roadmap for a realistic operation, so the principles become a plan you could adapt. Picture a mid-size operation localizing a software-plus-documentation product into a portfolio of languages, with a regulated component (the product ships in healthcare and financial-services contexts, so some content is high-liability), and a leadership team that has just funded a localization-AI program with the expectation of throughput gains and cost reduction. The operation handles five broad content types, UI strings, knowledge-base articles, marketing and brand content, technical documentation, and regulated medical and legal content, across a language portfolio with three mature languages (deep TM, strong assets, trusted linguists), four developing languages (moderate assets), and two new-market languages (thin assets, shallow linguist pools). Here is the roadmap, phase by phase, with the gate at each boundary.

Phase Zero (Pre-Quarter): Foundations and Governance

Before any content is machine-translated, phase zero stands up the dependencies the whole program rests on. The cross-functional governance group is formed, quality lead, terminology lead, a localization engineer, and a senior PM, with explicit authority to fail a file, hold a phase, and roll back. The MT-first tooling is built once and reused: the grounding pipeline, the severity-scored gate aligned to MQM/ISO 5060, the terminology and locale enforcement controls, and the quality-record system. The permanent human-only exclusion list is written and ratified: regulated medical content and legal contract content are named as never-MT-first, by policy, regardless of future volume pressure. And the quality baseline for the pilot content is established, so the crawl phase has something to score against. Phase zero ships nothing translated; it ships the foundation. The gate out of phase zero is simple and absolute: governance exists with stop authority, the tooling runs, the exclusion list is ratified, and the crawl-phase assets are confirmed clean. No crawl begins until that gate is green.

Phase One (Quarter 1): Crawl, the First Safe Win

The crawl phase is the first safe win, chosen by the three-axis method: the single most mature language, the highest-leverage lowest-risk content type, a small contained scale. Concretely, knowledge-base articles, the most mature language pair, the lowest risk tier, full human gate on every segment, run as a bounded batch with a real client deadline. The economics are deliberately not the point: the human reviews everything, so the throughput lift is modest, and that is correct, because crawl buys proof and knowledge, not savings. What crawl produces is a clean quality record on recoverable content: ISO 5060 scoring, zero unresolved Criticals, demonstrated terminology and locale conformance, and the operational learning of running the grounding, the gate, and the record on real work. The linguists running it are trained specifically for this content and tier and are surveyed for trust. The gate out of crawl: a clean record across the batch (not one lucky file), terminology and locale conformance demonstrated, the walk phase's assets and training confirmed ready, and linguist trust holding. If the error score is not clean, the gate says hold, and crawl continues until it is. The program does not advance on a date; it advances on proof.

Phase Two (Quarter 2): Walk by Adding Content, One Variable at a Time

With crawl proven, the walk phase widens on a single axis. The discipline here is one variable at a time, so the choice is to hold the language constant (still the most mature language) and add a content type: UI strings, the other high-leverage low-risk type. The language is unchanged, the risk tier is still low, only the content type is new, so when something regresses the cause is obvious. This is also where the economics begin to appear in earnest: with two proven content types in a mature language, the human gate can move from every-segment to risk-routed, capturing the throughput lift on content that has earned it, while the gate still fails any Critical. The dependency the roadmap scheduled for this phase is the UI-string tooling: placeholder and length-budget enforcement, the locale rules for UI, which are different from documentation. The gate out of phase two: clean record across both content types, the UI-specific enforcement (placeholders, length) demonstrated, and the dependencies for phase three (a second language's assets, or a higher content tier) confirmed. Hold if the UI placeholder errors are not driven to zero, because a broken string is a defect even when the prose is fluent.

Phase Three (Quarter 3): Walk by Adding a Language, Assets First

Now the roadmap changes a different single variable: it holds the proven content types constant and adds a second mature language. The critical move here is that the asset dependency was scheduled a phase ahead: building the second language's clean TM and approved termbase began during phase two, because asset-building takes time and a phase that discovers its assets are missing on day one is a phase that stalls. The new language's linguists are trained during this phase's ramp, and the language's engine quality was assessed in advance (a developing language with weaker engine output may need full-PE where the mature language used risk-routed PE, which the roadmap accounts for). The gate out of phase three: clean record on the new language across the proven content types, asset coverage confirmed adequate for the new language, the new linguist pool trusted and retained, and the higher-risk-tier dependencies for phase four confirmed. This is also the phase most likely to reveal that a language assumed mature is actually developing, in which case the gate says hold and the assets are deepened before advancing, a contained correction rather than a shipped failure.

Phase Four (Quarter 4): Walk Up the Risk Axis, Carefully

Only now, with the workflow proven across multiple content types and languages and a long clean track record, does the roadmap move up the risk axis: it adds technical documentation, a higher-consequence content type than knowledge-base articles, under full post-editing rather than risk-routed, with the slowest, most deliberate verification on the high-consequence elements. The language stays mature, the content type is the single new variable, and the risk tier rises by one deliberate step, not to the top. This phase depends on a technical termbase (scheduled and built in advance), on linguists trained for full-PE on technical content (a different discipline from the lighter tiers), and on a gate tuned for the higher consequence. The regulated medical and legal content is still excluded, by the phase-zero policy, and governance explicitly reaffirms the exclusion against any pressure to include it for its volume. The gate out of phase four: clean record on technical documentation under full PE, the full-PE discipline demonstrated and trusted, and a governance review of whether the program is ready to enter the run phase, where MT-first becomes the standing posture across qualified content and languages.

Phase Five (Quarter 5 and Beyond): Run, the Standing Posture

The run phase is not a finish line; it is a steady state with continuous governance. MT-first is now the default for the qualified content types and languages, the gates are routine, the quality record is automatic, and the human effort has migrated up to the judgment the engine cannot supply. Run expands deliberately and continuously: new languages enter through a repeatable mini-version of the crawl-walk sequence (assets, training, a contained pilot, a gate), new content types enter the same way, and the developing and new-market languages are brought up as their assets are built. The permanent exclusions hold: regulated medical and legal content remains human-only, and governance owns the line. Run also institutionalizes the leading and lagging signals into a standing dashboard, so a degradation in engine quality, a drift in termbase coverage, or a rise in linguist attrition triggers a contained hold before it becomes a shipped failure. The run phase is the program the bad board slide promised, but reached by sequencing rather than by switching everything on at once, which is the only way it could have been reached at all.

Reading the Worked Roadmap: What Each Phase Actually Bought

Step back and read the whole arc, because the logic is the lesson. Phase zero bought the brakes and the foundation. Crawl bought proof and knowledge on content nobody could be hurt by. Each walk phase changed exactly one variable, content, then language, then risk, so that every regression had an obvious cause and a bounded blast radius, and each one produced a defensible record that funded the next. Run is the standing posture the sequence earned. At no point did the program switch everything on at once; at every point it could have stopped, contained, and learned. The contrast with the opening's board slide is total: that slide promised all 19 languages and all content by end of quarter, and it delivered a shipped dosage error, two resignations, and a machine-translated indemnity clause. The roadmap delivers the same destination, MT-first as a standing posture, but reaches it with the credibility intact, the regulated content protected, the linguists retained, and a quality record at every gate that an auditor or a client could reconstruct. The destination is not the strategy. The sequence is the strategy.

The worked roadmap changes exactly one variable per phase, content then language then risk, so every regression has an obvious cause and a bounded blast radius, and every phase produces a defensible record that funds the next. The destination is not the strategy; the sequence is.

Defending the Roadmap to Leadership

The final skill of the strategist is selling the roadmap to a leadership team that funded the big-bang slide and expects speed, because a phased plan reads, to an impatient executive, like a slower plan, and the strategist has to reframe that. The reframe is that the roadmap is not slower; it is the only version that arrives at all. The big-bang launch is faster only on the slide. In reality it ships errors that trigger client losses, regulatory exposure, and rework that costs more time than the phasing ever would, plus the resignations that drain the linguist pool the whole program depends on. The phased roadmap front-loads a small amount of deliberate slowness in the crawl and early walk phases and saves the catastrophic slowness of a shipped Critical error in a regulated client's content. You are not trading speed for safety. You are trading a small, predictable, contained cost early for the avoidance of a large, unpredictable, program-ending cost later.

Frame the roadmap to leadership on the dual axis the whole program lives on: throughput and quality risk, together. The big-bang slide shows only throughput, which is why it is dangerous. The roadmap shows throughput rising phase by phase as content earns MT-first, and quality risk held flat or falling because every phase passes a gate and every high-liability exclusion holds. That dual-axis story is the one a CFO and a regulated client can both accept, because it does not ask them to trust that nothing will go wrong; it shows them the gate that stops it when it does. The sentence the strategist can say that the big-bang advocate cannot is this: here is the throughput we will capture, here is the sequence that captures it, here is the gate at every phase boundary, here is the content the machine will never touch, and here is the point at which we can stop without losing the program. That sentence is the roadmap's credential, and it is the difference between a localization-AI program that compounds quarter over quarter and one that ends in a quarterly review with a dosage error on the screen.

Key Takeaways

  • A big-bang localization-AI launch fails not because the technology is immature but because switching everything on at once removes the three things a strategist needs: a bounded blast radius when something breaks, an isolated signal about what broke, and a defensible place to stop. The roadmap exists to preserve all three, so sequencing is the strategy, not a project-management detail.
  • Every credible roadmap follows a crawl, walk, run spine: crawl proves the workflow on the safest content in the most mature language with a human gate on every segment (buying proof, not savings); walk widens one variable at a time and captures the throughput lift; run makes MT-first the standing posture across qualified content while permanent human-only exclusions hold.
  • You sequence along three axes at once: language maturity (TM depth, termbase, engine quality, linguist pool), content type (lead with high-leverage low-consequence content like knowledge-base articles, not high-volume regulated content), and content risk (consequence of a fluent error). Risk has veto power: high-volume mature-language content can still be permanently excluded because its consequence forbids MT-first.
  • The first safe win is the cell where all three axes point to safest (most mature language, highest-leverage content, lowest risk, small contained scale) and is optimized for the cleanest learning and most defensible proof, not for maximum savings. Anti-patterns that kill it: leading with highest-volume content regardless of risk, learning on the strategically important new-market language with thin assets, scoping the pilot too big to stop.
  • Every phase rests on four dependencies the roadmap must schedule as deliverables, not assume: assets (clean TM, approved termbase, style guide), tooling (grounding, the MQM/ISO 5060 gate, enforcement controls, the quality-record system), training (the linguists and PMs who run that phase's tier), and governance (the cross-functional group with authority to fail a file, hold a phase, and roll back).
  • Gates between phases are defined by quality, dependency, and human criteria together, never a date alone: a clean ISO 5060 record across a meaningful sample with zero unresolved Criticals, demonstrated terminology and locale conformance, the next phase's dependencies in place, and linguist trust holding. A gate has three outcomes (go, hold, roll back), and the roll-back outcome is what makes the right to stop real.
  • The worked multi-quarter roadmap changes exactly one variable per phase (phase zero foundations and governance, crawl on the first safe win, walk by adding a content type, then a language with assets built a phase ahead, then a deliberate step up the risk axis, then run as the standing posture), so every regression has an obvious cause and every phase produces a defensible record that funds the next.
  • Defend the roadmap to leadership on the dual axis of throughput and quality risk together: it is not slower than the big-bang launch, it is the only version that arrives, trading a small predictable contained cost early for the avoidance of a large program-ending cost later. The credential sentence is: here is the throughput, the sequence that captures it, the gate at every boundary, the content the machine will never touch, and the point at which we can stop without losing the program.