Assessing Your Localization-AI Readiness
The board approved the number on a Tuesday: forty percent off the localization budget inside eighteen months, delivered by going "MT-first" the way two competitors already had, and the deck that won the approval had exactly one slide of detail. You are the localization director who has to make that slide real, and the first thing you do is the thing nobody on the board did, which is open the actual cupboards. You pull the translation memories and find that of your nine target languages, three have memories that are large, clean, and segment-aligned, two have memories that are full of duplicated and contradictory entries nobody has cleaned since 2019, and four have almost nothing because that work used to go to a freelancer who kept the memory on her own laptop. You pull the termbases and discover one language has a maintained, approved glossary of four thousand terms and the rest have a spreadsheet last touched by someone who left. You ask how quality is measured today and the honest answer is "the reviewer reads it and if it looks fine it ships." You ask who would run the error scoring and realize nobody in the building has ever assigned a Critical against a typology in their life. The forty-percent number assumed an organization that does not exist yet. The job in front of you is not to launch the MT-first program. It is to find out, honestly and on paper, whether the organization is ready to launch it, and where exactly it will break if you launch it anyway. That honest finding is a readiness assessment, and this lesson is how you build one that a board, an auditor, and your own anxious linguists will all believe.
Why Readiness Is the Decision Before the Decision
Before anything else, fix the vocabulary, because a readiness assessment lives or dies on whether everyone in the room means the same thing by the same word. A translation memory (TM) is the organization's database of previously approved source-and-target sentence pairs, the asset that lets a new project reuse language a human already signed off on. A termbase is the controlled glossary of approved terms, the place where "the device is called the Lumera 400 and never 'the unit'" is written down as a rule. Machine translation (MT) is any system that renders text from one language to another with no human writing the words, and the production-grade flavor is neural MT; a large language model (LLM) is a general-purpose text predictor that translates as a side effect and tends to be more fluent and more confidently wrong than classic neural MT. Machine-translation post-editing (MTPE) is the workflow where a human edits machine output instead of translating from a blank target, and "MT-first" means the pipeline pre-populates every segment with machine output before a human opens the file. MQM (Multidimensional Quality Metrics) is the analytic error typology that scores translation errors by category and severity; ISO 5060 (ISO 5060:2024) is the standard that formalizes the MQM-aligned Critical, Major, and Minor scoring of translation output; ISO 18587 is the post-editing standard, currently being revised to cover AI and LLM output and to require the post-editor to hold full professional-translator competence. An LSP is a language-service provider, the vendor or in-house team that actually performs the translation and post-editing work. Hold those, and the rest of this lesson is legible.
Now the strategic point. Every failed MT-first program fails the same way, and it is never the way the post-mortem first claims. The post-mortem says "the engine was not good enough" or "the linguists resisted," but underneath both of those is a readiness gap that was present on day one and that nobody measured before committing. The engine was not the problem; the engine was asked to ground itself on a termbase that did not exist, in a language whose memory was contaminated, evaluated by a process that could not tell a Critical error from a stylistic preference, run by people who had never been trained to own quality inside a machine pipeline, with no governance to say which content the machine was forbidden to touch. The program did not fail at launch. It failed at the assessment that was never done.
An MT-first program does not fail at launch. It fails at the readiness assessment that was skipped, and every failure mode it later exhibits was a measurable gap on day one.
A readiness assessment is the deliberate act of measuring, before you commit budget and reputation, whether the organization can actually run an MT-first program without shipping the silent critical error that ends a client relationship. It is not a vibe and it is not optimism. It is a scored audit across the dimensions that determine whether the program will work, producing a single honest picture that says: here is where we are strong enough to go now, here is where we are weak enough that going now would be reckless, and here is the specific, sequenced work that closes the gap between the two. The decision-maker who runs this assessment first is buying something the board cannot see but desperately needs, which is the difference between a program that captures the speed the deck promised and a program that captures the liability the deck never mentioned.
Readiness Is Not Engine Quality (the Most Common Mistake)
The single most common error a localization leader makes here is to collapse "are we ready for MT-first" into "is the engine good," and then to spend the entire assessment running engine bake-offs while the real gaps sit unexamined. Engine quality is one dimension of readiness, and not the dimension that usually decides the outcome. The engine is a commodity you can swap; an excellent engine grounded on a contaminated memory and evaluated by an untrained reviewer produces fluent, confident, scored-nowhere output that is exactly as dangerous as a mediocre engine's. The dimensions that actually determine readiness are mostly about your own house: the state of your linguistic assets, the discipline of your quality process, the skills of your people, and the governance that decides what the machine may touch. The engine is the easiest of these to fix and the one leaders over-index on, because it is the only one that comes with a vendor demo. Resist that gravity. A readiness assessment that is mostly an engine bake-off is a readiness assessment that measured the easy thing and skipped the hard ones.
The Six Dimensions of Readiness
A readiness assessment is structured, not impressionistic, and the structure is a set of dimensions you score independently so that a strength in one cannot quietly hide a fatal weakness in another. Six dimensions cover the surface that determines whether an MT-first program will work in a localization operation. Each is a question with a measurable answer, and each maps to a specific way the program breaks if the answer is "not ready."
- Linguistic assets. Do you have translation memories and termbases that are large enough, clean enough, and current enough to ground the engine and seed the post-editing, per language and per content domain? This is the dimension that decides whether the engine answers from your approved record or guesses from the open web.
- Data and content readiness. Is your source content structured, segmentable, and clean enough to feed a pipeline, and do you know your volumes and content types well enough to route them? This is the dimension that decides whether the pipeline can even ingest your work without choking on broken files and unsegmentable formats.
- Engine and tooling fit. Do your candidate MT or LLM engines perform on your actual domain and language pairs, and does your CAT/TMS stack support grounding, severity scoring, and a quality record? This is the dimension leaders over-weight, and it still matters: the tool must be able to do the controls, not just the translating.
- Quality process maturity. Do you measure quality against an error typology with severities, or do you "read it and ship it"? This is the dimension that decides whether you can tell a Critical error from a stylistic preference, which is the entire safety case for MT-first.
- People and skills. Do your linguists, reviewers, PMs, and engineers have the skills to own quality inside a machine pipeline, and the willingness to do it? This is the dimension that decides whether the program has anyone who can actually run the controls the standards require.
- Governance and risk. Do you have a way to classify content by risk, decide what the machine is forbidden to touch, and respond when a Critical ships? This is the dimension that decides whether the cheap workflow ever lands on content that can kill or sue.
The discipline of scoring these separately is the whole point. An organization can have a world-class engine and a maintained termbase in its flagship language and still be catastrophically unready, because its quality process cannot detect a Critical error and its governance has no concept of MT-forbidden content. A single blended "readiness score" would average those together and let the strengths outvote the fatal weakness, which is exactly the averaging mistake that ships the lethal segment. You score each dimension on its own, and then you read the profile as a shape, not a sum, because in readiness as in error scoring, one fatal gap is not redeemed by four strong scores around it.
Score the six dimensions separately and read the profile as a shape, not a sum. One fatal gap, an untrained evaluator or no MT-forbidden governance, is not redeemed by a strong engine and a clean memory around it.
How to Score Each Dimension
A readiness score that means anything has to be defensible, which means each dimension needs a rubric: a small set of levels with concrete, observable criteria, so that two assessors looking at the same evidence land on the same level. The temptation is to score on a feeling ("our memories are pretty good"), and the entire value of the assessment evaporates the moment a number rests on a feeling. Use a four-level scale per dimension, because four levels force a real decision and avoid the lukewarm middle of a five-point scale: 1 Absent (the capability does not exist), 2 Emerging (it exists in patches, undocumented, inconsistent), 3 Established (it exists, is documented, and works in most cases), and 4 Optimized (it exists, is measured, maintained, and improving). The bar for green-lighting an MT-first program in a given language and content type is generally a 3 on every dimension and a hard 3-or-better on the two that gate safety, quality process and governance. Let us make each dimension concrete.
Scoring Linguistic Assets
Do not score "do we have TMs and termbases" as a yes or no, because the answer is almost always a misleading yes. Score them per language and per content domain, against evidence you can actually pull. For the translation memory, the observable criteria are size (how many approved segments, and is that enough to leverage usefully), cleanliness (the rate of duplicates, contradictory pairs where the same source maps to different approved targets, and untrusted entries of unknown provenance), alignment (are the segments properly aligned, or is the memory full of mismatched pairs that will poison leverage), and currency (when was it last maintained, and does it reflect current approved language or a product two versions ago). A language at Optimized has a large, deduplicated, aligned, recently maintained memory with known provenance. A language at Absent has the freelancer's laptop. Most real organizations are a patchwork, which is precisely why you score per language: the flagship language might be a 4 and four others a 1, and a blended "we have memories" hides the four that will sink the program.
For the termbase, the criteria are coverage (do the approved terms exist for the domains you translate), approval status (are the terms actually client-approved, or just someone's draft), structure (is it a real termbase with source, target, domain, and usage notes, or a flat spreadsheet), and maintenance (is there an owner and a process, or is it abandoned). The reason termbase readiness gates MT-first so hard is that the termbase is what you ground the engine on to stop it drifting off the approved term; a missing or unapproved termbase means the engine will confidently substitute synonyms across thousands of segments and you will have no automated way to catch it. A spreadsheet last touched by someone who left is an Emerging termbase at best, and it is a direct ceiling on how much of the post-editing the engine can offload.
Scoring Data and Content Readiness
This is the dimension leaders forget, and it is the one that breaks pilots in the first week. The criteria are whether your source content is structured and segmentable (clean XML or properly tagged files versus PDFs and flattened layouts the pipeline cannot parse), whether your content is clean at source (a badly written, inconsistent source produces worse MT and harder post-editing, and "garbage in, fluent garbage out" is a real failure mode), and whether you actually know your volumes and content types well enough to route them. An organization at Established can tell you, per content type, the annual word volume, the source format, and the risk tier. An organization at Absent discovers in week two that forty percent of its incoming content arrives as scanned PDFs that no engine can ingest cleanly. Score this honestly, because a strong engine and a clean memory are worthless against content the pipeline physically cannot read.
Scoring Engine and Tooling Fit
Score engine fit on your domain and your language pairs, never on a vendor's published benchmark, because a leaderboard score on generic news text tells you nothing about how the engine handles your medical dosing language or your legal boilerplate in your specific pairs. The criteria are domain performance (does the engine produce usefully accurate first drafts on your real content, measured by the post-editing effort it actually requires), terminology adherence (does it respect an injected termbase or drift off it), and critical-error behavior (how often does it produce the fluent, confident mistranslation on high-consequence segments). Note that the last of these is the criterion vendors never report and the one that matters most, and you can only measure it by running your own content through and scoring the output against a typology, which is why this dimension depends on the quality-process dimension being mature enough to do the scoring.
Tooling fit is the parallel question for your CAT and TMS stack: can it actually do the controls an MT-first program requires, or only the translating. The criteria are grounding support (can it feed the engine your TM, termbase, and style guide), severity-scoring support (can it capture an MQM/ISO 5060 error evaluation as a workflow step, or only an edit-distance number), and quality-record capability (can it produce the segment-level provenance an auditor needs). A tool that pre-translates beautifully but cannot capture a severity-scored quality record is a tool that will leave you unable to prove conformance, which means it caps your governance readiness no matter how good the translations look.
Scoring Quality-Process Maturity
This is the safety-gating dimension, and you score it with brutal honesty because it is the one organizations most want to flatter themselves about. The question is not "do we check quality." Everyone checks quality. The question is whether quality is measured against an analytic error typology with severities, reproducibly, by trained evaluators. The levels are stark. An organization at Absent has "the reviewer reads it and if it looks fine it ships," which is not a quality process at all; it is a vibe with a job title, and it is structurally incapable of reliably catching the fluent error that reads perfectly. An organization at Emerging has a checklist or an edit-distance metric but no severity-scored typology. An organization at Established scores output against MQM/ISO 5060 categories and severities, with the one-Critical-fails rule, by people who can do it consistently. An organization at Optimized does that and measures inter-evaluator agreement so the scores are reproducible across people.
The reason this dimension hard-gates the whole program is the cardinal logic of the discipline: MT output is fluent first and accurate second, so the entire safety case for MT-first is the ability to detect the silent critical error, and that ability is exactly what a severity-scored evaluation process provides and a "looks fine to me" review does not. An organization scoring 1 or 2 here cannot run MT-first safely on anything with consequences attached, full stop, because it has no instrument capable of seeing the failure mode the program creates. You can launch on the lowest-stakes content while you build this up, but you cannot launch on anything that can kill or sue until this dimension reaches 3.
Quality-process maturity hard-gates the program, because the entire safety case for MT-first is the ability to detect the silent critical error, and a "looks fine to me" review is structurally incapable of seeing it.
Scoring People and Skills
Score the people dimension on both capability and willingness, because a program needs linguists who can own quality inside a machine pipeline and who will. The capability criteria: can your linguists post-edit against the source rather than the smooth target, can your reviewers assign a defensible severity against a typology, can your PMs quote a quality tier instead of a race-to-the-bottom price, can your engineers handle grounding and the quality record. The revised ISO 18587 raises this bar deliberately by requiring the post-editor to hold full professional-translator competence, which means MT-first does not let you staff the work with cheaper, less-skilled people; it demands the same competence, redirected from translating to owning. The willingness criteria are just as real: a vendor pool quietly leaving because the work feels like cleaning up after a machine for less money is a people-readiness failure even if every linguist is technically capable. An organization at Absent has nobody trained in severity scoring and a demoralized pool; at Optimized it has trained quality owners with a career path that moves them up the value chain, not out of it.
Scoring Governance and Risk
The final dimension is the one that decides whether all the others are ever applied correctly. Governance readiness is your organization's ability to classify content by consequence, decide and enforce what the machine is forbidden to touch, and respond when a Critical error ships anyway. The criteria are risk-tiering (is there a defined, applied scheme that sorts a marketing string from a drug label from an indemnity clause before any post-editing begins), MT-forbidden rules (is there an enforced rule that routes regulated, life-safety, and high-liability legal content to full human translation or full post-editing, never to the cheap workflow), and incident response (is there a defined containment-and-root-cause process for a shipped Critical, or does the organization improvise). An organization at Absent treats every file as one homogeneous job and has no concept of MT-forbidden content, which is the governance posture that eventually lands the cheap workflow on a contraindication. At Established, risk-tiering is defined and applied at intake and MT-forbidden content is routed by rule, not by the judgment of whoever happens to open the file.
The Gaps That Block MT-First (and the Ones That Only Slow It)
Not all readiness gaps are equal, and the most important judgment in the entire assessment is distinguishing a blocking gap from a slowing one. A blocking gap is a weakness that makes MT-first unsafe or undeliverable on the content in question, full stop, until it is closed; launching anyway is not a slower program but a reckless one. A slowing gap is a weakness that costs you speed, leverage, or money but does not by itself put a Critical error on a ship path; you can launch with it present and close it as you go. Confusing the two in either direction is expensive: treat a slowing gap as blocking and you delay a program that was ready; treat a blocking gap as slowing and you ship the lethal segment.
The Blocking Gaps
Three gaps are blocking on any content with consequences attached, and the assessment should flag them in red the moment they appear:
- No severity-scored quality process (quality dimension at 1 or 2). This is the first blocking gap, because without the ability to detect and score the silent critical error, you have no instrument that can see the failure mode MT-first creates. You cannot launch on consequential content with a "looks fine to me" review, no matter how good the engine. The program is not slowed by this gap; it is unsafe under it.
- No MT-forbidden governance (governance dimension at 1). The second blocking gap is the absence of a rule and an enforcement mechanism that keeps the cheap workflow off regulated, life-safety, and high-liability legal content. Without it, risk-tiering is a suggestion, and the program will eventually route a drug label or an indemnity clause through light post-editing because nobody was empowered to stop it. This gap blocks the high-liability content specifically; you can launch on low-stakes content while closing it.
- Contaminated linguistic assets in a target language (assets dimension at 1 or 2 for that language). The third blocking gap is language-specific: a memory full of contradictory, untrusted pairs does not merely fail to help; it actively poisons the pipeline by grounding the engine on wrong approved language and propagating errors as if they were leverage. A dirty memory is worse than no memory, because it launders bad language through the authority of "previously approved." This blocks MT-first in that language until the memory is cleaned, though it does not block the languages whose assets are sound.
The Slowing Gaps
Other gaps cost you but do not block. A thin termbase slows the program by leaving more drift for the post-editor to catch by hand, but you can launch with heightened terminology review and build the termbase in parallel. A messy source-content situation slows you by requiring cleanup and format conversion before ingestion, but it does not put a Critical on a ship path. A tooling stack that grounds but cannot yet capture a full quality record slows your conformance story but can be supplemented manually while you upgrade. A people pool that is capable but undertrained in severity scoring slows you while you train them, and training is a tractable, scheduled fix. The discipline is to label each gap blocking or slowing in the assessment itself, because that label is what converts a list of weaknesses into a sequencing decision: blocking gaps must close before launch on the affected content; slowing gaps become the parallel workstream that runs alongside a phased rollout.
A dirty translation memory is worse than no memory, because it launders bad language through the authority of "previously approved" and grounds the engine on errors as if they were leverage.
A Worked Readiness Scorecard for One Organization
Abstraction is where readiness assessments go to die, so let us score a real organization end to end, the one from the opening: a mid-size company localizing a medical-device product line into nine languages, handed a board mandate to go MT-first and cut forty percent. We will assess it across the six dimensions, score each on the 1-to-4 scale, flag the blocking and slowing gaps, and arrive at a finding the director can take back to the board: not "yes" or "no," but "here, now, on this content, and here is the work that earns the rest."
The Evidence Gathered
The assessment is only as good as the evidence under it, so the director spent two weeks pulling the actual artifacts rather than asking people how they felt. The translation memories: three languages (German, French, Spanish) have large, deduplicated, aligned, currently maintained memories with known provenance; two (Italian, Dutch) have memories riddled with duplicates and contradictory pairs from years of uncontrolled imports; four (Polish, Czech, Swedish, Finnish) have almost nothing, the freelancer's-laptop situation. The termbases: German has a maintained, approved four-thousand-term glossary; the other eight have an abandoned spreadsheet. Source content: the structured product documentation is clean XML, but roughly a third of incoming content arrives as flattened PDFs from regulatory affairs. The engine: a candidate neural MT engine was run on real German and French content and produced usefully accurate drafts, but its critical-error behavior on dosing segments has not been scored because there is no scoring process to score it with. The tooling: the TMS grounds on TM and termbase but captures only edit-distance, not a severity-scored quality record. Quality process: "the reviewer reads it and if it looks fine it ships," with no typology, no severities, no trained evaluators. People: the in-house linguists are competent translators but none has run an MQM/ISO 5060 evaluation, and the contract pool is restive about MTPE rates. Governance: every file is treated as one job; there is no risk-tiering and no MT-forbidden rule, on a medical-device product line.
The Dimension Scores
Now score each dimension against the rubric, reading the evidence honestly rather than hopefully:
- Linguistic assets: 2 (Emerging), but read per language. German is a 4, French and Spanish are 3 to 4, Italian and Dutch are a 2 (contaminated memory, blocking for those languages), and the four CEE/Nordic languages are a 1 (Absent). The blended 2 is itself a warning: the program is ready on three languages and not on six.
- Data and content readiness: 2 (Emerging). The structured documentation is clean, but a third of content arriving as flattened PDFs is an ingestion problem that will break the pipeline on real intake. Slowing, not blocking, but real.
- Engine and tooling fit: 2 (Emerging). Engine performance on the strong languages is promising, but its critical-error behavior is unmeasured (gated by the quality process being absent), and the tooling cannot capture a quality record. Slowing on the record gap; the unmeasured critical-error behavior is a known unknown the pilot must close.
- Quality-process maturity: 1 (Absent). "Looks fine and it ships" is no process. This is a blocking gap on all consequential content, and on a medical-device line almost everything is consequential.
- People and skills: 2 (Emerging). Capable translators, zero severity-scoring experience, a restive pool. Slowing (training is tractable) but it cannot be hand-waved.
- Governance and risk: 1 (Absent). No risk-tiering and no MT-forbidden rule on a medical-device product line is the most dangerous score on the card. Blocking.
Reading the Shape, Not the Sum
If you averaged these six scores you would get something near 1.7 and a temptation to call the organization "about a third ready" and launch a third of the program. That averaging is exactly the mistake the assessment exists to prevent. Read the shape instead. Two dimensions, quality process and governance, are at 1, and both are hard safety gates on a medical-device line where a flipped dosage or a swapped contraindication can injure a patient and trigger a recall. Those two 1s are not outvoted by the strong German memory; they are the finding. The honest reading is: this organization is not ready to run MT-first on its consequential content, because it has no instrument to detect the silent critical error and no governance to keep the cheap workflow off the content that can kill. It is, however, ready to begin building toward readiness, and on the lowest-stakes content in its strongest languages it could even pilot now under heavy supervision.
The Finding and the Sequenced Close
The deliverable the director takes to the board is not a refusal and not a rubber stamp. It is a finding with a sequence attached, and the sequence is dictated by which gaps are blocking and which are slowing:
- Close the blocking gaps first, before any consequential-content launch. Stand up a severity-scored quality process (the quality dimension to 3): adopt an MQM/ISO 5060 typology, train evaluators, institute the one-Critical-fails gate. Stand up governance (the governance dimension to 3): define and apply risk-tiering at intake, write and enforce the MT-forbidden rule that routes dosing, contraindication, and indemnity content to full human translation or full post-editing. Clean the Italian and Dutch memories or freeze MT-first in those languages until they are clean. None of the forty-percent savings is real until these close, because all of it assumed a safety instrument that does not exist.
- Launch the phased pilot where readiness already exists. Begin MT-first on the lowest-consequence content (marketing and UI strings, not regulated documentation) in German, French, and Spanish, where the memories are clean and the termbase (German) is strong, under the new quality process as it comes online. This captures early, provable wins on content where a fluent error is recoverable, and it gives the people a real environment to build severity-scoring muscle.
- Run the slowing gaps as parallel workstreams. Build the eight missing termbases starting with the highest-volume domains; fix the PDF ingestion path with regulatory affairs; upgrade the tooling to capture a quality record; train the linguist and reviewer pool on severity scoring and give them the quality-owner career path that addresses the willingness gap. These run alongside the pilot, not before it.
- Sequence the remaining languages by asset readiness. Bring the four CEE/Nordic languages onto MT-first only after their memories and termbases reach 3, which is a build-the-asset project, not a switch-it-on decision. The board's "all nine, now" becomes "three now on low-stakes content, three more as assets and process mature, the riskiest content last and only behind full human review."
That finding is worth more to the organization than a yes would have been, because a yes would have launched the forty-percent number into a building with no instrument to catch the error that ends a medical-device client relationship. The assessment converted a reckless mandate into a defensible program: phased by readiness, gated by safety, and honest about the gap between the slide and the cupboards. The director did not say no to the board. The director said: here is exactly what "ready" costs, here is what we can capture now, and here is the order in which the rest becomes real, and that sentence is the one a leader who skipped the assessment can never say.
A readiness assessment converts a reckless mandate into a defensible program: phased by what is ready now, gated by the safety dimensions, and honest about the distance between the strategy slide and the actual assets in the cupboards.
Turning the Assessment Into a Living Instrument
The last move that separates a strategist from someone who ran a one-time audit is recognizing that readiness is not a state you reach and check off; it is a posture you maintain. The scorecard is a snapshot, and snapshots decay. A memory that was clean at assessment drifts as uncontrolled imports resume; a termbase that was maintained falls behind a product release; a quality process that reached 3 regresses when the trained evaluators leave and are replaced by people who never calibrated; an engine that fit the domain degrades when the content mix shifts. The strategist treats the readiness assessment as an instrument to re-run, tied to the events that change the underlying reality: a new language, a new content type, a new engine, a new regulatory regime, a vendor change, a team turnover. Each of those is a trigger to re-score the affected dimensions, because each can move a green dimension back to red without anyone noticing until a Critical ships.
There is also a quiet strategic benefit to running the assessment as a repeatable instrument rather than a one-off: it becomes the shared language that lets quality, terminology, engineering, and PM functions talk about readiness in the same terms, and it becomes the evidence base for every subsequent decision in the L4 program, the roadmap, the prioritization, the business case, the engine evaluation, the governance program. The roadmap you build next sequences exactly the gaps this assessment found. The business case quantifies exactly the savings these readiness levels make real and the savings the gaps make fictional. The engine evaluation runs precisely the domain-and-critical-error tests the engine-fit dimension said were unmeasured. The assessment is not a gate you pass and forget; it is the foundation the rest of the strategy is built on, and a strategist who keeps it current keeps the whole program honest.
So the readiness assessment is, in the end, the discipline of refusing to let a strategy slide outrun the organization that has to deliver it. It is the leader opening the cupboards before promising the meal. The forty-percent number on the board's deck was not wrong; it was unearned, and the assessment is how you earn it: by measuring the six dimensions honestly, scoring each against a defensible rubric, separating the gaps that block from the gaps that only slow, and producing a sequenced finding that captures the real speed where readiness exists and refuses the false speed where it does not. The organization that runs this assessment first ships an MT-first program that works. The organization that skips it ships the liability the deck never mentioned and writes the post-mortem that blames the engine for a gap that was measurable on day one.
Key Takeaways
- A readiness assessment is the decision before the MT-first decision: every failed MT-first program failed at an assessment that was skipped, and every failure mode it later showed was a measurable gap on day one, not an engine problem that appeared at launch.
- Readiness is not engine quality. The engine is a swappable commodity and the dimension leaders over-weight; the dimensions that actually decide the outcome are mostly about your own house, the assets, the quality process, the people, and the governance, and an excellent engine on a contaminated memory and an untrained reviewer is exactly as dangerous as a mediocre one.
- Score six dimensions independently, linguistic assets, data and content, engine and tooling fit, quality-process maturity, people and skills, and governance and risk, on a four-level scale (Absent, Emerging, Established, Optimized), and read the profile as a shape, not a sum, because one fatal gap is not redeemed by strong scores around it.
- Score linguistic assets per language and per content domain, not as a yes or no: size, cleanliness, alignment, and currency for TMs; coverage, approval status, structure, and maintenance for termbases. A blended "we have memories" hides the languages that will sink the program.
- Quality-process maturity and governance hard-gate the program. Without a severity-scored evaluation process you have no instrument to detect the silent critical error that is MT-first's defining failure mode; without an enforced MT-forbidden rule the cheap workflow eventually lands on a drug label or an indemnity clause. An organization scoring 1 or 2 on either cannot launch MT-first safely on consequential content.
- Distinguish blocking gaps from slowing gaps explicitly in the assessment. Blocking: no severity-scored quality process, no MT-forbidden governance, and a contaminated memory in a target language (a dirty memory is worse than none, because it launders bad language through "previously approved"). Slowing: thin termbases, messy source content, an incomplete quality-record tool, an undertrained but capable team, all of which run as parallel workstreams alongside a phased launch.
- The worked scorecard, a medical-device shop mandated to go MT-first and cut forty percent, scored 1 on quality process and 1 on governance: not ready on consequential content despite a strong German memory, because the two safety gates were at the floor and cannot be outvoted by the strengths.
- The deliverable is a sequenced finding, not a yes or no: close the blocking gaps before any consequential-content launch, pilot MT-first now only on low-stakes content in the strongest-asset languages, run the slowing gaps as parallel workstreams, and sequence the remaining languages by asset readiness. The assessment is a living instrument to re-run on every new language, content type, engine, or team change, and it is the evidence base for the roadmap, the business case, and the engine evaluation that follow.
Skill.re