Training Linguists, PMs, and Engineers
The strategy deck was flawless. A localization director at a medical-device company, eighteen months into building an MT-first pipeline, stood in front of her leadership team in the spring of 2026 and presented a roadmap that had everything: a risk-tiered intake model, a chosen engine that beat its rivals on a domain proof-of-concept, a defensible MTPE pricing tier, and an ISO 18587 and ISO 5060 governance program on paper. The board approved it. Six weeks later the first shipped file failed a client audit. A dosage had flipped in a fluent, grammatical, confident German sentence, and it had sailed through a post-editor who had spent fifteen years as a superb translator and had never once been taught how to read machine output against a source with the specific suspicion that the smooth sentence is the dangerous one. The PM who scoped the file had quoted light post-editing on content that legally required full human translation, because nobody had trained her on risk-tiering. The engineer who set up the project had let the engine touch a locked terminology set, because nobody had told him it mattered. The strategy was not wrong. It was inert, because it lived in a slide deck and not in the hands and habits of the people who touch the file. This lesson is about the single thing that turns a localization-AI strategy from an inert document into a working operation: a training program that builds, role by role, a workforce that can actually run the pipeline the strategy describes.
Why the Strategy Fails Without the Training
Every level of this curriculum before this one taught you to design something: the pipeline, the quality gate, the pricing tier, the governance program. This lesson confronts the uncomfortable truth that a design is not an operation. You can write the most defensible risk-tiering rulebook in the industry, and it does nothing until the intake coordinator can look at an incoming file and correctly place it in a tier. You can operationalize ISO 5060:2024, the standard that formalizes the MQM-aligned Critical, Major, and Minor error scoring, and it produces nothing until you have evaluators who can apply those severities consistently across ten thousand segments. The gap between a strategy and a capability is a trained workforce, and closing that gap is not a soft, nice-to-have HR exercise. It is the load-bearing wall of the whole build.
Here is the reframe that a localization leader has to internalize before anything else in this lesson makes sense. In an MT-first shop, your training program is a quality control. It is not adjacent to your quality system; it is part of it. The revised ISO 18587, the international standard for post-editing machine translation output (in DIS ballot with publication targeted for late 2025 into 2026, and now expanded to cover AI and LLM "non-human translation output"), does not merely suggest that your post-editors be competent. It requires the post-editor to hold the same linguistic competence as a professional translator, and it makes competence something you must be able to evidence, not assume. The moment competence becomes a documented requirement, training stops being a benefit you offer and becomes a control you operate. If you cannot show that the person who signed the delivery was trained and qualified to do so, your conformance claim is a hope, and a hope does not survive a client audit.
In an MT-first operation, the training program is not support for the quality system. It is part of the quality system. Untrained hands running a governed pipeline produce ungoverned output that merely looks governed.
The reason untrained output is so dangerous is the same reason the whole discipline is dangerous: the failure mode is silent. A machine-translation or LLM (large language model, the general-purpose text engine that translates as a side effect of predicting the next token) engine produces output that is fluent first and accurate second. Studies of LLM output on medical content found error rates around 59% on drug names, 60% on dates and times, and 66% on adverse events, every one of them delivered in grammatically perfect prose that does not trip the eye. An untrained post-editor does not fail loudly. They fail quietly, by approving a fluent sentence that means the opposite of the source, and their name goes on the delivery. The organization does not discover the training gap until the flipped dosage ships. That is why the training program is a control: it is the thing standing between a fluent hallucination and a delivered file, and if it is weak, nothing downstream can compensate.
The Competence You Are Actually Building
It helps to name precisely what the training is for, because it is easy to drift into "teach everyone to use the tool" and miss the point entirely. The competence you are building is not fluency in a CAT tool (the computer-assisted translation environment) or a TMS (the translation-management system that routes work and holds the assets). Tool fluency is table stakes and vendors teach it. The competence you are building is the judgment layer the engine cannot supply for itself: the ability to read fluent output skeptically against a source, to assign an error's severity by consequence rather than by size, to place a file in the correct risk tier, to hold an approved term against a drifting engine, to know which content the machine is forbidden to touch. That judgment is the product. The training program exists to manufacture it reliably, at scale, across a workforce, instead of relying on the two or three heroes who happen to have it already.
The Five Role Tracks
The central design decision of a localization-AI training program is this: you do not run one course. You run role-specific tracks, because the five people who touch an MT-first file need five different competencies, and a single generic "AI for localization" course teaches all of them a little and none of them enough. A linguist and a project manager both need to understand risk-tiering, but the linguist needs to execute post-editing effort against the tier while the PM needs to scope and price against it. Teaching them the same content wastes both of their time and under-serves both of their roles. The program is a matrix: a shared foundation everyone takes, then five tracks that go deep on what each role must actually own.
Below are the five tracks, each with the role's job in one line and the specific competencies the training must build. Read them as a specification for the curriculum you will build, not as a list of topics to gesture at.
Track 1: The Linguist and Post-Editor
This is the largest track and the one closest to the file. The post-editor's job is MTPE (machine-translation post-editing, the workflow where a human edits machine output rather than translating from a blank page), and the competencies the training must build are, in order of importance:
- Verification against the source. The single most important habit, and the one a career of translating from scratch does not teach. The post-editor must learn to read the machine output not for flow but against the source segment, hunting the specific high-consequence elements the engine flips: negations, numbers, dosages, dates, named entities, and obligations. This is a taught, drilled skill, not an instinct, and it is where most of the training hours should go.
- Effort calibration. Light post-editing (fix only what breaks meaning or usability, accept adequate prose) versus full post-editing (bring to publishable, human-quality standard), matched to the risk tier of the content, without over-editing a throwaway or under-editing a contract. Both over-editing and under-editing are failures: one burns the MTPE economics, the other ships a defect.
- Severity awareness. Enough fluency in the MQM dimensions and severities to recognize when they have just fixed a Critical error versus a Minor one, and to flag it into the record, because the post-editor is the first line of the quality gate even when a separate evaluator runs the formal scoring.
- Terminology discipline. Honoring the approved term against an engine that prefers a fluent synonym, and knowing that an unverified fix saved into the TM (translation memory, the database of approved source-target pairs the tool reuses) propagates the error across every future project.
- Tool and structure safety. Not breaking placeholders, tags, or length budgets while editing.
The ISO 18587 tie is direct and non-negotiable here: this track exists to produce and evidence the full professional-translator competence the standard requires of a post-editor. Your linguist track is, quite literally, your ISO 18587 competence program.
Track 2: The Terminologist
The terminologist is the specialist who builds, governs, and enforces the termbase (the controlled glossary of a client's approved terms with their correct equivalents in each target language and the rules for using them). This is a smaller, deeper track. The competencies:
- Term extraction and verification. Drafting candidate terms fast, increasingly with AI assistance, then verifying every candidate against the source domain and the subject-matter experts before it becomes a rule. The verification, not the extraction, is the skill.
- Termbase engineering. Structuring the termbase so the CAT tool flags departures, defining approved equivalents with rationale, and managing forbidden terms and context notes.
- Drift diagnosis. Understanding exactly where and why an engine reaches for a near-synonym over the approved term, so the enforcement mechanisms are aimed at the real failure points.
- TM hygiene as governance. Keeping the memory clean so terminology fidelity compounds instead of infecting future leverage.
Track 3: The QE Analyst and Evaluator
This track produces the people who run the quality gate. Two ideas anchor it, and the training must keep them rigorously separate. Quality estimation (QE) is an automatic, reference-free confidence number a model assigns to output with no human reading it, a routing signal. An MQM evaluation (Multidimensional Quality Metrics, the analytic framework that classifies each error by dimension and severity, formalized for translation output in ISO 5060) is a human verdict. The evaluator's competencies:
- Reading QE as routing, never ruling. Treating the automatic score as a triage tool that decides which segments deserve a slow read, and never as a verdict that ships, because the silent critical error is exactly the fluent segment a confidence score rates as safe.
- Applying the MQM/ISO 5060 typology consistently. Classifying each error by dimension (accuracy, terminology, locale, fluency) and severity (Critical, Major, Minor), the same way, across thousands of segments and across different evaluators.
- Severity by consequence. The hardest calibration: a one-word dropped negation is Critical; a stiff marketing paragraph is Minor. Size does not determine severity; consequence does.
- Producing the defensible record. A structured error report, with counts by dimension and severity and a clear pass or fail verdict, that survives a client audit.
The signature discipline of this track is inter-rater reliability: two trained evaluators scoring the same file should produce substantially the same result. If they do not, you do not have an evaluation program; you have two opinions. Calibration is the whole game, and it is a training outcome, not an accident.
Track 4: The Localization Engineer
The localization engineer owns the structural failure modes that are not linguistic at all: the ways an engine ships broken product without mistranslating a single word. The competencies:
- Placeholder and tag protection. Ensuring the engine treats a code placeholder (like
{count}or%s, the token a running program replaces with a live value) as untouchable structure, not text to be made fluent, so the application does not crash or display garbage. - Length and layout budgets. Handling the reality that a target language expands or contracts, so translated text does not overrun a fixed UI element or a subtitle's reading-rate budget.
- Encoding and format integrity. Preventing the corruption (the mojibake) that turns a language's accented characters into garbage when text is decoded wrong.
- Pipeline integration. Wiring the internationalization (i18n, building software so it can be adapted to any language without re-engineering) and localization checks into the build so quality holds at build time, with the quality gate enforced in CI/CD rather than hoped for.
Track 5: The Localization Project Manager
The localization PM owns the file from intake to delivery, and in an MT-first world the PM's hardest new competency is risk-tiering, the discipline of classifying content by consequence before a single segment is post-edited. The competencies:
- Risk-tiered intake and routing. Correctly placing a file (marketing string versus drug label versus indemnity clause) into a tier that dictates the workflow: MT with light post-editing, full post-editing, full human translation, or MT-forbidden.
- Workflow and effort scoping. Matching the post-editing effort and the human sign-off to the tier, and never quoting the cheap workflow on content that can harm.
- Defensible pricing. Quoting a provable quality tier instead of racing a client to the bottom on price, and holding the line when a client demands MTPE rates on content that legally requires full human translation.
- Standards literacy. Enough command of ISO 18587 and ISO 5060 to explain to a client what conformance means and why the quality record is part of the deliverable.
Five people touch an MT-first file, and they need five competencies. A single generic "AI for localization" course teaches each of them a little and none of them enough. The program is a shared foundation plus five deep tracks, and the depth is where the quality lives.
The Shared Foundation Everyone Takes
Before the tracks branch, everyone in the operation, all five roles, takes the same foundation, and the reason is cultural as much as technical. If the post-editor, the PM, and the engineer do not share a common mental model of the core danger, they will make locally reasonable decisions that combine into a shipped defect, exactly as they did in the opening story. The foundation is short, but it is mandatory, and it is the same for a freelance post-editor and a head of engineering.
The foundation carries four non-negotiables, the spine of the entire program:
- Fluent is not correct. The engine's output is fluent first and accurate second, and a fluent error is more dangerous than a clumsy one because it does not trip the eye. Accuracy is a relationship to the source segment and the approved terminology, not a property of the prose. Everyone must hold this, because everyone makes a decision that depends on it.
- The silent critical error is the killer failure mode. A flipped dosage, a dropped negation, or an inverted indemnity clause, delivered in perfect prose, is the error that costs a life or a lawsuit. The whole operation exists to catch it.
- Accountability stays human. "The engine wrote it" is never an answer when a Critical error ships. The person who signs the delivery owns the quality, whatever their role.
- Some content is MT-forbidden. Effort matches consequence, and some content the engine must never touch. Risk-tiered intake is the first control, and everyone needs to recognize high-consequence content when it crosses their desk, not only the PM who tiers it.
Keep the foundation deliberately short, a half-day at most, because its job is alignment, not depth. Depth belongs in the tracks. But do not skip it or make it optional, because it is the shared language that lets a linguist flag a mis-tiered file back to a PM, and an engineer flag a locked terminology set back to a terminologist. Without the common foundation, the handoffs between roles are where the silent errors slip through.
How to Build and Deliver the Program
Content is only half of a training program. The other half is delivery, and a leader who gets the curriculum right and the delivery wrong ends up with a beautifully designed course that nobody completes and that changes no behavior. The delivery model has to fit the reality that your workforce is busy, partly freelance, distributed across time zones, and skeptical of anything that smells like the machine coming for their jobs. Here is how to build and deliver it so it lands.
Build on Real Files, Not Toy Examples
The single highest-leverage decision in building the content is to teach on real, sanitized files from your own domain and your own engine's actual output, not on clean textbook examples. The reason is that the silent critical error is a domain-specific, engine-specific phenomenon. Your engine has particular ways it drifts on your content: the terms it substitutes, the negations it drops, the placeholders it mangles. A generic example of a flipped dosage teaches the concept; a real example of your engine flipping a dosage on your kind of content teaches the reflex. Build a library of curated failure cases, drawn from your own quality records and near-misses, sanitized of client-confidential data, and use them as the training's spine. Nothing you can buy off the shelf will teach your workforce your engine's specific failure signature, and that signature is what they need to learn to catch.
Teach by Doing, With a Scored Artifact
Localization competence is procedural, not declarative. Nobody learns to catch a fluent mistranslation by reading about it; they learn by post-editing a real file, missing a Critical error, having it caught in review, and feeling the specific sting of the smooth sentence that fooled them. Every track must culminate in a hands-on artifact that mirrors the real work and is scored the way the real work is scored:
- The linguist track ends with a post-edited file, scored against MQM/ISO 5060 severities, that the trainee must clear with zero undetected Criticals.
- The terminologist track ends with a verified termbase and a drift-enforcement setup on a real content set.
- The evaluator track ends with a scored evaluation that must reach a calibration threshold against a master scoring of the same file.
- The engineer track ends with a pipeline configuration that survives placeholder, length, and encoding stress tests.
- The PM track ends with a correctly tiered and scoped intake of a mixed-content project, including the correct MT-forbidden calls.
The artifact matters twice: it is how people learn, and it is your competence evidence. A scored artifact, retained, is the document that turns "we trained them" into "here is the file that proves the trainee could do the work to standard." That is the difference between a training claim and a training record, and only the record survives an audit.
Sequence It, and Refresh It
Do not deliver the program as a one-time event, because a one-time event decays. Engines change, content changes, and skills atrophy. Sequence the delivery: the shared foundation first, then the track for the role, then supervised production where a trainee's early real files are reviewed at a higher rate before they earn autonomy. Then refresh: when the engine is retuned or replaced, when a new content type or vertical enters the pipeline, and on a periodic cadence for the evaluators, whose calibration drifts silently over time exactly the way the engine's output does. Treat re-calibration of your evaluators as a recurring, scheduled control, not a thing you do when something breaks, because by the time an evaluator's drift causes a visible failure, uncounted files have already shipped under a miscalibrated gate.
Train on your own engine's real failure signature, teach by doing until the trainee has felt the sting of the fluent error, retain the scored artifact as competence evidence, and re-calibrate the evaluators on a schedule, because their drift is as silent as the engine's.
Certification and Competence Evidence
Here is where the training program stops being an internal nicety and becomes a strategic asset. The revised ISO 18587 requires the post-editor to hold full professional-translator competence, and it requires that competence to be something you can evidence. Your certification and competence-evidence system is how you meet that requirement, and if you build it well, it is also a thing you can sell.
What Competence Evidence Actually Is
Competence evidence is not a certificate on a wall. It is a defensible, retained record that a named person was qualified to perform a named role to a named standard, on a named date, with the artifact that proves it. For each person in each role, your evidence file should contain: the training they completed, the scored artifact they produced (the post-edited file with zero undetected Criticals, the calibrated evaluation, the tiered intake), the standard the artifact was scored against, and a re-qualification date. This is the same logic as the quality record for a delivered file, applied to the person instead of the file. When a client's auditor asks "how do you know the person who post-edited our regulated content was competent to do so," the answer is not "they are experienced." The answer is a retained evidence file, and the auditor can read it.
Tie this explicitly to ISO 18587's full-competence requirement. The standard asks you to establish that your post-editors hold professional-translator-level competence; your linguist track and its scored artifact are precisely how you establish it. Do not treat these as two separate things, an ISO obligation over here and a training program over there. The training program, properly recorded, is the mechanism that satisfies the ISO obligation. Design them as one system and you get conformance and capability from a single investment; design them separately and you do double the work and can prove neither.
The Evaluator Calibration Problem, Specifically
Evaluators need a stricter certification than any other role, for a reason worth understanding. The evaluator is the human who decides whether a file passes, so an evaluator who scores inconsistently or leniently corrupts the entire quality gate, and does so invisibly. Their certification must be based on calibration, not just knowledge: the trainee scores a set of master-scored files, and their scoring must agree with the master to a defined threshold before they are certified to score live work. And their re-certification must be periodic, because inter-rater reliability drifts. A leader who certifies an evaluator once and never re-checks is running a quality gate whose calibration is unknown, which is functionally the same as running no gate at all while believing you have one. The evaluator certification is the keystone of the whole competence system; get it wrong and every downstream quality claim is built on sand.
Competence Evidence as a Sellable Asset
The strategic payoff: an operation that can produce competence evidence on demand sells something its commodity competitors cannot. When a regulated client is choosing between an LSP (language-service provider, the agency that sells translation as a service) that says "our linguists are experienced" and one that says "here is the retained competence evidence for every person who will touch your content, tied to ISO 18587, with the scored artifacts on file," the second wins the regulated account and the price premium that comes with it. The training program you built as an internal control becomes, at the sales table, a differentiator. This is the recurring pattern of the entire strategist tier: the disciplines you build to be safe are the same disciplines you sell to be chosen. Competence evidence is the clearest example of it, because it is the one thing a client's auditor most wants to see and a race-to-the-bottom vendor most conspicuously cannot produce.
A Worked Training Plan
Abstract principles are inert until you see them shaped into a plan. Here is a worked example: a mid-size LSP with a mixed workforce (twelve staff linguists, a freelance pool of forty, four PMs, two localization engineers, and one part-time terminologist) standing up its training program over a first year to support an MT-first pipeline that already runs but is under-governed. The numbers are illustrative, but the structure is the deliverable.
Phase 1: Foundation and Triage (Weeks 1 to 4)
Everyone, all five roles, all fifty-nine people, takes the half-day shared foundation first. In parallel, leadership builds the failure-case library from the last year of quality records and near-misses, sanitized, because the tracks cannot be built without it. The output of Phase 1 is a workforce that shares the four non-negotiables and a curated library of the engine's real failure signature to teach on. Do not start a single track until the foundation is universal, because the tracks assume it.
Phase 2: The Critical-Path Tracks (Weeks 5 to 16)
Sequence the tracks by risk, not by headcount. The evaluator track goes first and gets the most rigor, because the quality gate is the control everything else depends on, and you need certified evaluators before you can score anyone else's training artifacts. Certify two to three evaluators to a calibration threshold against master-scored files. Then run the linguist track for the staff linguists and the most-used freelancers, culminating in the scored post-edited artifact with zero undetected Criticals, and retain each artifact as competence evidence tied to ISO 18587. The terminologist deepens on drift diagnosis and enforcement. The PM track runs in parallel, culminating in a correctly tiered mixed-content intake including the MT-forbidden calls. The engineer track hardens the pipeline against placeholder, length, and encoding failures. The output of Phase 2 is a first cohort of role-certified people with retained evidence.
Phase 3: Supervised Production and Record (Weeks 17 to 28)
Certification is not autonomy. Newly certified linguists and evaluators enter supervised production, where their early real files are reviewed at a higher rate before they earn full autonomy, and the reviews feed back as targeted coaching. This phase is where the training meets reality and where the near-misses that the library did not anticipate get caught, corrected, and added to the library. The output is a workforce operating the governed pipeline with a paper trail: every certified person, every scored artifact, every supervised-production review, retained as the competence-evidence system.
Phase 4: Refresh and Re-Calibrate (Ongoing, from Week 29)
The program becomes a standing function, not a project that ends. Schedule evaluator re-calibration on a fixed cadence (a set of master-scored files they must agree with to stay certified), refresh the whole workforce when the engine is retuned or replaced, and run the foundation plus the relevant track for every new hire and every newly onboarded freelancer before they touch live work. The failure-case library grows continuously from the operation's own quality records. The output is a self-sustaining competence system: a workforce that stays qualified as the engine, the content, and the standards move underneath it, with the evidence to prove it at any moment a client or a certifier asks.
Sequence the training by risk, not by headcount: certified evaluators first because the gate is the keystone, then linguists on the scored artifact, then supervised production, then a standing refresh-and-recalibrate function. The plan is not a course you deliver once; it is a control you operate forever.
The Cost Frame for Leadership
One last thing a decision-maker needs: the honest cost frame, because a training program has a J-curve like any investment. The costs land first (the build of the library and the tracks, the hours the workforce spends in training instead of production, the slower throughput during supervised production) and the gains arrive later (a governed pipeline that ships fewer Criticals, a competence-evidence system that wins regulated accounts, a workforce that runs the strategy instead of ignoring it). Present it to leadership as exactly that shape, and frame the training not as a cost to minimize but as the mechanism that makes every other investment in the localization-AI program actually pay off. The engine, the tooling, the governance framework, the pricing model: none of them produce value until a trained workforce operates them. The training program is the investment that unlocks the return on all the others, and that is the sentence that funds it.
Key Takeaways
- A localization-AI strategy is inert until a trained workforce can run it. In an MT-first shop the training program is not adjacent to the quality system; it is part of it, standing between a fluent hallucination and a delivered file. Untrained hands running a governed pipeline produce ungoverned output that only looks governed.
- Run five role-specific tracks, not one generic course: linguist/post-editor (verification against source, effort calibration, terminology discipline), terminologist (extraction, verification, drift diagnosis, TM hygiene), QE analyst/evaluator (QE as routing not ruling, consistent MQM/ISO 5060 scoring, severity by consequence), localization engineer (placeholders, length, encoding, pipeline integration), and localization PM (risk-tiering, scoping, defensible pricing, standards literacy).
- Everyone first takes a short shared foundation carrying the four non-negotiables (fluent is not correct, the silent critical error is the killer failure mode, accountability stays human, some content is MT-forbidden), because the handoffs between roles are where silent errors slip through and a common mental model is what catches them.
- Build the content on your own engine's real, sanitized failure cases, not textbook examples, because the silent critical error is a domain-specific and engine-specific phenomenon, and only your engine's actual failure signature teaches the reflex to catch it.
- Teach procedurally: every track ends in a hands-on artifact scored the way real work is scored (a post-edited file with zero undetected Criticals, a calibrated evaluation, a correctly tiered intake). The scored artifact is both how people learn and your retained competence evidence.
- Certification and competence evidence are how you satisfy ISO 18587's full-competence requirement. The evidence is a retained record that a named person was qualified for a named role to a named standard, with the scored artifact that proves it and a re-qualification date, applied to the person the way the quality record is applied to the file.
- Evaluators need the strictest certification, based on calibration against master-scored files, with periodic re-certification, because a miscalibrated evaluator corrupts the entire gate invisibly. Certifying an evaluator once and never re-checking is functionally the same as running no gate while believing you have one.
- The worked plan sequences by risk: universal foundation and a failure-case library first, then certified evaluators, then linguists on the scored artifact, then supervised production, then a standing refresh-and-recalibrate function. The training carries a J-curve (costs first, gains later) and is the investment that unlocks the return on the engine, the tooling, the governance, and the pricing model, none of which pay off until a trained workforce operates them. Competence evidence you can produce on demand wins the regulated account a commodity competitor cannot.
Skill.re