โ†
AI for Translation & Localization
Aware ยท M19 ยท lesson 19 of 19 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Your 90-Day On-Ramp
๐Ÿ“–
now learning

Your 90-Day On-Ramp

15 min

It is the first Monday of the quarter, and you have just told yourself a useful lie to get started: that you already know how to do this. You have read every lesson in this level. You can recite that fluent is not correct, that one Critical error fails the file, that ISO 18587 makes the post-editor accountable as a full-competence linguist, that some content the machine must never touch. You can talk about it at a conference. But this morning a real file is sitting in your translation-management system, the machine has already pre-translated every segment of it, and a project manager is waiting on a delivery date. Knowing the cardinal rule and being able to ship a defensible, severity-scored, post-edited file against that rule are two entirely different things, separated by roughly ninety days of deliberate, unglamorous practice. This lesson is the bridge. It is a concrete first-quarter plan that takes you from "I read machine translation skeptically" on day one to "I shipped one severity-scored post-edited file with a quality record attached" on day ninety, which is exactly the artifact this level asks you to produce as your capstone: an MT Opportunity and Risk Memo on one real content type. We are going to build that, week by week, slowly, with checklists, so that on day ninety the lie you told yourself on day one has quietly become true.

Why a Ninety-Day Plan and Not a Weekend Checklist

Before the calendar, the reasoning, because a plan you do not believe in is a plan you will abandon by week three. You might reasonably ask why this takes ninety days. The machine drafts a translation in seconds. The standards fit on a few pages. Why not a weekend crash course and a certificate? The answer is that the thing you are building is not knowledge. It is a habit of attention, and a habit is the one thing that cannot be installed quickly, because it has to survive contact with deadline pressure, with fatigue, with a file that reads so smoothly you stop wanting to check it. Knowledge is what you have after a weekend. A habit is what you have after you have done the slow, against-the-source verification a hundred times, on a hundred boring segments, until the verification stops feeling like extra work and starts feeling like the work.

Recall the asymmetry this whole level has been about. A machine-translation (MT) engine, meaning any system that converts text from a source language to a target language with no human writing the words, produces output that is fluent first and accurate second. The danger is the fluent error: the grammatical, confident, native-sounding sentence that means something different from the source, often the opposite. This is the silent critical error, and the reason it is silent is that your reading reflex skims smooth prose and only slows down on awkward seams. You cannot fix a reflex with a fact. You can only retrain it with repetition. Ninety days is roughly how long it takes a working linguist, on real files, to convert "I know I should read against the source" into "I cannot stop myself from reading against the source on a number, a negation, or a drug name." That involuntary flinch is the deliverable. The plan exists to build it.

You are not memorizing a rule in ninety days. You are retraining a reflex, and a reflex only changes through repetition under real conditions, not through one weekend of good intentions.

The Shape of the Quarter

The ninety days break into three thirty-day phases, and the phases are sequenced deliberately, because each one is the foundation the next stands on. Days 1 to 30 are read and skeptic: you build the vocabulary, internalize the standards, and learn to read MT output with disciplined suspicion, without yet being responsible for fixing it. Days 31 to 60 are post-edit against the source: you begin actually editing machine output, but only against the source segment, the termbase, and the translation memory, and you start applying the MQM error dimensions to what you find. Days 61 to 90 are score and ship: you take one real file, post-edit it to a defensible standard, score it against the Critical, Major, and Minor severities, and produce the quality record that is your capstone. You do not skip ahead. You do not score a file in week two, because you have not yet built the reading discipline that makes the score mean anything. The order is the method.

Days 1 to 30: Read and Skeptic

The first month has one job, and it is not to fix anything. It is to rebuild how you read a pre-translated segment, and to load the vocabulary and the standards so deeply that they become the lens you look through rather than facts you look up. You are going to read a great deal of machine output this month and change almost none of it, on purpose, because the discipline you are installing is noticing, and noticing has to come before fixing or you will fix the wrong things and miss the dangerous ones.

Week 1: Vocabulary as Muscle Memory

You cannot reason precisely about something you cannot name precisely. The first week makes the vocabulary of this level automatic, because every later judgment depends on it. By the end of the week you should be able to define each of these terms in working language, without hesitation, the way you know your own name. A segment is the unit the translation tool splits text into, usually a sentence, and the atom of everything that follows. Machine-translation post-editing (MTPE), sometimes shortened to PE, is the workflow in which a human edits machine output instead of translating from scratch. Quality estimation (QE) is an automatic confidence score a model assigns to output with no human reading. A translation memory (TM) is the database of previously approved source-and-target segment pairs your tool reuses. A termbase is the controlled glossary of the client's approved terms. Locale is the specific language-and-region convention set: dates, units, currency, formality. MQM is the Multidimensional Quality Metrics framework, an analytic error typology that classifies translation errors by dimension and severity. A large language model (LLM) is a general-purpose text-prediction system that translates as a side effect of its general competence, and is more fluent and more confidently wrong than classic neural machine translation (NMT), a network trained only on the translation task.

Your Week 1 checklist:

  • Write each term above on one side of a card and its working definition on the other. Drill them until you can go both directions without pausing. This is not busywork; the precision is load-bearing later.
  • Open three recent files in your CAT tool, the computer-assisted translation environment where you work, and label five segments in each by hand: which are pre-populated from MT, which from a TM match, which are exact or fuzzy matches. A fuzzy match is a TM segment that is similar but not identical to the current source. Knowing where each segment came from changes how much you trust it.
  • For every acronym you used out loud this week, force yourself to expand it once. MTPE, QE, TM, MQM, NMT, LLM. The expansion keeps the meaning live instead of letting the initials go hollow.
  • Write one paragraph, for yourself, explaining the difference between fluency and accuracy to an imaginary client. If you cannot do it cleanly, you do not yet own the distinction.

Week 2: The Standards as Questions You Ask

The three standards this level taught are not trivia to recite. They are a set of questions you learn to ask of any AI-touched file, and the second week converts each standard from a number into a question. ISO 17100 is the human-translation services baseline: it defines what professional translation competence and process actually mean, and it is the foundation the post-editing standards build on. The question it gives you is: does the person and the process behind this file meet the baseline a professional translation requires, or is the machine being treated as a substitute for that baseline rather than a draft within it? ISO 18587 is the machine-translation post-editing standard, and its revision, in DIS ballot with publication targeted for late 2025 into 2026, expands its scope from machine translation to non-human translation output, meaning it explicitly now covers AI and LLM output, retires the rigid light-versus-full split for an effort spectrum, aligns with ISO 17100, and insists the post-editor hold the same full linguistic competence as a professional translator. The question it gives you is: does the effort applied to this content match its consequence, and is the human who signs it actually qualified to catch what the machine missed? ISO 5060:2024 formalizes the MQM-aligned analytic evaluation: each error is classified by dimension (accuracy, terminology, locale, fluency) and by severity (Critical, Major, Minor). The question it gives you is: scored honestly against this typology, does the file pass, or does it carry even one Critical?

Your Week 2 checklist:

  • Write the three standards on a single page, each reduced to the one question it makes you ask. Pin it where you work. You will refer to it for the rest of the quarter.
  • Take one file you have already delivered in the past and ask all three questions of it retrospectively. You are not redoing the work; you are practicing the interrogation on something low-stakes.
  • Memorize the four MQM dimensions and the three severities cold. Accuracy, terminology, locale, fluency. Critical, Major, Minor. These are the coordinates of every score you will ever produce.
  • Write, in one sentence each, an example from your own domain of a Critical, a Major, and a Minor error. Concrete and specific. If your Critical example is not something that could harm a reader or expose someone legally, you have not understood the severity yet.

Week 3: Reading Against the Source, Not the Flow

This is the pivotal week of the first month, because it installs the core operation. You stop reading the target for flow and start reading it against the source, element by element, on the high-consequence categories. The point this week is to do this without responsibility for fixing anything: you are only noticing, only flagging, only building the flinch. A comparison cannot be fooled by fluency, because fluency is a property of only one of the two things being compared, so the entire defense against the silent critical error is to compare rather than to read.

The never-trust-the-fluency categories, the ones you check against the source every single time regardless of how smooth the sentence sounds, are these:

  • Negations. Every "not," every prohibition, every "must not," "shall not," "contraindicated." A dropped or added negation is the single most common silent Critical, and it is one tiny word.
  • Numbers, dosages, and units. Every figure, character by character: 2.5 is not 25, mg is not mcg, twice daily is not twice weekly.
  • Names and approved terms. Drug names, proper names, and the client's approved termbase entry, not a fluent synonym the engine preferred.
  • Dates, times, and durations. Against the source and against the locale convention.
  • Obligations and parties. In legal content, who must do what to whom: which party indemnifies, who is liable.
  • Adverse events and warnings. The description of harm and the line between expected and dangerous.

Your Week 3 checklist:

  • On three live files, run a noticing-only pass: read each target segment against its source on the six categories above, and mark, do not fix, every segment where the fluency would have carried you past a discrepancy. Count them at the end. The count is your baseline.
  • For every flagged segment, write one word: the category it belongs to. You are building a felt sense of where the machine fails in your specific content.
  • Deliberately catch yourself reading for flow. When you notice your eye skimming a smooth segment without checking a number in it, stop, back up, and check. The catching-yourself is the training.
  • At week's end, write down the one category where you flagged the most discrepancies. That is your personal high-alert zone, and it tells you where your attention leaks.

Week 4: Where the Machine Is Forbidden, and Where It Genuinely Helps

The first month closes by teaching judgment about the engine itself: where it is a near-free gift and where it must never be allowed near the content. This is the seed of the risk triage your capstone depends on. The machine genuinely helps on high-volume, lower-liability content: product descriptions, internal knowledge bases, user-generated content, first drafts of marketing copy, subtitling at scale. Here a fluent error is recoverable, a light post-editing pass suffices, and the speed is a genuine gift. The machine is forbidden, or at minimum demands full human translation or full post-editing with the most senior qualification, on regulated and life-safety content: drug labels and patient information leaflets, clinical-trial protocols, contracts and indemnity clauses, financial disclosures, evacuation and hazard warnings. Here a fluent error reaches a body, a balance sheet, or a regulator, and the asymmetry between the engine's speed and the cost of its silent mistake is at its most extreme.

Your Week 4 checklist:

  • List the content types you actually work on. For each, write one of three labels: MT-helps (light PE acceptable), MT-with-full-PE (high consequence, machine draft allowed but full verification required), or MT-forbidden (machine must not touch). Defend each label in one sentence tied to consequence.
  • For one MT-forbidden type on your list, write the specific harm a silent critical error would cause. Make it concrete: a body, a sum of money, a regulatory event.
  • Review the four MQM dimensions against one MT-helps file and one high-consequence file, and notice how the same dimension carries wildly different stakes depending on content. A locale error on a blog date is Minor; a locale error on a dosing schedule can be Critical.
  • Write your day-30 self-assessment: can you read a segment against its source on the six categories without being asked? Can you name the standard that applies and the question it raises? If yes, you are ready to start fixing. If not, repeat Week 3 before advancing. Honesty here protects you later.

Days 31 to 60: Post-Edit Against the Source and the Termbase

The second month is where you start changing the file, and the discipline shifts from noticing to editing without over-editing. The trap of the second month is the opposite of the first: now that you are allowed to fix, you will want to fix everything, including things that do not matter, which burns the budget that makes MTPE economically sane and erases the very speed advantage the workflow exists to capture. The skill of this month is editing against the source, the termbase, and the translation memory, and only against those, fixing what affects meaning, terminology, locale, and genuine usability while leaving alone the merely-preferential rewrites that an unsupervised reviewer's ego produces.

Week 5: Your First Post-Edits, Against the Source Only

You begin editing, but on lower-consequence content, the MT-helps tier, so that the cost of an early mistake is small while the habit forms. The rule of this week is simple and strict: every edit you make must be justified by the source segment, the termbase, or the TM, never by your own sense of how you would have phrased it. If you cannot point to the source meaning, the approved term, or the locale rule that your edit serves, you are over-editing, and you stop.

Your Week 5 checklist:

  • Post-edit one MT-helps file. Before each edit, ask: what does this fix, source accuracy, an approved term, a locale rule, or a genuine fluency problem that impedes understanding? If the answer is "I just prefer it," revert.
  • Keep a running tally of your edits in two columns: necessary (meaning, term, locale, blocking fluency) and preferential (style, taste, equivalent rephrasing). At week's end, the preferential column tells you how much budget your ego is spending.
  • On every segment containing one of the six high-consequence categories, do the against-the-source check from Week 3 before you accept or edit. The noticing habit now becomes a gate before action.
  • Time one file. Note your words-per-hour. You are establishing a personal baseline you will watch as quality discipline and speed learn to coexist.

Week 6: The Termbase as Law, Against a Drifting Engine

This week you confront the specific failure of an engine that "prefers" a more common synonym over the client's approved term. The engine is optimizing for probable, fluent text, and the client's mandated term for a device, a feature, or a legal concept is often not the most statistically common word, so the engine quietly substitutes a near-neighbor that reads beautifully and violates the termbase. Term drift is a Major error at minimum and, on a regulated term, can be Critical, and it propagates: once the wrong term enters the TM, it infects every future leverage from that memory.

Your Week 6 checklist:

  • On one file with a real termbase, check every termbase entry's appearance in the target by hand. Mark each place the engine drifted off the approved term, even when the substitute reads perfectly. Fluency is not compliance.
  • For each drift, classify the severity: is this a Minor inconsistency, a Major term violation, or, on a regulated or safety term, a Critical? Tie the severity to consequence, not to how different the words look.
  • Notice whether any drifted terms already live in the TM. If they do, you have found a propagation vector, and you note it for the quality record. A clean TM compounds quality; a dirty one compounds error.
  • Practice writing a one-line query for the client on any term where the approved entry seems wrong for the context. The terminologist's instinct is to ask, not to guess.

Week 7: Applying the MQM Dimensions to What You Find

Now you connect your editing to the analytic typology, so that every problem you catch gets a name from the ISO 5060 vocabulary rather than a vague "this is off." This is the week your noticing becomes scoring-ready. Every error you find belongs to exactly one of four dimensions, and naming it correctly is the foundation of a defensible record. Accuracy: does the target convey the source meaning? Mistranslations, omissions, additions, dropped negations all live here. Terminology: does it use the approved terms from the termbase? Term drift lives here. Locale conventions: dates, units, currency, formality, encoding. Fluency: is the target itself well-formed, grammar, spelling, punctuation, register, independent of the source?

Your Week 7 checklist:

  • On one post-edited file, label every error you caught with its MQM dimension. Force yourself to choose exactly one. The discipline of choosing trains the distinction.
  • Watch for the most common confusion: an error that feels like fluency but is actually accuracy. A dropped negation reads as a fluency-clean sentence but is an accuracy error of the highest severity. Dimension is about what is wrong, not about how it reads.
  • Tally your errors by dimension. The shape of the tally tells you where this engine, on this content, fails most, which is intelligence you carry into the capstone.
  • Pair each error with a severity. You now have, for the first time, error records with both a dimension and a severity, which is the raw material of a quality score.

Week 8: Matching Effort to Consequence

The second month closes by operationalizing the ISO 18587 effort spectrum: not every file deserves the same depth of post-editing, and applying full verification to a throwaway is as much a failure of judgment as applying a light pass to a contract. The revised standard retired the rigid light-versus-full binary precisely because real content sits on a spectrum, and the post-editor's judgment about where a given file falls is itself a professional skill. The principle: effort rises with consequence. Low-consequence, high-volume content gets a light pass focused on blocking errors. High-consequence content gets full post-editing with the complete against-the-source verification on every high-alert category, and the most senior qualification.

Your Week 8 checklist:

  • Take three files of genuinely different consequence and assign each an effort level: light PE, full PE, or human-only. Write the one-sentence justification tied to what happens if a fluent error survives.
  • Post-edit one of them at the matched effort level, and deliberately do not over-edit the light one or under-verify the full one. Discipline runs in both directions.
  • Re-read your day-30 content-type labels from Week 4 and refine them now that you have edited real files. Your judgment about where the machine helps and where it is forbidden should be sharper.
  • Write your day-60 self-assessment: can you edit against the source without over-editing? Can you name every error's dimension and severity? Can you match effort to consequence and defend it? If yes, you are ready to score and ship. If your preferential-edit column is still large, spend extra days on restraint before advancing.

Days 61 to 90: Score, Ship, and Write the Memo

The final month produces the artifact. You take one real file, post-edit it to a defensible standard, score it against the Critical, Major, and Minor severities, and assemble the quality record and the MT Opportunity and Risk Memo that is this level's capstone. Everything in the first two months was preparation for the thing you build now: a deliverable a client or an auditor could read and a go/no-go decision you could defend out loud.

Week 9: Choose the File and Tier the Risk

The capstone begins with selection and triage, because the memo is fundamentally an argument about where MT fits for one real content type and what risk it carries. Choose a single, real content type you actually work on, ideally one with genuine stakes so the exercise is honest, but not so high-stakes that you cannot responsibly post-edit it yet. Then build the risk argument the memo is named for.

Your Week 9 checklist:

  • Pick one real content type and one representative file. Name it precisely: not "marketing" but "in-app onboarding microcopy for the German market."
  • Write the risk tier: is this MT-helps, MT-with-full-PE, or MT-forbidden, and why, in terms of the consequence of a surviving fluent error? This is the spine of the memo.
  • State which standards apply and the question each raises for this file: the ISO 17100 baseline, the ISO 18587 effort and competence requirement, the ISO 5060 scoring obligation.
  • Decide the post-editing effort level the content warrants, and justify it against the risk tier. The effort decision must follow from the consequence, not from the deadline.

Week 10: Post-Edit the File to a Defensible Standard

This week you do the actual work, applying everything from the first two months at full strength on the chosen file. You post-edit against the source, the termbase, and the TM; you run the against-the-source verification on every high-consequence category; you edit what matters and leave the preferential alone; and you keep notes as you go, because the record is being built in real time, not reconstructed afterward.

Your Week 10 checklist:

  • Post-edit the full file at the matched effort level, with the six high-alert categories checked against the source on every segment that contains one.
  • As you work, log each error you find: the segment, the MT output, your edit, the MQM dimension, and a provisional severity. This log is the seed of your quality record.
  • Enforce the termbase by hand against engine drift, and flag any drifted term already in the TM as a propagation note.
  • Verify locale conventions explicitly: dates, units, currency, formality. Do not assume the engine got the locale right because the prose reads naturally.

Week 11: Score the File, Critical, Major, Minor

Now you convert your error log into a severity-scored evaluation, the heart of the capstone and the thing that turns "I post-edited this" into "here is the proof." You apply the ISO 5060-aligned model: each error already has a dimension; now you finalize its severity, remembering that severity tracks consequence, not edit size. A one-word dropped negation in a high-stakes segment is Critical; a stiff but accurate phrasing on a low-stakes line is Minor.

Severity is about what happens when a real reader acts on the sentence, not about how many characters changed. One inverted contraindication is Critical even though it is one word; a clumsy but accurate paragraph may be only Minor.

Your Week 11 checklist:

  • Finalize a severity for every logged error: Critical, Major, or Minor, each justified by consequence in one line.
  • Apply the gate: is there even one Critical? If yes, the file does not pass, and your memo must say so honestly. One Critical fails the file regardless of how clean the rest looks, because harm does not average out.
  • Summarize the score: counts by severity and by dimension. This summary table is the readable face of your evaluation.
  • Sanity-check your own severities against a peer or against the lesson's examples. Calibration is part of the skill; a severity no one else would agree with is not yet defensible.

Week 12: Assemble the MT Opportunity and Risk Memo

The final week produces the deliverable. The MT Opportunity and Risk Memo is a short, defensible document about one content type that brings together everything the quarter built: where MT fits, the risk tier, the post-editing effort it warrants, the standards that apply, the severity score with the Critical gate, and a clear go or no-go, including the honest recognition of any MT-forbidden cases. It is the sentence the goldmine of this whole program is organized around, written down and backed by evidence: here is the throughput, here is the risk tier this content got, here is the 5060 error score with the Critical count, here is the terminology conformance, here is the post-editing record, all defensible under the revised ISO 18587.

Your Week 12 checklist:

  • Opportunity. State where MT genuinely helps for this content type and what speed it captures. Be specific and honest; do not oversell.
  • Risk tier and standards. State the risk classification, the consequence of a surviving fluent error, and the three standards' questions answered for this file.
  • Effort and score. State the post-editing effort applied, and attach the severity-scored evaluation: counts by severity and dimension, with the Critical gate result stated plainly.
  • Go or no-go. Make the recommendation. Ship, ship-with-conditions, or do-not-machine-translate. Tie it to the score and the risk tier, not to the deadline. If the content is MT-forbidden, say so and say why.
  • The record. Attach the segment-level log, MT output, edit, dimension, severity, so a client or auditor can reconstruct your reasoning. The record is what makes the memo a credential rather than an opinion.

When that memo is written and the record is attached, you are done, and you have done the thing the lie on day one assumed you already could. You did not learn it. You built it, slowly, one boring verified segment at a time, until the reflex was yours.

Staying on the Ramp After Day Ninety

A ninety-day plan has a quiet danger built into it: the assumption that on day ninety-one you are finished. You are not, and the linguists whose careers move up in the MT era are the ones who understand that the on-ramp merges into a road, not a finish line. The habit you built is perishable. The against-the-source flinch decays if you stop using it, the same way the reading-for-flow reflex will try to reassert itself the moment a deadline gets tight and a file reads smoothly. The work after day ninety is maintenance and deepening: you keep running the verification on the high-consequence categories on every file, you keep scoring honestly even when the deadline argues against it, and you keep your content-type risk tiers current as the work changes.

This is also where the next level begins. Everything you built here was awareness and disciplined manual practice. The next level, the AI-assisted linguist, teaches you to direct the engine deliberately rather than only react to it: prompting it against a source, grounding it on your termbase and locale rules, and building the verification into a repeatable system rather than a manual pass. But none of that is safe to learn until the reflex this quarter built is real, because directing an engine you do not read skeptically only lets you ship fluent errors faster. The ninety days were the prerequisite, not the destination. You earned the right to go faster by first proving you can go carefully.

The on-ramp merges into a road. The flinch you built is perishable, and the linguist whose role moves up is the one who keeps running the verification long after the ninety days, when the deadline argues hardest against it.

Key Takeaways

  • The ninety-day on-ramp builds a habit of attention, not a body of knowledge: a reflex that reads a pre-translated segment against its source automatically. A reflex changes only through repetition under real deadline pressure, which is why it takes a quarter and not a weekend.
  • The quarter runs in three sequenced phases that must not be reordered: days 1 to 30 read and skeptic (vocabulary, standards, against-the-source reading with no responsibility to fix), days 31 to 60 post-edit against source, termbase, and TM while applying the MQM dimensions, and days 61 to 90 score one real file and ship it with a quality record.
  • Month one installs the noticing habit before any editing: drill the vocabulary (segment, MTPE, QE, TM, termbase, locale, MQM, NMT, LLM), convert ISO 17100, ISO 18587, and ISO 5060 into the question each makes you ask, and practice reading the six high-consequence categories against the source without fixing anything yet.
  • The six never-trust-the-fluency categories, checked against the source every time regardless of how smooth the prose reads, are negations, numbers and dosages and units, names and approved terms, dates and times, legal obligations and parties, and adverse events and warnings.
  • Month two adds editing with restraint: every edit must be justified by the source, the termbase, or the TM, never by preference, because over-editing burns the budget that makes MTPE viable. You enforce the termbase by hand against a drifting engine, label every error by MQM dimension (accuracy, terminology, locale, fluency), and match effort to consequence under the ISO 18587 effort spectrum.
  • Month three produces the capstone, an MT Opportunity and Risk Memo on one real content type: choose and risk-tier the file, post-edit it to a defensible standard while logging every error, score it against Critical, Major, and Minor severities where one Critical fails the file, and assemble the memo plus a segment-level record an auditor could reconstruct.
  • Severity tracks consequence, not edit size: a one-word dropped negation is Critical, a stiff but accurate paragraph is Minor. The go/no-go in the memo follows from the score and the risk tier, not from the deadline, and includes the honest recognition of MT-forbidden cases.
  • Day ninety is a merge, not a finish line. The flinch is perishable and decays without use, so the work continues as maintenance and deepening, and only once the reflex is real is it safe to advance to directing the engine deliberately in the next level, because directing an engine you do not read skeptically just ships fluent errors faster.