AI for Translation & Localization
Proficient · M16 · lesson 16 of 19 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
The MT-First Post-Editing Workflow
📖
now learning

The MT-First Post-Editing Workflow

15 min

At 8:40 on a Thursday morning a single file lands in the queue, and by the time it ships at 4:15 that afternoon it will have passed through your hands, the engine's, a termbase, a translation memory, an error-scoring grid, and a record that an auditor could reconstruct line by line. The file is real in the way every file you will ever touch is real: it is a 412-segment German-to-English pharmaceutical product document, a mix of a marketing summary at the front, a clinical dosing section in the middle, and a supplier indemnity clause stapled to the back, all submitted as one job by a client who wants it back the same day at machine-translation-post-editing prices. Five years ago that request was a contradiction. Today it is Thursday. The whole reason it is possible, and the whole reason it is dangerous, is that an engine already pre-translated every one of those 412 segments before you opened the project, so you are not starting from an empty grid; you are starting from a grid that is already full, already fluent, and silently wrong in three or four places that will glide past your eye at reading speed unless something in the workflow stops them. This lesson is the backbone of the entire program: the MT-first post-editing workflow, walked end to end on this one file, with the verification built into every handoff so that speed and defensibility are not a trade-off but the same single pass. We are going to follow the file from the moment it arrives to the moment it is delivered, naming every stage, showing exactly what is checked at each one, and proving that the place where AI's biggest upside collides with the discipline's biggest liability is the place a workflow, not a hero, has to own.

The Stakes and the Shape of the Pipeline

Before we touch the file, fix the vocabulary, because we will use it precisely for the rest of the lesson and the whole point of a workflow is that the words mean the same thing every time. Machine translation (MT) is any system that converts text from a source language to a target language with no human writing the words; the production-grade neural flavor is neural machine translation (NMT), and a large language model (LLM) is a general-purpose text predictor that translates as a side effect of its broad competence and tends to be even more fluent, and more confidently wrong, than classic NMT. Machine-translation post-editing (MTPE), often shortened to post-editing (PE), is the workflow you are in: a human editing machine output instead of translating from a blank target. Quality estimation (QE) is an automatic, reference-free score the engine attaches to each segment to guess how good its own output is; it is a signal that routes your effort, never a verdict that ships a segment. A translation memory (TM) is the client's database of previously approved source-target pairs you can leverage. A termbase is the client's controlled glossary of approved terms. MQM (Multidimensional Quality Metrics) is the analytic error typology, formalized for translation output by ISO 5060:2024, that scores errors by category (accuracy, terminology, locale, fluency) and severity (Critical, Major, Minor). A locale is the specific language-and-region convention bundle: en-US dates and units are not en-GB and are not de-DE. Hold those, and the pipeline becomes legible.

Now the shape. An MT-first post-editing pipeline is not one act of editing; it is a chain of handoffs, and a Critical error can be injected, missed, or caught at any one of them. The discipline of the workflow is that every handoff has a verification built into it, so that quality is not a final inspection bolted on at the end but a property accumulated at each stage. Here are the five stages this file will pass through, and the load-bearing idea each one carries:

  • Risk-tiered intake. Before a single segment is post-edited, the content is classified by consequence so the workflow can decide what the machine may touch and how hard a human must look. The concept: match post-editing effort to consequence, the spirit of the ISO 18587 effort spectrum.
  • Grounded pre-translation. The engine drafts every segment, but grounded on the client's own TM, termbase, and style guide rather than the open web, so the first draft already speaks the approved language where it can. The concept: retrieval over your linguistic assets narrows the gap the human has to close.
  • Post-editing against source, TM, and termbase. The human works the MT output against the source segment, the approved terms, and the memory, never against the smooth surface. The concept: accuracy is a relationship to the source and the terminology, not a property of the prose.
  • The severity-scored QE gate. The post-edited output is scored against the MQM/ISO 5060 typology, and a single Critical error blocks delivery regardless of how clean the rest looks. The concept: error typology operationalized into a go/no-go, not a vibe.
  • Terminology and locale enforcement plus the quality record. Approved terms and locale conventions are enforced as controls, and every MT source, edit, term decision, and score is captured into a record a client or auditor can reconstruct. The concept: provenance and conformance under the revised ISO 18587 and ISO 5060.
A pipeline is not a sequence of steps that produce a translation. It is a sequence of handoffs, each with a verification built in, so that quality is accumulated stage by stage and a single Critical error has nowhere to hide between intake and delivery.

Two numbers anchor why this is worth doing carefully. The upside is real: a hybrid MT-first workflow lifts a linguist from roughly 2,000 words a day to 5,000 or more, MTPE adoption rose from 26% in 2022 to about 46% in 2024, 81.1% of language-service providers now offer it, and MTPE prices land at roughly 50 to 75% of full human translation. The liability is just as real: studies of LLM output on medical content found error rates around 59% on drug names, 60% on dates and times, and 66% on adverse events, every one of them delivered in grammatically perfect prose. Our 412-segment file contains drug names, dosing numbers, and an indemnity clause. The upside and the liability are sitting in the same file, and the workflow is how you take the first without shipping the second.

Stage One: Risk-Tiered Intake and the Decision the Machine Cannot Make

The file arrives as a single job, but treating it as a single job is the first and most expensive mistake available to you. Risk-tiered intake is the act of looking at the content before any post-editing begins and classifying it by what a fluent error in it would actually do to someone, because the depth of human verification a segment warrants is a function of its consequence, not its word count. The engine cannot make this decision. It has no concept of liability; it renders a drug label and a tagline with the same uniform confidence. The decision of what the machine may touch, and how hard a human must look at what it touched, is a human judgment exercised at the door, and it is the first place your value as a quality owner shows up.

Open the file and read its structure, not its sentences. Our 412 segments fall into three blocks, and they are not the same kind of content at all:

Three Blocks, Three Tiers

The marketing summary (segments 1 to 96). Product positioning, benefit statements, a brand promise. A fluent error here is recoverable: a slightly off phrasing costs a little polish, not a life or a lawsuit. This is the low-consequence tier. It is a candidate for light post-editing, where the bar is accurate and clear rather than publishable-perfect, and the source check concentrates on meaning and the few approved brand terms.

The clinical dosing section (segments 97 to 318). Indications, contraindications, dosages, administration schedules, adverse-event descriptions. This is the high-consequence tier. A flipped negation in a contraindication, a moved decimal in a dose, a swapped unit, or a scrambled conditional in an adverse-event warning is a Critical error that can injure a patient and trigger a recall. This block gets full post-editing at minimum, and the segments that carry a dosing figure or a safety instruction get the slowest, most deliberate point-and-match verification against the source.

The supplier indemnity clause (segments 319 to 412). Contractual obligations, liability allocation, warranties, the direction of who indemnifies whom. This is also high-consequence, but in the legal rather than the clinical sense: a fluent engine that swaps which party bears liability produces a clean, lawyerly clause that allocates risk to the wrong party, and both versions read with identical authority. This block gets full post-editing, and on a strict reading of the program's own rule it sits at the edge of MT-forbidden: high-liability legal content where the safest routing is full human translation or full post-editing by a competent legal linguist, never light PE.

The single most important output of intake is not a translation; it is a routing decision, recorded, that says: light PE on 1 to 96, full PE on 97 to 318, full PE with legal-grade scrutiny on 319 to 412, and a flag that the indemnity block is high-liability content the cheap workflow must never land on. That recorded decision is the first entry in the quality record, and it is the thing that lets you tell a client later, defensibly, exactly which scrutiny each part of their file received.

Intake classifies content by what a fluent error in it would do, not by how many words it contains. The depth of verification a segment earns is a function of consequence, and that judgment is the first thing the machine cannot make for you.

Why One File Holds Three Tiers (and Why That Is Normal)

The instinct of a tired linguist at 8:40 in the morning is to treat the file as homogeneous: one job, one rate, one pass. Resist it, because the consequence is not homogeneous, and the workflow's entire economic and safety logic depends on spending your slow attention precisely where a flip is catastrophic and reading faster where it is not. If you full-edit the marketing summary with the same intensity as the dosing section, you burn budget the MTPE economics did not include and you will not finish by 4:15. If you light-edit the dosing section to save time, you ship a Critical. The tiering is what lets the same file be both fast and safe: the budget you save by reading the recoverable marketing block lightly is the budget you spend reading the contraindications deliberately. Risk-tiering is not bureaucracy. It is the mechanism by which speed and defensibility stop competing and start funding each other.

Stage Two: Grounded Pre-Translation, So the First Draft Already Knows Your Language

With the routing decided, the engine drafts. But a workflow that lets the engine draft from nothing but its own training is a workflow that hands the post-editor a harder problem than necessary, because an ungrounded engine does not know the client's approved term for the device, does not know the previously approved translation of the boilerplate, and does not know the style guide's rule on formality or units. Grounded pre-translation is the practice of feeding the engine the client's own linguistic assets, the translation memory, the termbase, and the style guide, so that the first draft already speaks the approved language wherever those assets can reach. It is the difference between an engine guessing and an engine answering from your record.

What Grounding Actually Changes in the Draft

On our file, grounding does three concrete things before the human ever opens the grid. First, TM leverage: the indemnity clause and several dosing-section warnings are boilerplate the client has translated and approved before, so the TM returns exact and high-fuzzy matches that pre-populate those segments with previously human-approved language rather than fresh MT. A 100% TM match on an approved indemnity sentence is worth more than the most fluent machine rendering of it, because it has already passed a human and a client. Second, termbase injection: the engine is given the client's approved terms, so where the source says "Messsonde" the grounded engine is steered toward the approved "measuring sensor" instead of the fluent-but-wrong "probe," and where the source names the device the engine uses the mandated product name rather than a synonym. Third, style-guide constraint: the engine is told the target locale is en-US with its date, unit, and formality conventions, so it is biased toward the right separators and formats from the start.

Grounding does not make the human unnecessary. It narrows the gap the human has to close. The verification built into this handoff is a quiet but real one: you confirm the grounding actually fired. A grounded pipeline can silently fail, the termbase can be misconfigured, the TM can be the wrong version, the locale can default to the wrong region, and the result is an engine that looks grounded and is not. So the post-editor's first move on opening the grid is not to start editing; it is to spot-check that the approved terms appear where they should and the TM matches are the right ones. Grounding that you did not verify is grounding you cannot rely on.

A grounded engine answers from your approved record instead of guessing from the open web. It does not replace the human; it shrinks the gap the human must close, and the first verification is confirming the grounding actually fired.

The QE Signal Rides Along With the Draft

Alongside the grounded draft, the engine attaches a quality-estimation score to each segment: a confidence number guessing how good it thinks its own output is. Read it as a triage signal and nothing more. A low QE score is a useful flag that says look here first, this segment may need more of your attention. A high QE score is not a clearance: the engine is exactly as confident about the fluent contraindication it inverted as about the segment it got right, because QE measures the model's confidence, not the truth, and the silent Critical error is precisely the one the engine is confident about. So you use QE to order your attention, starting with the low-confidence segments and the high-consequence blocks, but you never let a high QE score talk you out of the source check on a dosing number or a negation. The score routes effort; it does not grant a pass.

Stage Three: Post-Editing Against Source, TM, and Termbase

Now the human works the file, and this is the stage where the program's cardinal distinction does all its work: the post-editor edits against the source segment, the approved termbase, and the translation memory, never against the smooth target. The reason is the one fact that governs the entire discipline: MT output is fluent first and accurate second, and a fluent error does not trip the eye, which is exactly why it is the dangerous one. Accuracy is a relationship between the target and the source meaning and the approved terms; it is not a property of the prose. You cannot read the target for flow and catch an error that is itself perfectly fluent. You have to compare the two cells.

The mechanics, applied at the intensity each block's tier earned at intake: read the source segment first and form your own understanding of what it claims before the fluent target anchors you, then point-and-match the high-consequence elements element by element. On the dosing and indemnity blocks, that means slowing down deliberately on every negation, number, named entity, obligation, date, scope word, and warning, and confirming each one survived into the target with the same polarity, the same magnitude, the same party, the same direction. On the marketing block, the source check is lighter and faster because a fluent error there is recoverable. Let us walk a handful of segments from each block exactly as you would in the grid, so you see the verification working rather than just described.

Marketing Block, Segment 12: The Light-PE Pass

Source (DE): "Vertrauen Sie auf eine Lösung, die seit über zwanzig Jahren in Kliniken weltweit eingesetzt wird."

Target (EN), grounded MT: "Trust a solution used in clinics worldwide for over twenty years."

Read the source: trust a solution that has been used in clinics worldwide for over twenty years. Point-and-match the one high-consequence element a marketing sentence carries, the number: "über zwanzig Jahren" is "over twenty years," and the target says "over twenty years." Correct. The named entity, the approved product positioning, is consistent with the brand termbase. The prose is clean and on-brand. This is the light-PE outcome the tier intended: the meaning matches the source, the one figure is verified, and you do not burn an hour transcreating a benefit statement that is already accurate and clear. You confirm and move on. The discipline here is restraint as much as scrutiny: over-editing a clean low-stakes segment is its own failure mode, because it spends budget the MTPE economics did not include and erases the speed the workflow exists to capture.

Dosing Block, Segment 148: The Moved Decimal

Source (DE): "Die empfohlene Anfangsdosis betragt 2,5 mg pro Tag."

Target (EN), grounded MT: "The recommended starting dose is 25 mg per day."

Read the source: the recommended starting dose is two-and-a-half milligrams per day. The German decimal comma in "2,5" means two point five. Point-and-match the number, digit by digit, separator by separator. The target says "25." The decimal comma was read as a thousands separator or simply dropped, and 2.5 became 25, a tenfold error in a starting dose. The unit, mg, is correct, which is exactly the partial correctness that makes the segment feel trustworthy. The sentence is a flawless English imperative and it is off by a factor of ten on a clinical dosing parameter. Caught only because this block earned the full point-and-match treatment at intake and numbers in it are verified rather than read. This is a Critical error. Flag it, fix to "2.5 mg per day," and note the locale dimension: the decimal-mark convention inverts between de-DE and en-US, which is a locale failure as much as an accuracy one, and we will return to it at the enforcement stage.

Dosing Block, Segment 203: The Inverted Conditional

Source (DE): "Wenn schwere Hautausschlage auftreten, setzen Sie die Behandlung nicht fort und suchen Sie sofort arztliche Hilfe."

Target (EN), grounded MT: "If severe skin rashes occur, continue treatment and seek medical help immediately."

Read the source carefully as a unit, because it is a conditional with two halves. If severe skin rashes occur, then do not continue treatment and seek medical help immediately. The trigger is "severe skin rashes," and the instructed action has a prohibition, "setzen Sie die Behandlung nicht fort," do not continue treatment, and a direction, seek medical help. Point-and-match the whole conditional. The trigger survived. "Seek medical help immediately" survived. But the prohibition evaporated: "nicht fort" became "continue," so the target tells a patient experiencing a severe adverse skin reaction to keep taking the drug. The surrounding correct material, the right trigger and the right second instruction, makes the segment feel even more trustworthy than a bare flipped sentence would, which is how a dropped negation hides inside a conditional. This is a Critical safety error wearing the costume of a correctly translated instruction. Caught because the conditional was read against the source as a unit and the negation inside the action half was checked for polarity. Flag Critical, fix to "do not continue treatment."

Indemnity Block, Segment 357: The Flipped Obligation

Source (DE): "Der Lieferant haftet nicht fur mittelbare Schaden, die aus der Nutzung des Produkts entstehen."

Target (EN), grounded MT: "The supplier shall be liable for indirect damages arising from use of the product."

Read the source: the supplier shall not be liable for indirect damages arising from use of the product. The polarity-bearing core is "haftet nicht," is not liable. Point-and-match the obligation, which in legal content means checking who must do what to whom and in which direction. The target says "shall be liable," dropping the negation and inverting the entire risk allocation: a clause that protects the supplier from liability for indirect damages now imposes it. Both versions read as clean, authoritative legal English; the only thing that distinguishes them is the negation, and a flipped indemnity is the kind of fluent error that loses a lawsuit rather than a recall. This is why the indemnity block was tiered high-liability at intake and why a strict shop routes it to full human translation or full PE by a competent legal linguist. Flag Critical, fix to "The supplier shall not be liable." Note also that this was not a 100% TM match; had the TM returned an approved version of this exact clause, the grounded draft would have carried the correct negation from the start, which is grounding earning its keep.

What the Post-Editing Pass Produced

Across the file the human pass found three Criticals (segments 148, 203, 357) and a scatter of Major terminology and Minor fluency issues, every Critical delivered in flawless English and every one invisible to a flow read. The marketing block confirmed quickly because its segments carried few high-consequence elements; the dosing and indemnity blocks consumed the slow attention because that is where a flip is catastrophic. The output of this stage is not yet a deliverable. It is a post-edited file with a set of caught and fixed errors, each flagged, ready to be scored. The post-editing did not certify the file. It prepared it for the gate.

Post-editing edits against the source, the termbase, and the memory, never against the smooth target, because a fluent error is invisible to a flow read and accuracy is a relationship to the source, not a property of the prose. The intensity scales with the tier; the principle never does.

Stage Four: The Severity-Scored QE Gate

A post-edited file is not a delivered file. Between them sits the gate, and the gate is what turns "I looked at it and it seems fine" into a defensible go/no-go a client and an auditor can read. The gate scores the post-edited output against the MQM/ISO 5060 error typology: every error is classified by category (accuracy, terminology, locale, fluency) and by severity (Critical, Major, Minor), and the scoring follows the one rule that makes the gate a gate rather than a suggestion: a single Critical error fails the file, regardless of how clean the rest of it looks.

The Error Typology as a Workflow Step, Not a Feeling

Scoring is the act of replacing a vibe with a record. For each error caught in stage three, you assign a category and a severity, and the assignment is defensible because it follows the typology rather than your mood. The moved decimal in segment 148 is an accuracy error, Critical, because it can injure a patient. The inverted conditional in 203 is an accuracy error, Critical, for the same reason. The flipped indemnity in 357 is an accuracy error, Critical, because it inverts a legal obligation. The "probe" versus approved "measuring sensor" miss is a terminology error, Major, because it sends a reader to the wrong component without inverting a safety claim. A target that wrote a date as 03/04 where the en-US locale wanted 04/03 is a locale error, severity depending on whether the ambiguity is consequential. A slightly awkward but accurate marketing sentence is a fluency error, Minor. The discipline is that each label is reproducible: another competent evaluator looking at the same segment against the same typology would assign the same category and severity. That reproducibility is what makes the score a record instead of an opinion.

The One-Critical-Fails Rule and Why It Is Not Negotiable

Here is the rule that gives the gate its teeth, and it is worth dwelling on because it is the most counterintuitive thing in the workflow for someone trained to think in averages. The gate does not pass a file because most of it is clean. It fails the file on a single Critical. A 412-segment file with 411 flawless segments and one inverted contraindication is a failed file, because the one inverted contraindication is the segment that injures a patient, and an average that lets 411 good segments outvote one lethal one is an average that ships the lethal one. Severity does not average; it gates. One Critical, no matter how clean the surrounding 411 segments, blocks delivery until the Critical is resolved and re-verified.

On our file, the gate caught three Criticals, all fixed in stage three. The gate's job now is to confirm those fixes hold and that no new Critical was introduced by the fix, and only then does the file pass. Notice what the gate gives you that a final read-through never could: a number a client can audit. You do not tell the client "it looks good." You tell them: 412 segments, scored against ISO 5060, three Criticals found and resolved, zero Criticals remaining, the Major and Minor counts here, and a record of each. That sentence is the difference between a quality you assert and a quality you can prove.

The gate scores the file against the MQM/ISO 5060 typology and fails it on a single Critical, no matter how clean the rest looks, because severity gates rather than averages and the one lethal segment cannot be outvoted by the clean ones.

Stage Five: Terminology and Locale Enforcement, and the Quality Record

The file has passed the gate, but two controls run across the whole pipeline rather than at a single point, and the delivery is not defensible until both are closed: terminology and locale enforcement, and the quality record that captures the provenance of every decision. These are the controls that turn a one-time clean file into a conformant, auditable delivery.

Terminology Enforcement as a Control, Not a Preference

Terminology fidelity is the requirement that the client's approved term appears, exactly, on every segment that should carry it, from intake to delivery, instead of drifting segment by segment as a fluent engine substitutes synonyms it prefers. Grounding at stage two steered the engine toward the approved terms; post-editing at stage three caught the ones that drifted anyway; enforcement at stage five is the final sweep that confirms the approved term holds across all 412 segments as a consistency check, not a preference. If the termbase says the device is "the Lumera 400" and one segment in the marketing block calls it "the unit," that inconsistency is a terminology error even though "the unit" reads fine, because terminology is about identity and conformance, not about whether the synonym is plausible. The enforcement sweep is a control: it is run by rule across the file, it produces a conformance result you can report, and it does not rely on the post-editor having happened to notice every drift while reading. Approved terms that hold across the whole pipeline are the difference between a translation and a conformant one.

Locale Enforcement: The Conventions an Engine Gets Subtly and Expensively Wrong

Locale enforcement is the parallel control for the convention bundle: dates, times, numbers, units, currency, and formality must match the target locale, en-US in our case, not the source locale and not some default. The moved-decimal error in segment 148 was an accuracy failure, but it was also a locale failure, because the de-DE decimal comma and the en-US decimal point invert, and an engine that does not respect the target locale will flatten one into the other. The enforcement sweep checks the file's locale dimension as a category: every date in the right format and unambiguous, every number with the right separators, every unit in the convention the target market expects, formality consistent with the style guide. Locale errors are the ones that are subtle enough to pass a flow read and expensive enough to require a re-delivery, which is exactly why they get a dedicated control rather than relying on a tired post-editor at 3:30 in the afternoon to catch every separator by eye.

The Quality Record: Provenance That Survives an Audit

The last artifact, and the one that makes the whole pipeline defensible rather than merely good, is the quality record. For every segment it captures the provenance of the work: the MT or TM source that pre-populated it, the edit the human made, the term decisions enforced, the locale dimension checked, and the error score with its category and severity. The record is what lets a client or an auditor reconstruct, line by line, exactly what happened to their file, and it is what turns the human accountability the revised ISO 18587 requires from an assertion into a demonstration.

This matters because of the cardinal rule the whole program is built on: accountability stays human, and "the engine wrote it" is never an answer when a Critical ships. The revised ISO 18587, the post-editing standard expanded to cover AI and LLM "non-human translation output," in DIS ballot with publication targeted for late 2025 into 2026, retires the rigid light-versus-full split for an effort spectrum and requires the post-editor to hold the same full professional-translator competence as a human translator, precisely because catching the fluent error in high-stakes content is a translator's judgment, not a button-pusher's reflex. ISO 5060:2024 supplies the Critical/Major/Minor scoring the gate uses. The quality record is how you stand behind both: it shows the risk tier each block received, the grounding that fired, the source-verified edits, the severity-scored gate result with zero Criticals remaining, and the terminology and locale conformance, all reconstructable. It is the artifact that lets you say the sentence a raw MT vendor can never say.

The quality record captures, per segment, the MT or TM source, the human edit, the term and locale decisions, and the error score, turning the human accountability the standard requires from an assertion into a demonstration an auditor can reconstruct.

Why Speed and Defensibility Are the Same Pass, Not a Trade-Off

Step back and look at what the file's day actually proves, because it overturns the assumption that quality is the tax you pay on speed. The naive picture is that you can have the file fast or you can have it safe, and the MTPE rate buys you fast at the cost of safe. The pipeline shows that picture is wrong. The grounding at stage two made the file faster by handing the post-editor a draft already speaking the approved language, and it made the file safer by carrying the approved terms and the previously approved TM segments into the draft. The risk-tiering at stage one made the file faster by letting the marketing block be read lightly, and it made the file safer by concentrating the slow attention on the dosing and indemnity blocks where a flip is catastrophic. The same move did both. Speed and defensibility were not in tension; they were produced by the same structure.

This is the rare win the goldmine is named for: the place where AI's biggest upside, speed and volume, collides head-on with the discipline's biggest liability, the silent critical mistranslation in regulated content, and the linguist who owns that collision with a workflow is the one whose role moves up instead of away. The engine handed you fluency, which it produces for free on every segment and guarantees on none of them for accuracy. The pipeline is how you take the fluency as a gift and supply the one thing the engine structurally cannot: a verified relationship between the smooth output and what the source actually said, scored against a typology, enforced for terminology and locale, and recorded so it can be proven. That verified relationship is not cleanup. It is the entire reason a human is in the loop, and the workflow is what makes it repeatable instead of heroic.

The Sentence That Is the Credential

At 4:15 the file ships, and here is what you can tell the client, which is the whole payoff of building the workflow instead of just editing the file: here is the throughput, here is the risk tier each content type received, here is the ISO 5060 error score with three Criticals found and zero remaining, here is the terminology conformance, here is the locale conformance, and here is the post-editing record, all defensible under the revised ISO 18587. That sentence is the credential. It is the one a raw MT vendor selling unverified machine output can never say, and it is the difference between a price you raced to the bottom on and a quality tier you can prove. The 412-segment file that arrived at 8:40 as a same-day MTPE request became, by 4:15, a fast delivery and a defensible one, because the verification was built into every handoff instead of bolted on at the end.

Key Takeaways

  • The MT-first post-editing workflow is not one act of editing but a chain of five handoffs, risk-tiered intake, grounded pre-translation, post-editing against source and TM and termbase, the severity-scored QE gate, and terminology/locale enforcement with a quality record, each carrying a verification so quality is accumulated stage by stage rather than inspected at the end.
  • Risk-tiered intake classifies content by what a fluent error would do, not by word count: the same 412-segment file held a low-consequence marketing block (light PE), a high-consequence clinical dosing block (full PE), and a high-liability indemnity clause (full PE or human-only). The tiering is what lets speed and defensibility fund each other instead of competing.
  • Grounded pre-translation feeds the engine the client's own TM, termbase, and style guide so the first draft already speaks the approved language; it shrinks the gap the human must close rather than replacing the human, and the first verification is confirming the grounding actually fired. QE scores ride along as a triage signal that routes effort, never a clearance that ships a segment.
  • Post-editing edits against the source segment, the approved termbase, and the memory, never against the smooth target, because MT output is fluent first and accurate second and a fluent error is invisible to a flow read. The walkthrough caught three Criticals in flawless English: a 2.5-to-25 mg moved decimal, an inverted adverse-event conditional, and a flipped supplier-indemnity negation.
  • The severity-scored gate scores the file against the MQM/ISO 5060 typology by category (accuracy, terminology, locale, fluency) and severity (Critical, Major, Minor), and fails the file on a single Critical no matter how clean the rest looks, because severity gates rather than averages and one lethal segment cannot be outvoted by 411 clean ones.
  • Terminology and locale enforcement run as controls across the whole pipeline, not as one-time reads: the approved term must hold on every segment regardless of how plausible a synonym reads, and dates, numbers, units, and formality must match the target locale (the de-DE decimal comma versus the en-US point is both an accuracy and a locale failure).
  • The quality record captures, per segment, the MT or TM source, the human edit, the term and locale decisions, and the error score, turning the human accountability the revised ISO 18587 requires into a demonstration an auditor can reconstruct, with ISO 5060:2024 supplying the Critical/Major/Minor scoring the gate uses.
  • Speed and defensibility are the same pass: grounding and risk-tiering each made the file both faster and safer in one move, which is the rare win where AI's upside (a hybrid workflow lifts a linguist past 5,000 words a day at 50 to 75% of human rates) and the discipline's liability (LLM medical error rates around 59% on drug names, 60% on dates, 66% on adverse events, all fluent) are owned by a workflow rather than a hero.