Where MT Is Forbidden
The arbitration ran for fourteen months, and the entire dispute turned on a single word that a machine had quietly added. A European equipment manufacturer had sold a packaging line to a buyer in another country, and the warranty section of the supply contract had been localized like everything else in the deal: the engine pre-translated every segment, a generalist post-editor cleared the file against the clock, and the contract was signed in two languages, both deemed equally authentic. The source clause limited the manufacturer's liability for consequential damages. The target clause, in fluent, lawyerly, native-sounding prose, did the opposite. A negation had been dropped in one place and an obligation inverted in another, so that the localized warranty made the manufacturer liable for exactly the losses the original had excluded. Nobody noticed for two years, because the clause read beautifully, until a conveyor failed, a production line stopped, and a lawyer finally read the foreign-language version word against word. The manufacturer's exposure ran into seven figures. The post-editor who cleared the file had done nothing lazy: they had post-edited a contract the way you post-edit a product catalogue, and that was the whole problem. This lesson is about the content that should never have reached a post-editor's desk in that workflow at all, the categories where machine translation is not a risky shortcut but a forbidden one, and how to recognize that content before a single segment is touched.
What MT-Forbidden Actually Means
Begin with the vocabulary, because the whole lesson hangs on a few precise terms. Machine translation (MT) is any system that converts text from a source language to a target language with no human writing the words: the neural engine or the large language model that pre-populates your segments before you open the file. Machine-translation post-editing (MTPE), sometimes shortened to PE, is the workflow in which a human edits that machine output rather than translating from scratch. A locale is the specific language-plus-region-plus-convention target you are producing for, French for France versus French for Canada, with its own rules for dates, units, currency, and formality. And an IFU is an instruction for use, the document, printed or digital, that tells a clinician or a patient how to use a drug or a medical device safely. Hold those four; we will use them throughout.
Now the claim. In an MT-first pipeline, the default assumption is that every segment arrives pre-translated and your job is to post-edit it. MT-forbidden content breaks that default. It is content where the correct first move is not to post-edit the machine output but to refuse the machine output entirely, route the work to full human translation by a qualified specialist, and in the highest-stakes cases route it to a sworn or certified translator whose signature carries legal weight. The machine may not be allowed to produce the draft at all, and where it is, that draft must never be the thing the human edits and trusts. This is not a quality preference. It is a category boundary.
The reason the boundary exists comes straight from the failure mode that haunts the entire field: the silent critical error. An MT or LLM engine produces output that is fluent first and accurate second. It is structurally excellent at generating grammatical, idiomatic, confident prose, and structurally unable to verify that the prose means what the source means, because accuracy is a relationship to the source segment and the approved terminology, not a property the engine can inspect inside itself. On most content, a fluent error is recoverable: someone notices, someone fixes it, the cost is rework. On MT-forbidden content, a single fluent error is not recoverable. It reaches a body, a balance sheet, or a court, and it reaches them disguised as the truth, because the prose that carries the error reads exactly as well as the prose that carries the fact.
MT-forbidden does not mean the machine produces worse prose here. It means a single fluent error here is catastrophic and irreversible, so the workflow that trusts the machine's surface is the wrong workflow no matter how good the surface looks.
The Difference Between High-Consequence and Forgiving Content
It helps to picture two extremes on a single axis and then locate everything else between them. At the forgiving end sits a product description on an e-commerce page, a help-centre article, an internal knowledge-base note. If the engine renders one of these with a clumsy phrase or even a small meaning slip, a reader is mildly inconvenienced, support handles a confused ticket, an editor patches it next sprint. The content is high-volume, low-stakes, and self-correcting, which is exactly why MTPE economics work there: the speed is a near-free gift and a light pass suffices.
At the high-consequence end sits a contraindication on a drug label, a dosing instruction in an IFU, an indemnity clause in a supply contract, an evacuation procedure on a hazard sign. If the engine renders one of these wrong, there is no confused ticket and no next sprint. A patient combines two drugs that must never be combined. A worker reads "do" where the source said "do not." A company signs away a liability it meant to exclude. The content is often lower in raw word count than the marketing site, but each word is load-bearing, and the cost of one fluent error is measured in harm, recalls, and lawsuits rather than in rework hours. The skill this lesson teaches is recognizing, at intake, which end of the axis a given file sits on, before the engine and the cheap workflow get anywhere near it.
Risk-Tiering at Intake: The First Control
The recognition has to happen at intake, which is the moment a job arrives and before anyone opens a CAT tool. This is the first control in the goldmine pipeline, and it has a name: liability triage, the discipline of classifying content by consequence before a single segment is post-edited. The mental model is a hospital emergency department. A triage nurse does not treat patients in the order they arrive; they assess severity first and route the chest-pain case ahead of the sprained ankle, because consequence, not arrival order, decides priority. Risk-tiered intake does the same to content. It asks, of every file, one question before any other: what happens in the real world if a fluent error survives into this delivery?
The answer sorts content into tiers, and the tiers determine the workflow. A workable scheme, the spirit of the ISO 18587 effort spectrum, runs like this:
- Low risk, MT-ready. High-volume, low-consequence content (product catalogues, UI strings of no safety relevance, internal documentation). MT pre-translates; a light post-edit cleans it. The engine is doing its best work here.
- Elevated risk, full post-editing. Customer-facing content where errors damage trust or brand but not bodies (websites, support content with some product specifics). MT drafts; a full post-edit against the source and the termbase is required.
- High risk, full human translation. Content with real legal, financial, or safety consequence where the value is in the verified relationship to the source. MT may assist as a reference at most, but the deliverable is produced and owned by a qualified human translator working from the source.
- MT-forbidden. Regulated, life-safety, and legally binding content where a single fluent error is catastrophic or where the law or the client contract simply prohibits machine processing. The engine does not produce the trusted draft. The work goes to a qualified specialist, and the highest-stakes cases go to a sworn or certified translator.
The crucial move, the one that separates a defensible shop from a careless one, is that the tier is assigned at intake by a human who understands the content, not discovered halfway through post-editing when a linguist stumbles on a contraindication in a file that was quoted as light PE. By then the cheap workflow is already running, the deadline already assumes machine speed, and the pressure is to keep going. Triage at the door is cheap. Triage in the middle of a file is a renegotiation nobody wants to have, which is exactly why it gets skipped, which is exactly how the leaflet and the warranty shipped.
Liability triage is an emergency department for content: assess consequence first, route by severity, and never let arrival order or a cheap quote decide whether a drug label gets the same workflow as a product catalogue.
The Signals That Flag MT-Forbidden Content
Recognition is a skill you can systematize. Certain signals, visible at intake without reading the whole file, mark content as high-risk or forbidden. Train yourself to scan for them the way the triage nurse scans vital signs:
- The audience acts on it physically. If a human will swallow, inject, operate, evacuate, or administer based on the text, it is life-safety content. Drug labels, IFUs, device instructions, hazard warnings.
- The text allocates legal rights or obligations. If it says who must do what, who is liable, who indemnifies whom, who warrants what, it is binding legal content. Contracts, warranties, terms of service, informed-consent forms.
- A regulator will read it. If a government agency reviews, approves, or files the document, it is a regulatory submission. Clinical-trial documents, marketing-authorization dossiers, financial prospectuses, patent filings.
- A signature certifies it. If the deliverable must carry a translator's sworn or certified attestation to be valid in court or before an authority, it is certified-translation content. Birth certificates, judgments, sworn statements.
- The client or the law says so. Some clients contractually forbid machine processing of their content for confidentiality reasons; some jurisdictions require a named human translator for official documents. The prohibition can be external to the content's risk entirely.
Any one of these signals is enough to pull a file out of the default MTPE lane and into a conversation about the right tier. Two of them together (a regulator-reviewed document that allocates legal obligations, say, like an informed-consent form) should be treated as MT-forbidden until a qualified person proves otherwise. The rest of this lesson walks the specific categories, the consequence of one silent error in each, and why the boundary is non-negotiable.
The Medical Categories: Labels, IFUs, Dosing, and Devices
Medical content is the canonical MT-forbidden territory, and it is worth understanding why with the actual evidence rather than as an article of faith. When researchers evaluated large-language-model output on medical content, they did not find a machine that stumbled visibly. They found one that produced clean, confident, grammatical prose and got the load-bearing facts wrong at rates no professional would tolerate: error rates of roughly 59% on drug names, roughly 60% on dates and times, and roughly 66% on adverse events, every one delivered in fluent prose with no hedge and no change in tone between the right outputs and the catastrophic ones. Sit with the categories. A drug name is the most identity-critical token in a pharmaceutical document; substitute one and you have prescribed the wrong medicine. A date governs when a dose is taken. An adverse event is the documented harm a treatment can cause, the exact content a patient reads to tell a normal side effect from an emergency. These are precisely the elements an engine optimizing for smooth prose smooths over, because they carry the least statistical weight and the most consequence.
Drug Labels, Patient Leaflets, and IFUs
A drug label, a patient information leaflet, a summary of product characteristics, and an IFU all share one property: the reader trusts them absolutely because they are the official document in the box, and that trust is the delivery mechanism for any error. A flipped negation on a contraindication tells a patient to combine two drugs that must never be combined. A corrupted dosage tells them to take ten times the safe amount or a tenth of the effective amount. A mistranslated adverse-event description tells them a stroke symptom is a side effect to wait out. The leaflet reads fluently, so the post-editor's eye, trained to read smooth prose as competence, slides past the inverted clause. The cost is not rework. It is a recall, a regulatory investigation, a reportable safety event, and in the worst case a death, every one of which has the localization vendor's delivery somewhere in the chain of accountability. This is why drug labelling is full human translation at minimum, frequently with an independent back-translation and reconciliation step, and never raw post-edited machine output trusted on its surface.
Medical-Device Instructions and Dosing
Medical-device instructions raise the same stakes with an added wrinkle: the reader is often a clinician operating equipment on a patient in real time, where a misread step is immediate. An instruction that inverts a setting, transposes a calibration figure, or drops a "not" in a sterilization step does not produce a confused support ticket; it produces a harmed patient. Dosing instructions deserve their own line because they concentrate every dangerous element into a few characters: a number, a unit, a frequency, and often a negation, all of which the engine treats as low-statistical-weight tokens it can smooth without disturbing grammar. 2.5 mg is not 25 mg. Twice daily is not twice weekly. Milligrams are not micrograms. Fluent prose around a wrong number does nothing to protect the number, and the sentence reads identically whether the dose is right or lethal. There is no version of "light post-editing" that is appropriate to a dosing instruction. The verification against the source, character by character, is the entire value, and it is human work.
Clinical-Trial Documents and Informed Consent
Clinical-trial protocols, investigator brochures, and above all informed-consent forms sit at the intersection of three of our intake signals at once: a regulator reads them, they carry legal weight, and a human acts on them physically. An informed-consent form is the document a trial participant reads to understand the risks they are agreeing to bear. If a machine fluently softens a described risk, drops a side effect, or inverts an eligibility criterion, the participant has consented to something other than what the source described, and the consent is no longer informed. The harm is both physical and legal, and the regulatory consequence can halt a trial. Content of this kind is MT-forbidden in the strict sense: the engine does not produce the trusted draft, qualified medical translators produce it, and the process typically includes independent review precisely because the asymmetry between fluent and correct is too dangerous to leave to a single pass.
In medical content the fluent error is not a quality slip; it is the official document in the box telling a patient to do the thing that will hurt them, in prose so smooth that no eye reading for flow will ever catch it.
The Legal Categories: Contracts, Liability, and Sworn Translation
If medical content is where the fluent error reaches a body, legal content is where it reaches a balance sheet, and the warranty dispute that opened this lesson is the archetype. Legal language is built from a small set of words that carry enormous weight relative to their size: "shall" and "shall not," "indemnify" and "be indemnified," "including" versus "including without limitation," "warrant," "in no event," "notwithstanding." These are exactly the low-information, high-consequence tokens an engine optimizing for fluent continuation can drop or invert without disturbing the grammar of the clause. A fluent engine that renders "the supplier shall not be liable" as "the supplier shall be liable," or that swaps which party indemnifies which, has not produced an awkward sentence. It has produced a clean, lawyerly, professional-sounding clause that allocates risk to the wrong party, and the two versions read with identical authority.
Contracts, Indemnities, and Warranties
The trap in a contract is timing. The inverted clause does no visible damage on the day it ships, because nobody reads a warranty closely when everything is going well. The clause is read closely at exactly one moment: when a dispute arises and a lawyer goes through the foreign-language version word against word, often years later, looking for advantage. That is the worst possible moment to discover that the localized version says the opposite of the original, because by then a counterparty's lawyer has found it first and is relying on it. Where two language versions are both deemed authentic, the localized error is not a translation defect to be quietly corrected; it is a binding term. An inverted indemnity clause is a fluent error wearing a suit, and it is dormant until it is expensive. Contracts, indemnities, liability and limitation clauses, warranties, and terms of service are full human translation by a legal specialist, and the highest-stakes instruments are reviewed by counsel in both languages.
Safety Warnings and Regulatory Filings
Two adjacent categories follow the same logic. A safety warning, an operating instruction for machinery, an evacuation procedure, a hazard label, reaches whoever is standing in front of the equipment when the warning that should have said "do not" said "do." The consequence is immediate and physical, and the text is often short, which paradoxically raises the risk, because a one-word file feels trivial and gets the least scrutiny. A regulatory filing, a marketing-authorization dossier, a financial prospectus, a patent application, reaches a government agency that reads it adversarially and can reject the submission, levy a penalty, or invalidate the right being claimed if a figure is transposed or a condition inverted. Both belong outside the MTPE default. The engine's speed is worthless when the cost of a single fluent error is a halted product launch or a worker's injury.
Sworn and Certified Translation
The most absolute category is sworn or certified translation, and here the prohibition is not only about consequence; it is about authority. A sworn or certified translation is one a qualified, often officially appointed, translator attests to be a complete and accurate rendering, signing and stamping it so that it carries legal validity before a court, a registry, or an immigration authority. Birth and marriage certificates, court judgments, sworn statements, academic records used for official purposes, all require this. The signature is the product. A machine cannot swear an oath, cannot be held professionally accountable, and cannot lawfully certify anything, so by definition this content cannot be machine-translated and delivered as certified. Even using MT as a hidden first draft is professionally fraught here, because the certifying translator is attesting to a rendering they personally produced and verified, not to a machine output they tidied. This is the cleanest case of MT-forbidden there is: the law itself names a human as the only valid author.
The Linguist's Duty to Escalate, Not Post-Edit
Here is the part that turns recognition into professional behaviour, and it is the most important habit in this lesson. Suppose intake failed. Suppose a file was quoted and routed as ordinary post-editing, and you, the linguist, open it and realize three segments in that you are looking at a dosing instruction, an informed-consent passage, or an indemnity clause. The wrong move, the move that the deadline and the quote and the open file all silently pressure you toward, is to keep post-editing it carefully and hope your care is enough. The right move is to stop and escalate.
Escalation means flagging to the project manager that the content has been misclassified, that it belongs in a higher tier or is MT-forbidden, and that the workflow and the quote need to change before the work continues. This feels like raising a problem, and culturally it can feel like slowing down a job everyone wants finished. It is the opposite. It is the single most valuable thing a linguist does, because the alternative is a silent critical error shipping with your name on the delivery. "The engine wrote it" is never an answer when a Critical error ships; the accountability sits with the human who signed the file. The revised ISO 18587, the post-editing standard expanded to cover AI and LLM output and in DIS ballot with publication targeted for late 2025 into 2026, makes this explicit by requiring the post-editor to hold the same full linguistic competence as a professional translator, precisely because recognizing forbidden content and refusing to post-edit it is a translator's judgment, not a button-pusher's reflex.
When you open a file and find a contraindication, a consent risk, or an indemnity clause that was routed as cheap post-editing, your job is not to post-edit it more carefully. Your job is to stop, flag the misclassification, and refuse to let the wrong workflow proceed.
Why Escalation Protects Everyone
It is worth seeing escalation not as self-protection but as the control that protects the whole chain. When you escalate a misclassified file, you protect the patient who would have read the inverted contraindication, the company that would have signed the flipped warranty, the client whose recall you prevented, the project manager who would otherwise have unknowingly released a liability, and yourself, whose name and qualification are on the line. You also protect the pipeline's integrity: every escalation is a data point that intake missed something, and a shop that listens to its linguists tightens its triage over time. A linguist who quietly post-edits forbidden content to avoid friction is not being efficient; they are absorbing a catastrophic risk on behalf of an organization that did not even ask them to. The professional, and the standard, expect the opposite. You are the last human who can see the content clearly before it ships, and seeing clearly includes seeing that it should never have reached you in this lane.
How This Connects to the Pipeline
This lesson is the awareness layer of the first stage of the goldmine pipeline, liability triage, the risk-tiered intake that decides what the machine may touch before any post-editing begins. At later levels you will build this hands-on: a risk-tiered intake-and-routing step that classifies every file by consequence, a source-to-delivery map that marks each step MT-ready, human-only, or MT-forbidden, and a high-liability routing rule that sends medical, legal, and financial content to full human translation or full post-editing by rule rather than by guess. For now, the load-bearing skill is simpler and earlier: the ability to look at a file at intake, or to recognize three segments into a post-edit, that you are holding content the machine must not own, and to act on that recognition by routing or escalating rather than by trusting the fluent surface. Everything downstream depends on this one judgment being made correctly and early, because no quality gate, however good, can fully protect content that was put in the wrong lane at the door.
Key Takeaways
- MT-forbidden content is not content the engine renders worse; it is content where a single fluent error is catastrophic and irreversible (it reaches a body, a balance sheet, or a court), so the default post-edit-the-machine workflow is the wrong workflow no matter how clean the output reads.
- The boundary comes from the field's core asymmetry: MT and LLMs are fluent first and accurate second, structurally able to guarantee grammatical prose and structurally unable to verify it means what the source means, because accuracy is a relationship to the source segment and the approved terms, not a property the engine can inspect.
- Recognition happens at intake through liability triage, the first stage of the pipeline: an emergency-department model that assesses consequence first and routes by severity into tiers (MT-ready light PE, full PE, full human translation, MT-forbidden), assigned by a human who understands the content before any segment is post-edited.
- Five intake signals flag high-risk or forbidden content: the audience acts on it physically, the text allocates legal rights or obligations, a regulator reads it, a signature must certify it, or the client or the law explicitly prohibits machine processing. Any one pulls the file out of the default MTPE lane.
- Medical content is canonical MT-forbidden territory, backed by measured LLM error rates of roughly 59% on drug names, 60% on dates and times, and 66% on adverse events, all in fluent prose: drug labels and IFUs, device instructions, dosing (2.5 mg is not 25 mg, twice daily is not twice weekly), and clinical-trial and informed-consent documents.
- Legal content is where the fluent error reaches a balance sheet: contracts, indemnities, liability and warranty clauses built from tiny high-consequence words (shall, shall not, indemnify, in no event) that an engine can invert without disturbing grammar, plus safety warnings, regulatory filings, and the absolute case of sworn or certified translation where the law names a human as the only valid author.
- The legal trap is timing: an inverted clause does no visible damage until a dispute, when a lawyer reads the foreign-language version word against word, often years later, and where both language versions are authentic the localized error is a binding term, not a defect to quietly fix.
- The linguist's duty when forbidden content surfaces mid-file is to stop and escalate, not to post-edit more carefully. "The engine wrote it" is never an answer when a Critical ships; the revised ISO 18587 requires the post-editor to hold full professional-translator competence precisely so they can recognize forbidden content and refuse the wrong workflow, protecting the patient, the client, the PM, and themselves.
Skill.re