Matching Post-Editing Effort to Risk Tier
The audit started with a sentence nobody in the room wanted to hear. A pharmaceutical client's compliance officer slid a single printed segment across the table and asked the localization PM, Dario, one question: "Who decided this got light post-editing, and where is that written down?" The segment was from a patient-information leaflet. It read smoothly in the target language, perfectly grammatical, and it told a patient to take the medication with food when the source said to take it on an empty stomach. The negation had flipped somewhere between the engine and the delivery, and nobody had caught it, because the file had been processed under a light-PE brief that told the linguist to make it readable and stop. Dario had a risk-tiered intake system. The content had been correctly classified, on paper, as Tier 1, the highest-consequence tier his shop ran. But somewhere between that classification at intake and the work that actually happened on the file, the connection broke. The drug leaflet got the cheap workflow. And when the compliance officer asked Dario to show the rule that linked the tier to the effort, and the record proving that rule was followed, Dario had a tier label and nothing else. This lesson is about the join that failed in Dario's shop: the assignment that connects a risk tier to a post-editing effort level, defensibly, in writing, so the cheap workflow can never land on the content that can kill or sue, and so that when an auditor asks who decided and why, the answer is a document and not a shrug.
The Join That Actually Fails
You learned in the previous lessons how to build the pipeline and how to classify content at intake. Risk-tiered intake is the control that sorts incoming content by consequence before any engine touches it: a marketing string, a drug label, and an indemnity clause are not the same risk and must not get the same treatment. A risk tier is a label, usually Tier 1 through Tier 3 or named bands like high, medium, and low, that records how much harm a silent error in that content could cause. That part you have. What this lesson dissects is the next link in the chain, the one Dario's shop got wrong: the rule that takes a risk tier and assigns it a specific post-editing (PE) effort level, the human work of editing machine output rather than translating from a blank page.
It helps to name the three effort levels precisely, because the whole assignment turns on the differences between them. Light post-editing (light PE) aims at a single target: understandable. The linguist fixes outright mistranslations and anything that misleads, then takes their hands off the keyboard, leaving clunky-but-accurate prose alone and making no preferential edits. Full post-editing (full PE) aims at indistinguishable from a human translation: the linguist fixes everything light PE fixes and then keeps going, correcting style, register, terminology, locale conventions, and flow, with rigorous source-against-target verification on every high-consequence element. And the third level is not a post-editing level at all: human-only translation, where the content never enters the machine translation (MT) workflow in the first place, MT being any system that renders text from one language to another with no human writing the words, whether a dedicated neural machine translation (NMT) engine or a general-purpose large language model (LLM) that translates as a side effect of predicting plausible text. Human-only means a qualified translator produces the target from the source from scratch, with no machine draft in the loop, because the consequence of a machine-introduced error is too high to risk at any effort level.
Here is the trap that swallowed Dario. The tier and the effort level feel like the same thing, so people assume the classification does double duty: classify the content and the workflow follows automatically. It does not. A tier is a statement about the content ("this is high-consequence"). An effort level is a statement about the work ("apply full PE with maximum verification, or route to human-only"). Between them sits a rule, and if the rule is unwritten, lives in one PM's head, or is overridden file by file under deadline pressure, then the tier is decoration. Dario's drug leaflet was correctly tiered and incorrectly worked, and the gap between those two facts is the exact space this lesson exists to close.
A risk tier describes the content. A post-editing effort level describes the work. The assignment rule is the bridge between them, and an unwritten bridge is the one that collapses under audit.
Why the Bridge Must Be a Rule, Not a Judgment
You might object that a competent linguist can simply look at a drug leaflet and know to verify it carefully. That is true and it is not enough, for three reasons that an audit will expose every time. First, the person who classifies content at intake is frequently not the person who post-edits it, and the only thing that travels reliably between them is the metadata on the file, not the intuition in someone's head. Second, intuition does not scale: a shop processing two hundred files a week across nine languages cannot run on every linguist independently re-deriving the correct effort from the content, because they will derive it differently, and the variance is precisely the risk. Third, and decisively, an intuition cannot be audited. When the compliance officer asks "who decided this got light PE," the answer "an experienced linguist used their judgment" is not a control, it is the absence of one. A rule that says Tier 1 always gets full PE or human-only, never light PE, with no exceptions outside this named override process can be shown, checked, and defended. The judgment of one busy person on one busy afternoon cannot. The revised ISO 18587, the standard that defines post-editing requirements and is in DIS (Draft International Standard) ballot with publication targeted for late 2025 into 2026, makes this concrete by requiring the post-editor to hold the same linguistic competence as a professional translator and by replacing the rigid two-box light-versus-full model with an effort spectrum matched to content and consequence. The spectrum gives you the freedom to assign effort precisely; the assignment rule is what turns that freedom into something you can prove.
The Assignment Rules: Tier to Effort
So let us build the bridge. The assignment rule is a small, fixed table that maps each risk tier to a default post-editing effort level and a default verification depth, plus the conditions under which the default can change. It must be written down, agreed with the client where the client cares, and applied without per-file improvisation. Here is a defensible three-tier mapping you can adapt, stated as rules rather than suggestions.
Tier 1: The Content That Can Kill or Sue
Tier 1 is regulated, life-safety, or binding-legal content: drug labels and patient leaflets, dosage and contraindication text, medical-device instructions, safety warnings on equipment, contracts and indemnity clauses, financial disclosures, and anything where a silent error harms a person or creates legal liability. The rule for Tier 1 is the strictest and the least negotiable: Tier 1 content gets full PE at maximum verification, or it gets routed to human-only translation, and it never gets light PE under any circumstance. The choice between full PE and human-only is itself governed by a sub-rule: if the client, the regulator, or your own risk appetite permits a machine draft to exist at all for this content, full PE with segment-by-segment source verification applies; if a machine draft is forbidden by regulation or contract, or if the domain is one where the consequence of any machine-introduced error is unacceptable, the content is human-only and leaves the MT pipeline entirely. The reason light PE is categorically excluded is structural, not cautious. Light PE instructs the linguist to verify for comprehension and stop. A flipped negation in a dosage instruction is perfectly comprehensible; it just means the opposite of the source. The discipline that defines light PE, do not over-edit, do not verify past understandability, is exactly the discipline that cannot catch the Tier 1 error. You do not apply a tool that is structurally blind to the failure mode that defines the tier.
The evidence for why this is not theoretical caution is blunt. Studies of LLM output on medical content found error rates of roughly 59% on drug names, roughly 60% on dates and times, and roughly 66% on adverse events, every one of those errors delivered in grammatically perfect, confident prose. If your assignment rule lets any Tier 1 content reach a light-PE brief, you have built a pipeline that processes the most dangerous content with the workflow least able to catch the most common machine errors in exactly that content. That is the shape of Dario's failure, drawn as a system.
Tier 2: The Consequential Middle
Tier 2 is consequential but not life-safety or binding-legal content: customer-facing product UI and onboarding flows, marketing with factual claims, support documentation that guides a user through a real task, help-center articles, terms-of-service summaries that are not the binding instrument itself. The rule for Tier 2 is: full PE is the default, with verification depth scaled to the specific high-consequence elements the content contains. Tier 2 is the home of the spectrum's most important move, the decoupling of polish from verification. A banking app's onboarding flow needs high surface polish because it carries the brand and is the first thing a customer sees, and it needs targeted high verification on the specific elements that bite, a permissions prompt where a flipped negation becomes a privacy breach, a placeholder the engine helpfully translated and thereby broke, a legal disclosure that must say what the source said. So the Tier 2 rule does not collapse to one effort number. It says full PE by default, and then it points the verification at named element types: negations, numbers and currency, instructions a user will act on, legal disclosures short of the binding instrument, and format integrity for placeholders and string length. Light PE can appear inside Tier 2 only by explicit exception for a clearly low-consequence sub-batch, and that exception is documented, not assumed.
Tier 3: The High-Volume, Low-Consequence Floor
Tier 3 is high-volume, low-shelf-life, low-consequence content: bulk product descriptions in a large catalog, internal knowledge-base content, user-generated content summaries, throwaway promotional copy where a stiff sentence harms nobody. The rule for Tier 3 is: light PE is the default, with pointed verification reserved for the thin islands of consequence inside otherwise harmless prose. This is the tier light PE was invented for, and on Tier 3 the cheap workflow is not reckless, it is correct, because the economic logic of the whole MT-first model depends on low-consequence content getting the cheapest pass that preserves meaning. But even Tier 3 carries small threads of risk: a material-composition claim ("flame-resistant"), a capacity figure, a compatibility statement ("fits all 2024 models"), a safety instruction on a product a child will use. The Tier 3 rule keeps polish at the floor and lifts verification on exactly those islands. The distinction the rule encodes, light polish everywhere, pointed verification on the few claims a customer could act on and be harmed or refunded over, is the difference between a Tier 3 pass that is fast, cheap, and defensible and one that is fast, cheap, and lucky.
Tier 1: full PE or human-only, never light, maximum verification. Tier 2: full PE by default, verification aimed at named high-consequence elements. Tier 3: light PE, with pointed verification on the islands of consequence. Write it as a table. Apply it without improvisation.
Documenting the Assignment So It Survives an Audit
Dario had the right tier on the file and lost the audit anyway, because a tier label by itself proves nothing about the work that happened. The assignment has to be documented as a record, not just a setting, and the difference between a setting and a record is whether you can reconstruct, after the fact, who decided what and on what basis. An auditor, a regulated client, or a courtroom does not accept "the file was Tier 1" as evidence that Tier 1 handling occurred. They want the decision trail. So the assignment record for any file should capture, at minimum, six things, and you should be able to produce them for any segment in any delivery.
- The tier and who assigned it. The risk tier the content received at intake, the person or rule that assigned it, and the date. This is the classification half of the bridge, and it must be timestamped and attributed, not just present as a dropdown value with no provenance.
- The effort level the tier produced, and the rule that produced it. Not just "full PE" but "full PE, assigned by the Tier 1 rule in the version of the assignment table in force on this date." The point is to show the effort was not a per-file judgment but the deterministic output of a written rule, so the same content would have received the same effort no matter which PM touched it.
- The verification scope that was applied. Which element types were verified source-against-target, negations, numbers, obligations, terms, dates, placeholders, and that the verification actually happened rather than being nominally required. A checklist completed per file is the cheapest form of this evidence.
- Any exception, and its authorization. If the default effort was changed for this file, the record must show what it was changed to, why, and who had the authority to approve the change. An unauthorized downgrade is precisely the failure mode that ships a Tier 1 file at light PE; the record makes that downgrade visible instead of silent.
- The error outcome. The result of the quality gate, ideally an MQM/ISO 5060 severity-scored record, MQM being the Multidimensional Quality Metrics error typology and ISO 5060:2024 the standard that formalizes its Critical, Major, and Minor severities. A clean gate with zero Criticals on a Tier 1 file is the evidence that the assigned effort did its job.
- The human who signed the delivery. Because accountability never transfers to the engine. "The engine wrote it" is not an answer when a Critical error ships, and the record names the full-competence human who owns the quality of the file.
The unifying principle is provenance: every file should answer the question "why did this content get this much human effort?" with a trail that a stranger can follow without talking to anyone. When Dario was asked "who decided this got light PE," the correct system would have answered for him: nobody decided, the rule decided, the rule says Tier 1 never gets light PE, and here is the exception record showing whether anyone overrode it and on whose authority. Dario could not produce that trail, so the absence of the record became the finding. The lesson is that the record is not paperwork bolted on after the work; the record is part of how you prove the work was the right work. A defensible assignment that is undocumented is, to an auditor, indistinguishable from no assignment at all.
Where the Record Actually Lives
This does not have to be a separate document anyone maintains by hand, and if it is, it will rot. The durable version lives in the metadata of the file as it moves through the translation-management system (TMS), the platform that orchestrates the localization workflow, so that the tier, the assigned effort, the verification checklist, the exception log, and the gate result travel with the content automatically. The goal is that the record is a byproduct of doing the work correctly, not an extra task competing with the deadline. When the assignment rule is encoded in the TMS so that classifying a file as Tier 1 mechanically sets its effort level to full PE and attaches the Tier 1 verification checklist, the record builds itself, and the human override becomes a deliberate, logged action rather than a quiet click. Dario's shop had the tier field and nothing downstream of it; the fix is to make the field do work, to wire the tier to the effort and the checklist so the bridge is structural and not a habit.
The Edge Cases That Break Naive Rules
A clean three-tier table handles the easy 80% of files. The cases that actually cause incidents are the ones where a single file does not have a single tier, or where the tier you assigned turns out to be wrong partway through the job. A rule that cannot handle these is a rule that will be quietly violated under deadline pressure, which returns you to Dario's problem by a different road.
Mixed-Tier Files
The most common edge case is the file that contains more than one tier of content. A medical-device manual is mostly Tier 2 procedural prose, get-started steps, navigation, descriptions of the interface, but it contains Tier 1 islands: the contraindications, the dosage-equivalent settings, the safety warnings, the "do not use if" conditions. A software product's release notes are mostly Tier 3 marketing, but one paragraph describes a security fix whose mistranslation has real consequence. If your assignment rule operates at the file level only, you are forced into a bad binary: tier the whole file up and pay full PE on the throwaway prose, or tier it down and apply light PE to the contraindication. Both are wrong. The first burns the budget; the second is Dario's failure exactly.
The defensible rule for mixed-tier files has two parts. First, the file inherits the highest tier of any content it contains, as a floor: a file with any Tier 1 segment is, at the file level, never eligible for a light-PE-only workflow, because the cheap pass must never be the default on a file that contains content that can kill or sue. Second, and this is what keeps the first part economical, the assignment is applied at the segment or section level wherever your tooling allows it: the Tier 1 islands get full PE and maximum verification, the Tier 2 procedural body gets full PE with targeted verification, the Tier 3 prose gets light PE. The record then documents not one effort level for the file but the effort applied to each section, with the tier that drove it. The key discipline is that splitting a file by tier is a deliberate, documented act, not a linguist silently deciding mid-file to "go easy on this part." If your tooling cannot segment effort within a file, the safe fallback is to treat the whole file at its highest tier and accept the cost, because over-paying on prose is recoverable and under-editing a contraindication is not.
Reclassification Mid-Project
The second edge case is the tier that changes after work has started. You classified a batch of help-center articles as Tier 3 at intake. Three days in, the linguist post-editing them flags that one article is actually a regulatory compliance notice that was mislabeled as general help content, or the client adds a new product whose documentation includes dosage guidance the original scope did not, or a marketing batch turns out to contain claims that legal now says are binding. The content's tier was wrong, or it changed, and some of it has already been processed under the lower-tier effort. A naive system has no answer for this and simply ships what was already done at the wrong effort, which is how a Tier 1 segment reaches delivery with a Tier 3 pass behind it.
The defensible reclassification rule has three moves. First, reclassification is a first-class event, not an informal correction: when anyone, the linguist, the PM, the client, identifies that content is in the wrong tier, that triggers a logged reclassification with a reason and an authorizer, exactly like an exception. Second, reclassification upward forces rework of everything already processed at the lower effort. If a Tier 3 batch is reclassified to Tier 1, the segments already light-PE'd do not get grandfathered in; they are re-queued for full PE and maximum verification, because the whole point of the tier is that the lower effort is structurally incapable of catching the higher tier's errors, so the prior work cannot be trusted no matter how clean it looks. This is expensive and it is non-negotiable, and the cost of rework is the argument for getting the intake classification right the first time. Third, reclassification downward is allowed but must be authorized and logged, because a downward reclassification is the most dangerous action in the system, it is the mechanism by which Tier 1 content can legitimately become a lower effort, and so it requires the strongest authorization and the clearest reason. The asymmetry is the safety property: upgrading a tier is easy and triggers rework; downgrading a tier is hard and requires sign-off. A system that makes downgrades easy is a system that will downgrade under deadline pressure, which is the original disease.
A mixed-tier file inherits its highest tier as a floor, then splits effort by section. A reclassification is a logged event: upgrading forces rework of the cheap pass; downgrading requires authority, because the easy downgrade is how the dangerous workflow reaches dangerous content.
A Worked Assignment Across a Multi-Tier Project
Abstraction earns its keep only when it changes what you do with a real queue, so let us run the rules across one project end to end, the kind that lands on a PM's desk on a Monday. The client is a medical-device company launching a new insulin pump in three European markets. The localization package contains five deliverables, and the temptation, the Dario temptation, is to give the whole package one label because it came in one purchase order. We will give it five assignments instead, and document each.
Deliverable one: the patient-facing instructions-for-use (IFU), including dosing and contraindications. This is unambiguous Tier 1: a flipped negation or a corrupted number here harms a patient and creates regulatory liability. The assignment rule produces full PE at maximum verification, and the sub-rule on machine drafts kicks in: because this is regulated medical-device labeling, you check whether the client's regulatory submission permits a machine draft to exist at all. If it does, full PE with segment-by-segment source verification on every number, negation, dose, and warning, by a linguist with medical-domain competence, scored at the gate with a hard zero-Critical requirement. If the submission forbids a machine draft, this deliverable is human-only and leaves the MT pipeline. Either way, light PE is categorically impossible here, and the record names the rule, the verification scope, the domain-qualified linguist, the gate result, and the human signer.
Deliverable two: the device's on-screen UI strings. These are mostly Tier 2, customer-facing, brand-carrying, needing high polish so the device does not feel cheap and foreign, but they contain Tier 1 islands: any string that displays a dose, a warning, an alarm message, or a "do not" instruction is life-safety content even though it sits in a UI file. So deliverable two is a mixed-tier file. The assignment: the file inherits Tier 1 as its floor (no light-PE-only workflow on this file), full PE applies across the body for brand and polish, and the dose/alarm/warning strings get the Tier 1 treatment, maximum verification, plus a format check on placeholders and string length so a correct German alarm does not overflow the pump's tiny screen and truncate the word "not." The record documents the per-section effort, not one number for the file.
Deliverable three: the quick-start guide. Procedural Tier 2: it walks a user through setup. High polish because it is customer-facing, full PE by default, verification aimed at the instructions the user will physically act on (insertion steps, priming, the sequence that matters), but no binding-legal or dosing content, so it does not carry the Tier 1 floor that deliverable two does. Full PE, targeted verification, standard gate.
Deliverable four: the marketing landing pages for the three markets. Mostly Tier 3 promotional copy where stiff prose harms nobody, with thin islands of consequence: any efficacy or safety claim ("clinically proven," "reduces hypoglycemia risk") is a regulated claim in a medical context and is actually Tier 1 hiding in marketing. So this is another mixed-tier file. The body gets light PE; the claims get pulled up to Tier 1 treatment with full verification and, because medical claims are involved, very likely a legal and regulatory review outside the linguistic workflow entirely. The discipline here is to resist tiering the whole landing page down to Tier 3 because "it's just marketing," which is precisely how a regulated efficacy claim ships with a light pass.
Deliverable five: the bulk catalog entries for the company's accessory store, replacement reservoirs, carrying cases, skin adhesives. Genuine Tier 3: high-volume, low-consequence, where the cheap workflow is correct. Light PE, pointed verification only on the few specs that carry a consequence (sterility claims, sizing, compatibility with the pump model), and the budget you saved here is the budget that funds the maximum-verification full PE on deliverable one. That is the whole economic point of tiering: you spend the most expensive resource in the pipeline, careful human attention, on the content that earns it, and you stop spending it where it does not.
Now watch what the reclassification rule does mid-project. On day four, the linguist on deliverable four flags that one landing-page section describes off-label use in a way the client's regulatory team has not cleared. That is a logged reclassification: the section jumps from the Tier 3 body to Tier 1, the light-PE pass already done on it is discarded and re-queued for full PE plus regulatory review, and the reason and authorizer are recorded. Nobody quietly "fixes it up." The event is visible, the rework happens, and six months later when a regulator asks how that claim was handled, the record answers. That visible, logged, rule-driven trail across all five deliverables is the exact thing Dario could not produce, and producing it is the difference between owning the quality and hoping for it.
The Failure Modes the Rule Is Built to Prevent
It is worth naming, sharply, the specific ways an assignment system fails, because each one is a clause in the rule you now know how to write. Recognizing the failure coming is most of the skill.
The silent downgrade. A Tier 1 file gets light PE because a PM under deadline pressure relabeled it, or because the tier field was set but never wired to the effort, or because a linguist decided the content "looked fine." This is Dario's failure and the most dangerous one, because the content most able to harm gets the workflow least able to catch the harm. The defenses are the categorical exclusion (Tier 1 never gets light PE, full stop), the authorized-and-logged requirement on any downward move, and the structural wiring of tier to effort in the TMS so the downgrade cannot be a quiet click.
The blanket upgrade. The opposite failure, where a nervous shop tiers everything up to be safe and full-PEs the entire catalog, is not safe, it is broke. It throws away the throughput that makes the MT-first model viable, lifting a linguist from roughly 2,000 words a day toward 5,000 or more only if light content actually gets light effort. A shop that full-edits Tier 3 has quietly converted its hybrid workflow back into manual translation and given away the economics. The defense is the discipline to assign Tier 3 honestly and let the cheap pass be cheap.
The undocumented assignment. The effort was correct but unprovable, because no record links the tier to the effort to the verification to the signer. This passes until the audit, and then it fails the way Dario failed, with correct work and no evidence of it. The defense is provenance: the record is a byproduct of the workflow, not an afterthought.
The file-level blindness. The system tiers whole files and cannot see the Tier 1 island inside the Tier 2 manual or the regulated claim inside the Tier 3 landing page, so it ships the dangerous segment under the file's average effort. The defense is the highest-tier-floor rule plus section-level assignment wherever tooling allows, and the willingness to split a file deliberately rather than average across it.
Every one of these is a way of breaking the bridge between the tier and the effort. The assignment rule, written down, wired into the tooling, and recorded per file, is the single control that holds all four shut. It is not glamorous. It is a small table and a logged trail. But it is the difference between a localization shop that can stand in front of a compliance officer and answer "who decided, and why" with a document, and a shop that, like Dario's, has a tier on the file and a flipped negation in a patient's hands.
Key Takeaways
- The dangerous gap is not classifying content; it is the unwritten join between the risk tier and the post-editing effort level. A risk tier describes the content's consequence; a PE effort level (light PE, full PE, or human-only) describes the work. The assignment rule is the bridge, and an unwritten bridge collapses under audit because intuition cannot be shown, scaled, or defended.
- The defensible mapping: Tier 1 (regulated, life-safety, binding-legal) gets full PE at maximum verification or human-only, and never light PE, because light PE verifies for comprehension and a flipped negation is perfectly comprehensible. Tier 2 (consequential, customer-facing) gets full PE by default with verification aimed at named high-consequence elements. Tier 3 (high-volume, low-consequence) gets light PE with pointed verification on the thin islands of consequence.
- Light PE is categorically excluded from Tier 1 for a structural reason, not caution: it is blind to exactly the failure mode that defines the tier, evidenced by LLM medical error rates of roughly 59% on drug names, 60% on dates and times, and 66% on adverse events, all in fluent prose.
- Document the assignment as a record, not a setting: the tier and who assigned it, the effort level and the rule that produced it, the verification scope applied, any exception and its authorization, the MQM/ISO 5060 gate result, and the human who signed. Provenance means a stranger can answer "why did this content get this much effort?" without talking to anyone. An undocumented-but-correct assignment is, to an auditor, indistinguishable from no assignment.
- Wire the rule into the TMS so classifying a file as Tier 1 mechanically sets full PE and attaches the verification checklist, making the record a byproduct of the work and the downgrade a deliberate logged action rather than a quiet click.
- Mixed-tier files inherit their highest tier as a floor (no light-PE-only workflow on a file containing Tier 1 content), then split effort by section so the Tier 1 islands get full PE while the Tier 3 body stays cheap. If tooling cannot segment effort, treat the whole file at its highest tier, because over-paying on prose is recoverable and under-editing a contraindication is not.
- Reclassification is a logged, authorized event: upgrading a tier forces rework of everything already processed at the lower effort (no grandfathering the cheap pass), while downgrading requires the strongest authorization because the easy downgrade is the exact mechanism that lands the dangerous workflow on dangerous content.
- The four failure modes the rule holds shut are the silent downgrade (Tier 1 at light PE), the blanket upgrade (full-PE everything, killing the 2,000-to-5,000-words-a-day throughput that makes MT-first viable), the undocumented assignment (correct work, no evidence), and file-level blindness (shipping a Tier 1 island under a file's average effort). Each is a broken bridge between tier and effort, and the written, wired, recorded assignment rule is the single control that prevents all four.
Skill.re