Risk-Tiered Intake and Routing
The batch landed in the queue at 4:42 on a Thursday, named the way these things always are: ACME_Q3_DE_mixed_47k.zip. Forty-seven thousand words, German into English, due Monday, quoted as machine-translation post-editing at the standard rate. The coordinator who accepted it did what coordinators do under a Friday deadline: glanced at the file count, saw "marketing refresh" in the client email, and pushed the whole zip into the light-post-editing lane where the engine had already pre-translated every segment and a fast linguist could clear five thousand words a day. What nobody opened was folder 09. Inside folder 09, between a banner headline and a set of product blurbs, sat eleven segments of a dosing table for an over-the-counter analgesic the client also happened to manufacture, and one segment of an indemnification clause from the distribution agreement that had been pasted into the same export by mistake. The marketing rate, the marketing deadline, and the marketing workflow were now sitting on top of a drug dose and a liability term, and not one human had looked at the content closely enough to know it. This lesson is about the control that was missing on that Thursday: the discipline of sorting content by what it can do to a person, a company, or a court before the cheap workflow ever touches it. It is called risk-tiered intake, and it is the first stage of a defensible localization pipeline.
Why Intake Is the First Control That Matters
Start with the vocabulary, because the lesson stands or falls on a handful of precise terms. Machine translation (MT) is any system that turns source text into target text with no human writing the words: the neural engine or the large language model (LLM) that pre-populates your segments before you open the file. Post-editing (PE), and its full name machine-translation post-editing (MTPE), is the workflow where a human edits that machine output instead of translating from scratch. A risk tier is a label assigned to a piece of content that says how much harm a single undetected error in it could cause, and therefore which workflow it is allowed to receive. And MT-forbidden content is the most extreme tier: content where the machine must never produce the trusted draft at all, because one fluent error in it is catastrophic and irreversible. Hold those four. We will turn each one into a working tool.
Here is the claim that organizes everything else. In an MT-first shop, the default assumption is that every file is a post-editing job, because the engine has already drafted it and the economics only work if the human is editing rather than writing. That default is correct for most content and lethal for some. The job of intake, the moment a file arrives and before anyone opens a computer-assisted-translation (CAT) tool, is to decide which kind of content this is. Intake is not paperwork. It is the single decision that determines whether a drug dose gets the workflow it needs or the workflow the marketing site got, and it is made before a single segment is edited, which is exactly why it is the first control and not the last.
The reason a misjudgment at intake is so dangerous is that every downstream control assumes the file is already in the right lane. A severity-scored quality gate, a terminology check, a careful post-editor: all of them work on the file as it sits in front of them, in the tier it was assigned. None of them can rescue content that was put in the wrong lane at the door, because they were never told to give it the scrutiny it needed. The dosing table in folder 09 will get a light pass not because the post-editor is careless but because the file is flagged "light PE" and the deadline assumes machine speed. The control that should have caught it ran out before the post-editor ever opened the file. It ran at intake, and it did not run at all.
Intake is the only control that runs before the workflow is chosen. Every control after it inherits the lane intake assigned, which is why a triage miss at the door cannot be fixed by diligence downstream.
The Asymmetry That Makes This Necessary
Risk-tiering would be unnecessary if machine output failed visibly. If the engine produced garbled, obviously broken prose on the dangerous segments, any post-editor would catch it on sight and no triage would be needed. The opposite is true, and that is the whole problem. MT and LLM output is fluent first and accurate second. The engine is structurally excellent at producing grammatical, idiomatic, confident prose, and structurally unable to verify that the prose means what the source means, because accuracy is a relationship between the target segment and the source segment and the approved terminology, not a property the engine can inspect inside its own output. A fluent error does not trip the eye. The sentence that says "take 25 mg" reads exactly as smoothly as the sentence that says "take 2.5 mg," and the post-editor scanning for flow slides past both at the same speed.
This means the danger of a piece of content is not visible in the machine output. You cannot tell a high-risk file from a low-risk one by looking at how good the translation looks, because the high-risk file looks just as good and is hiding a worse error. The danger lives in the consequence of the content, not in the quality of the draft. That is the insight that forces triage to the front of the pipeline. You have to classify content by what it can do in the world, because you cannot classify it by how the machine rendered it, because the machine renders the lethal segment and the harmless one with identical polish.
The Four-Tier Rubric: Low, Medium, High, Forbidden
A rubric is only useful if it produces a different action at each level, so we will define the four tiers by the workflow each one triggers, not by an abstract sense of importance. The scheme below is the working spirit of the ISO 18587 effort spectrum, the post-editing standard now revised to cover AI and LLM output, which retired the rigid light-versus-full split in favor of matching effort to consequence. Four tiers, four routing decisions.
Tier 1: Low Risk, Light Post-Editing
What it is. High-volume, low-consequence content where an error costs an inconvenience and the system self-corrects: product catalogue descriptions, user-interface strings with no safety role, internal knowledge-base notes, help-centre articles about non-critical features. The test. If the worst realistic outcome of an undetected fluent error is a confused reader, a support ticket, or an editor patching it next sprint, the content is low risk. The routing decision. Light post-editing. The engine pre-translates, and a fast pass fixes anything that blocks comprehension or embarrasses the brand, leaving stylistic imperfections alone because over-editing here burns the budget that justifies the workflow. This is where MTPE earns its reputation: the speed is close to free and the risk is genuinely low. The mistake to avoid is not under-serving this tier; it is letting its comfortable economics leak upward onto content that does not belong in it.
Tier 2: Medium Risk, Full Post-Editing
What it is. Customer-facing content where an error damages trust, brand, or a purchasing decision but not a body or a balance sheet: marketing pages with product specifics, support content that describes how a paid feature behaves, transactional emails, contractual-adjacent copy like a returns policy that is annoying but not litigated. The test. If an undetected error would cost reputation, a sale, or a customer-service escalation, but nobody is harmed and nothing is legally binding, the content is medium risk. The routing decision. Full post-editing. The engine still drafts, but the human edits every segment against the source and the approved termbase, not just against fluency, treating the machine output as a draft to be verified rather than a result to be tidied. Full PE costs more time per word than light PE and produces a delivery the linguist can stand behind to a client, which is exactly what medium-risk content needs.
Tier 3: High Risk, Full Human Translation
What it is. Content with real legal, financial, or safety consequence where the entire value is the verified relationship to the source: most contractual language, financial reporting, safety-relevant instructions that are not strictly life-critical, regulated marketing claims, anything a regulator may read. The test. If an undetected error could trigger a lawsuit, a financial loss, a regulatory finding, or physical risk, the content is high risk. The routing decision. Full human translation. A qualified specialist translates from the source and owns the deliverable. MT may sit beside the linguist as a reference or a productivity aid at most, but it does not produce the draft the human trusts, because on this content the human's verified judgment against the source is the product, and a machine draft is a contaminant that invites the eye to confirm rather than to verify. The cost is higher and the throughput lower, and on this content that trade is the correct one.
Tier 4: MT-Forbidden
What it is. Regulated, life-safety, and legally binding content where a single fluent error is catastrophic and irreversible, or where the law or the client contract simply prohibits machine processing: drug labels and patient leaflets, instructions for use of a medical device, dosing tables, clinical-trial and informed-consent documents, indemnity and liability clauses, sworn and certified documents, and any content a client has contractually ring-fenced from machine processing for confidentiality. The test. If one undetected fluent error reaches a body, a court, or a regulatory authority and cannot be recalled, or if processing the content by machine is itself prohibited, the content is MT-forbidden. The routing decision. The machine does not produce the trusted draft. The work goes to a qualified specialist, often with independent review or back-translation, and the highest-stakes legal instruments go to a sworn or certified translator whose signature carries legal weight. The distinction from Tier 3 is not a matter of degree but of category: in Tier 3 a machine may assist; in Tier 4 it may not touch the trusted draft at all.
The four tiers are not four levels of importance. They are four different workflows: light PE, full PE, full human translation, and machine-off. The tier is the routing decision, and assigning it is the whole job of intake.
The Rubric Is Segment-Aware, Not Just File-Aware
The Thursday batch teaches the subtlety the four tiers hide. A "file" is rarely one tier all the way through. The Q3 batch was overwhelmingly Tier 1 and Tier 2 marketing content with eleven Tier 4 segments and one Tier 4 clause buried inside it. A rubric that tags whole files and stops there will miss exactly the case that matters, because the dangerous content hides inside a benign container. Real triage therefore works at two grains: it tiers the file to set the default lane, and it scans for the higher-tier islands inside it, because the worst outcomes come not from a file that is obviously a drug label but from eleven dosing segments smuggled inside something labelled "marketing." The rule is that the tier of a file is the tier of its highest-risk segment, not the average. One Tier 4 segment makes the file a Tier 4 problem until that segment is pulled out and routed correctly, exactly as one Critical error fails an entire file regardless of how clean the rest reads.
The Signals That Classify Content at Intake
Tiering is not a feeling. It is a scan for specific signals that are visible at intake without reading the whole file, the way a triage nurse reads vital signs before knowing the diagnosis. Train yourself to look for these, because they are the operational definition of the tiers above. Any one of them pulls a file or a segment out of the comfortable default lane.
- The audience acts on it physically. If a human will swallow, inject, operate, calibrate, evacuate, or administer based on the text, it is life-safety content and at minimum high risk, usually MT-forbidden. Signals: drug names, doses, units, frequencies, contraindications, "do not," warnings, device steps.
- The text allocates legal rights or obligations. If it says who must do what, who is liable, who indemnifies whom, who warrants what, it is binding legal content. Signals: "shall," "shall not," "indemnify," "warrant," "in no event," "liable," "notwithstanding," party names, monetary caps.
- A regulator will read it. If a government agency reviews, approves, or files the document, it is a regulatory submission. Signals: marketing-authorization language, clinical-trial terminology, prospectus and disclosure language, patent claims.
- A signature certifies it. If the deliverable must carry a translator's sworn or certified attestation to be valid before a court, registry, or authority, it is certified-translation content and cannot be machine-produced. Signals: certificates, judgments, sworn statements, official records.
- The content is itself sensitive or contractually restricted. Some clients forbid machine processing for confidentiality or data-residency reasons regardless of the content's risk; some jurisdictions require a named human translator. Signals: a confidentiality clause in the client contract, personal data, unreleased financials.
Two of these signals appearing together should be treated as MT-forbidden until a qualified person proves otherwise. An informed-consent form trips three at once: a regulator reads it, it allocates legal obligations, and a human acts on it physically. The dosing table in folder 09 tripped the first signal alone, and that was enough; the indemnity clause tripped the second, and that was enough too. The discipline is not subtle reasoning about each file. It is a fast, repeatable scan for these unmistakable markers, run on every job at the door, treated as a hard stop rather than a suggestion.
Who Runs the Scan, and When
The scan has to be run by a human who understands the content, at intake, before the quote and the lane are locked. This is the move that separates a defensible shop from a careless one. The failure on Thursday was not that the post-editor was bad; it was that no qualified human looked at the content before it was classified, so the classification defaulted to the client's email subject line, which said "marketing." Risk tiers cannot be assigned by file name, by client habit, or by the deadline. They are assigned by someone who can recognize a dosing table when they see one and who has the authority to say "this batch cannot all be light PE." Where volume makes a full human read impossible, automated pre-scans for the signal words above can flag candidate segments for human review, but the flag routes to a human decision; it does not replace it. The pre-scan is a smoke detector, not a fire marshal.
A risk tier is assigned by a human who can recognize the content, at intake, before the lane is locked. A tier inferred from a file name or a client's email subject is not triage; it is a guess wearing triage's clothes.
Routing: The Decision Each Tier Triggers
Tiering is worthless without routing, and routing is just the act of sending each tier to its workflow and refusing to let a tier drift into a cheaper lane than it earned. Walk the four routes concretely, because the value of the rubric is entirely in the action it forces.
Tier 1 routes to light post-editing. The file goes into the high-throughput lane. The engine's draft is trusted on its surface and a fast pass fixes only what blocks comprehension. The quote reflects light-PE rates, the deadline assumes machine speed, and the linguist is measured on throughput because the risk genuinely permits it. The only governance here is keeping the lane clean: nothing higher-tier is allowed to ride in on Tier 1 economics.
Tier 2 routes to full post-editing. The file goes to a linguist who edits every segment against the source and the termbase. The quote reflects full-PE rates, the deadline budgets verification time, and the delivery carries a quality record the linguist can defend to the client. The engine drafts; the human verifies; the surface is never trusted on its own.
Tier 3 routes to full human translation. The file leaves the post-editing lanes entirely and goes to a qualified specialist who translates from the source and owns every word. MT, if present at all, is a reference the linguist may consult and must not lean on. The quote reflects human-translation rates and the deadline reflects human throughput, and both are correct because the verified relationship to the source is the deliverable.
Tier 4 routes to a qualified human and, where required, a sworn or certified translator, with the machine switched off for the trusted draft. This is the route that most often gets fought, because it is the most expensive and the client most wants the cheap rate. Routing it correctly means saying no to the cheap rate on this content and explaining why: a single fluent error here is irreversible, the law or the contract may prohibit machine processing, and the signature or the verified human authorship is the product the client is actually buying. The route frequently adds independent review or back-translation and reconciliation, because on this content one pass is not enough margin against the fluent-error asymmetry.
The Renegotiation Trap
The reason triage must happen at the door and not mid-file is economic as much as it is professional. Once a batch is quoted as light PE, accepted at a marketing rate, and scheduled against a Monday deadline that assumes machine speed, discovering a Tier 4 island in the middle of it forces a renegotiation nobody wants to have. The deadline is already committed, the rate is already agreed, the client already believes the job is "marketing," and the pressure on the linguist who found the dosing table is to keep going and not be the person who reopens a closed deal on a Friday afternoon. That pressure is precisely how the leaflet ships. Triage at intake is cheap because nothing is committed yet; triage in the middle of a running job is a confrontation with a deadline, a quote, and a client expectation all pulling the other way. The whole point of moving the control to the front is to make the correct decision when it is still free to make.
A Worked Intake Triage of the Thursday Batch
Take the Q3 batch through the discipline as it should have run, segment grain and all, so the rubric stops being abstract. The coordinator opens the zip and inventories the folders rather than the file name. There are nine folders. Folders 01 through 06 are website and campaign copy: headlines, product blurbs, a landing page, a set of feature descriptions. Folder 07 is an internal release note. Folder 08 is a batch of help-centre articles. Folder 09 is labelled "product_inserts," and that label alone is a signal worth a human's attention.
The triage runs the scan on each. Folders 01 through 06 and 08 trip no high-risk signal. Nobody acts on a headline physically, nothing allocates a legal obligation, no regulator reads a feature blurb, no signature certifies a help article. These are Tier 1 and Tier 2: the campaign copy with brand-specific claims goes to full PE because a wrong product claim damages trust and may even brush regulated-claim territory, and the plain catalogue and help content goes to light PE. Folder 07, the internal release note, is the lowest stakes in the batch and routes to light PE. So far the marketing lane holds, and the bulk of the forty-seven thousand words moves at speed exactly as quoted.
Folder 09 is where the triage earns its keep. The coordinator opens "product_inserts" and finds, among genuine marketing inserts, the eleven-segment dosing table for the analgesic. The first signal fires immediately: a human will swallow a dose based on this text. Numbers, units, and a frequency are present, the exact low-statistical-weight tokens an engine smooths without disturbing grammar, where 2.5 mg and 25 mg read identically fluent and only one is safe. This is not Tier 1 marketing despite living in a marketing batch. It is Tier 4, MT-forbidden, and it is pulled out of the batch, routed to a qualified medical specialist, and flagged for the independent verification that dosing content demands. The eleven segments are not post-edited at the marketing rate. They are removed from the lane entirely.
Then the triage finds the stray indemnification clause, pasted into folder 09's export by whoever assembled it. The second signal fires: it allocates liability, it contains "indemnify" and "in no event," it names parties. This is binding legal content, Tier 4, and it goes to a legal specialist, not a marketing post-editor. The fact that it arrived by accident changes nothing about the routing; the content is what it is regardless of how it was packaged.
The outcome of correct triage is not that the whole batch slows down. It is that forty-seven thousand words minus twelve segments run fast and cheap exactly as the client wanted, while the twelve segments that could kill or sue are surgically removed and sent where they belong, with their own rate, their own deadline, and their own qualified human. The client gets the speed on the content that can take it and the safety on the content that cannot, and the shop can prove which content got which workflow and why. That proof, the record of the tier each segment received and the signal that triggered it, is the artifact that survives a client audit and is the reason intake triage is a control and not a chore. Without the triage, the same batch ships with a fluent dosing error and an inverted liability term inside a delivery stamped "light PE," and the first time anyone reads the dose closely is in a pharmacy or a courtroom.
Correct triage does not make the fast work slow. It makes the dangerous twelve segments visible so they can be removed, and lets the safe forty-six thousand words run exactly as quoted. The speed and the safety are not in tension once the content is in the right lane.
Escalation When Intake Already Failed
Intake will sometimes fail anyway, and the linguist is the last human who can catch it. Suppose the triage above did not happen and you, the post-editor, open the batch already in the light-PE lane and three segments into folder 09 you realize you are editing a dosing table. The wrong move, the one the deadline and the quote and the open file all silently pressure you toward, is to post-edit it extra carefully and hope your care substitutes for the right workflow. It does not. Care is not a tier. The right move is to stop and escalate: flag to the project manager that the content is misclassified, that it belongs in Tier 4, and that the lane, the rate, and the deadline have to change before the work continues. "The engine wrote it" is never an answer when a Critical error ships, and neither is "I was told it was light PE." The accountability sits with the human who signs the delivery, and the revised ISO 18587 requires the post-editor to hold full professional-translator competence precisely so they can recognize misclassified content and refuse the wrong workflow rather than diligently executing it. Escalation feels like slowing a job down. It is the opposite: it is the one move that stops a silent critical error from shipping with your name on it.
Key Takeaways
- Risk-tiered intake is the first control in a localization pipeline: it classifies content by consequence before any segment is post-edited, because every downstream control inherits the lane intake assigns and none can rescue content put in the wrong lane at the door.
- The danger of content is not visible in the machine output. MT and LLMs are fluent first and accurate second, so a lethal segment reads exactly as smoothly as a harmless one; you must classify content by what it can do in the world, not by how good the draft looks.
- The four-tier rubric is defined by the workflow each tier triggers: low risk routes to light post-editing, medium risk to full post-editing, high risk to full human translation, and MT-forbidden to a qualified human (or a sworn translator) with the machine switched off for the trusted draft.
- Triage is segment-aware, not just file-aware: the tier of a file is the tier of its highest-risk segment, because the worst outcomes come from a few dangerous segments (a dosing table, an indemnity clause) smuggled inside a benign-looking batch.
- Five signals classify content at intake: the audience acts on it physically, the text allocates legal rights or obligations, a regulator reads it, a signature must certify it, or the content is contractually or legally restricted. Any one signal pulls the file out of the default lane; two together mean MT-forbidden until proven otherwise.
- The tier must be assigned by a human who can recognize the content, at intake, before the quote and lane are locked. A tier inferred from a file name, a client email subject, or a deadline is a guess, not triage; automated pre-scans flag candidates but route to a human decision.
- Triage must run at the door because mid-file triage forces a renegotiation against a committed deadline, an agreed rate, and a client expectation, and that pressure is exactly how dangerous content ships at the cheap rate. Correct triage at intake is free; correct triage mid-job is a confrontation.
- When intake has already failed and forbidden content surfaces mid-file, the duty is to stop and escalate, not to post-edit more carefully. Care is not a tier; the human who signs owns the delivery, and ISO 18587 requires the post-editor's full competence precisely so they can refuse the wrong workflow.
Skill.re