AI for Translation & Localization
Proficient · M3 · lesson 3 of 19 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Catching the Silent Critical Error Systematically
📖
now learning

Catching the Silent Critical Error Systematically

15 min

The recall notice arrived on a Tuesday, and it ran to four languages because the error had shipped in four languages. A mid-size medical-device firm had localized the instructions for use of a home injection pen across eleven markets, MT-first, full post-editing on the regulated content, a competent vendor pool, a clean delivery record. In the Brazilian Portuguese file, one segment of the dosing table had carried the source phrase "do not exceed 40 units in a single injection" into a fluent, natural, grammatically immaculate target that said, in effect, "administer 40 units in a single injection." The engine had dissolved the prohibition into an instruction, the post-editor had read the Portuguese, found it flawless, and confirmed it, and the reviewer had spot-checked the file the way reviewers spot-check, reading for the rough cells and trusting the smooth ones. The segment was the smoothest cell in the file. It took a pharmacist in São Paulo, reading the leaflet against the English original a patient had also been handed, to notice that the two documents told the patient opposite things about the most dangerous number on the page. Nobody on the localization chain had been careless. Every one of them had been skilled, fast, and looking. What none of them had was a procedure that forced a polarity check on that exact segment regardless of how good the Portuguese looked. This lesson is about that procedure: the small, ruthless, by-design scan you run on four specific categories of content, on every file, not because you suspect a problem but precisely because you cannot suspect this kind of problem, because it does not announce itself. The four categories are negations, numbers, dosages, and obligations, and they are where the Critical errors live.

Why These Four Categories Carry the Most Critical Risk

Start with the word that governs everything in this lesson. A Critical error, in the MQM (Multidimensional Quality Metrics) and ISO 5060:2024 framework that decides whether a localized file ships, is an error that can cause real harm or render the content unfit for its purpose: a flipped dosage, an inverted contractual obligation, a reversed safety instruction, a corrupted financial figure. It is the severity class that, under the one-Critical-fails rule, blocks delivery of an entire file no matter how clean the rest is. Most errors a post-editor catches are Minor (a slightly awkward phrasing) or Major (a real fidelity problem short of catastrophe). The Critical is rarer, and it is the only one that ends relationships, triggers recalls, and shows up in litigation. The whole economic and professional weight of post-editing in 2026 rests on one asymmetry: you can ship a hundred Minor errors and survive, you cannot ship one Critical and survive, and the Critical is the one engineered to be invisible.

Now ask the question this lesson exists to answer: if Critical errors can in principle appear anywhere, why scan four specific categories rather than everything with equal suspicion? Because Critical errors are not evenly distributed. They cluster, heavily and predictably, in content that carries a small set of semantic features, and those features are the four categories. The reason is structural, not accidental, and worth understanding rather than memorizing, because understanding it is what lets you recognize a fifth high-risk element when your domain produces one.

The Feature They Share: Meaning That Inverts Without a Seam

What negations, numbers, dosages, and obligations have in common is that each carries meaning that can flip to its opposite, or to a dangerously wrong value, while the surrounding sentence stays perfectly grammatical. This is the precise property that defeats the human eye and the precise property an MT or LLM (large language model) engine, being a fluency machine first and an accuracy machine second, is most prone to corrupt. Consider what makes ordinary prose self-correcting. If an engine mistranslates a verb in a descriptive sentence, the result usually reads slightly off, a collocation no native speaker would choose, a register that wobbles, a seam your trained eye catches. The fluency itself is the alarm. But the four high-risk categories have no seam to expose, because the corruption happens at the level of a single token or a single logical operator, and the engine fluently rebuilds the sentence around the change.

Take a negation, which we define broadly as any element that establishes polarity, whether to forbid, permit, deny, or condition: the words "not," "never," "no," "unless," "except," "without," "fails to," and the grammatical constructions that carry the same logical force. When an engine drops a "not" or collapses an "unless," it does not leave a hole. It produces a complete, natural sentence that happens to mean the reverse. "Do not exceed two tablets" becomes "Take two tablets," and the target is internally consistent, fluent, and confident. There is nothing for fluency-based detection to catch, because the target's fluency is real; it is the meaning that is wrong, and meaning is exactly what reading the target alone cannot verify, because the target is a coherent instruction that announces nothing about the source having said the opposite.

Take a number, the rare element supposed to survive translation completely unchanged: "0.075%" is "0.075%" in every language, a language-independent fact the translation must carry across intact. This makes numbers verifiable without deep target-language competence, but it also makes their corruption invisible to a reader, because a transposed digit, a vanished decimal point, or a swapped unit produces a number that is still perfectly plausible. "25 mg" reads exactly as confidently as "2.5 mg." The sentence does not wobble. The believability of the wrong value is what kills you: a corrupted number that read as obviously absurd would be safe, and the dangerous ones land inside the plausible range.

A dosage is the most lethal special case of a number, and it earns its own category because it combines the invisibility of a corrupted number with the highest possible consequence and a uniquely treacherous set of failure shapes. A dosage is a quantity of a substance bound to a unit, a frequency, and often a route and a maximum: "2.5 mg, twice daily, not to exceed 10 mg in 24 hours." Every element in that string is a separate corruption opportunity, and several combine the number failure with the negation failure (the "not to exceed" is a polarity element guarding a number). When an engine flips the decimal, swaps mg for ml, drops the frequency, or dissolves the maximum, the result is a fluent dosing instruction off by a factor of ten or inverted at the ceiling, and it reads as authoritative as the correct one. The medical-content error rates the industry keeps citing, around 59% on drug names, 60% on dates and times, 66% on adverse events, all in grammatically perfect prose, are the empirical shadow of exactly this: the categories where the engine fails most are the categories where the failure is least visible.

An obligation is the legal analog of a dosage, and it inverts with the same seamlessness. An obligation is any element that establishes who must, may, or must not do something: the modal verbs "shall," "must," "may," "will," "may not," and the deontic constructions that carry legal force. The difference between "the Supplier shall indemnify the Client" and "the Supplier may indemnify the Client" is the difference between a binding duty and an optional courtesy, and an engine that renders "shall" as a permissive form, or flips the party an obligation binds, produces a clause that reads like flawless legal prose and reallocates millions in liability. Obligations also carry a polarity dimension: "shall not disclose" guards a prohibition, and dropping the "not" inverts a confidentiality duty into a disclosure permission. Like dosages, obligations are a compound category, a meeting place of polarity, modality, and party assignment, each independently corruptible and none visible by reading the fluent target.

You scan four categories because Critical errors cluster where meaning can invert without leaving a grammatical seam, and negations, numbers, dosages, and obligations are precisely the elements whose corruption an engine hides inside perfect prose.

Why Not Just Scan Everything Equally

The obvious objection is that if Critical errors are so dangerous, you should scrutinize every element of every segment with maximum intensity. The objection fails for the same reason "just be more careful" fails: attention is finite, fade-prone, and competitively allocated, and spreading it uniformly across a file means spreading it thin, which means the high-risk elements get the same diluted glance as the throwaway prose. The discipline of the systematic scan is not more attention; it is concentrated attention, deliberately withdrawn from where Criticals do not live and aimed at where they do. You are not lowering your guard on the rest of the file; your normal post-editing pass already covers fluency, terminology drift, locale, and the rest. The systematic scan is a targeted layer added on top, a narrower pass whose entire job is to find and verify the four categories, by design, on every file, regardless of how it looks.

This is the move that separates this lesson from general post-editing. Normal post-editing is broad and shallow-to-medium across the whole file: you read every segment, fix what is rough, confirm what is clean. The systematic high-risk scan is narrow and deep on four categories: you ignore everything that is not a negation, number, dosage, or obligation, and verify every instance of those four against the source as a separate, ruthless pass with its own attention. The two passes catch different things. The broad pass catches the Major and Minor errors and the obvious Criticals. The targeted scan catches the silent Critical, the one the broad pass skims past precisely because it reads perfectly. You need both, and you must not let the targeted scan dissolve into the broad one, because the moment you check polarity while smoothing style, the polarity check loses to the fluency it is supposed to distrust.

The Systematic Scan as a Procedure, Not a Habit

The difference between a linguist who occasionally catches the silent Critical and one who catches it by design is the difference between a habit and a procedure. A habit is reliable until the one time it is not: the late Thursday, the file you trusted because the vendor is good, the segment that looked so clean you did not bother. A procedure does not depend on you remembering to perform it or feeling sharp enough to perform it well; it is a defined set of steps you run the same way every time, on every applicable file. The systematic scan is a procedure, and like every good procedure it has four properties: a trigger, a scope, an order, and an output.

The Trigger and the Scope

The trigger is the rule that decides the scan runs at all, and the correct rule for the four high-risk categories is the strongest possible: it always runs, on every file, every time, with no exception for how good the file looks, how trusted the source, or how tight the deadline. This absoluteness is the entire point. A procedure beats a habit because it removes the decision of whether to check, and that decision is exactly where the silent Critical gets through, because the files that carry it are disproportionately the ones that looked clean enough to skip the check. A scan you run only when suspicious is a scan you do not run on the dangerous file, because the dangerous file is the one that did not make you suspicious. Treat the four-category scan the way aviation treats the killer items on a pre-flight checklist: run without exception, every time, regardless of how routine the flight feels, because the cost of the rare miss is catastrophic and the cost of the check is small.

The scope is the rule that decides which segments the scan touches, and here the procedure earns back the time the trigger spent. The scan does not deeply verify every segment; it verifies every instance of the four categories, and most segments contain none. The first move is therefore an isolation step: sweep the file and flag every segment that contains a polarity operator, a number, a dosage element, or an obligation, and ignore the rest for this pass. On a typical file this isolates a minority of segments and concentrates the deep verification exactly there. This is why the systematic scan is affordable even under MTPE (machine-translation post-editing) throughput pressure: it is not a second full read of the file but a deep verification of the fraction where Criticals live, and the isolation step is largely mechanical and partly automatable.

The Order and the Output

The order matters because the four categories are not equally fast or equally deadly, and the procedure should front-load the checks that are cheapest to run and most catastrophic to miss. Numbers and dosages go first because they are the most mechanical: you reconcile a digit against a digit without reading the sentence, so the check is fast, language-light, and partly handed to a tool. Negations and obligations come next because they require a semantic judgment, reducing each operator to its polarity or modal force and comparing source to target, which is slower and cannot be fully automated. Running the mechanical checks first anchors the scan before fatigue sets in and clears the fast-to-verify elements, leaving concentrated attention for the slower semantic checks.

The output is what makes the scan defensible rather than merely diligent, and it is the property that most distinguishes a procedure from a habit. The scan does not end with a feeling that the file is clean; it ends with a record. Every high-risk element it isolated produces a logged line: the segment, the category, the source value, the target value, the verdict (match or defect), and, for defects, the severity. A scan that found nothing produces a record that it found nothing, which is itself evidence the verification happened. This record is the artifact the whole program keeps returning to: the severity-scored quality record that turns "I checked the numbers" into "here is every number I checked, source and target value for each, and the one that did not match." It is what a client audit reads, what a reviewer inherits, and what stands between your name on the delivery and the recall notice.

A habit catches the silent Critical on a good day. A procedure catches it every day, because it runs on every file regardless of how the file looks, isolates only the high-risk elements, verifies them in a fixed order, and produces a record that proves what it checked.

The Four-Category Scan in Detail

Now the substance of each category: the failure shapes the engine produces, and the fast technique for catching them. These are the four checks that constitute the systematic scan, stated as actions you perform against the source rather than goals you hold in mind.

Scanning Numbers

The number check is the most mechanical and therefore the one to run first and to lean on hardest, because it is the rare Critical-class verification a tool and a non-fluent reader can largely perform. The action is: isolate every number and unit in the source segment, find its twin in the target, and compare them as bare tokens, ignoring the surrounding prose. Do not read the number in context, because reading the sentence invites your brain to process the meaning and skim the digit; instead list the source numbers, list the target numbers, and line them up like a clerk reconciling an invoice. The failure shapes are specific and worth holding in mind so you recognize them: a transposed or altered digit (0.075 becoming 0.75, a tenfold corruption that reads perfectly), a vanished or added decimal point (2.5 becoming 25), a silently swapped unit (ml becoming mg, a one-letter shift that turns a volume into a mass), flipped decimal and thousands separators across locales (the European "1.000" meaning one thousand mistaken for the English one point zero), and invented precision where the source's hedged "about 30%" hardens into a false "30.0%." The strength of the bare-token technique is exactly that it bypasses fluency: a bilingual ten-year-old could line up the digits, which is the point, because the digit comparison is the one Critical check that does not require the expert judgment fluency fools.

Scanning Dosages

A dosage is a number wearing armor, and the dosage check is the number check plus a structured decomposition, because a dosage is not one value but a bundle. The action is: decompose every dosage in the source into its components, amount, unit, frequency, route, and maximum, and verify each component independently against the target. A dosage like "2.5 mg, subcutaneously, twice daily, not to exceed 10 mg per day" is five separate Critical opportunities: the amount (2.5), the unit (mg), the route (subcutaneous), the frequency (twice daily), and the ceiling (10 mg per day, guarded by the polarity operator "not to exceed"). Verify them one at a time, because the engine can carry four of them perfectly and corrupt the fifth, and the fifth is the one that hospitalizes someone. The dosage check is where the number check and the negation check meet, because the maximum is a number guarded by a prohibition, and dropping the "not to exceed" while keeping the number turns a ceiling into a target. Treat every dosage as the highest-consequence element in the file: it is the single place where a corruption that reads as confident, plausible prose can directly cause physical harm, which is why dosages get the structured decomposition rather than the quick token reconciliation a plain number gets.

Scanning Negations and Polarity

The negation check is the deadliest in the sense that its failure leaves the smallest trace, because a dropped negation leaves no gap at all; the engine fluently rebuilds the sentence around the absence. The action is: reduce every segment carrying a logical operator to its polarity skeleton, forbid, permit, require, or condition, and confirm the source and target agree. You cannot catch an inverted polarity by reading the target alone, because the target is internally consistent: "Take two tablets every twenty-four hours" is a coherent instruction that announces nothing about the source having said "do not exceed two." The only reliable method is to read the source first, register its polarity as an abstract logical fact (this segment forbids), then read the target and confirm it carries the same polarity (it must also forbid). The polarity family is wider than the word "not," and the scan must recognize all its members: the vanished prohibition ("do not" becoming a neutral instruction), the flipped conditional ("unless" becoming "when," which turns an exception into a trigger), the double-negative collapse ("not uncommon" becoming "common," a subtle shift of degree), the scope error where the negation lands on the wrong clause, and the implicit negation carried by words like "except," "without," "fails to," and "absent." Run this check on every segment that contains any polarity operator, every time, with no exception; it is the killer item among killer items.

Scanning Obligations and Modality

The obligation check is the negation check's legal cousin, and it verifies three things at once because an obligation has three corruptible dimensions: its modal force, its polarity, and the party it binds. The action is: for every segment carrying a deontic operator, identify the modal force (must, may, must not), the polarity, and the bound party, and confirm all three survive into the target. The modal force is the strength of the duty: "shall" and "must" create a binding obligation, "may" creates a permission, and an engine that renders "shall" with a permissive form, common when the target's modal system does not map cleanly onto English, downgrades a duty to an option. The polarity is the same dimension the negation check tracks, because "shall not compete" is an obligation guarding a prohibition and dropping the "not" inverts a non-compete into a license to compete. The bound party is who carries the duty: "the Supplier shall indemnify the Client" and "the Client shall indemnify the Supplier" differ only in word order in some languages, and an engine can swap them while producing flawless legal prose, reallocating the entire risk of the contract. Reduce each obligation to a triple, who, must-or-may-or-must-not, do-what, and confirm the triple matches the source. On legal content, this check is a killer item, and the modal verbs (shall, must, may, will) deserve their own dedicated line because a flipped "shall" to "may" is a Critical that costs a contract.

Automated Assists and Their Hard Limits

A reasonable hope at this point is that the whole scan could be handed to a tool. Part of it can, and you should automate every part that automates well, because every check you offload from attention is attention freed for the checks only judgment can make. But the boundary between what tooling does and what the human must do is sharp, principled, and worth understanding exactly, because misunderstanding it in either direction is dangerous: over-trusting the tool ships the Critical it cannot see, and under-trusting it wastes human attention on clerical work a machine does better.

What the Machine Catches: Presence, Form, and Consistency

Tooling excels at the mechanical, form-level checks, the ones that compare tokens rather than meaning. A QA filter in a CAT tool (computer-assisted translation environment) or a TMS (translation-management system, the platform that orchestrates the localization workflow) can be configured to do several of the number-check's jobs automatically and reliably. It can verify number presence and consistency: extract every numeric token from the source and confirm the same set appears in the target, flagging any segment where a number was added, dropped, or altered. This single automated check catches a large fraction of the transposed-digit and vanished-decimal failures, because most of them change the set of numeric tokens. A regex-based check (a pattern-matching rule) can flag unit and separator anomalies: a segment where the source has "mg" and the target has "ml," or where decimal and thousands separators do not match the target locale. Tooling can flag negation-word presence: a configurable check that counts polarity words (not, never, no, unless, and their target-language equivalents) in source and target and flags when the counts diverge, a crude but useful signal that a "not" may have been dropped. And it can enforce placeholder and tag integrity and length-ratio outliers, both relevant because an omission that drops an obligation or a caveat often shows up as a length anomaly. Configure every one of these. They are fast, tireless, and run on every segment without fatigue, and they convert a portion of the four-category scan from a human task into an automatic flag the human merely adjudicates.

What the Machine Cannot Do: Judge Whether Meaning Inverted

Here is the hard limit, and it is not a temporary limitation of current tools but a structural boundary. The automated checks verify form, the presence, count, and shape of tokens. They cannot reliably verify meaning, whether the polarity actually inverted, whether the modal force actually downgraded, whether the party an obligation binds actually swapped. Consider the failure modes the form checks miss. A negation-word count check flags a segment where the count of "not" diverges, but it is silent when the engine inverts polarity without changing the count: rendering "unless the indicator shows red" as "when the indicator shows red" keeps zero explicit negations on both sides and flips the conditional, and no token-counting tool sees it, because the inversion is semantic, not lexical. The number-presence check confirms 2.5 and mg both appear in the target, but it cannot tell you that the engine attached the 2.5 to the wrong drug in a two-drug segment, or that the "twice daily" frequency silently became "once daily," when "once" and "twice" are single content words a presence check does not police. The obligation triple, who-must-do-what, is entirely beyond form checking: swapping the Supplier and the Client in an indemnity clause changes no token counts, no numbers, no negation words, and produces flawless prose, and only a human reading the clause against the source and asking "who carries this duty" catches it.

The deeper statement of the limit is this: the automated assists reduce the haystack, they do not find the needle. They flag the segments where a form-level anomaly suggests a problem, directing human attention, and they clear the segments where the form matches, letting the human skip clerical confirmation. But the residue, the polarity inversion that kept the token count, the frequency that flipped between two valid words, the party swap that changed nothing countable, is exactly the silent Critical, and it lands in the human's lap by construction. This is the operational meaning of the rule the revised ISO 18587 (the post-editing standard, in DIS ballot with publication targeted into 2026, which now covers AI and LLM output and insists the post-editor hold the full competence of a professional translator) encodes when it requires the post-editor to hold full professional-translator competence: the human in the loop is not there to do faster what the tool does, but to do at all what the tool cannot, the meaning-level verification of the four categories and the severity judgment that follows. Automate the form checks, trust them for what they cover, and never mistake a clean automated pass for a verified file, because the automated pass is silent on precisely the errors that matter most.

Automated checks reduce the haystack; they never find the needle. They verify the presence, count, and shape of tokens, and they are structurally blind to whether meaning inverted, which is exactly where the silent Critical lives and exactly what the human must verify.

A Worked Systematic Scan of a File

Abstraction is comfortable and useless until you run it on a real file, so let us run the four-category scan, in full, on a concrete batch. Picture a ten-segment file: patient-and-prescriber-facing content for the same injection pen from the opening, English source, Brazilian Portuguese target, pre-translated by an LLM, full post-editing on regulated content, with a loaded termbase, a brief specifying the formal register and the pt-BR locale, and a deadline. Assume the normal post-editing pass is done, fluency edited, terminology checked, locale verified, and we are now running the targeted high-risk scan as a separate, ruthless layer on top. The automated QA filter has already flagged what it can. We isolate the high-risk segments, then verify them in order: numbers, dosages, negations, obligations.

The isolation step. Sweeping for the four categories, segments 2, 3, 4, 6, 7, 8, and 9 contain at least one high-risk element; segments 1, 5, and 10 contain none and are excluded from this pass (handled by the normal post-editing pass). Seven of ten segments carry a high-risk element, high because this is regulated dosing-and-terms content; on a marketing file it would be one or two. The automated filter has raised three flags: a number-set divergence on segment 3, a separator anomaly on segment 6, and a length-ratio outlier on segment 4. We verify all seven regardless, because the automated flags reduce the haystack but do not define it.

Running the Numbers and Dosages

Segment 3 (number, flagged). Source: "Each pen contains 0.075 mg per unit." Bare tokens, source: 0.075, mg. Bare tokens, target: 0.75, mg. The decimal shifted one place; the engine produced a tenfold corruption that reads as a perfectly believable concentration. The automated filter flagged this as a number-set divergence, correctly, and the human confirms it is a real corruption rather than a locale-format artifact. This is structurally the fund-fee error from the program's earlier lessons and the dosing error from this lesson's opening, the same vanished-decimal failure, the same tenfold magnitude, the same fluent plausibility. Logged: Segment 3, number, source 0.075 mg, target 0.75 mg, decimal shifted, Critical candidate.

Segment 6 (number, flagged). Source: "Store between 2 and 8 degrees Celsius." Target: "Armazenar entre 2 e 8 graus Celsius." Bare tokens match: 2, 8. The separator flag the filter raised turns out to be a false positive, the filter mistook the range conjunction for a separator anomaly. Numbers verified clean. Logged: Segment 6, number, 2 and 8 match, pass. Recording the pass matters as much as recording the defect, because the record must show the check ran.

Segment 8 (dosage). Source: "Inject 2.5 mg subcutaneously, twice daily, not to exceed 10 mg in any 24-hour period." Decompose: amount 2.5, unit mg, route subcutaneous, frequency twice daily, ceiling 10 mg per 24 hours guarded by "not to exceed." Verify each. Amount 2.5, match. Unit mg, match. Route subcutaneous, match. Frequency: the target reads "uma vez ao dia," once daily. The engine halved the frequency, "twice" became "once," a single-content-word flip the automated number-presence check is blind to because "once" and "twice" carry no numeral it tokenizes. Ceiling: the target carries "10 mg" but its guarding polarity, examined under the negation check below, is also corrupted. The decomposition caught the frequency flip the form check missed. Logged: Segment 8, dosage, frequency source "twice daily" target "once daily," Critical candidate; ceiling deferred to negation check.

Running the Negations and Obligations

Segment 2 (negation). Source: "Do not use the pen if the solution is cloudy." Polarity skeleton, source: forbid (do not use, under a condition). Reduce the target, "Use a caneta se a solução estiver turva," to permit-under-condition, "use the pen if the solution is cloudy." The "não" dropped; the engine fluently rebuilt a coherent instruction that inverts a safety prohibition into a safety hazard. The Portuguese is flawless, and the target read alone is a perfectly sensible instruction. Caught only by registering the source polarity first. Logged: Segment 2, negation, source forbid, target permit, dropped negation inverting safety instruction, Critical candidate.

Segment 8 ceiling (negation, continued). Returning to the dosage ceiling: source "not to exceed 10 mg" is a number (10 mg) guarded by a prohibition (not to exceed). Reduce the target's guarding polarity: it reads "administrar 10 mg," administer 10 mg, the "not to exceed" dissolved. The ceiling became a target. The engine turned a maximum into an instruction, so the corrupted segment now reads "inject once daily, administer 10 mg," compounding the frequency flip with a ceiling inversion. This is the exact compound failure that makes dosages their own category: a number, a frequency, and a polarity-guarded maximum, three corruptions in one segment, none visible in the fluent target. Logged: Segment 8, negation, ceiling "not to exceed 10 mg" became "administer 10 mg," prohibition dropped, Critical candidate.

Segment 7 (negation). Source: "Replace the needle after each injection unless the device indicates otherwise." Polarity: default replace, with an "unless" exception. Reduce the target: "Substitua a agulha após cada injeção quando o dispositivo indicar," replace the needle when the device indicates, which collapses the "unless" exception into a "when" trigger and changes "always replace, except when told not to" into "only replace when told to." The conditional flipped. The number-and-token level is clean (no numerals corrupted), and a form check sees nothing, because the inversion is purely semantic. Logged: Segment 7, negation, "unless" collapsed to "when," conditional inverted, Critical candidate.

Segment 4 (obligation, flagged for length). Source: "The prescriber must confirm the patient's renal function before the first dose." Obligation triple, source: who = prescriber, force = must (binding), do-what = confirm renal function before first dose. Target reduces to: who = prescriber, force = "pode" (may, permissive), do-what = confirm renal function. The modal force downgraded from a binding "must" to a permissive "may," turning a mandatory safety step into an optional one. The length flag the filter raised was a near-miss signal, the permissive construction ran shorter, but the flag did not and could not diagnose the modal downgrade; only the obligation-triple check did. Logged: Segment 4, obligation, force source "must" target "may," binding duty downgraded to permission, Critical candidate.

Segment 9 (obligation). Source: "The clinic shall retain these records and shall not disclose them to third parties." Two obligations. First triple: clinic, shall (binding), retain records, target matches, pass. Second triple: clinic, shall not (binding prohibition), disclose to third parties. Target reduces to: clinic, "pode" (may), disclose to third parties, the "shall not" became a permissive "may," inverting a confidentiality obligation into a disclosure permission, a polarity inversion riding on a modal downgrade. The first obligation was clean and the second catastrophically wrong, which is why the check verifies every obligation in a segment, not the segment as a whole. Logged: Segment 9, obligation, second clause "shall not disclose" became "may disclose," confidentiality duty inverted, Critical candidate.

Reading the Verdict and the Record Off the Scan

Now the severity gate, run over the logged candidates. Segment 3: tenfold concentration error, Critical. Segment 8: frequency halved and dosage ceiling inverted, Critical (arguably two in one segment). Segment 2: dropped negation inverting a safety prohibition, Critical. Segment 7: conditional collapsed, inverting a maintenance instruction, Critical. Segment 4: mandatory safety step downgraded to optional, Critical. Segment 9: confidentiality obligation inverted into a disclosure permission, Critical. The gate's rule is unambiguous: six segments carry Critical errors, the file fails, and it would fail on any one of them. Recall that the normal post-editing pass had already run, fluency edited, terminology checked, locale verified, and had cleared this file, because every one of these six segments reads as flawless Portuguese. The broad pass is not negligent; it is structurally incapable of catching meaning inversions that leave no seam. The targeted scan caught all six, not because the person running it was more talented, but because the procedure isolated the four categories, verified every instance against the source, decomposed the compound segments, and refused to trust the fluent surface on a single one.

And the run is itself the record. Each logged line, "Segment 8, dosage, frequency source twice daily target once daily, Critical; ceiling not-to-exceed dropped, Critical," is a row a client, a reviewer, or an auditor can read and reconstruct. The scan did not only catch the six Criticals; it produced the evidence of what it checked across all seven high-risk segments, including the two that passed, which is the proof that the verification happened and was not a hopeful glance. This is the artifact that would have stopped the recall in the opening: not a more careful post-editor, but a procedure that forced a polarity check on the one segment whose flawless prose said the opposite of the source, and that left behind a record proving the check was run. The systematic scan is how the silent Critical stops being the error you catch on a good day and becomes the error you catch by design, on every file, whether or not it ever announces itself.

Building the Scan Into Your Pipeline and Keeping It Alive

A procedure that lives only in one linguist's discipline dies when that linguist is busy, so the final move is to build the four-category scan into the pipeline as a defined step rather than a personal habit, and to keep it sharp by feeding it your own failures. Two disciplines do this.

The first is tailoring the scan to where the Criticals live in your domain, because the four categories are the universal spine but their weight shifts by content type. A financial file elevates numbers, percentages, currency, and separators to maximum scrutiny, where negation and dosage are nearly absent. A pharmaceutical file makes dosages the highest-priority category and runs the decomposition twice on every dosing instruction, with numbers and negations close behind. A legal file makes obligations the lead category, gives modal verbs (shall, must, may, will) their own dedicated line, and treats every party-assignment as a triple to verify, with negations second because a dropped "not" in a contract is as deadly as in a label. A software UI file, lower in Critical density, still runs numbers and negations but folds them into a placeholder-and-length-hardened pass. The principle is to keep all four categories on the scan always, and to set order and intensity by where your domain's Criticals cluster, so the deepest verification lands where the consequence is highest.

The second discipline is updating the scan from your own escape log, which turns a static procedure into a learning system. Every time a Critical escapes the scan and is caught downstream, by a reviewer, by the client, by a pharmacist in São Paulo, the response is an entry, not embarrassment: what was the failure, which category did it belong to, why did the scan miss it, and what sharpened technique or added automated flag would have caught it. The injection-pen recall would become a permanent line in the scan's design: every dosing maximum is a polarity-guarded number, decompose it into the number and the prohibition and verify both. A scan that absorbs every escape this way improves monotonically, accumulating the institutional memory of every Critical the team has ever shipped and handing it to the next linguist as a step in a procedure rather than a war story they must live through themselves. The silent Critical is silent by nature, but a procedure that has learned from every one it ever missed grows steadily louder, until the category of error that once ended relationships becomes the category you catch first, on every file, by design.

Key Takeaways

  • A Critical error, in the MQM and ISO 5060 model, is one that can cause real harm or render content unfit for use, and under the one-Critical-fails rule a single one blocks delivery of the whole file. The systematic scan exists because Critical errors are not evenly distributed; they cluster in four categories where meaning can invert without leaving a grammatical seam.
  • The four high-risk categories are negations (any polarity operator that forbids, permits, or conditions), numbers (the rare element meant to survive translation unchanged), dosages (a number wearing armor, bundling amount, unit, frequency, route, and a polarity-guarded maximum), and obligations (modal force, polarity, and bound party in legal content). All four share the property that an engine corrupts them inside fluent, seamless prose the eye cannot catch by reading the target alone.
  • The systematic scan is a targeted layer added on top of normal post-editing, not a replacement for it. The broad pass catches Major and Minor errors and obvious Criticals; the narrow, deep four-category scan catches the silent Critical the broad pass skims past precisely because it reads perfectly. Keep the two passes separate so the polarity check never loses to the fluency it must distrust.
  • Run the scan as a procedure, not a habit: it triggers on every file every time with no exception (because the dangerous file is the one that looked clean enough to skip), it scopes to only the segments containing the four categories (most segments contain none), it verifies in a fixed order (mechanical number and dosage checks first, semantic negation and obligation checks second), and it outputs a logged record of every element checked.
  • The number check reconciles bare tokens ignoring the prose; the dosage check decomposes the bundle and verifies each component independently; the negation check reduces each segment to its polarity skeleton and compares source to target; the obligation check verifies the triple of who, must-or-may-or-must-not, and do-what. Each is stated as an action against the source, never a judgment of the target's surface.
  • Automated assists reduce the haystack but never find the needle. Tooling reliably verifies form, number presence and consistency, unit and separator anomalies, negation-word counts, placeholder integrity, and length-ratio outliers, and you should configure all of them. They are structurally blind to whether meaning inverted: the polarity flip that kept the token count, the frequency that swapped between two valid words, the party swap that changed nothing countable. That residue is the silent Critical, and it is the human's by construction.
  • This boundary is the operational meaning of the revised ISO 18587's requirement that the post-editor hold full professional-translator competence: the human is not there to do faster what the tool does, but to perform the meaning-level verification and severity judgment the tool cannot. Never mistake a clean automated pass for a verified file.
  • Build the scan into the pipeline as a defined step, tailor its order and intensity to where your domain's Criticals cluster (numbers for finance, dosages for pharma, obligations for legal), keep all four categories always on, and feed the scan every downstream escape so it improves monotonically. The run is itself the severity-scored quality record that turns "I checked" into auditable evidence of exactly what was checked and what was found.