AI for Translation & Localization
Capable · M5 · lesson 5 of 21 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Building Verification Checklists for Localization AI
📖
now learning

Building Verification Checklists for Localization AI

15 min

On a Thursday in March, Marek had three files to clear before a five o'clock handoff, and he was good, which is exactly how the trouble started. He was a senior English-to-Polish post-editor on a financial-services account, the kind of linguist a project manager hands the hard batch because he never misses. The engine had pre-translated all three files, the renderings read like a careful native had written them, and he cleared the first two on instinct: scan, trust the smooth ones, fix the few rough cells, confirm, deliver. The third file was a set of fund fact sheets. Segment 14 said the management fee was 0.75% in fluent Polish, except the source said 0.075%, a tenfold error the engine had introduced and Marek had skimmed past because the sentence read beautifully and 0.75% is a perfectly believable fee. He caught it only because a colleague, reviewing on a hunch, ran a number reconciliation she did because she always did, on every file, no matter how good the linguist. That is the whole difference between Marek and his colleague, and it is the whole subject of this lesson. He was relying on vigilance. She was relying on a checklist. Vigilance fails under deadline because it is a mood. A checklist does not fail under deadline, because it is a procedure, and this lesson is about how to build one for localization AI that catches the invented number, the dropped negation, and the drifted term every single time, instead of only on the days you happen to feel sharp.

Why Vigilance Is the Wrong Tool

Begin with an uncomfortable admission, because everything that follows depends on accepting it: your attention is not a reliable instrument, and the better you are, the less reliable it feels, because skill disguises the failures. A post-editor clearing machine-translation output is asked to do something the human mind is genuinely bad at, which is to sustain uniform, high-resolution scrutiny across hundreds of segments, most of which are fine, in search of the rare few that are catastrophically wrong while looking completely fine. That task is a vigilance task, and decades of research into vigilance, from radar operators in the war to baggage screeners today, converge on one finding: human detection of rare signals degrades over time, fast, and the degradation is invisible to the person experiencing it. You do not feel yourself getting worse. You feel exactly as sharp as you did an hour ago, and you are catching less.

Now stack the localization-specific conditions on top of that general human limit, because they make it worse. MTPE (machine-translation post-editing, the work of editing engine output rather than translating from a blank cell) is paid by throughput, so there is a deadline and a word count pressing on every pass. The output is fluent, so the errors do not announce themselves; they hide inside grammatical, natural sentences. And the base rate of errors is low enough that most segments really are fine, which trains your expectation toward "this is fine" and against the suspicion you need. You are doing a vigilance task, under time pressure, against a fluent adversary, with an expectation set tuned to trust. Of course the tenfold fee slips through. The surprise would be if it did not.

The first principle of this lesson, then, is not "be more careful." That is the advice that fails, and it fails because it has no shape: it asks you to do more of the exact thing that is already failing, which is to apply general attention to a problem general attention cannot solve. The fix is to stop relying on attention for the parts of the job that can be turned into procedure. This is the core insight behind every checklist in every high-consequence field, and it is worth stating plainly.

A checklist does not make you more vigilant. It makes vigilance unnecessary for the things on the list, by converting a judgment you might forget to make into a step you cannot skip.

What a Checklist Actually Does

The most useful way to understand a verification checklist is to be precise about the cognitive trade it makes, because that trade is the entire value. Memory and attention are scarce, expensive, and fade-prone. A written, ordered list is cheap, durable, and does not get tired at four-thirty on a Thursday. A checklist moves the burden of remembering-to-check off your working memory, where it competes with everything else you are doing, and onto a piece of paper or a screen, where it sits unchanged regardless of how stressed, rushed, or confident you are. The genius of the surgical checklist, the aviation checklist, the pre-flight checklist, is not that it contains secret knowledge experts do not already have. The surgeon knows to confirm the patient's name. The pilot knows to check the flaps. The whole point is that they know it and still, under pressure, sometimes skip it, and the list closes that exact gap between knowing and doing.

This distinction matters because linguists sometimes resist checklists as beneath their expertise, as if a list of things to verify were an insult to a professional who obviously knows to verify them. That resistance gets the function exactly backwards. The checklist is not for the knowledge; it is for the discipline. Marek knew to check numbers. He had checked numbers ten thousand times. What he lacked on that Thursday was a mechanism that made him check numbers on the file where it mattered, in the specific conditions, deadline plus fluency plus a believable wrong answer, that were engineered to make him not. The checklist supplies the mechanism. It is the difference between a habit, which is reliable until the one time it is not, and a procedure, which is reliable because it does not depend on you remembering to perform it.

There is a second, subtler thing a checklist does, and it is why this lesson exists as its own topic rather than as a footnote to a skills lesson. A checklist makes verification auditable and transferable. A linguist who verifies by instinct cannot prove what they checked, cannot hand the method to a junior, and cannot defend a delivery when a client asks how the error was missed. A linguist who verifies against a written checklist, and logs what the checklist found, produces a record. That record is the bridge from "trust me, I was careful" to "here is the procedure I ran and here is what it caught," and in an MT-first industry being reshaped by ISO 5060:2024 (the standard formalizing the MQM-aligned Critical, Major, and Minor error scoring that decides whether output ships) and the revised ISO 18587 (the post-editing standard, in DIS ballot with publication targeted into 2026, which now covers AI and LLM output and insists the post-editor hold the full competence of a professional translator), that record is increasingly the thing that distinguishes a defensible deliverable from a hopeful one.

The Anatomy of a Good Checklist Item

Before listing what belongs on a localization-AI verification checklist, it is worth slowing down on what makes a checklist item good, because a badly written checklist is worse than none. It produces the feeling of diligence without the substance, and a linguist who has ticked a box marked "check accuracy" believes they have verified accuracy when they have done nothing of the kind. The literature on checklists, much of it from aviation and surgery, is unambiguous about the properties that separate a checklist that works from a ceremony that does not.

A good checklist item is specific. "Check the translation is accurate" is not a checklist item; it is a wish. It names the goal, not the action, and because it names no concrete action it can be ticked without doing anything verifiable. "Isolate every number and unit in the source segment and confirm each appears unchanged in the target" is a checklist item, because it names a discrete action whose performance is either done or not done, with no room for the comforting self-deception that you "basically checked." The test of specificity is simple: could two different linguists, handed the same item and the same segment, disagree about whether the item was performed? If yes, the item is too vague.

A good checklist item is binary and observable. It resolves to a clear pass or fail, not a feeling. "Does every approved term in this segment match the termbase?" resolves to yes or no, and the no is a concrete defect you can point to. "Does this read naturally?" resolves to a vibe, and a vibe is exactly the fluency-trusting judgment the checklist exists to bypass. Wherever a checklist item can be made to resolve against an external reference, the source segment, the termbase, the locale rules, the brief, it should be, because the external reference is what the fluent target invites you to ignore.

A good checklist item is ordered for the way work actually flows, and the order is not arbitrary. The cheapest, highest-consequence checks come first, both because catching a Critical error early saves the effort of polishing a segment you are going to reject, and because the discipline of starting with the mechanical checks anchors the whole pass before fatigue sets in. And a good checklist is short enough to actually run. A checklist with forty items is a checklist no one runs, because running it costs more than the throughput the work is paid for. The art is to put on the list only the checks that are both high-consequence and prone to being skipped, and to leave off the things that either do not matter much or that you reliably do anyway.

A checklist item you can tick without performing a verifiable action is not a safeguard. It is theater, and theater is more dangerous than nothing because it manufactures the confidence of having checked.

The Pause Point and the Killer Item

Two design ideas borrowed from aviation make a verification checklist dramatically more effective, and both are worth importing into localization. The first is the pause point: a deliberate moment in the workflow where you stop and run the checklist, rather than trying to check continuously as you edit. Continuous checking sounds diligent but it is how things get missed, because while you are editing fluency you are not in the frame of mind to be reconciling numbers, and the two tasks interfere. The better pattern is to edit the segment for fluency and style as you normally would, and then, at a defined pause, run the verification checklist as a separate, distinct pass with its own attention. The pause point is what keeps the verification from dissolving into the editing.

The second idea is the killer item: the small number of checks where the failure is so consequential that they get special status, run without exception on every applicable segment regardless of time pressure. In aviation a killer item is something like "flaps set for takeoff," the kind of check that, skipped, kills everyone. On a localization-AI checklist the killer items are the ones that map to Critical errors under ISO 5060: the dropped negation that inverts a safety instruction, the corrupted number in a dosage or a financial figure, the reversed obligation in a contract. These get a heavier line on the checklist, not because the other items do not matter, but because these are the ones where "I was rushing" is never an acceptable account of why one shipped. The severity gate, which we come to at the end, is built on exactly this idea: a single Critical-class failure on a killer item fails the file, full stop.

What Belongs on a Localization-AI Verification Checklist

Now the substance. A verification checklist for localization AI should cover the specific categories where confident engines fail, and only those categories, in an order that puts the cheapest and deadliest checks first. What follows is the canonical set of eight, each stated as an actionable item rather than a goal, with the failure it catches and the fast technique for running it. These eight are not a menu to pick from; they are the spine of any serious checklist, and a checklist missing any of them has a hole where a known failure mode lives.

Item One: Numbers and Units

The first and most important item, because it is fast, language-light, and catches the deadliest mechanical error: isolate every number and unit in the source segment, find its twin in the target, and compare them as bare tokens, ignoring the surrounding prose. A number is the rare element that is supposed to survive translation completely unchanged. "0.075%" is "0.075%" in every language; the digit is a language-independent fact the translation must carry across intact. This makes numbers the one category you can verify without deep target-language competence, by clerical reconciliation rather than by reading. The tells are transposed or altered digits (Marek's 0.075 becoming 0.75), a silently swapped unit (ml becoming mg, a one-letter shift that turns a volume into a mass), flipped decimal and thousands separators across locales, and invented precision where "about 30%" hardens into a false "30.0%." The technique is mechanical on purpose: do not read the number in context, because reading the sentence invites your brain to process the meaning and skim the digit. List the source numbers, list the target numbers, line them up. A bilingual ten-year-old could do it, which is exactly its strength, because it bypasses the fluency that fools the expert.

Item Two: Negations and Polarity

The second item, and the deadliest of all because a dropped negation leaves no gap: reduce every segment carrying a logical operator to its polarity skeleton, forbid, permit, or require, and confirm the source and target agree. When an engine drops a "not," "never," "unless," or "do not," it fluently rewrites the sentence around the absence, producing a grammatical, natural target that means the exact reverse of the source. There is no awkwardness to snag your eye, because the prose closed seamlessly over the missing word. You cannot catch an inverted meaning by reading the target alone; the target is internally consistent. "Take two tablets every twenty-four hours" is a coherent instruction that announces nothing about the source having said "do not exceed two." The only way to catch it is to read the source first, register its polarity (this is a prohibition), and confirm the target carries the same polarity (it should also forbid). The polarity family is wider than the word "not": the vanished prohibition, the flipped conditional ("unless" becoming "when"), the double-negative collapse ("not uncommon" becoming "common"), and the scope error where the negation lands on the wrong clause. This item is a killer item. Run it on every segment with a logical operator, every time.

Item Three: Named Entities

The third item: confirm every proper name, product name, organization, person, place, drug name, and brand survives exactly, neither mistranslated, transliterated when it should stay, nor left untranslated when it should localize. Named entities are a distinct category from terminology because the failure shape differs. An engine, being a fluent author, will helpfully translate a brand name that must stay in English, or invent a plausible transliteration of a person's name, or render a product name as its literal components. Studies of LLM output on medical content found error rates around 59% on drug names alone, and a drug name is a named entity where a substitution can be lethal. The tell is any proper noun that has changed form between source and target without a documented reason. The technique is to treat the project's rules for names as a constraint to verify against: which names stay, which localize, which transliterate, and then confirm each named entity in the segment obeys its rule. Where there is no documented rule, the safe default is that proper names do not change unless the project says so.

Item Four: Approved Terms and Term Drift

The fourth item: confirm every controlled term in the segment matches the termbase entry, not a fluent near-synonym. A termbase is the database of approved terms, the client's mandated word for each concept, and the central fact about an engine is that it does not know your termbase exists unless explicitly given it, and even then it drifts, because it was trained on a world where every concept is expressed a dozen ways and at generation time it reaches for the statistically common phrasing, not the one phrasing your client mandated. The failure is term drift: the same source term rendered as the approved term in some segments and as plausible synonyms in others, each individually fluent, the error visible only as inconsistency against the termbase. The writerly instinct that defeats post-editors here is that varied vocabulary reads as more elegant; in terminology-controlled work, variation is a defect, because consistency carries meaning a regulator, a search index, or a downstream translation memory depends on. The technique is to let the termbase do the looking: a CAT tool (computer-assisted translation environment) with the termbase loaded flags, in real time, when the target lacks the approved term. Where tooling is absent, run a file-wide search keyed to each approved term. Catching drift segment by segment is hard; catching it with a list-keyed search is fast and near-total.

Item Five: Locale Correctness

The fifth item: confirm dates, currency, units, number formats, typographic conventions, and the language variant match the documented target locale. A locale is the full set of regional conventions a target audience expects, not just the language but the country variant and everything it carries. A locale slip is when the engine produces the right language but the wrong conventions, easy to miss because the language is correct and only the conventions underneath are off. The classic slip is the numeric date: "03/04/2026" is ambiguous, and an engine may carry it unchanged into a locale that reads it the other way, shifting the date by a month. Then currency and units (a dollar figure dropped into a Eurozone document, Fahrenheit left for a Celsius audience), then typography (German quotation marks differ from English, French inserts a space before certain punctuation). The most damaging slip is between variants of one language: Spanish for Spain is not Spanish for Mexico, Portuguese for Portugal is not Portuguese for Brazil. An engine asked for "Spanish" without a variant defaults to a blend, producing fluent Spanish that is persistently wrong for the actual market. The technique is to confirm the engine was targeted at the exact locale (es-MX, pt-BR, fr-CA, not "Spanish" or "Portuguese" or "French") and to verify the locale-bound elements against the project's locale rules rather than against the fluent surface.

Item Six: Tags and Placeholders

The sixth item: confirm every tag, code placeholder, variable, and markup token in the source is present, intact, and correctly positioned in the target. Localization content is rarely plain prose; it carries formatting tags, inline markup, and placeholders the running software replaces at runtime, things like a variable that becomes the user's name, an ICU placeholder (the message-formatting syntax used in software localization) that becomes a count, or an HTML tag that styles a word. An engine treating its input as text to make fluent may drop a placeholder, translate the variable name inside it, reorder tags so they wrap the wrong words, or corrupt the syntax so the software cannot parse it. The consequences run from a broken layout to a string that crashes the build to a message that displays raw code to a user. The tell is any difference in the set, count, or syntax of tags and placeholders between source and target. The technique is mechanical reconciliation, identical in spirit to the number check: list the tokens in the source, list them in the target, confirm the sets match and the syntax is intact. This is a check tooling does well, and a QA filter in the CAT tool or TMS (translation-management system, the platform that orchestrates the localization workflow) should be configured to flag placeholder mismatches automatically.

Item Seven: Omissions and Additions

The seventh item: run a two-directional mapping pass, source-to-target to catch omissions and target-to-source to catch additions, counting the units of meaning in and out. An LLM does not merely mistranslate; it edits. It will smooth a terse source by adding an explanatory clause that "should" be there, or quietly drop a phrase it judged redundant, both in seamless prose. Addition, also called hallucination, shows up as target material with no source: "Store below 25C" becoming "Store below 25 degrees Celsius, away from direct sunlight and moisture," with the last clause invented. Omission is the mirror: a qualifying condition, a caveat, a cross-reference that simply does not appear, the fluent target closing over the gap with no seam. The cheap trigger is gross length mismatch: languages expand and contract by predictable ratios, so a target dramatically shorter or longer than the expansion ratio predicts is a flag worth the slower check. The slower check, reserved for flagged and high-consequence segments, is the mapping pass: walk the source units of meaning and confirm each has a target counterpart, then walk the target units and confirm each has a source. One direction catches only one failure; you must do both. The discipline is to count meaning, not just to read it.

Item Eight: The Severity Gate

The eighth item is not a content check but a decision rule, and it is what turns the checklist from a list of observations into a go or no-go: classify every defect the checklist found by severity, and apply the rule that a single Critical error fails the file regardless of how clean everything else is. The first seven items find defects. The severity gate decides what the defects mean for delivery. Under the MQM (Multidimensional Quality Metrics) and ISO 5060 model, every error is scored Critical, Major, or Minor: a Critical error is one that can cause real harm or render the content unfit for use (a flipped dosage, an inverted contractual obligation, a corrupted financial figure, a safety instruction reversed), a Major error seriously degrades quality without that catastrophic edge, and a Minor error is a small blemish. The non-negotiable rule the gate enforces is that one Critical fails the file, no matter how fluent and polished the other ninety-nine segments are, because "looks fine overall" is precisely the judgment that ships the silent critical error. The severity gate is where the killer items earn their name: a failure on numbers, negations, named entities, or placeholders in high-consequence content is a Critical, and a Critical is a stop. This item is what makes the checklist defensible, because it produces a clear, recordable verdict a client and an auditor can read.

The first seven items find the defects. The eighth decides whether the file ships, and it enforces the one rule that distinguishes a quality gate from an opinion: a single Critical error fails the file, no matter how clean the rest looks.

A Worked Checklist Applied to a File

Abstraction is comfortable and useless until you run it on a real file, so let us build the reusable checklist as a concrete artifact and apply it, segment by segment, to a small batch. Picture an eight-segment file: patient-facing instructions for a medical device, English source, German target, pre-translated by an LLM, with a loaded termbase, a documented brief specifying the formal "Sie" register and the de-DE locale, and a deadline. Here is the checklist as Marek's colleague would run it, written as the ordered, pausing, killer-item-marked procedure it should be.

The reusable checklist (run at the pause point, after editing each segment for fluency):

  • 1. Numbers and units [KILLER]. List every source number and unit; confirm each appears unchanged in the target as a bare token. Decimal and thousands separators match the de-DE locale.
  • 2. Negations and polarity [KILLER]. Reduce every logical operator to forbid, permit, or require; confirm source and target agree.
  • 3. Named entities [KILLER on regulated content]. Every proper name, product name, and drug name survives per the project's name rules.
  • 4. Approved terms. Every controlled term matches the termbase entry, not a synonym; termbase QA flag clear.
  • 5. Locale. Dates, currency, units, formats, and typography match de-DE; register is "Sie" and consistent.
  • 6. Tags and placeholders [KILLER]. Every tag and placeholder present, intact, correctly positioned; TMS placeholder QA clear.
  • 7. Omissions and additions. Length within the German expansion ratio; two-directional mapping pass on flagged and high-consequence segments.
  • 8. Severity gate. Classify every defect Critical, Major, or Minor; a single Critical fails the file.

Running It Segment by Segment

Segment 1. Source: "Insert the {device_name} fully before activation." The target reads fluently and addresses the user with "Sie." Item 1: no numbers, pass. Item 2: no polarity operator, pass. Item 3: device name is a named entity, present per rule. Item 4: "activation" is a controlled term, termbase rendering present, flag clear. Item 5: register "Sie," correct. Item 6: the placeholder. The source carries {device_name}; the target carries {Gerätename}, the engine having helpfully translated the variable name inside the placeholder. That is a placeholder corruption: at runtime the software will not find a variable called Gerätename and the string will break or display raw. This is a killer item and a defect. Logged, with severity to be set at item 8.

Segment 2. Source: "Do not reuse the cartridge after the expiry date." Item 1: no numbers. Item 2, the killer: source polarity is forbid (do not reuse). Reduce the target to its polarity skeleton: it permits reuse. The negation dropped, the engine fluently rewrote "reuse the cartridge after the expiry date" as a neutral instruction, and the meaning inverted from a prohibition to a permission on a safety-critical instruction. This is exactly the failure that nearly ended Farah's contract in the previous lesson, and exactly the failure vigilance misses because the German is flawless. Logged as a Critical candidate.

Segment 3. Source: "Each dose delivers 2.5 mg of the active ingredient." Item 1, the killer: source numbers are 2.5 and mg. Target numbers, listed as bare tokens: 25 and mg. The decimal point vanished. "25 mg" is ten times the dose, fluent, believable, and lethal in the wrong direction. Caught not by reading the sentence, which reads perfectly, but by lining up the tokens. Logged as a Critical candidate. Note that this is structurally identical to Marek's fund-fee error: a vanished decimal, a tenfold corruption, a perfectly believable wrong number. The check that catches it is the same regardless of domain.

Segment 4. Source: "Store below 25C." Target, mapped: "Lagern Sie unter 25 °C, geschützt vor Sonnenlicht und Feuchtigkeit." Item 7: walk the units. Source units: store, below 25C. Target units: store, below 25C, protected from sunlight, protected from moisture. Two units became four; the engine added "away from sunlight and moisture" out of its general knowledge of storage. On a regulated medical document, content not in the source is content the client did not approve. Logged as a Major (an unapproved addition that degrades fidelity but does not, here, create a safety inversion).

Segment 5. Source: "Contact our support team or visit the help center." The target reads fluently but addresses the user as "du." Item 5: the brief mandates "Sie," and this segment lurched into the informal register. A consistency search across the file confirms the file mixes "Sie" and "du." Logged as a Major (a register defect that violates the documented brief and undermines the brand's chosen voice).

Segment 6. Source: "The device is compatible with the AccuFlow Mini." Item 3, named entity: "AccuFlow Mini" is a product name that, per the project's name rules, stays in English. The target rendered it "AccuFlow Mini," unchanged. Pass. Every other item: clean. This segment is fully clean, which matters to record too, because the checklist documents what passed, not only what failed.

Segment 7. Source: "Replace the filter every 30 days, unless the indicator shows red sooner." Item 2, the killer: this carries a conditional operator, "unless." Source polarity: replace every 30 days is the default, with an exception that pulls it earlier. Reduce the target: it renders "replace every 30 days when the indicator shows red," collapsing the "unless" exception into a "when" condition, which changes the instruction from "every 30 days, or earlier if red" to "only when red." The conditional flipped. Item 1 also applies: 30 appears in source and target, unchanged, pass on numbers; the defect is purely the polarity collapse. Logged as a Critical candidate (a flipped maintenance condition on a medical device).

Segment 8. Source: "For more information, see Section 4.2." Item 1: numbers 4 and 2; target shows "Abschnitt 4.2," unchanged, pass. Item 7: "see Section 4.2" maps cleanly, no omission, no addition. All items clean.

Reading the Verdict Off the Gate

Now item 8, the severity gate, run over the logged defects. Segment 1: placeholder corruption, a Critical, because at runtime it breaks the string. Segment 2: dropped negation inverting a safety instruction, Critical. Segment 3: tenfold dosage error, Critical. Segment 5: register defect, Major. Segment 4: unapproved addition, Major. Segment 7: flipped maintenance conditional, Critical. The gate's rule is unambiguous: four Critical errors, therefore the file fails, and it would fail on any one of them. Note what just happened. A linguist clearing this file on vigilance alone, under deadline, would very likely have caught the rough cells and confirmed the smooth ones, and segments 2, 3, and 7 are the smooth ones, fluent, natural, confident, and catastrophically wrong. The checklist caught all four Criticals not because the person running it was more talented than Marek, but because the procedure does not get tired and does not trust fluency, and it ran the killer items on every applicable segment regardless of how good the German looked.

And here is the move that makes this lesson reusable beyond a single file: that run is itself the source-verification record. Each line, "Segment 3, item 1, source 2.5 mg, target 25 mg, decimal dropped, Critical," is a row in a log a client and an auditor can read. The checklist did not only catch the errors; it produced the evidence that the verification happened and the evidence of what it found. This is the same artifact the program's standards keep circling back to: a defensible quality record, severity-scored, that turns "I was careful" into "here is the procedure I ran and here is what it caught."

Building Your Own and Keeping It Alive

The eight-item checklist above is a strong default, but the best checklist is one tuned to your actual content and kept alive rather than printed once and forgotten. Two disciplines keep it useful. The first is tailoring by content type, because a uniform checklist applied uniformly wastes attention on segments that do not need it and under-checks segments that do. A throwaway marketing tagline does not need the two-directional mapping pass; a dosing instruction needs every killer item run twice. Build the core eight as the spine, then add content-specific items: a financial file gets a dedicated check on figures, percentages, and currency that elevates them all to killer status; a software UI file gets a hardened placeholder and length-budget check; a legal file gets a killer item on obligations and modal verbs (shall, must, may), where a flipped "shall" to "may" is a Critical that the polarity item should already catch but that deserves its own line on legal content. The principle is to make the killer items match where the Criticals actually live in your domain.

The second discipline is updating the checklist from your own error log, which is the practice that turns a static list into a learning system. Every time a defect escapes the checklist and is caught downstream, by a reviewer, by the client, by a reader, the right response is not embarrassment but an entry: what was the failure, why did the checklist not catch it, and what item, sharpened or added, would have? Marek's colleague did not invent her number-reconciliation habit from a textbook; she built it the day a vanished decimal cost someone a re-delivery, and she never ran a file without it again. A checklist that absorbs every escape this way gets monotonically better, because it accumulates the institutional memory of every error your team has ever shipped, and it hands that memory to the next linguist as a procedure rather than as a war story they have to live through themselves.

Where Tooling Ends and the Human Begins

A fair question at this point: if the checklist is a procedure, why not automate it entirely and free the human from running it? Part of it should be automated, and the lesson has said so repeatedly. Placeholder mismatches, number presence, term-base non-matches, register-marker consistency, and length-ratio outliers are all checks a CAT tool or TMS QA filter runs faster and more reliably than a human, and you should configure every automatable item as an automatic flag, because every check you offload from attention is attention freed for the checks only judgment can make. The checklist should explicitly mark which items the tooling runs and which the human must run, so that the human's pass concentrates on what is left.

But the items at the heart of the checklist resist full automation precisely because they require comparing meaning to meaning. A QA filter can tell you the number 2.5 appears in the source and 25 in the target; it cannot reliably tell you that the dropped negation in segment 2 inverted the meaning, because judging that the polarity reversed is a semantic judgment, and judging whether an added clause in segment 4 is an approved elaboration or a hallucination requires knowing the regulatory context. This is the deeper reason the revised ISO 18587 insists the post-editor hold full professional-translator competence: the human in the loop is not there to do what the tool does faster, but to do what the tool cannot do at all, the meaning-level verification the checklist organizes and the severity judgment the gate requires. The checklist does not replace the linguist's expertise; it directs it, ensuring the expensive human judgment lands on exactly the segments and exactly the questions where it is irreplaceable, and is never quietly skipped on the smooth cell where the silent critical error is hiding.

This lesson closes Level 2's work on assisted oversight, and it points directly at the L2 capstone, where everything assembles into one deliverable. The capstone asks you to produce a verified AI-assisted deliverable, a post-edited file, a transcreated set, or a terminology-enforced batch, with a source-verification log and an MQM/ISO 5060 error score attached. The checklist you have just learned to build and run is the engine of that capstone: it is the procedure that produces the source-verification log, the killer items are what surface the Criticals, and the severity gate is what produces the score that decides whether the deliverable ships. Marek's colleague was not more vigilant than Marek. She had a checklist, she ran it on every file, and she could prove what it caught. By the time you finish the capstone, so can you.

Key Takeaways

  • Vigilance is the wrong tool for catching the silent critical error because it is a mood that fades under deadline, and human detection of rare signals degrades over time invisibly. A checklist does not make you more vigilant; it makes vigilance unnecessary for the things on the list by converting a judgment you might forget into a step you cannot skip.
  • A checklist's value is a cognitive trade: it moves the burden of remembering-to-check off scarce, fade-prone attention and onto a cheap, durable, ordered list. The checklist is not for the knowledge (you already know to check numbers); it is for the discipline of checking them on the exact file where deadline plus fluency plus a believable wrong answer conspire to make you not.
  • A good checklist item is specific (a discrete action, not a goal), binary and observable (resolves to pass or fail against an external reference, not a vibe), ordered cheapest-and-deadliest-first, and short enough to actually run. An item you can tick without performing a verifiable action is theater, and theater is more dangerous than nothing because it manufactures false confidence.
  • Import two ideas from aviation: the pause point (run verification as a separate pass after editing, so the two tasks do not interfere) and the killer item (the small set of checks, numbers, negations, named entities, placeholders, that run on every applicable segment without exception because a single failure is catastrophic).
  • The canonical eight items are numbers and units, negations and polarity, named entities, approved terms, locale, tags and placeholders, omissions and additions, and the severity gate. Each is stated as an action and targets a specific engine failure mode; a checklist missing any of them has a hole where a known failure lives.
  • The severity gate is the eighth item and the decision rule that turns observations into a go or no-go: classify every defect Critical, Major, or Minor under MQM/ISO 5060, and enforce that a single Critical error fails the file no matter how clean the other segments look. This is what distinguishes a quality gate from an opinion.
  • Running the checklist produces the source-verification record as a byproduct: each logged line ("Segment 3, source 2.5 mg, target 25 mg, decimal dropped, Critical") is auditable evidence of what was checked and what was found, turning "trust me, I was careful" into a defensible quality record under ISO 18587 and 5060.
  • Tailor the spine to your content type (financial files elevate every figure to killer status; legal files add an obligation-and-modal-verb item), keep it alive by absorbing every downstream escape into a sharpened item, and automate what tooling does well so the human's pass concentrates on the meaning-level verification only judgment can perform. The checklist directs expertise; it does not replace it.