Keeping the TM Clean
Two years after the project shipped, nobody on the new team had heard the word "réservoir" said in anger. The Lyon medical-device account had long since fixed the wrong French term for the insulin-pump cartridge, run the global find-and-replace, reprinted the guide, re-recorded the subtitle, retrained the helpline. Case closed, or so everyone believed. Then a junior post-editor in a different city, working a completely unrelated firmware-update file for the same client, opened a segment the translation memory had pre-filled at a confident 100% match. The source said "Insert the cartridge fully before priming." The target, leveraged automatically from the memory, said "réservoir." The post-editor, new to the account and trusting a green 100% match the way everyone is taught to trust one, accepted it without a second look. The poison had survived the cleanup. It had been hiding inside a stored segment that the find-and-replace missed, because that segment had been written back into the TM from a slightly different sentence and never got swept. The error was not in the file anyone fixed. It was in the memory everyone reused. This lesson is about that memory, the translation memory, and the single discipline that decides whether it makes a localization operation better every year or quietly worse: keeping it clean.
What a TM Actually Is, and Why It Compounds
Start with the asset itself, because the whole argument turns on understanding what it does. A translation memory (TM) is a database of previously translated and approved sentence pairs. Each time a human finalizes a segment (a sentence or sentence-like unit, the atomic chunk the pipeline operates on), the source text and its confirmed target text are stored together as one record. The next time an identical or similar source segment appears, the TM offers the stored target back, so the same sentence is never solved twice. The CAT tool (the computer-assisted translation environment where linguists actually work) reaches into the TM before the machine-translation engine ever runs, and fills what it can from memory first.
That reuse is called leverage, and leverage is the entire economic engine of a localization shop. A document with high internal repetition, or strong overlap with past work, can be 40, 60, even 80% leveraged, meaning most of it is filled from memory at a fraction of the cost of fresh translation. There are two flavors you must tell apart instantly. An exact match (a 100% match) means the incoming source is character-for-character identical to a stored source, and the tool drops in the stored target with full confidence. A fuzzy match is a partial match, scored as a percentage of similarity, typically from a 99% match down to a 75% threshold; the source is similar but not identical, and the tool offers the stored target while highlighting what changed so the linguist can adapt it.
Here is the property that makes a TM unlike almost any other asset a business owns: it compounds. Every approved segment you commit becomes a building block for every future project that shares vocabulary or phrasing with it. A clean TM is therefore not a static file; it is a flywheel. Each correct segment you add makes the next project faster and more consistent, which produces more correct segments, which makes the project after that faster still. Over years, a well-kept TM becomes the single most valuable thing a localization team owns, more valuable than any tool, because the tools are replaceable and the institutional memory of how this client's product is correctly expressed in this language is not.
A translation memory is the only asset in the building that gets more valuable every time you use it, or less valuable every time you use it. The difference is entirely a matter of what you let into it.
The Flywheel Spins Both Ways
The uncomfortable corollary of compounding is that the flywheel is indifferent to direction. The exact mechanism that makes a clean TM compound quality, that the system faithfully reuses past decisions at scale, makes a dirty TM compound mistakes at the very same scale. A correct segment, leveraged a hundred times, spreads correctness a hundred times. A poisoned segment, leveraged a hundred times, spreads the poison a hundred times, and it does so silently, because leverage is automatic and trusted. The TM does not know the difference between a good decision and a bad one. It reuses both with identical, mechanical fidelity.
This is why "keeping the TM clean" is not housekeeping. It is the difference between an asset that appreciates and a liability that depreciates, and the two look exactly alike on the day you commit the segment. The cost of a bad segment is not paid when you write it. It is paid, with interest, every future time the TM hands it back. In an MT-first 2026 pipeline, where engines pre-fill faster and post-editors move faster than ever, the rate at which segments enter the TM has gone up, which means the rate at which a bad one can enter has gone up too. The flywheel spins faster in both directions now.
How a TM Gets Poisoned
A TM does not rot on its own. It is poisoned, segment by segment, by specific acts of committing bad data into it. If you can name the ways it happens, you can build a guard against each one. There are three that account for nearly all of it, and a working linguist needs to recognize all three on sight.
Poison One: Committing Unverified MT Output
This is the dominant failure mode of the MT-first era, and it is exactly what happened in Lyon. The pipeline pre-fills a segment with raw machine-translation output (from a neural MT engine or an LLM, it makes no difference here). The output is fluent, grammatical, and confident. A post-editor, working at speed on a per-word rate that did not budget for deep scrutiny of every clean-looking segment, reads a smooth sentence, sees nothing obviously wrong, and confirms it. The instant they confirm, the CAT tool writes that segment, raw machine guess and all, back into the translation memory as if it were an approved human translation.
That single act is the poisoning. The TM has no field that says "this one was a rushed accept of an unverified machine guess." Once committed, it sits in the memory with exactly the same status as a segment a senior linguist agonized over for ten minutes. The next project that leverages it inherits a 100% match that looks like gold and is actually an unverified machine rendering that was never checked against the source. The danger is precisely that MT output is fluent first and accurate second: the segments most likely to be waved through are the smooth ones, and a fluent mistranslation is the one that reads perfectly while meaning something the source never said.
Poison Two: Bad Alignments
The second route is more technical and easy to underestimate. Much of the content in a mature TM did not come from live translation at all. It was created by alignment, the process of taking a pair of documents that were translated in the past (a source file and its finished target file, translated before the team used a CAT tool) and machine-matching each source sentence to its corresponding target sentence to manufacture TM segments in bulk. Alignment is how a shop "imports" years of legacy translation into a memory in an afternoon instead of retranslating it.
The problem is that alignment is a guessing game whenever the two documents do not line up sentence for sentence. If the source has three sentences where the target has two (because the translator merged them), or the target added a sentence the source did not have (a localized example, a regulatory note), the aligner can pair the wrong source with the wrong target. The result is a TM segment whose source says one thing and whose target says something from the neighboring sentence: a wrong-content pair that is internally fluent on both sides but is not a translation of each other at all. These are insidious because each half reads perfectly. Only the relationship between them is broken, and a leverage match will hand you the mismatched target with full 100% confidence the next time that source appears.
Poison Three: Wrong-Context Matches
The third route is the subtlest, because nothing in the segment is "wrong" in isolation. A short source segment can be perfectly translated one way in one context and need a completely different translation in another, and a plain TM stores only the sentence pair, not the world around it. Consider the English string "Open." On a button it might be a command, "Ouvrir." As a status label it might be an adjective, "Ouvert." As a store-hours line it might be "Ouvert" again but in a different register. A TM that stored "Open" → "Ouvrir" from the button context will cheerfully leverage that target into the status-label segment as a 100% match, and it will be wrong, not because anyone mistranslated anything, but because the right translation depended on a context the TM never recorded.
This is the wrong-context match: a leverage that is a faithful record of a real, correct past translation, surfacing into a new segment where that translation no longer fits. It is the reason a 100% match is a strong signal, not a guarantee. The percentage measures source-text similarity. It says nothing about whether the surrounding context, the field, the screen, the legal frame, is the same. A linguist who treats every 100% as untouchable will ship the wrong-context error every time, because the tool will never flag it: by its own measure, it found a perfect match.
A TM gets poisoned three ways: you commit an unverified machine guess, you import a mismatched alignment, or you leverage a correct translation into a context where it is no longer correct. The first is a discipline failure, the second is an import failure, and the third is a trust failure. All three end the same way: a bad segment wearing a green match's clothing.
The Anatomy of a Propagated Error: A Worked Example
Abstract failure modes do not teach the way a single error walked all the way through a system does. So follow one, slowly, from the keystroke that created it to the project two years later where it surfaced again. This is the réservoir error, traced as a TM-poisoning event rather than a terminology event, because the lesson is in the propagation.
Day zero, segment 12. The MT engine, optimizing for the most probable fluent French, renders the approved term "cartouche" as the more common "réservoir." A post-editor confirms the fluent segment. The CAT tool writes segment 12 into the TM. The poison is now in the memory. Crucially, nothing visible has gone wrong yet. The file reads fine. The post-editor moves on, unaware they have just deposited a defective building block into the asset every future project will draw from.
Day zero, later in the same file. Segments 47, 88, 134, and 201 contain similar sentences. The CAT tool, now finding the freshly stored segment 12 as a high fuzzy or exact match, leverages "réservoir" into each of them. The post-editor sees green and high-percentage matches, the interface highlights only the words that changed (a port number, a step order), and the unchanged term sits quietly beneath the highlight where the design of the fuzzy-match view actively steers the eye away from it. Five more poisoned segments enter the file, and as each is confirmed, each is written back into the TM. The single drift has become six TM records, all mutually reinforcing.
Day three, a different deliverable. The quick-start guide, the training-video subtitle file, and the helpline script are all generated from projects that share this TM. Each leverages the poisoned segments. The error crosses from the manual into print, into video, into a call-center script, without any human ever re-introducing it. By now "réservoir" appears in twenty-four segments across three channels, every instance a faithful leverage of a stored segment, every instance reading perfectly.
The cleanup that did not clean. The error is caught, eventually, in a print proof. The team runs a global find-and-replace across the active project files, reprints, re-records, retrains. They believe it is over. But find-and-replace operates on the files in front of them, and the TM is a database behind the files. One poisoned segment had been written back from a sentence whose surrounding text differed enough that the cleanup query, written to match the printed strings, never selected it. It stayed in the memory, dormant.
Two years later. A new post-editor on a new firmware file opens a segment pre-filled at 100% from that surviving record. "réservoir." They accept it, because a 100% match from an established client TM is exactly the thing every linguist is trained to trust. The error, declared dead, has reproduced itself into a new generation of deliverables from a single stored segment that outlived the entire cleanup. That is the anatomy of a propagated error: cheap to create, expensive to contain, and capable of surviving its own funeral because it lives in the asset, not the file.
Why the Cost Curve Is So Cruel
The reason TM hygiene pays for itself is the shape of the cost curve. Catching the error at the moment of the keystroke, before segment 12 is ever committed, costs one verification: a glance at the source, a corrected word, three seconds. Catching it after it is in the TM but before it leverages costs a cleanup pass. Catching it after it has propagated into the file costs a re-review of the file. Catching it after it has crossed into print, video, and the helpline costs a reprint, a re-record, and a retraining. And catching it two years later costs all of that again, plus the credibility of a team that told the client it was fixed. The cost roughly multiplies at each stage it escapes. Every control in TM hygiene is, at bottom, an attempt to catch the error as early on that curve as possible, ideally before the commit that makes it institutional memory.
TM Hygiene Practices That Actually Hold
Knowing how a TM gets poisoned tells you exactly where to build the guards. TM hygiene is not one habit; it is a set of controls, each closing a gap the others leave open. Here are the ones that earn their place, with the failure each one stops.
Metadata: The Provenance of Every Segment
The foundational practice is to record, on every segment, where it came from and who stands behind it. A bare TM stores only the source and target. A governed TM stores metadata alongside each pair: the origin (human translation, MT post-edited, aligned import), the date, the linguist or reviewer who confirmed it, the project, the domain, and ideally a status (draft, reviewed, approved). This sounds like bureaucracy until you need it. Metadata is what lets you ask the TM a question it otherwise cannot answer: "show me every segment that was committed as raw MT and never independently reviewed," or "show me everything imported by alignment in that legacy batch." Without provenance, every segment looks equally trustworthy, which means the poisoned ones are invisible. With provenance, you can find and quarantine an entire class of suspect segments in one query, which is the only way the Lyon cleanup could have been complete.
Penalties for MT-Origin Segments
Most CAT tools and translation-management systems let you apply a penalty to a class of TM matches: a configured percentage subtracted from the displayed match score for segments with a given attribute. The classic and most important use is the MT-origin penalty. If a segment entered the TM as machine-translated or lightly post-edited, you apply a penalty so that a "100%" match from that segment displays as, say, 85% instead. The number is not cosmetic. It changes behavior: a 100% match is one a linguist is trained to wave through, while an 85% fuzzy match is one the workflow forces them to open, read, and confirm against the source. The penalty deliberately demotes a less-trustworthy segment out of the auto-trust zone and into the must-verify zone. It is the system encoding the truth that a machine-origin segment has not earned the confidence a fully reviewed human segment has, and refusing to let it impersonate one.
A penalty is how you stop an unverified segment from wearing a trusted segment's badge. It does not delete the leverage; it removes the leverage's disguise, forcing a human to look where they would otherwise have trusted.
Context and ICE Matches
The defense against the wrong-context error is the context match, often called an ICE match (in-context exact match). An ICE match is a 100% match that is also confirmed to share the same surrounding context as the stored segment: the tool checks not only that the source string is identical but that the segment before it (and sometimes after) is identical too, meaning the segment is appearing in the same place in the same kind of document. An ordinary 100% match guarantees the sentence matched. An ICE match guarantees the sentence matched in the same context, which is the thing that actually determines whether the stored target still fits. The practical discipline is to treat ICE matches as genuinely safe to leverage with minimal review, and to treat plain 100% matches as a strong suggestion that still requires a context check, exactly because the "Open" → "Ouvrir" trap lives in the gap between the two. A tool that distinguishes them is handing you the difference between "the words are the same" and "the situation is the same," and only the second one is safe.
Not Auto-Confirming
The single most consequential hygiene practice is also the simplest to describe and the hardest to hold under deadline: do not auto-confirm. Many tools offer to automatically confirm and lock all 100% matches, or all matches above a threshold, so the linguist never has to touch them. This is the feature that built the surviving Lyon segment. Auto-confirmation treats the match percentage as a proxy for correctness, and we have seen three separate ways that proxy lies: the 100% can be an unverified MT commit, a bad alignment, or a wrong-context match. Every auto-confirmed segment is a segment no human looked at, which means every poison that already lives in the TM gets relaundered into the new deliverable untouched, and any new poison sails straight through. Turning auto-confirm off, at least for any content with consequence, is the choice to keep a human in the loop at the exact point the loop matters most: the moment a stored segment becomes a shipped one.
Cleanup Passes and TM Maintenance
Finally, a TM needs scheduled maintenance the way any compounding asset does, because no commit-time control is perfect and errors accumulate. A cleanup pass (also called TM maintenance) is a deliberate, periodic audit of the memory itself rather than of a live project. The valuable ones include: a duplicates-and-conflicts pass that finds segments with the same source but different targets (a sign of inconsistency or a buried error), a terminology consistency pass that checks stored targets against the current termbase, an orphan pass that finds aligned segments whose source and target lengths suggest a mismatch, and a provenance pass that uses metadata to re-review or quarantine everything from a suspect batch. The point of maintaining the TM as an asset in its own right, on a schedule, is that the poison that survives the file-level controls is precisely the poison that will surface years later in an unrelated project, exactly as réservoir did. The cleanup pass is the only control that looks at the memory itself, which is the only place that surviving segment could have been found before a new post-editor accepted it.
The Discipline of the TM Boundary
Step back from the individual controls and see the single principle underneath all of them. Every one of these practices guards the same line: the boundary between a segment in a live file and a segment written back into the translation memory. That boundary is the most important checkpoint in the entire pipeline, and most linguists cross it dozens of times an hour without noticing it exists.
On the file side of the boundary, an error is local and cheap. It lives in one deliverable, it can be caught in review, and if it ships it damages one project. On the memory side of the boundary, the same error is institutional and expensive. It becomes a building block, it leverages into everything, and it can outlive every attempt to kill it. The entire discipline of keeping a TM clean reduces to a single instinct: treat the write-back into the TM as a publishing decision, not a save. When you confirm a segment, you are not saving your work; you are publishing a sentence into the asset your whole team and your future self will reuse without re-checking. That reframing is the whole lesson. A segment you would happily ship in today's file is not automatically a segment you should commit to the memory forever, because today's file has a context the memory will forget.
Who Owns the Boundary
In a mature operation, somebody owns this boundary explicitly. It is part of what the terminologist and the TM manager do, but in an MT-first shop the responsibility has spread to every post-editor, because every post-editor commits segments. The job is no longer "translate the sentence." It is "decide whether this sentence is good enough to become permanent reusable memory," which is a higher bar and a different kind of judgment. The engine cannot make that judgment; it has no concept of "this is good enough to keep." The TM cannot make it; it reuses whatever it is fed. Only the human at the boundary can look at a fluent machine guess and decide it is not yet trustworthy enough to outlive this project. That decision, made hundreds of times a day, is what separates a TM that compounds quality from one that compounds mistakes.
Confirming a segment is not a save. It is a publishing decision into an asset that will reuse your sentence without ever re-checking it. The linguist who feels the weight of that boundary keeps the TM clean. The one who treats it as a keystroke poisons it, fluently, one green match at a time.
Why This Is Leverage Worth Protecting
It would be easy to read all of this as a catalog of dangers and conclude that a TM is a hazard to be managed. It is not. A clean TM is the closest thing a localization operation has to a genuine compounding return. It means a regulated client's product is expressed in exactly the approved way, every time, across every future update, for years, at a fraction of the cost of redoing the work, with consistency no human could maintain by memory alone. The leverage is real and it is enormous, and the hygiene practices are not a tax on that leverage; they are the thing that keeps the leverage worth having. A poisoned TM still leverages just as fast. It just leverages the wrong thing.
The marvel and the danger are, once again, the same mechanism. Faithful reuse at scale is the most powerful tool the discipline has, and it amplifies whatever you feed it. Keeping the TM clean is simply the decision to be careful about what you feed an amplifier this powerful. Done well, it turns the memory into an asset that makes every linguist who touches it faster and more accurate than the last. Done carelessly, it turns the same memory into a machine for replicating a single tired keystroke across a product family for years. The mechanism does not choose. You do, every time you cross the boundary.
Key Takeaways
- A translation memory (TM) is a database of approved source-target sentence pairs that the CAT tool reuses (leverage) before the MT engine runs. It is the one asset that compounds: each clean segment makes future projects faster and more consistent, and each poisoned segment compounds mistakes at the same scale, because leverage reuses good and bad decisions with identical, mechanical fidelity.
- A TM gets poisoned three ways: committing unverified MT output (a fluent machine guess confirmed at speed and written back as if approved), bad alignments (legacy import that mismatches source and target sentences into a wrong-content pair), and wrong-context matches (a correct past translation leveraged into a new context where it no longer fits, like "Open" as a button versus a status label).
- The match percentage measures source-text similarity, not correctness. A 100% match can still be an unverified MT commit, a bad alignment, or a context mismatch, so a green match is a strong signal, never a guarantee.
- A single drift compounds because confirming a segment writes it back into the TM as institutional memory; it then leverages forward into more segments, across manuals, subtitles, and scripts, and can survive a file-level find-and-replace because the TM is a database behind the files. The réservoir error resurfaced two years later from one stored segment the cleanup missed.
- The cost curve is cruel: catching the error at the keystroke costs three seconds, after the commit costs a cleanup, after propagation costs a re-review, after it crosses channels costs a reprint and re-record, and after years costs all of that plus client trust. Every hygiene control aims to catch the error as early on that curve as possible.
- The hygiene stack: record metadata (provenance) on every segment so you can query and quarantine suspect classes; apply penalties to MT-origin segments so they display below the auto-trust threshold and force a human to look; prefer context/ICE matches (in-context exact matches) over plain 100% matches because they confirm the situation, not just the words; do not auto-confirm matches, because every auto-confirmed segment is one no human checked; and run scheduled cleanup passes that audit the memory itself for duplicates, conflicts, terminology drift, and bad alignments.
- The single principle beneath every control is the TM boundary: the line between a segment in a live file and a segment written back into the memory. On the file side an error is local and cheap; on the memory side it is institutional and expensive. Treat the write-back as a publishing decision, not a save.
- The human at the boundary is the only part of the system that can decide a fluent segment is not yet trustworthy enough to become permanent reusable memory. The engine has no concept of "good enough to keep," and the TM reuses whatever it is fed. Keeping the TM clean is the decision to be careful about what you feed an amplifier this powerful, so it compounds quality instead of mistakes.
Skill.re