โ†
AI for Translation & Localization
Aware ยท M11 ยท lesson 11 of 19 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
MT and LLMs in Translation and Post-Editing
๐Ÿ“–
now learning

MT and LLMs in Translation and Post-Editing

15 min

It is 9:14 on a Tuesday and Omar has just double-clicked a file called es-ES_release_notes_v7.sdlxliff. His CAT tool, the computer-assisted translation environment where he edits target text segment by segment, takes a second to open it, and then the grid fills the screen: source English on the left, target Spanish on the right, a status column in the middle. He expects what he has expected for two decades, an empty right column waiting for him to type. That is not what he gets. Every single target cell is already populated. Two hundred and eleven segments, top to bottom, filled with Spanish before his hands touched the keyboard. Some cells carry a small green icon and a percentage. Some carry a different icon and a label that reads MT. None of them are blank. The file did not arrive as a translation job in the sense he learned the word. It arrived as a draft that something else already wrote, and the project manager's instruction in the email said, in so many words, just clean it up. This lesson is about that screen: what put the words there, what each of those icons means, how the machine and the memory and the human stack on top of one another inside the CAT tool and the translation-management system, and why the moment every segment is pre-filled, the most important thing about your job quietly changes from writing to deciding whether to trust.

The Stack That Fills the Segment

To understand Omar's screen you have to understand that the Spanish in those cells did not come from one place. It came from a stack of systems, each one taking a pass at the file before he opened it, each one leaving its mark in a different color. Before we walk the stack, we need a shared vocabulary, because the words in this part of the industry are dense and most of them are acronyms.

A segment is the unit the whole machine works in. It is usually a sentence, sometimes a heading or a list item or a cell, the chunk of source text that the tool treats as one translatable thing. The file is not a flowing document to the software; it is a numbered list of segments, and every system in the stack operates segment by segment. When Omar's colleagues say "I did four thousand segments yesterday," they are describing units of work, not pages.

A CAT tool (computer-assisted translation) is the editing environment itself: the two-column grid, the status flags, the keyboard shortcuts that confirm a segment and jump to the next. It is where a human does the linguistic work. A TMS (translation-management system) is the layer above it, the platform that receives the source files from the client, splits them into segments, decides which linguist gets which file, applies the automated steps before the human ever sees them, and collects the finished work for delivery. In a modern shop the CAT tool and the TMS are often the same product or are wired tightly together, and the pre-translation that fills Omar's grid is something the TMS does automatically the instant the file is created, long before 9:14 on Tuesday.

A TM (translation memory) is a database of every source-and-target segment pair the team has translated and approved before. Translate "Insert the battery with the positive terminal facing up" into Spanish once, confirm it, and the pair is stored. The next time that exact sentence, or one close to it, appears in any future file, the TM can offer the stored Spanish back. The TM is the team's accumulated, human-verified memory, and it is the most trusted thing in the stack because a human already signed off on every pair in it.

A fuzzy match is what happens when the new source segment is close to something in the TM but not identical. The tool measures the similarity and reports it as a percentage. A 100% match means the source segment is word-for-word identical to a stored pair. A 95% match means almost identical, perhaps one word changed. A 75% match means recognizably related but meaningfully different. The percentage is not a guess about quality; it is a mechanical measure of how much of the source overlaps. That distinction matters enormously, and we will come back to it, because a high fuzzy percentage tells you the source is similar, not that the offered target is right for the new context.

MT (machine translation) is the engine that fills the segments the TM could not. NMT (neural machine translation) is the specific kind of MT that runs most production pipelines: a neural network trained only to map source sentences to target sentences. Increasingly an LLM (large language model), a general text-prediction system that translates as a side effect of its broad training, is wired into the same slot or layered on top of it. When a segment has no TM match, the TMS sends the source to the MT or LLM engine and drops the returned target into the cell. That is the MT icon Omar sees.

And MTPE (machine-translation post-editing) is the name for what Omar is now being asked to do: edit machine-produced output rather than translate from a blank cell. The human inside an MTPE workflow is a post-editor, and the act of editing is PE (post-editing). The order of operations across the whole stack is the thing to hold onto.

The target text in a pre-translated file is a layer cake: TM matches where the team has been before, MT or LLM output where it has not, and a human who must decide, segment by segment, which layer to trust.

How Pre-Translation Decides What Goes Where

The TMS follows a priority order when it pre-fills a file, and that order is the reason Omar's cells carry different icons. For each segment, the system asks a sequence of questions:

  • Is there a 100% or context match in the TM? If the exact source segment, ideally with the same surrounding segments, exists in the memory, the stored human-approved target is dropped in and the segment is often locked or marked as needing only a glance. This is the cheapest and safest fill.
  • Is there a high fuzzy match in the TM? If the source is, say, 85% similar to a stored pair, the tool offers the stored target with the differing words highlighted, so the human can adjust the small delta rather than retranslate the whole sentence.
  • Is there any usable fuzzy match at all, above the project threshold? Shops set a floor, commonly somewhere around 70 to 75%, below which a fuzzy match is considered more trouble than help.
  • No match? Send it to the engine. Every segment the TM cannot cover is routed to the MT or LLM engine, and the raw machine output fills the cell.

The result is a single file where some Spanish was written and approved by a human months ago, some was assembled by adjusting a near-match, and some was generated by an engine seconds before delivery, and all of it sits in the same column in the same font looking exactly alike. That uniformity is the central fact of the post-editor's day, and the central danger. The grid does not visually shout "this one is human-trusted and that one is raw machine guesswork." It shows you a metadata flag, a percentage, a small icon, and otherwise presents two hundred and eleven cells of equally clean, equally confident Spanish.

What MT-First Actually Means Operationally

The industry phrase for Omar's reality is MT-first, and it is worth being precise about what it does and does not mean. It does not mean the machine replaced the human. It means the machine goes first in the order of operations. The default starting state of a file is no longer blank; it is pre-populated. The human enters the process after the engine, not before it. That single reordering changes the economics, the deadline, the pricing, and most of all the cognitive posture of the person doing the work.

This is not an emerging trend you have time to prepare for. It is the established 2026 baseline. According to Nimdzi industry research, MTPE adoption rose from 26% in 2022 to roughly 46% in 2024, and 81.1% of language-service providers (LSPs) now offer MTPE, with roughly 70% offering subtitling, much of it machine-first as well. The pre-filled grid is not a special engagement Omar signed up for. It is the shape of mainstream work, and a linguist who treats it as exotic is a linguist who has not opened a representative file lately.

The economics explain both why the shift happened and why it presses on you the way it does. MTPE typically prices at 50 to 75% of full human translation, around $0.05 to $0.15 per word, with light post-editing sometimes as low as $0.02 per word. A hybrid MT-plus-human workflow lifts a linguist's throughput from roughly 2,000 words a day to 5,000 or more. From the client's seat, that reads as cheaper and faster. From the post-editor's seat, it reads as a file that arrived pre-filled, a deadline that assumes the machine already did the thinking, and a per-word rate that quietly assumes you are correcting rather than creating. Every one of those assumptions is a force pushing you to move fast and trust the clean cells, and every one of them is the precise mechanism by which a fluent error reaches the delivery.

The Blank Cell and the Filled Cell Are Different Jobs

It is tempting to think pre-translation just gives you a head start on the same task. It does not. Translating a blank cell and post-editing a filled one are different cognitive activities, and confusing them is where post-editors get hurt.

When the cell is blank, you read the source, build the meaning in your head, and produce the target. You are the author. You know exactly what you put in the sentence because you constructed it word by word, weighing each choice against the source you just read. Your attention is naturally anchored to the source, because the source is the thing you are converting.

When the cell is filled, the task inverts. A fluent, finished, confident target sentence is already sitting there, and your job is to decide whether it is right. The path of least resistance is to read the target, find it grammatical and natural, and accept it. But reading the target tells you only that the target is fluent. It tells you nothing about whether the target means what the source means. To know that, you have to do the thing the filled cell quietly discourages: drag your eyes back to the source, hold the source meaning in your head, and compare. The filled cell is engineered, by its very fluency, to make you skip that step. That is not a flaw in any one tool. It is the structural hazard of post-editing, and naming it is the first defense against it.

Walking Omar's Screen

Let us go back to the grid and read it the way a disciplined post-editor reads it, icon by icon, because the icons are the map of where to spend attention and where the engine has left a trap.

The 100% Match: The Comfortable Lie

Segment 3 carries a 100% match icon. The source is "Press and hold the power button for five seconds," and the TM returns the Spanish a colleague confirmed last quarter. A 100% match feels like the safest cell on the screen, and usually it is. But "safe" is doing work it has not entirely earned, for two reasons. First, a 100% match means the source segment is identical to a stored one; it does not mean the stored target was correct, only that a human once confirmed it. If an error was banked into the TM months ago, a 100% match propagates that error confidently into every future file forever. The match is only as trustworthy as the segment that seeded it. Second, a 100% match can be a context failure: the same source sentence can require different target wording depending on what surrounds it. "Open the cover" in a section about a battery is not necessarily the same imperative, with the same gendered article and noun, as "Open the cover" in a section about a paper tray. A bare 100% match without a context match can drop the right words for the wrong place. Most CAT tools distinguish a context match (often shown as 101% or with a dedicated icon) from a plain 100%, and the difference is exactly the surrounding-segment check. Read the flag, not just the number.

The Fuzzy Match: The Highlighted Delta

Segment 14 carries an 88% fuzzy match. The tool shows the stored source and target with the changed words highlighted: the previous segment said "Wait ten seconds," this one says "Wait thirty seconds." The tool has helpfully offered the old Spanish with "diez" highlighted so Omar can change it to "treinta." This is leverage at its best and its most treacherous. At its best, it saves him from retranslating an entire sentence to fix one word. At its most treacherous, the highlighted delta narrows his attention to exactly the words that changed and away from everything else, and the thing that changed in the source, the number, is precisely the kind of high-consequence element a tired post-editor can mis-key while focused on the highlight. The fuzzy percentage measures source overlap. It does not promise the target is right, and a high fuzzy match plus split attention is a classic place for a wrong number to slip through clean.

The MT Segment: The Confident Stranger

Segment 27 carries the MT icon. There was no TM match, so the source went to the engine and the engine returned Spanish that reads beautifully. This is the cell that most deserves and least invites scrutiny. It deserves scrutiny because it was produced by a system that optimizes for fluent, probable target text and only approximates faithfulness to the specific source in front of it, which is why machine output fails in characteristic ways: a dropped negation, a corrupted number, a swapped approved term, a silently omitted clause, an invented phrase on thin input. It least invites scrutiny because it is, by construction, the smoothest text on the screen. The engine's entire competence is fluency. Fluency is exactly the quality you have spent your career reading as a sign that a sentence is fine. On the MT segment, your professional instinct and the engine's failure mode point in the same wrong direction, and only a deliberate habit, going back to the source, pulls you off the trap.

The icons are a heat map of risk. A confirmed context match earns a glance; a raw MT segment earns the source-comparison you are most tempted to skip precisely because it reads the best.

The LLM Layer: When the Draft Is Even Smoother

Increasingly the engine in that slot is not a narrow NMT system but an LLM, or an LLM layered after the NMT pass to "improve" the draft. The operational effect for the post-editor is that the machine output gets even more fluent, even more context-aware, even more native-sounding, and therefore even harder to distrust. An LLM is more willing than a narrow NMT engine to smooth a confusing source by guessing intent, to add a clause that "should" be there, to render something it has effectively invented in authoritative prose. The smoother the draft, the higher the cost of the fluency illusion, and the LLM is the smoothest drafter in the stack. The icon may still just say MT. The behavior underneath it has changed, and the direction of the change is toward output that is more persuasive and not necessarily more correct.

The Trust Trap

We can now name the central psychological mechanism of the MT-first file directly, because Omar's whole morning runs on it. Call it the trust trap: the better the draft looks, the less you check it, and the machine's drafts now look very good. The trap has a specific anatomy worth dissecting, because understanding the mechanism is how you defeat it.

The first element is fluency as a false signal. For your entire career, fluent target text correlated with competent translation, because the only thing producing fluent target text was a competent human who had understood the source. That correlation is now broken. A machine produces fluent target text whether or not it understood anything, because fluency is what it optimizes for and meaning-faithfulness is not something it verifies. Your trained instinct, "this reads well, so it is probably fine," was reliable for thirty years and is now a liability, because the thing that reads well may have understood nothing.

The second element is volume fatigue. The economics that made MTPE attractive also raised the daily quota. A post-editor moving toward 5,000 words a day cannot give every segment the deliberation a from-scratch translator gives 2,000. Attention is finite and the file is long, and the cells that read smoothly are the natural place to economize attention, which is exactly where the fluent error hides. The workflow's speed and the error's camouflage are not separate problems. They are the same problem.

The third element is diffusion of responsibility. There is a quiet, almost unconscious sense that the engine, or the TM, or the project manager who set up the pre-translation, has already taken some responsibility for the text. The cell is filled, after all. Someone or something decided it was good enough to put there. But that sense is false. Nothing in the stack is accountable for the meaning of segment 27. The engine has no accountability; it is a probability machine. The TM only certifies that some past human approved some past segment. The PM certified nothing about the linguistic content. When the file ships, the only party who reviewed segment 27 against its source and let it through is the post-editor. The accountability did not diffuse. It concentrated, on the one human in the loop, and the filled cell disguises that concentration as shared comfort.

What Changes When the Draft Is Already There

Put the three elements together and you can describe precisely what the pre-filled file changes about the work. It does not change the standard the output is held to; a contraindication must still never be inverted, an approved term must still hold, a number must still match. What it changes is the path of least resistance. In a blank file, the path of least resistance is to read the source, because you cannot produce a target without it. In a pre-filled file, the path of least resistance is to read the target and accept it, because a plausible target is already there and the source is one extra glance away. The MT-first workflow quietly relocates the cheapest action from "engage the source" to "trust the surface," and it does so on a per-word rate and a deadline that reward the cheap action. The entire discipline of post-editing is the deliberate refusal of that relocation. It is the trained habit of treating every filled cell, especially the smooth ones, as a claim to be checked against the source rather than a result to be accepted.

The Post-Editor's Real Workflow

So what does disciplined work on Omar's file actually look like, segment by segment, once you have named the trap? It is not slower-is-always-better paranoia, which would erase the economics that make MTPE viable and earn him a reputation for over-editing. It is a structured allocation of a finite attention budget, spent where the risk is, governed by the icons and by the content.

  • Read the source first, deliberately, on every segment that carries consequence. The single habit that defeats the trust trap is reversing the natural reading order: source before target, meaning before surface. On low-stakes segments a glance suffices; on any segment carrying a number, a negation, a dosage, an obligation, a safety instruction, or an approved term, the source reading is non-negotiable and comes first.
  • Treat the MT and LLM cells as the high-risk zone. They had no human-verified TM behind them. Their fluency is manufactured, not earned. They are where the dropped negation and the corrupted number live. Spend disproportionate attention there, in inverse proportion to how trustworthy they look.
  • Treat fuzzy matches as a delta to verify, not just a delta to edit. Fix the highlighted change, then widen your eyes back out to the whole segment and confirm the unchanged part is actually right for this context. The highlight tells you where the source differs, not where the target might be wrong.
  • Do not over-edit. Post-editing is not retranslation. If the machine's rendering is accurate, on-term, and natural, changing it to your personal preferred phrasing burns budget you were not paid for and risks introducing an error into text that was fine. The job is to make it correct, not to make it yours. Preferential rewriting is the opposite failure from the trust trap, and it is just as expensive.
  • Check the elements the engine corrupts, every time, by design. Numbers, units, dates, negations, named entities, approved terms, and placeholders are the high-consequence, easily-corrupted elements. A reliable post-editor checks these on every segment as a fixed ritual, not as a function of how suspicious the cell happens to look, because the dangerous cell looks innocent.
  • Honor the termbase and the TM, but do not worship the match percentage. A 100% match can carry a banked error or a context mismatch; a 70% fuzzy can be exactly right after the delta is fixed. The percentage measures source overlap, not target correctness. Read it as a hint about where the leverage is, not as a verdict that the cell is done.

None of this is glamorous and all of it is the job. The post-editor who internalizes this stops being the machine's cleanup crew and becomes the one indispensable verifier in a pipeline full of confident systems that cannot verify themselves. The engine handed Omar speed. It also handed him a new failure mode and labeled it his responsibility, and the workflow above is how he discharges that responsibility without throwing the speed away.

Why This Is a Marvel, and Not a Demotion

It is easy to read all of this as a story of decline: the proud translator reduced to cleaning up after a machine for half the rate. That reading misses what actually happened to the value. The machine took over the part of the work that was always the most mechanical, the production of a plausible first draft of routine text, and left in human hands the part that was always the most valuable and is now the most exposed: the judgment about whether a fluent sentence is true to its source, safe for its reader, correct on its terms, and fit to ship with a name attached. The pre-filled grid did not lower the ceiling on the linguist's value. It raised the floor of speed and concentrated the human's contribution onto pure judgment. The post-editor who sees segment 27's confident, beautiful, subtly wrong Spanish and drags their eyes back to the English to catch it is doing something the entire stack of systems above them cannot do. That is not cleanup. That is the one job in the room that matters, and the MT-first workflow, for all its pressure, is what made that job the center of the work rather than its afterthought.

Key Takeaways

  • A pre-translated file is a layer cake assembled by a stack of systems before the human opens it: the TMS (translation-management system) splits the source into segments, fills exact and fuzzy matches from the TM (translation memory), routes every no-match segment to the MT (machine translation) or LLM engine, and presents all of it in one uniform column inside the CAT tool (computer-assisted translation environment).
  • MT-first means the machine goes first in the order of operations, not that it replaced the human. The default state of a file is now pre-populated, and the human enters the workflow after the engine as a post-editor doing PE (post-editing), the workflow named MTPE (machine-translation post-editing).
  • MT-first is the established 2026 baseline: MTPE adoption rose from 26% in 2022 to roughly 46% in 2024, 81.1% of LSPs (language-service providers) now offer it, it prices at 50 to 75% of full human translation, and it lifts throughput from about 2,000 to 5,000-plus words a day.
  • Translating a blank cell and post-editing a filled one are different jobs. The blank cell anchors your attention to the source because you must convert it; the filled cell tempts you to read only the target, find it fluent, and accept it, which tells you nothing about whether it means what the source means.
  • Read the icons as a risk heat map. A context match (often 101%) earns a glance, a plain 100% match can carry a banked error or a context mismatch, a fuzzy match is a delta to verify rather than just edit, and a raw MT or LLM segment is the highest-risk cell precisely because its manufactured fluency makes it look the most trustworthy.
  • The fuzzy-match percentage measures how much of the source overlaps a stored segment, not whether the offered target is correct for the new context. A high percentage tells you where the leverage is; it is not a verdict that the cell is done.
  • The trust trap has three parts: fluency is now a false signal because machines produce fluent text without understanding, volume fatigue thins attention on the cells that read smoothly, and responsibility feels diffused even though accountability actually concentrates on the single human who lets the segment ship.
  • The discipline of post-editing is the deliberate refusal of the path of least resistance: read the source first on every consequential segment, spend disproportionate attention on the smooth MT and LLM cells, check numbers, negations, dosages, terms, and placeholders every time by ritual, and fix what is wrong without over-editing what is merely not your preference.