Post-Editing Productivity Without Quality Loss
Theo took the rush job because the rate was good and the deadline was insane: 28,000 words of consumer-electronics support content, German into English, machine-translated already, due in five working days. The project manager had done the arithmetic out loud on the call. "It's MTPE, the engine's already drafted everything, you just clean it. Call it five thousand six hundred words a day. You've done four thousand before, this is a stretch but it's doable." Theo said yes, because the math was real and the money was real, and machine-translation post-editing (MTPE, the workflow where an engine drafts every segment and a human revises it) is supposed to make exactly this kind of throughput possible. For four days it went beautifully. He was flying, clearing segments faster than he ever had, and the file looked clean. On day five, at 4:40 p.m., twenty minutes from delivery and roughly 26,000 words in, he hit a segment that read: "If the battery overheats, you may continue charging the device." He nearly let it pass. It was fluent, it was grammatical, it sat in a paragraph of plausible troubleshooting advice, and his eyes were doing 5,600-words-a-day skimming. Something snagged. He opened the source. The German said the opposite: do not continue charging. A flipped negation, in a battery-safety instruction, in flawless prose, at the exact moment fatigue and speed had hollowed out his attention. He fixed it, delivered on time, and then sat very still for a while, because he understood that he had not caught that error through skill. He had caught it through luck, and luck is not a quality system. This lesson is about how to do what Theo did on purpose: push throughput past 5,000 words a day and still catch the one fluent, silent, critical error, by building speed out of triage and discipline instead of borrowing it from your own vigilance. It is honest about the tension, because the tension is real, and it gives you a worked productivity setup that resolves it.
The Tension Nobody Wants to Name
Let us put the uncomfortable thing on the table first, because every productivity technique in this lesson is an answer to it, and if you pretend it does not exist you will build the wrong habits. Speed and quality are, at the level of raw attention, in genuine competition. A human reading carefully catches more than a human reading fast. That is not a motivational poster you can override with willpower; it is how attention works. The faster you move through segments, the more your eyes do pattern-matching instead of comprehension, and pattern-matching is exactly the failure mode that lets a fluent mistranslation through, because a fluent mistranslation matches the pattern of correct text. The danger is not that you will read fast and the text will look obviously broken. The danger is that you will read fast and the text will look perfect, because the engine's most dangerous output is the sentence that is grammatically flawless and means the opposite of the source.
So when a project manager says "MTPE lifts you from 2,000 words a day to 5,000 or more," they are stating a real industry benchmark, and they are also, if you are not careful, describing the precise conditions under which you ship the error that ends a relationship. The 2024 through 2026 numbers are consistent: a disciplined hybrid workflow, machine translation plus a human post-editor, is supposed to push a linguist from roughly 2,000 words a day to 5,000 or more, and MTPE prices at roughly 50 to 75% of full human translation precisely because that throughput is assumed. The throughput is the premise of the rate. But the throughput is also the thing that, done wrong, manufactures the silent critical error. This is the tension. You cannot wish it away, and you should be suspicious of anyone who tells you that going faster is simply free.
Speed and quality compete for the same scarce resource: your attention. Productivity in post-editing is not the art of caring less. It is the art of spending your finite attention only where an error can actually live.
Here is the reframe that makes the rest of this lesson possible, and it is the single most important idea in it. The naive view of productivity is that you go faster by reading everything faster, by lowering your standard uniformly across the file so each segment gets less care. That view is a trap, because it lowers your care on the dangerous segments by exactly as much as it lowers care on the safe ones, and the dangerous segments are where the Critical hides. The professional view of productivity is the opposite: you go faster by reading unevenly. You spend almost no attention on the segments that cannot hurt you and you concentrate the attention you saved onto the segments that can. Speed is not a uniform dimming of the lights. It is a spotlight. The whole craft is learning where to point it.
One acronym before we go further, because the spotlight runs on it. Quality estimation (QE) is the automatic confidence score a system assigns to each machine-translated segment, an algorithm's guess at how likely that segment is to be wrong. QE is not a verdict and it never clears a segment for delivery on its own. But as a way of deciding where to point the spotlight first, it is the most powerful single lever you have, and most of this lesson is about using it well without trusting it blindly.
Triage by QE: Pointing the Spotlight
The first and largest source of honest speed is to stop treating every segment as equally deserving of your attention, because they are not. On a typical MT-first file, the engine produces a distribution: a large mass of segments that are genuinely fine, a smaller band that are subtly off, and a thin, deadly slice that are confidently and completely wrong. If you give all three the same read, you waste most of your attention on the safe mass and arrive at the deadly slice already exhausted. That is the Theo failure mode. Triage is how you avoid it. Triage means deciding, before you read a segment for meaning, how much scrutiny it deserves, and the best available input to that decision is the QE score the system already computed.
Read the QE band, not the number
QE typically surfaces as a per-segment confidence, often a number or a color band: high confidence, medium, low. The mistake beginners make is treating the number as precise, as if 0.91 and 0.89 mean something different. They do not. QE is a noisy signal, so use it the way it is actually informative, which is in coarse bands.
- Low-confidence band. The engine itself is unsure. These segments get your fullest attention, every time, source open, no exceptions. The engine waving a flag is the cheapest warning you will ever get; ignoring it is malpractice.
- Medium-confidence band. The treacherous middle. The engine thinks it is probably fine. Sometimes it is, sometimes it is hiding a fluent error. These get a real source check, faster than the low band but never skimmed.
- High-confidence band. The mass of the file. The engine is confident, and it is usually right. These get a fast read, but, and this is the load-bearing caveat, never a blind pass. High QE confidence reduces the probability of an error; it does not eliminate it, and the errors that survive in the high band are exactly the fluent, confident ones that QE is worst at flagging.
The productivity gain is enormous and it is legitimate. Instead of spending equal time on 28,000 words, you spend most of your time on the low and medium bands, which might be a quarter of the file, and you move quickly through the high-confidence mass. That is where the leap from 2,000 to 5,000 words a day actually comes from. Not from reading everything faster, but from reading the safe majority faster so you can read the risky minority slower.
Why QE can never be the whole answer
Now the warning, because triage by QE has a structural blind spot that, untreated, will eventually ship a Critical. QE estimates how confident the model is, and the model's confidence is correlated with fluency. A segment that reads beautifully tends to get a high QE score. But the silent critical error is, by definition, the one that reads beautifully and means the wrong thing. So the exact failure mode that hurts you most, the fluent mistranslation, is the one QE is structurally worst at catching, because it looks to the model exactly like a confident, correct rendering. Theo's flipped negation almost certainly carried a respectable QE score. The engine was not unsure; it was confidently wrong.
This is why QE routes attention but never clears a segment, and why certain categories get checked regardless of what the QE band says. You triage by QE to decide where to start and how to allocate the bulk of your time, and then you overlay a set of content-based checks that run on every file no matter how green the QE looks. We will get to those non-negotiable checks; for now, hold the rule: QE points the spotlight, but some things you check in the dark, by hand, every time, because the engine's confidence is precisely what cannot be trusted on them.
QE tells you where the engine is unsure. It cannot tell you where the engine is confidently wrong, and confidently wrong is the failure mode that costs a life or a lawsuit. Route by QE; verify the dangerous categories regardless of it.
Fix-What-Matters: The Discipline That Funds Speed
Triage decides where you look. The next discipline decides what you do once you are looking, and it is the one that most often separates a 5,000-word day from a 2,500-word day on identical content. It is the discipline of fixing only what is actually defective and leaving acceptable output alone, and it is the direct subject of the previous lesson on over-editing, so we will be brief on the principle and concrete on its productivity consequences.
The principle: there are two kinds of edit. A necessary edit fixes something genuinely wrong, an accuracy defect, a wrong number, a terminology violation, a locale failure, a real grammatical error, meaning-destroying garble. A preferential edit replaces correct, clear, acceptable output with output you simply like better. The international standard for post-editing, ISO 18587, directs the post-editor to fix errors and make no purely preferential changes, using as much of the raw machine output as possible. The revised standard, in DIS ballot with publication targeted for late 2025 into 2026, raises the bar by requiring the post-editor to hold full professional-translator competence, but it does not relax the restraint; it asks you to be a full translator who chooses not to rewrite acceptable text.
Here is why this is a productivity lesson and not just an economics lesson. Every preferential edit costs you twice. It costs the seconds you spend making it, which is the obvious cost. And it costs the attention you spend deciding to make it, which is the hidden and larger cost, because attention is the scarce resource you need for catching the Critical. When Theo was flying for four days, part of what made him fast was that he was not stopping to improve acceptable segments. The moment a post-editor starts polishing, two things happen at once: the clock burns and the attention budget drains. The over-editor does not just run slow; they run slow and arrive at the dangerous segments depleted, which is the worst of both worlds. Restraint is not a virtue you practice instead of speed. Restraint is where a large part of the speed comes from, and it is also what preserves the attention that speed would otherwise spend.
The half-second test, at throughput
At 5,000 words a day you do not have time to deliberate over each edit, so the decision has to be a reflex. Train one question to fire automatically on every segment: would this, untouched, mislead, harm, embarrass, or fail a defined requirement? Yes means edit, it is necessary. No means leave it, even if you can imagine something better. The test is about consequence, not taste, and at throughput its real job is to kill the reflexive reach toward acceptable text before it starts, so the seconds and the attention both stay in your pocket for the segments that need them. A post-editor who has internalized this test is not making the same number of edits faster. They are making far fewer edits, and that is the speed.
Batching: Fixing the Same Error Once Instead of Forty Times
The next gear of honest throughput is batching, and it is the technique most underused by linguists who think of speed as a per-segment property. It is not. A large fraction of MT errors are not unique; they are systematic. The engine that mistranslated the client's approved device name once mistranslated it everywhere it appeared. The engine that used the British "colour" did it across the whole file. The engine that rendered a recurring UI label inconsistently did so the same wrong way in forty segments. If you fix these one segment at a time as you encounter them, you pay the cognitive cost of recognizing and deciding the same fix forty times. If you batch them, you pay it once.
The systematic errors worth batching
- Terminology drift. The engine consistently prefers a synonym over the client's approved term. This is the highest-value batch, because terminology is a defined requirement and the fix is mechanical. Run a find-and-replace, verified, across the whole file in one pass, rather than catching the wrong term forty separate times.
- Locale conventions. A spelling convention, a date format, a decimal separator, a quotation-mark style that the engine got systematically wrong for the target locale. One pass, one rule, the whole file.
- Recurring phrases and boilerplate. Repeated UI strings, standard warnings, legal boilerplate, navigation labels. Decide the correct rendering once, then enforce it everywhere it recurs.
- A consistent engine tic. Many engines have a signature habit on a given language pair: over-formalizing, misrendering a particular construction, mishandling a specific connector. Once you have seen it twice in a file, assume it is everywhere and hunt it deliberately rather than waiting to stumble on each instance.
Batching does two things for productivity, and the second matters more than the first. Obviously it saves time, one decision instead of forty. Less obviously, it improves consistency in a way that protects the translation memory (TM, the database that stores your approved segments for reuse). If you fix a recurring error inconsistently across forty segments because you were tired by segment thirty, you have just seeded the TM with variants that will degrade future leverage. Batching enforces a single correct decision uniformly, which keeps the file consistent and keeps the asset clean. Speed and quality align here too: the faster method is also the more consistent one.
When not to batch, and the trap inside batching
One caution, because batching has a sharp edge. A find-and-replace is a blunt instrument, and the same word can be right in one context and wrong in another. The approved term might genuinely not apply in a particular segment; a global replace can introduce errors as fast as it removes them. So batch the decision but verify the application: when you replace across the file, review each replacement rather than trusting the count. The productivity still holds, because reviewing a flagged list of forty changes is far faster than discovering and fixing forty errors blind, but the discipline of checking each application is what keeps batching from becoming a new error source. Batch boldly; apply carefully.
Most MT errors are not unique, they are systematic. Recognize a defect once, fix it everywhere in a single verified pass, and you convert forty decisions into one without lowering your care on any of them.
Keyboard and CAT Efficiency: The Free Speed
There is a category of throughput that costs you nothing in attention, which makes it the only free lunch in this lesson: mechanical efficiency in your computer-assisted-translation (CAT) tool, the editing environment where you work the segments. Every second you spend reaching for the mouse, hunting through menus, or retyping something the tool could insert is a second stolen from the file with zero quality benefit. Unlike reading faster, getting faster at the mechanics has no downside, because it does not touch your attention budget at all. It is pure gain, and most linguists leave a great deal of it on the table.
The mechanics that actually move the needle
- Learn the segment-navigation and confirm shortcuts cold. Moving to the next segment, confirming a segment, jumping to the next untranslated or next low-QE segment: these are the motions you make thousands of times a day. Doing them by keyboard instead of mouse, without thinking, is the largest single mechanical saving available.
- Use terminology and TM insertion shortcuts. When the termbase suggests the approved term or the TM offers a match, insert it with a keystroke rather than retyping. This is faster and it also prevents the typo that retyping introduces.
- Filter and sort the file by QE band. If your tool can sort or filter segments by QE score, you can work the low-confidence band as a group, then the medium, then sweep the high. This operationalizes triage mechanically instead of leaving it to your memory.
- Master find-and-replace and the QA-check panel. These are the engines of batching and of the never-skipped checks. Knowing them cold turns a ten-minute manual hunt into a thirty-second filtered pass.
- Use auto-propagation deliberately. Many CAT tools auto-propagate a confirmed segment to identical segments. This is batching for free, but it carries the same trap: verify that the propagated context is genuinely identical, not just textually identical.
The reason mechanical efficiency belongs in a lesson about quality, not just speed, is that it is the only kind of speed that does not borrow against your attention. When you save a second by knowing a shortcut, that second is genuinely free. When you save a second by reading a segment less carefully, that second is a loan against the moment you hit a flipped negation. Maximize the free speed first and hardest, because every second you claw back mechanically is a second you do not have to claw back by cutting corners on comprehension. The linguists who hit real throughput without quality loss are almost always the ones whose hands have stopped touching the mouse.
Fatigue Management: Protecting the Attention You Run On
Everything so far assumes a steady supply of the scarce resource. But attention is not steady. It degrades over a session, sharply and predictably, and the degradation is invisible from the inside, which is what makes it dangerous. Theo's error happened at 4:40 p.m. on day five, 26,000 words in, and that is not a coincidence. The flipped negation was probably no harder to catch than a dozen he had caught earlier in the week; what had changed was him. Fatigue does not announce itself. It does not feel like "I can no longer catch errors." It feels like "I am cruising, the file is clean, I am almost done," which is the exact subjective state in which a tired post-editor skims past the Critical. Managing fatigue is not wellness advice bolted onto a productivity lesson. It is the productivity lesson, because your throughput is only as good as the attention you bring to the dangerous segments, and fatigue is what hollows that attention out.
The specific shape of post-editing fatigue
Post-editing has a particular fatigue profile that differs from translation from scratch. Translation is generative and varied; it keeps you engaged because you are producing. Post-editing is evaluative and repetitive: read, judge, mostly accept, occasionally fix, advance. That rhythm, accept, accept, accept, advance, advance, is hypnotic, and hypnosis is the enemy of catching the anomaly. The long stretches of acceptable segments, which are good for speed, lull you into a rhythm of approval, so that when the one bad segment arrives wearing the same clothes as the good ones, your judgment is set to "accept" by momentum. The high-confidence mass that makes you fast is the same mass that lulls you, and the Critical is most likely to slip through right after a long run of clean segments, when your guard is lowest precisely because nothing has been wrong for a while.
The countermeasures that actually work
- Break before you think you need to. The break that protects you is the one taken before fatigue is conscious, because by the time you feel it, you have already been missing things for a while. Short, regular breaks, on a timer, not on a feeling, keep the attention floor higher than long marathon stretches followed by collapse. The total throughput over a day is higher with breaks, not lower, because the post-break segments are read at full attention instead of half.
- Put the hardest content where your attention is highest. If you know a file contains a high-stakes section, a safety warning block, a section with many numbers, a legally sensitive passage, schedule it for early in your session or right after a break, not for the exhausted tail. Theo's battery instruction landed in the worst possible slot by accident; you can place your high-stakes work in the best slot on purpose.
- Cap the day honestly. Five thousand words a day is a sustainable benchmark; it is not a floor to exceed indefinitely on willpower. The eleventh thousand words is read at a fraction of the attention of the first thousand. A linguist who routinely pushes to 8,000 by grinding late is not more productive; they are manufacturing the conditions for a shipped Critical and calling it dedication.
- Change modes to break the trance. When you feel the accept-accept rhythm setting in, deliberately switch tasks for a few minutes: run a batch find-and-replace pass, do a numbers-only sweep, run the QA panel. Changing the cognitive mode disrupts the hypnosis and resets your anomaly detection.
Fatigue does not feel like fatigue; it feels like cruising. The danger is not that you will know your attention has dropped, it is that you will not. Protect attention with structure, breaks, placement, and honest caps, because you cannot protect it with willpower you cannot feel running out.
The Checks That Never Get Skipped, No Matter the Deadline
Now the spine of the whole system, the part that makes speed safe rather than reckless. Triage, restraint, batching, mechanical efficiency, and fatigue management all increase throughput, and all of them, pushed hard, increase the risk of skimming past the silent critical error. The thing that holds the line is a small set of checks that run on every file, by rule, regardless of QE band, regardless of deadline, regardless of how clean the file looks and how tired you are. These are the floor. Everything else flexes; these do not. They are deliberately few, because a check that is too long to run under deadline pressure is a check that gets skipped under deadline pressure, and a skipped check is worse than no check because it breeds false confidence.
The non-negotiable list
- Numbers. Every number in the target verified against the source: dosages, prices, quantities, measurements, percentages, dates, times, model numbers. Numbers are where the fluent error is most catastrophic and most invisible, because a wrong number reads exactly as plausibly as a right one. This check is mechanical and fast and it is never skipped. A numbers-only sweep across the whole file is one of the highest-value passes you can run.
- Negations. Every "not," "no," "never," "without," "do not," and every inverted condition checked against the source for a flip. This is Theo's error. A dropped or inverted negation reverses meaning while keeping the prose flawless, and it is the single most dangerous pattern in MT output. You hunt negations deliberately; you do not wait to notice them.
- Approved terminology. Every client-mandated term confirmed present and correct, every forbidden term confirmed absent. This is a defined requirement, not a preference, and it is exactly where engines drift. The termbase and the QA panel make this fast; running it is non-negotiable.
- Obligation and prohibition language. "Must," "shall," "may," "should," "is required to," "is prohibited from": the modal verbs that carry legal and safety weight. An engine that softens a "must" to a "should" or flips a permission to a prohibition has changed the meaning that matters most, in content where meaning is liability.
- Placeholders, tags, and length where applicable. Code placeholders intact, tags balanced, length within the UI budget for software content. A broken placeholder ships a defect that no amount of linguistic quality redeems.
The discipline is not that you run these checks when you have time. It is that you run them especially when you do not have time, because the deadline pressure that tempts you to skip them is exactly the condition under which errors ship. The checks are short by design so that "no time" is never a real excuse. A numbers sweep, a negation hunt, and a terminology QA pass on a 5,000-word day cost minutes, not hours, and they are the difference between throughput that is fast and defensible and throughput that is fast and lucky. Theo was lucky. A system does not run on luck.
Why the floor is what licenses the speed
Here is the synthesis that ties the lesson together. The reason you can move fast through the high-confidence mass without dread is that the never-skipped checks form a safety net underneath the whole file. You triage aggressively by QE, you restrain your edits, you batch the systematic errors, you fly through the clean segments, and the thing that lets you do all of that without gambling is the knowledge that, regardless of how the per-segment work went, a numbers sweep and a negation hunt and a terminology check will pass over the entire file before it ships. The floor is what converts aggressive speed from recklessness into method. Without it, every fast pass is a roll of the dice. With it, the speed is bounded by a guarantee: no matter how fast I went, these specific catastrophic categories were verified by rule. That is how you get past 5,000 words a day and still sleep, because the silent critical error has to survive both your attention and the floor, and the floor does not get tired.
Theo's Second Rush Job: The Setup That Works
Months later Theo took a similar job, 25,000 words, four days, and this time he did not run on luck. Watch the setup, because it assembles every technique in this lesson into a single working method, and the throughput it produces is real, not bravado.
Before he touched a segment, he read the brief and the risk tier. This was consumer support content, full post-editing, not a regulated drug label, so the acceptable bar was "correct, clear, on-term, on-locale," not "rewritten to my taste." He confirmed the approved-terms list and the target locale, and he noted that the file contained a safety-warning section, which he flagged for a high-attention slot.
He set up triage. He sorted the file by QE band so the low-confidence segments were grouped. He decided up front that the low band got full source-open scrutiny, the medium band got a real but faster source check, and the high band got a fast read plus the floor checks at the end. He did not pretend the high band was safe; he decided to cover it with the never-skipped passes rather than with per-segment vigilance.
He worked the bands, restraining edits. Through the high-confidence mass he moved fast, making only necessary edits, leaving acceptable segments alone, naming each edit "necessary" or "preferential" in the half-second before his hands moved. When he noticed the engine consistently using "switch off" where the client's term was "power down," he stopped and batched it: one verified find-and-replace across the file, each application reviewed, forty decisions collapsed into one.
He managed his attention deliberately. He took the safety-warning section first thing in the morning, at full attention, not at the exhausted tail. He broke on a timer, not on feeling. When the accept-accept trance set in around mid-afternoon, he switched modes and ran a batch pass to reset. He capped each day at a sustainable number rather than grinding to exhaustion to bank words early.
Before delivery, regardless of the clock, he ran the floor. A numbers-only sweep against the source. A negation hunt for every "not" and inverted condition, which is where the battery instruction would have been hiding had this file contained it. A terminology QA pass against the approved list. An obligation-language check on the safety section. A placeholder and tag check. The floor took under half an hour and it ran because the deadline was tight, not in spite of it.
He delivered on time, past 5,000 words a day, with a method that did not depend on a lucky snag at 4:40 p.m. The throughput was the same as the rush job that scared him. The difference was that the speed now came from triage, restraint, batching, and mechanics, and the safety came from a floor that ran no matter what, rather than from his own fraying attention. That is the rare win the lesson promised: the speed move and the quality move turned out to be the same move, because the disciplines that make you fast, spending attention only where errors live, are the same disciplines that make you safe.
Throughput past 5,000 words a day is not built from reading everything faster. It is built from triage that points attention, restraint that conserves it, batching that multiplies it, mechanics that come free, fatigue management that protects it, and a floor of never-skipped checks that catches what slips. Assemble those and speed stops costing quality.
Key Takeaways
- Speed and attention compete; productivity is allocation, not uniform haste. You do not go faster by reading every segment less carefully. You go faster by reading the safe majority quickly so you can read the dangerous minority slowly. Speed is a spotlight, not a dimmer.
- Triage by QE points the spotlight. Quality estimation (QE), the engine's per-segment confidence score, routes your attention: low-confidence segments get full scrutiny, medium get a real source check, the high-confidence mass gets a fast read. Most of the leap from 2,000 to 5,000 words a day comes from this reallocation.
- QE has a fatal blind spot. It is structurally worst at catching the fluent, confident mistranslation, because that error looks to the model exactly like correct output. Route by QE, but verify the catastrophic categories by hand on every file regardless of the QE band.
- Fix-what-matters funds the speed. Necessary edits fix real defects; preferential edits rewrite acceptable text. Every preferential edit costs seconds and, worse, drains the attention you need for the Critical. Restraint (ISO 18587's "no preferential changes") is where much of the throughput comes from.
- Batch systematic errors. Most MT errors recur: terminology drift, locale conventions, boilerplate, engine tics. Decide the fix once and apply it across the file in a single verified pass, but review each application, because find-and-replace is blunt. Batching saves time and protects TM consistency at once.
- Mechanical efficiency is free speed. Keyboard shortcuts, TM and termbase insertion, QE sorting, find-and-replace, and the QA panel buy throughput without spending attention. Claw back speed mechanically first, because every second saved on the mouse is a second you do not steal from comprehension.
- Manage fatigue as a quality control. Attention degrades invisibly and feels like cruising. Break on a timer not a feeling, place high-stakes content where attention is highest, cap the day honestly, and switch modes to break the accept-accept trance. The Critical slips through right after a long run of clean segments.
- The floor never gets skipped. Numbers, negations, approved terminology, obligation language, and placeholders get checked on every file by rule, especially under deadline pressure, because that pressure is when errors ship. The floor is short by design and it is what converts aggressive speed from a gamble into a method.
Skill.re