ISO 18587 and Post-Editing - Including the Revision
In 2017 a standards committee in Geneva published a quiet little document, twelve pages of dry requirements, that almost nobody outside the language industry ever read. It was called ISO 18587, and it did something the translation world had been doing badly for years: it wrote down, in plain procedural language, what it actually means for a human to "fix" a machine translation. It gave the industry two words that linguists, project managers, and clients would argue about for the next eight years: light and full. Light post-editing meant "make it understandable, leave the rest." Full post-editing meant "make it good enough that nobody can tell a machine started it." For most of a decade those two words were the whole map. Then the ground moved. The machines stopped being clumsy neural engines that produced obviously broken sentences and became fluent large language models that produced beautiful, confident, native-sounding prose that was sometimes completely wrong. The old map no longer matched the territory. And so, in the run-up to 2026, the same committee did something rare for a standards body: it admitted the map was out of date and redrew it. This lesson is about that redrawing. It is about the old light-versus-full split and exactly what each meant, the revision now in DIS ballot with publication targeted for late 2025 into 2026, and the single sentence in that revision that should change how you price your work, defend your deliveries, and think about your own career: the post-editor must now hold the same competence as a professional translator.
What ISO 18587 Actually Is
Before we touch the revision, you need to know what the standard is, because most of the people who invoke it have never read it and use it as a vague badge rather than a real requirement. ISO 18587 is the international standard that defines the requirements for the human post-editing of machine-translation output. That phrase is doing a lot of work, so let us unpack every part of it.
Machine translation (MT) is any system that converts text from a source language to a target language with no human writing the words. Neural machine translation (NMT) is the production-grade flavor that has dominated since around 2016: a neural network trained specifically and only on the translation task. A large language model (LLM) is a general-purpose text-prediction system, trained to continue text plausibly across every domain, that translates as a side effect of that general competence. Post-editing (PE) is the act of a human editing that machine output instead of translating from a blank page. Machine-translation post-editing (MTPE) is the named workflow built around that act: the engine drafts every segment, the human revises. A segment is the unit a CAT tool (computer-assisted translation tool) breaks the text into, usually a sentence, and it is the atom the whole pipeline operates on.
ISO 18587, then, is the rulebook for what a human must do, and what an organization must guarantee, when the workflow is "the machine drafts, the human cleans." It does not tell you the machine is good. It does not endorse any engine. It assumes the machine has produced something and specifies the human discipline that turns that something into a deliverable a client can rely on. It sits inside a small family of related standards, and you should hold the family in your head as a single picture:
- ISO 17100 is the baseline for human translation services. It defines what a professional translation process looks like: a qualified translator, a separate qualified reviser, defined competences, a documented process. It is the reference point for "what good human translation work is."
- ISO 18587 is the standard for post-editing machine output. Historically it was the lighter, MT-specific cousin of 17100, written for a world where the machine did the rough draft and the human tidied.
- ISO 5060:2024 is the newer standard for evaluating translation output: the analytic, MQM-aligned model that classifies each error by dimension and severity (Critical, Major, Minor) and decides whether output ships. MQM is the Multidimensional Quality Metrics framework, the error typology that scoring sits on.
Keep that triangle in mind: 17100 is the human baseline, 18587 is the post-editing process, 5060 is the scorecard. The revision we are about to discuss is, in large part, the story of 18587 moving closer to 17100, until the gap between "post-editor" and "translator" nearly closes.
ISO 18587 is not a quality certificate for the machine. It is a process requirement for the human who is accountable for the machine's output.
The Old Light-vs-Full Split
For most of its life, the single most quoted thing about ISO 18587 was its distinction between two levels of post-editing effort. This split was genuinely useful, and it shaped how the entire industry priced and scoped work, so you need to understand exactly what each level meant before you can understand why the revision is dismantling it.
Light Post-Editing
Light post-editing (light PE) was the cheaper, faster pass. Its goal was a single word: understandable. The post-editor read the machine output and intervened only enough to make the meaning correct and comprehensible. The brief, stated plainly, was: do not chase style, do not impose your preferences, do not polish. If the machine produced a clunky but accurate sentence, you left it clunky. You fixed outright mistranslations, you fixed anything that confused the meaning, you fixed errors that would mislead a reader, and you stopped there. The output was allowed to read like a machine produced it, as long as a reader could understand it correctly. The classic use case was high-volume, low-shelf-life, internal-facing content: a knowledge-base article nobody would frame, a user-generated review, a product description in a catalog of fifty thousand SKUs. The asymmetry that made light PE rational was that the content's consequence was low, so the cheapest pass that preserved meaning was the economically correct one.
Full Post-Editing
Full post-editing (full PE) was the higher-effort pass, and its goal was a different word: indistinguishable. The output had to read as though a competent human translator had produced it from scratch. The post-editor fixed not only the mistranslations and the comprehension errors but the style, the register, the terminology, the locale conventions, the flow, everything. The brief was: deliver a translation that meets the quality a client would expect from full human translation, with no tell that a machine started it. The use case was client-facing, brand-sensitive, or higher-stakes content: a website, marketing copy, documentation a customer reads, anything where the surface quality was part of the product. Full PE cost more and took longer because the post-editor was doing nearly the full job of a translator, just starting from a draft instead of a blank page.
Why the Split Was Useful, and Where It Cracked
The light-versus-full split was useful because it gave everyone a shared vocabulary for matching effort to consequence. A project manager could quote a client "light PE on the catalog, full PE on the homepage" and both sides knew roughly what they were buying. It connected to pricing: light PE could be priced as low as about two cents a word, full PE landed in the range of full MTPE, roughly five to fifteen cents a word, against full human translation that might be double that. It gave linguists a way to scope their own effort and not over-invest in throwaway content. For a decade, that two-level map was the working tool of the field.
But the map had a crack running through it, and the crack widened as the machines changed. The split assumed a clean correlation between effort and risk: light effort for low-risk content, full effort for high-risk content. In practice, the two levels were defined by how much polishing the output got, not by how dangerous a hidden error would be. Those are not the same axis. A piece of content could need only "understandable" surface quality and yet carry a catastrophic risk if a single fact flipped. And the new machines made the problem acute, because an LLM produces output that is already fully fluent, already indistinguishable on the surface, while being silently wrong underneath. The light-versus-full question, "how much do I polish," became almost meaningless when the machine had already polished everything to a mirror shine and the only real work left was verifying that the gorgeous prose was true. The two-level map was built for a world of visibly rough machine output. It did not fit a world of fluent machine output that needed not polishing but interrogation.
Why the Machine Changed the Rules
To understand the revision you have to feel the shift in the machines, because the standard is reacting to it. The post-editing standard of 2017 was written in the age of neural machine translation. NMT engines were good, far better than the statistical systems before them, but their failures still tended to look like machine failures: an awkward construction, a word in the wrong place, a stiffness that signaled "an engine made this." A post-editor reading NMT output could often spot the machine's fingerprints, and the light-versus-full distinction made sense against that texture: light PE smoothed the worst of it, full PE erased it.
Then large language models arrived in production translation workflows, and the texture changed. An LLM is trained to continue text plausibly, which means it is optimized, above all, for fluency. Its output is grammatical, idiomatic, in the right register, and confident, even when it is wrong. The fingerprints disappeared. The machine stopped producing prose that looked machine-made and started producing prose that looked like a careful native speaker wrote it. This is the deep shift: fluency stopped being a signal of correctness. In the NMT era, smooth often correlated with right and rough often correlated with wrong, so a post-editor reading for flow was a half-reasonable heuristic. In the LLM era, smooth correlates with nothing about correctness. The engine writes the wrong drug name with exactly the same calm authority as the right one. The measured evidence is stark: studies of LLM output on medical content found error rates of roughly 59% on drug names, roughly 60% on dates and times, and roughly 66% on adverse events, every one delivered in grammatically perfect prose.
This is why the old map broke. Light versus full was a question about surface effort. The new machines made surface effort nearly free and made the dangerous work entirely about verification against the source, an activity the two-level model never named. A post-editor in 2026 is rarely deciding "how much do I polish this rough draft." They are deciding "this draft is already polished, so where among these beautiful sentences is the one that lies, and is this content the kind where a lie is a recall or a lawsuit." That is not a light-or-full question. It is a question about effort distributed along a spectrum and matched to consequence, and it demands the judgment of a full translator, not the lighter touch the standard once implied a post-editor could get away with.
In the NMT era, fluent often meant right. In the LLM era, fluent means nothing about right. The standard had to stop measuring polish and start measuring the verification only a full translator can do.
The Revision: What Actually Changed
The revised ISO 18587 is, as of this writing, in DIS ballot, the Draft International Standard stage, with publication targeted for late 2025 into 2026. It is not a tweak. It is a redrawing of the standard to fit the world the machines created. There are five changes that matter to you, and we will take each one slowly, because each one changes something concrete about how you work, price, and defend your output.
Change One: Scope Expands From MT to "Non-Human Translation Output"
The original standard's scope was explicitly machine translation: the output of an MT engine. That framing made sense in 2017, but it left a loophole that grew into a chasm. When an LLM produces a translation, is that "machine translation" in the standard's sense? The engines, the workflows, and the failure modes are different enough that the question was real, and a vendor could argue an LLM-based workflow fell outside 18587's scope entirely. The revision closes that loophole by expanding the scope from "machine translation" to non-human translation output, with AI and LLMs named explicitly. The standard now covers any output not written by a human, regardless of which kind of engine produced it. This matters because it pulls the entire wave of LLM-first and generative translation workflows squarely under a post-editing standard with real human-accountability requirements, instead of letting them float in an unregulated space where "the AI did it" could be a defense. If an engine drafted it and a human is cleaning it, you are in 18587 territory, whatever the engine is called.
Change Two: Multimodal, Hybrid, and Human-in-the-Loop Terminology
The revision introduces vocabulary the 2017 standard did not have, because the workflows did not exist yet. It brings in multimodal terminology, acknowledging that "translation output" is no longer only text-to-text but can involve speech, images, and mixed media. It brings in hybrid workflows, where MT, LLMs, translation memory, and human work interleave rather than running in a clean machine-then-human sequence. And it brings in human-in-the-loop terminology, the idea that the human is not a final cleanup stage bolted on at the end but a participant woven through the process, verifying, steering, and correcting at multiple points. The practical effect is that the standard now has language for the messy, interleaved, AI-saturated pipeline you actually work in, rather than the tidy two-stage assembly line it assumed in 2017. It can describe your real workflow instead of a simplified cartoon of it.
Change Three: The Rigid Light-vs-Full Split Becomes an Effort Spectrum
This is the change that retires the old map. Instead of two discrete buckets, light and full, the revision moves toward an effort spectrum: post-editing effort is understood as a continuum, matched to the content, the risk, and the client's quality target, rather than forced into one of two boxes. Why does this matter? Because the two-box model created a false choice. Real content does not divide neatly into "make it understandable" and "make it indistinguishable." A piece of high-volume content might need minimal stylistic work but rigorous verification of a handful of safety-critical facts, a combination the light-versus-full vocabulary could not express. Light PE said "don't verify too hard, just make it readable." Full PE said "polish everything." Neither said "polish lightly but verify the dosages with your life," which is exactly what a great deal of real content needs. The spectrum lets you place effort where consequence is, distributing attention rather than choosing a tier. It is the standard catching up to the truth that effort should track risk, not a label.
Change Four: Alignment With ISO 17100
The revision deliberately aligns ISO 18587 more closely with ISO 17100, the human-translation services baseline. Remember the triangle: 17100 defines what professional human translation work is. By pulling 18587 toward it, the revision is narrowing the historical gap between "post-editing" as a lighter, cheaper discipline and "translation" as the full professional craft. The two processes converge. This alignment is the structural reason for the fifth and most important change, the one that follows directly from it.
Change Five: The Post-Editor Must Hold Full Translator Competence
Here is the sentence that should reorganize how you think about your work. The revised ISO 18587 requires the post-editor to hold the same linguistic competence as a professional translator. Not a lighter, cheaper, machine-cleanup competence. The full thing. The competence ISO 17100 demands of a translator working from a blank page.
Sit with why this is the inevitable consequence of everything above. If the machine's output is now fully fluent, the surface polishing that a lesser-skilled post-editor could do has become nearly worthless, because the machine already did it. The remaining work, the work that actually protects the client, is catching the silent error in beautiful prose: the flipped negation, the swapped drug name, the inverted obligation, the corrupted dosage, the mistranslated adverse event. That work is not a cleanup reflex. It is a translator's deepest judgment, the ability to read the target against the source and know, with professional certainty, whether the meaning survived. A button-pusher cannot do it. Only someone who could have produced the translation themselves can reliably tell whether the machine produced the right one. The competence requirement is the standard formally recognizing that in the LLM era, the post-editor's job is the translator's job, just starting from a draft. The cheap-cleanup framing is dead, and the standard has buried it.
Why This Matters for Your Career and Your Price
This is not abstract standards trivia. The revision reshapes the economics and the defensibility of your work, and understanding it is the difference between being commoditized and being indispensable.
The Pricing Trap the Standard Closes
Here is the squeeze that has been crushing post-editors. Clients read that machine translation is "good enough" now, and they want MTPE priced at a fraction of full human translation: roughly 50 to 75% of full human rates, about five to fifteen cents a word, with light PE pushed as low as two cents. The implicit logic was: the machine does most of the work, so the human does less, so the human is worth less. The old light-versus-full vocabulary fed this logic, because "light PE" sounded like genuinely lighter work that deserved a genuinely lighter price.
The revision detonates that logic for any content where a silent error has real consequence. If the standard requires you to hold full translator competence and to verify fluent output against the source with a translator's judgment, then on high-stakes content you are not doing "less" work, you are doing the most demanding work in the field: finding the one fluent lie among a thousand beautiful true sentences. The standard hands you the argument. You stop selling "cheap machine cleanup" and start selling "ISO 18587-conformant post-editing, full translator competence, matched effort along the spectrum, verified against the source." That is a defensible premium, not a race to the bottom. The post-editors who get commoditized are the ones who accept the "you just tidy the machine" framing. The ones who thrive cite the revised standard and price the judgment.
Accountability Never Transfers to the Engine
The competence requirement also nails down accountability, and this protects you as much as it binds you. When a Critical error ships, "the engine wrote it" is never an answer, because the standard makes the human post-editor the accountable, full-competence professional who cleared the file. That sounds like a burden, and it is, but it is also the entire source of your value. The reason your role moves up the value chain in the MT era, instead of away, is precisely that the machine cannot be accountable and you can. The client is not paying for the engine, which they could license themselves. They are paying for the human who owns the quality, who can stand behind the delivery, who holds the competence to catch what the engine cannot. The revised standard codifies that this person must be a full translator. That codification is your job security written into an international standard.
The revision turns "you just clean up the machine" into "you hold full translator competence and you own the quality." One of those gets commoditized. The other gets paid.
How to Work Under the Revised Standard
Awareness becomes useful only when it changes what you do on the next file. Here is how the revised ISO 18587 translates into working habits, even at the beginner level where this lesson sits.
Stop asking "light or full" and start asking "where on the spectrum, and where is the risk." Before you touch a file, separate two questions the old model fused: how much surface polish does the client need, and where in this content would a silent error cause real harm. Those answers can diverge, and the spectrum lets you honor both. You might apply light stylistic effort and maximum verification effort on the same file. The two-box model could not express that. The spectrum can, and so should your scoping conversation with the PM.
Verify against the source on the high-consequence elements, regardless of effort tier. The competence requirement exists because fluent output hides errors in specific places. Treat these as never-trust-the-fluency zones and check each against the source independently of how the sentence reads: negations and prohibitions, numbers and dosages and units, drug names and proper names and approved terms, dates and times and durations, legal obligations and which party owes what, and adverse-event and safety-warning descriptions. This is the verification the standard now assumes you are competent to perform. It is the work that the machine's fluency made invisible and that your full translator competence makes catchable.
Connect your post-editing to the scorecard. The revision aligns 18587 with the broader standards family, and the file's fate is decided by ISO 5060-style evaluation: each error classified by dimension and severity, with a single Critical error failing the file regardless of how clean the rest is. Work knowing that a fluent flip you miss is not "one small error in a good file." It is the Critical that fails the whole delivery and that you, the full-competence post-editor, were accountable for catching.
Name the standard when you scope and price. When a client pushes for light-PE pricing on content that carries real risk, you now have a precise, authoritative answer: the revised ISO 18587 requires full translator competence and effort matched to consequence, and this content sits where consequence is high. You are not being difficult. You are being conformant. The standard is the backbone of the conversation that moves you from commodity to professional.
Key Takeaways
- ISO 18587 is the international standard for the human post-editing of machine-translation output: not a quality certificate for the engine, but a process and competence requirement for the accountable human. It sits in a triangle with ISO 17100 (the human-translation baseline) and ISO 5060 (the analytic error-scoring scorecard).
- For a decade the standard's defining feature was the light-versus-full split: light post-editing aimed only at "understandable" and left the prose rough, while full post-editing aimed at "indistinguishable from human translation" and polished everything. The split priced and scoped the entire industry, with light PE as low as about two cents a word and full PE in the roughly five-to-fifteen-cents MTPE range.
- The light-versus-full model measured surface effort (how much polishing) rather than consequence (how dangerous a hidden error is), and it broke when LLMs replaced clumsy NMT with fluent output that is polished on the surface and silently wrong underneath, evidenced by LLM medical error rates of roughly 59% on drug names, 60% on dates and times, and 66% on adverse events, all in perfect prose.
- The revised ISO 18587, in DIS ballot with publication targeted for late 2025 into 2026, makes five changes: it expands scope from machine translation to "non-human translation output" with AI and LLMs named explicitly; introduces multimodal, hybrid, and human-in-the-loop terminology; retires the rigid light-versus-full split for an effort spectrum; aligns more closely with ISO 17100; and requires the post-editor to hold full professional-translator competence.
- The effort-spectrum change matters because real content does not divide into "make it understandable" and "make it indistinguishable." Effort should track risk, so you can apply light stylistic polish and maximum verification on the same file, a combination the two-box model could not express.
- The competence requirement is the inevitable consequence of fluent machines: when the engine already polishes the surface, the remaining work is catching the silent fluent error against the source, which is a full translator's judgment and not a button-pusher's reflex. The cheap-cleanup framing of post-editing is dead.
- This reshapes your price and your defense: the standard hands you the argument to refuse commodity light-PE pricing on high-consequence content and to sell ISO 18587-conformant post-editing with full translator competence and effort matched along the spectrum. The post-editors who accept "you just tidy the machine" get commoditized; the ones who cite the revised standard get paid.
- Accountability never transfers to the engine: "the engine wrote it" is no defense when a Critical error ships, because the revised standard makes the full-competence human post-editor the accountable professional. That accountability is not only a burden, it is the source of value that moves the linguist's role up the chain in the MT era instead of away.
Skill.re