Operationalizing ISO 18587 and 5060
The board of a mid-size language-service provider gave its new head of quality a single sentence of a mandate, and it sounded easy until she tried to execute it: "Make us ISO 18587 and 5060 compliant across the whole company, not just on the deals where the client asks." She had come up as a linguist, so she knew the standards cold. She could post-edit a regulated file and score it against a severity typology in her sleep. What she discovered in her first month was that the company did not have a quality problem. It had a program problem. On the medical account, a brilliant senior reviser ran a flawless post-editing process entirely out of her own head, kept immaculate records in a personal spreadsheet, and would be leaving in April. On the consumer-electronics account, three project managers each interpreted "post-editing" differently and none of them scored errors at all. On the two newest accounts, nobody could have told you who was qualified to touch a drug label. The standards were being followed, in patches, by heroes, in ways that lived and died with the individual. What the board wanted was not conformance on one file. It was a running machine that produced conformance on every file, by every person, on every account, whether or not the hero was in the building that day. That machine is a program, and this lesson is about how you build one. It is the difference between "we follow the standard" and "we operate a standard," and at the org level that difference is the whole job.
From a Conformant File to a Conformant Organization
There is a level shift hiding in the word "operationalize," and if you miss it you will build the wrong thing. At the practitioner level, the question is "can I make this one file conformant." At the strategist level, the question is "can I make conformance an emergent property of how the organization runs, so that it happens automatically, repeatably, and independently of any one person." Those are not the same problem scaled up. They are different problems. A hero can make any single file conformant through sheer competence and diligence. Only a program can make ten thousand files conformant across forty linguists and six accounts without a hero at the center of each one.
Let us define the load-bearing terms in working language before we build anything, because the rest of the lesson leans on them. ISO 18587 is the international standard that defines the requirements for the human post-editing of machine-translation output: the competences the post-editor must hold, the process that must be followed, and the records that must be kept when the workflow is "the machine drafts, the human revises." Its revision, in DIS ballot (the Draft International Standard stage) with publication targeted for late 2025 into 2026, widens the scope from machine translation to "non-human translation output" with AI and large language models named explicitly, retires the rigid light-versus-full split in favor of an effort spectrum, aligns more tightly with ISO 17100, and requires the post-editor to hold the same linguistic competence as a professional translator. ISO 5060:2024 is the companion standard for evaluating translation output: the analytic, MQM-aligned model that classifies each error by dimension and by severity (Critical, Major, Minor) and decides whether output ships or fails. MQM is the Multidimensional Quality Metrics framework, the error typology the scoring rests on. ISO 17100 is the baseline standard for professional human translation services, which defines what "a qualified translator" and "a defined process" mean in the first place. Conformance is the demonstrable, third-party-verifiable state of meeting every applicable requirement of a standard, backed by evidence you already hold. And a program, the word this whole lesson turns on, is the org-wide system of roles, documented processes, evidence capture, a repeatable gate, and periodic calibration and audit that makes conformance happen the same way every time, by design rather than by luck.
You will also need the working vocabulary this program runs on: MT is machine translation, NMT is neural machine translation, LLM is large language model, PE is post-editing, MTPE is machine-translation post-editing, QE is automatic quality estimation, TM is translation memory, TMS is the translation-management system, CAT is the computer-assisted translation tool, a segment is the sentence-sized unit the CAT tool breaks text into, locale is the language-plus-region target like de-DE, and LSP is a language-service provider.
A conformant file is an achievement. A conformant organization is a system. The strategist's job is to build the system so the achievement stops depending on who happens to be on shift.
The Hero Problem and Why It Fails an Audit
The single most dangerous thing a growing localization operation can have is a conformance record that depends on one exceptional person. It feels like strength. It is actually your largest hidden liability. The board's medical-account hero is the perfect illustration: she produces genuinely conformant work, but the conformance lives in her head and her personal spreadsheet, which means the moment she takes leave, changes accounts, or resigns in April, the organization's conformance on that account evaporates and nobody notices until a client audit lands on a different reviser's desk. A standard is not satisfied by a person being good. It is satisfied by a system that stays good when the good person is gone. Auditors know this, which is why a mature audit does not ask "is your best linguist competent." It asks "show me the process that guarantees any linguist on this account is competent, and show me it works on a file picked at random, including one your hero never touched." A program passes that test. A hero cannot, because the hero is a single point of failure wearing a cape.
The Three Roles a Program Is Built On
A program is not a document. It is people with named responsibilities that do not overlap and do not have gaps. The revised ISO 18587 and ISO 5060 together imply three distinct roles, and the most common failure in an immature operation is to collapse them into one person, which quietly destroys the independence the standards depend on. We will take each role slowly, because the separation between them is not bureaucratic tidiness. It is the structural reason the program can be trusted.
Role One: The Qualified Post-Editor
The post-editor is the person who works the machine output against the source, the approved terminology, and the client's specifications, correcting to the agreed quality target. Under the revised standard this person must hold the full linguistic competence of a professional translator, the same bar ISO 17100 sets for a human translating from a blank page. At the file level, competence is an attribute. At the program level, competence is a managed pool. The strategist's obligation is not "hire good post-editors." It is to run a system that answers, at any moment, for any account: who is qualified to post-edit this content, in this language pair, at this risk tier, and where is the dated evidence of that qualification. That means a maintained register of linguist qualifications, a defined onboarding qualification step, a mechanism that prevents an unqualified person from being assigned high-liability work, and a re-qualification cadence so a credential does not silently go stale. The program treats post-editor competence as inventory it must always be able to prove it has on hand, not as a hope that the person assigned happened to be good.
Role Two: The Evaluator
The evaluator is the person who scores the output against the ISO 5060 typology, classifying each error by dimension and severity and producing the structured record that decides pass or fail. Here is the rule that immature operations violate constantly and that the standards care about deeply: the evaluator should not be the same person as the post-editor for the same file. An evaluation is a check on the post-editing, and a check performed by the person being checked is not a check, it is a self-assessment with a conflict of interest baked in. Independence is the entire point of the evaluation step. At the program level this means the strategist must design the operation so that scoring is a distinct, staffed function with its own qualified people, not a box the post-editor ticks on their own work at the end of a long day. The evaluator role also carries a second competence requirement that people forget: scoring reliably against a severity typology is a skill in its own right, and an evaluator who has never been calibrated against other evaluators will drift. The program qualifies, calibrates, and monitors its evaluators as deliberately as it qualifies its post-editors.
Role Three: The Quality Owner
The third role is the one that only exists at the org level, and it is the role the board handed the new head of quality. The quality owner owns the program itself: the documented processes, the role definitions, the evidence architecture, the gate, the calibration cadence, and the audit readiness. The post-editor owns a file. The evaluator owns a score. The quality owner owns the machine that makes files and scores conformant regardless of who is holding them. This is not a linguist role and it is not a project-management role, though it borders both. It is a systems-ownership role. Its core deliverable is not a translation or a score. It is the answer to the question "if a client audits any file we shipped last quarter, can we produce the conformance evidence in minutes, and is that true whether or not the person who worked it still works here." When that answer is a confident yes, the quality owner has done the job. When it depends on which file or which linguist, the program has a hole the quality owner has not yet closed.
Post-editor, evaluator, quality owner: three roles, deliberately separated. Collapse the evaluator into the post-editor and you lose independence. Collapse the quality owner into a project manager and you lose the program.
The Segregation-of-Duties Principle
Borrow a concept from financial controls, because it is exactly right here. Segregation of duties means the person who does the work is not the person who checks the work is not the person who owns the controls that govern both. In an auditor's mind, this separation is what makes evidence credible: a score is trustworthy precisely because the scorer had no incentive to pass the file. When a growing LSP is under margin pressure, the tempting economy is to have the post-editor "self-check" to save an evaluator's hours. That economy does not save money. It converts your evaluation evidence from proof into assertion, because a self-scored file is a file scored by someone who benefits from it passing. The strategist who understands segregation of duties will protect the evaluator role from being absorbed the way an accountant protects the separation between the person who writes checks and the person who reconciles the account. It is not overhead. It is the thing that makes the number believable.
The Documented Process That Outlives the Hero
A program's second pillar, after roles, is documented process. The word "documented" is doing heavy lifting and it is worth being precise about what it buys you. A process that lives in a senior linguist's head is not a defined process in the standard's sense, because it cannot be audited, taught, transferred, or proven. The moment you write it down, three things become possible that were impossible before: a second person can run it identically, an auditor can verify the project followed it, and the organization survives the loss of the person who originally invented it. Documentation is how you extract the program from the hero's skull and make it property of the company.
The documented process for an MTPE-first operation has to specify, at minimum, the following, and the discipline is to write each one as an instruction a competent stranger could follow, not as a description of what the good people already do:
- Intake and risk-tiering: how content is classified by consequence at the door (marketing string versus drug label versus indemnity clause), and how the tier determines routing (MT-ready, full post-editing, human-only, or MT-forbidden). This is the first control and the one that prevents a cheap workflow from ever landing on content that can harm.
- Engine and grounding configuration: which engine drafts which content, grounded on which termbase and TM, and how that configuration is captured automatically at pre-population.
- Assignment: the rule that routes a task only to a post-editor whose qualification, for that language pair and risk tier, is on file and current before the task opens.
- The post-editing pass: the inputs (source, MT output, termbase, style guide, reference TM), the expected output, and the quality target, so two qualified post-editors handling the same content go through the same steps and land in the same place.
- The evaluation: who scores, against which typology and weights, on what sample, producing which record.
- The gate: the pass/fail rule, including the absolute rule that a single Critical error blocks delivery.
- Sign-off and delivery: who is accountable, how the named, dated release is recorded, and how it is bound to the exact version shipped.
Notice that this documented process is not new work invented for the audit. It is a written specification of the pipeline the operation already runs. The act of documenting does not add steps. It makes the existing steps repeatable by anyone and provable to a third party. The hero was doing all of this already. The program writes it down so the hero's departure in April does not take the process with her.
The Repeatability Test
There is a simple, brutal test for whether your process is genuinely documented or merely described: hand it to a competent linguist who has never worked your accounts, and see whether they can run a file to the same outcome your best person would, using only the written process and the linked assets. If they can, you have a program. If they cannot without someone tapping the hero on the shoulder to explain "what we actually do here," then what you have is folklore with a document stapled to it, and it will fail the moment the folklore-keeper is unavailable. The strategist should run this test deliberately, on real files, as a standing check on the program's health, because a process degrades silently: people add undocumented tweaks, the written version drifts from the practiced version, and one day the two have diverged so far that the document is fiction. Repeatability is not a one-time property. It is a thing you re-verify.
The Evidence Trail as Org Architecture
The third pillar is evidence, and at the org level the crucial insight is that evidence is an architecture decision, not a documentation task. A single conformant file can have its evidence assembled by hand after the fact if someone remembers what happened. An organization shipping thousands of files a month cannot reconstruct evidence on demand, because the labor of reconstruction scales with volume and the memory required does not exist. So the program's evidence must fall out of the tools as a byproduct of doing the work, automatically, for every file, without anyone deciding to capture it. The distinction is between an operation that logs and an operation that remembers. A remembering operation relies on people recalling what they did and hunting through email when an auditor asks. A logging operation has the trail emitted by the TMS, the CAT tool, the evaluation step, and the sign-off action as an inevitable consequence of the pipeline running. When the audit letter arrives, the logging operation does not produce new evidence. It exports evidence that already existed.
For each delivery, the evidence architecture must reliably capture, without human intervention, a pack that a client or certifier can read end to end:
- The intake and risk-tier record: content type, assigned tier, routing decision, and the rule or named person that set it.
- The specification the project ran against: brief, quality target, termbase version, style-guide version, all version-stamped.
- The engine and grounding record: which engine drafted the segments and against which TM and glossary.
- The post-editor qualification and assignment: the named linguist, their dated qualification, the assignment timestamp, all showing the qualification predated the work.
- The edit evidence: the CAT edit log showing what changed from raw MT to delivered text.
- The severity-scored evaluation: the MQM/ISO 5060 report with per-error detail and the score against threshold, attributed to a named evaluator who was not the post-editor.
- The gate decision and sign-off: the go/no-go outcome, the stated fail-rule, and the named, dated, version-bound release.
The test for whether any single record is evidence rather than a story is the same at the org level as at the file level: evidence is attributed, dated, and reconstructable. A named human or system action stands behind it. It is fixed in time, and the order matters (qualification predates work; sign-off follows evaluation). A skeptic who was not present can rebuild the claim from the artifact alone. The program's job is to make sure that every one of these records, for every file, is captured with all three properties automatically, so that the quality owner never has to say "I'm sure we did it, let me try to find proof."
The Random-File Standard
Here is the standard the quality owner should hold the whole program to, and it is deliberately harsh because auditors are: pick any file the organization shipped in the last quarter at random, and produce its complete evidence pack in minutes. Not the flagship file the account team is proud of. A random one. The one worked by the new hire on the account nobody thinks about, on a Friday, when the hero was on leave. If the program can do this for the random file, it can do it for all of them, and conformance is a genuine org property. If it can only do it for the files the good people worked, then the program is really a hero with better paperwork, and the random file is where the client audit will, by bad luck or by design, land. Designing for the random file is what turns conformance from an anecdote into a guarantee.
The Repeatable Severity-Scored Gate
Every program needs a single, unambiguous decision point where output either earns delivery or does not, and it must produce the same verdict regardless of which qualified evaluator runs it or which account it sits on. This is the gate, and its repeatability is what distinguishes a program from a collection of individual judgment calls. The gate has three declared components, and all three must be written down and identical across the organization, because a gate that means different things on different accounts is not a gate, it is a mood.
First, the scoring model: the MQM/ISO 5060 error typology with its dimensions (accuracy, terminology, locale convention, style and fluency, and so on) and the severity weights the organization assigns to Critical, Major, and Minor. Second, the threshold: a numeric pass/fail line, typically expressed as a normalized error score per thousand words, so "good enough" becomes a number an auditor and a client can both see, not a feeling. Third, and non-negotiable, the absolute Critical rule: a single Critical error fails the file regardless of how clean the average looks, because a Critical is a categorically different event from a quality-of-degree problem. A flipped dosage or an inverted indemnity clause is not offset by a thousand beautiful sentences. The average forgives; the Critical rule refuses to.
The reason the Critical rule sits at the org level and not just the file level is that averages are how a program quietly poisons itself. If the gate were nothing but a normalized score, an evaluator under deadline pressure could let a single catastrophic mistranslation through on the strength of an otherwise pristine file, and the number would defend the decision. The absolute rule removes that discretion by design. It tells every evaluator, on every account, that the one thing they may never do is average away a Critical. Making the rule identical everywhere is precisely what makes the gate a program control rather than a personal standard that varies with who is scoring.
Calibration: The Thing That Keeps the Gate Honest
Here is the subtle failure that undoes even a well-designed gate, and the org-level control that prevents it. Two qualified evaluators, handed the identical file, will not necessarily assign the same severities. One reads a terminology slip as Major; the other calls it Minor. One flags a locale-convention error the other waves through. This is evaluator drift, and left alone it means your gate does not actually produce a repeatable verdict, because the verdict depends on who happened to score. A gate that gives different answers depending on the evaluator is not a program control at all. It is the illusion of one.
The cure is calibration: a recurring exercise in which the program's evaluators independently score the same reference files and then compare, discuss the discordant severities, and converge on a shared interpretation of the typology. Calibration is not a training event you do once. It is a standing cadence, monthly or quarterly, that keeps the evaluator pool aligned as people join, leave, and slowly drift apart. The strategist runs calibration for the same reason an instrument is re-calibrated against a known reference: not because the instrument is broken, but because instruments drift, and a measurement you cannot trust to be consistent is not a measurement. A program that scores errors but never calibrates its scorers is measuring with rulers that have all been stretched slightly differently, and reporting the results as if they were comparable. The calibration record, showing that the evaluator pool was aligned to a shared standard on a recent date, is itself a piece of audit evidence: it proves the gate is repeatable, not merely present.
A gate without calibration is a gate whose verdict depends on the evaluator. Calibration is the org-level control that makes "the file passed" mean the same thing no matter who scored it.
A Worked Program Build for an LSP
Abstractions become real when you build the thing, so let us take the board's mid-size LSP from patchwork to program, the way the new head of quality actually would. The company has forty linguists, six accounts spanning consumer electronics, medical devices, financial services, and marketing, an MT-first pipeline already in place, and the specific pathologies we opened with: a hero on the medical account, inconsistent post-editing definitions on consumer electronics, no evaluation at all on two accounts, and no clear map of who is qualified for what. The mandate is org-wide ISO 18587 and 5060 conformance that does not depend on any individual. Here is the build, in the order that actually works.
Phase One: Establish the Roles and the Qualification Register
The quality owner starts with people, not process, because roles are the foundation everything else attaches to. She formally names the three roles, and critically, she separates evaluator from post-editor as a hard rule: no linguist evaluates their own file, on any account, ever. She stands up a qualification register: for every one of the forty linguists, a dated record of their competence by language pair and by domain, and a rule in the TMS that a task tagged at the medical or financial risk tier can only be assigned to a linguist whose register entry qualifies them for it. This single control fixes the "nobody could tell you who was qualified to touch a drug label" problem structurally: the system now refuses to assign a drug label to an unqualified linguist, rather than relying on a PM to remember. The hero's tacit competence is entered into the register as a formal, dated qualification, so that when she leaves in April her qualifications leave a documented standard behind, and her replacement is qualified against the same recorded bar.
Phase Two: Write the One Process, Retire the Six
Next, she replaces the six accidental processes with one documented process, applied everywhere, with risk-tiering as the branch that adapts it to content rather than six unrelated workflows. The consumer-electronics team, where three PMs each meant something different by "post-editing," now runs the same intake, assignment, post-editing, evaluation, gate, and sign-off steps as the medical team, differing only in the risk tier and therefore the effort spectrum applied. She runs the repeatability test deliberately: she hands the written process to a linguist from the marketing account and asks them to run a consumer-electronics file, and where they get stuck she finds the folklore that was never written down and writes it. The two accounts that did no evaluation at all now have the evaluation step in their process by default, staffed by evaluators drawn from the pool and separated from their post-editors.
Phase Three: Instrument the Evidence to Log, Not Remember
Now she attacks the evidence architecture. The hero's personal spreadsheet is retired, not because it was bad but because it was personal and therefore a single point of failure. In its place, the TMS captures engine, grounding, and assignment automatically; the CAT tool keeps the edit log by default; the evaluation step writes a structured MQM/5060 report by construction; and sign-off becomes a recorded action bound to the delivered version. The quality owner then applies the random-file standard as an acceptance test for the whole phase: she picks three files shipped that week, at random, across three accounts, and tries to export the complete evidence pack for each in minutes. Where she cannot, she has found an instrumentation gap, and she closes it before declaring the phase done. The goal state is that no human ever again reconstructs evidence after the fact on any account.
Phase Four: Standardize and Calibrate the Gate
She writes one gate for the whole company: the shared scoring model, a threshold expressed per thousand words that can be set tighter for higher risk tiers, and the absolute Critical rule identical on every account. Then she institutes calibration: once a quarter, every evaluator in the pool scores the same set of reference files independently, the team compares severities, discusses every disagreement, and converges. The first calibration is revealing, because it surfaces exactly how much the evaluators had been drifting apart while everyone assumed "we all score the same way." The calibration record becomes a standing piece of the program's audit evidence, proving that the gate produces a repeatable verdict rather than an evaluator-dependent one. From this point the gate means the same thing whether the file was scored by the medical hero's successor or the newest evaluator on the marketing account.
Phase Five: Audit Readiness as a Standing State
Finally, she converts audit readiness from an event into a state. Instead of scrambling when a client audit letter arrives, she runs a quarterly internal audit that behaves exactly like an external one: it picks random files across all six accounts, demands the full evidence pack for each within minutes, checks that qualifications predated the work, that evaluators were independent of post-editors, that the gate applied its Critical rule, and that the calibration cadence is current. Any file that cannot pass the internal audit is a defect in the program, not in the file, and it gets fixed at the program level. When a real client audit eventually lands, on any account, on any random file, the answer is already a ten-minute export, because the organization has been auditing itself the whole time. The board's one-sentence mandate is now satisfied not by a hero but by a machine, and the machine keeps producing conformance whether or not the hero is in the building, which was the entire point.
The worked build has an order for a reason: roles first, one process second, logged evidence third, a calibrated gate fourth, standing audit readiness fifth. Skip roles and you document folklore; skip calibration and your gate lies; skip the internal audit and you learn your gaps from the client instead of before them.
What the Program Lets the LSP Say
The payoff is a sentence the company could not truthfully say before and can now say to any client, any auditor, any prospect, about any file: "Show us the file, and we will show you who was qualified to work it and when they qualified, the process they followed, the independent severity-scored evaluation with zero Criticals, the gate decision that would have failed it on a single Critical, and the named, dated sign-off, exported in minutes, and this is true of every file we ship, not just the ones our best people touched." That sentence is a commercial weapon. It moves the LSP off the commodity price war (where the pitch is "we clean up the machine cheaply") and onto governance ground competitors cannot occupy, because the competitor selling raw machine output and hope cannot say it, and the competitor relying on a hero can only say it until the hero leaves. The program is the credential, and it is durable in a way that individual talent never is.
Key Takeaways
- Operationalizing the standards is an org-level shift, not a scaled-up file task. A conformant file is an achievement any competent hero can produce; a conformant organization is a system that produces conformance repeatably, independently of any individual. A program, not a person, is what an org-level audit actually tests.
- A program is built on three deliberately separated roles. The qualified post-editor (with full professional-translator competence, managed as a proven pool), the evaluator (who must be independent of the post-editor for the same file, or the evaluation is a self-assessment), and the quality owner (who owns the program itself: process, evidence, gate, calibration, audit). Segregation of duties is what makes the evidence credible.
- Documented process is how you extract the program from the hero's head. Writing the process down makes it repeatable by any competent stranger, provable to an auditor, and survivable when the person who invented it leaves. The repeatability test (can an outsider run a file to the same outcome from the document alone) is the standing check that it is genuinely documented, not folklore with a document stapled on.
- Evidence is an architecture decision: log, do not remember. At org scale you cannot reconstruct evidence on demand, so the TMS, CAT tool, evaluation step, and sign-off must emit an attributed, dated, reconstructable record for every file automatically. The random-file standard (produce the full evidence pack for any file shipped last quarter, in minutes) is the harsh test that conformance is a real org property and not a hero with better paperwork.
- The gate must be repeatable, with three declared components identical across the whole org: the MQM/ISO 5060 scoring model, a numeric threshold per thousand words, and the absolute rule that one Critical error fails the file regardless of the average. The Critical rule sits at the org level precisely to remove the discretion that would let an evaluator average away a catastrophic mistranslation.
- Calibration is the control that keeps the gate honest. Qualified evaluators drift, assigning different severities to the same file, so a gate without a recurring calibration cadence gives an evaluator-dependent verdict, which is no verdict at all. The calibration record is itself audit evidence that the gate is repeatable, not merely present.
- The worked LSP build follows a deliberate order: establish roles and a qualification register, retire the accidental processes for one documented process, instrument the evidence to log rather than remember, standardize and calibrate the gate, and make audit readiness a standing state through quarterly internal audits that behave like external ones. Each phase closes a specific pathology of the patchwork operation.
- The program is a durable commercial asset, not just a compliance chore. It lets the organization say, about any random file, "here is who was qualified, the process, the independent severity-scored evaluation with zero Criticals, the gate, and the dated sign-off, in minutes," a sentence a raw-MT competitor cannot say and a hero-dependent competitor can say only until the hero leaves.
Skill.re