Conformance in Practice - ISO 18587 and 5060
A medical-device client sent your team an audit letter on a Tuesday. Not an angry one, just a routine one, the kind a regulated company sends when its own auditors come knocking. The letter had one line that mattered: "Please provide evidence that the German user manual for the infusion pump, delivered on 14 March, was post-edited in conformance with ISO 18587 and evaluated in conformance with ISO 5060." Your PM forwarded it to you with a single word in the subject line: "Help." And here is the quiet horror of that moment, the one this lesson exists to prevent. You did follow the standards. The file went through machine translation, a qualified linguist post-edited it, someone scored the errors, the file shipped clean. You know you did it right. But the auditor did not ask whether you followed ISO 18587. The auditor asked you to prove it. And as you open the project folder you realize, with a sinking feeling, that "we follow the standard" and "we can produce the records that prove we followed the standard" are two completely different things, and you only ever built the first one. This lesson is about closing that gap. It is about turning ISO 18587 and ISO 5060 from documents you cite into a workflow you actually run, clause by clause, with the evidence falling out of the process as a byproduct rather than being reconstructed in a panic on a Tuesday afternoon.
What Conformance Actually Means
Start with the word, because almost everyone in the industry uses it loosely and the looseness is exactly where the trouble lives. Conformance is the state of demonstrably meeting every applicable requirement of a standard, in a way a third party can verify from evidence you hold. Read that definition again and notice the two halves people forget. The first half is "meeting every applicable requirement," which is about what you do. The second half is "in a way a third party can verify from evidence," which is about what you can show. Conformance is not a feeling. It is not "we work to a high standard." It is the specific, auditable claim that you met named clauses of a named standard, backed by records that a person who was not in the room can read and confirm.
Let us anchor the two standards in plain working terms before we go further, because the rest of this lesson maps a real pipeline onto their clauses. ISO 18587 is the international standard that defines the requirements for the human post-editing of machine-translation output: the qualifications the post-editor must hold, the process that must be followed, and the records that must be kept when the workflow is "the machine drafts, the human revises." Its revision, in DIS ballot (the Draft International Standard stage) with publication targeted for late 2025 into 2026, expands scope from machine translation to "non-human translation output" with AI and large language models named explicitly, retires the rigid light-versus-full split for an effort spectrum, aligns more closely with ISO 17100, and requires the post-editor to hold the same linguistic competence as a professional translator. ISO 5060:2024 is the companion standard for evaluating translation output: the analytic, MQM-aligned model that classifies each error by dimension and severity (Critical, Major, Minor) and decides whether the output ships or fails. MQM is the Multidimensional Quality Metrics framework, the error typology the scoring sits on. And underneath both sits ISO 17100, the baseline standard for professional human translation services, which defines what "a qualified translator" and "a defined process" mean in the first place.
Hold the relationship in your head as a sentence: 18587 tells you how to run the post-editing process, 5060 tells you how to evaluate what came out of it, and 17100 defines the human competence both of them lean on. Conformance in practice means running a pipeline where every requirement of those standards has a corresponding step you actually perform and a record you actually keep. Acronyms you will need throughout: MT is machine translation, NMT is neural machine translation, LLM is large language model, PE is post-editing, MTPE is machine-translation post-editing, QE is automatic quality estimation, TM is translation memory, TMS is translation-management system, CAT is the computer-assisted translation tool, a segment is the sentence-sized unit the CAT tool breaks text into, locale is the language-plus-region target like de-DE, and LSP is a language-service provider.
Conformance is not "we followed the standard." Conformance is "we can prove we followed the standard, to someone who was not there, from records we already hold."
The Two Claims That Separate Pros From Amateurs
There are two sentences a localization shop can say about a standard, and they sound almost identical, which is why the difference is so dangerous. The first is "we follow ISO 18587." The second is "we can prove we conform to ISO 18587." The first is a claim about your intentions and your craft. The second is a claim about your records. A shop can be entirely sincere about the first and completely unable to deliver the second, and when that happens the auditor does not give you credit for the quality work you genuinely did. From the auditor's chair, work that cannot be evidenced did not happen. This is the single most important reframe in the whole lesson: conformance is an evidence discipline, not a quality discipline. The quality is necessary but it is not the thing being audited. The thing being audited is whether you can hand over the trail.
What ISO 18587 Concretely Requires
Most people who invoke ISO 18587 have never read it, and they treat it as a vague badge of seriousness. To run it as a workflow you have to know what it actually demands, and the demands fall into five concrete buckets. We will take each one slowly, because each one becomes a step in your pipeline and a record in your folder.
One: A Qualified Post-Editor
The standard does not let just anyone post-edit. It specifies that the post-editor must hold defined competences, and the revision sharpens this to the full linguistic competence of a professional translator, the same bar ISO 17100 sets for a human translator working from a blank page. In concrete terms this means a documented qualification: a relevant degree, or a recognized credential, or proven professional experience in the language pair and domain, plus demonstrated post-editing competence specifically. The reason the standard cares is the reason this whole program exists. In the LLM era the output is already fluent, so surface tidying is nearly worthless, and the work that protects the client is catching the silent fluent error against the source. That is a translator's deepest judgment, not a button-pusher's reflex. Only someone who could have produced the translation themselves can reliably tell whether the machine produced the right one. So the qualification requirement is not bureaucratic box-ticking. It is the standard insisting that the accountable human is competent to catch what the engine cannot.
For conformance, the concrete obligation is this: you must be able to show, for the specific person who post-edited the specific file, that their qualification was on record before the work began. A CV in someone's inbox is not evidence. A dated qualification record in your vendor-management system, linked to the project, is.
Two: A Defined, Documented Process
ISO 18587 requires that post-editing follow a defined process rather than being improvised file by file. The process must specify what the post-editor does: work the MT output against the source segment, against the approved terminology, and against the client's specifications, correcting errors to the agreed quality target. It must specify the inputs (source, MT output, termbase, style guide, any reference TM) and the expected output. The point of "defined" is repeatability: the same content handled by two different qualified post-editors should go through the same steps and land in the same place. A process that lives only in a senior linguist's head is not a defined process, because it cannot be audited, taught, or proven. Conformance requires the process to exist as a written specification that the project demonstrably followed.
Three: Records and Traceability
This is the bucket that ambushes the unprepared shop. The standard requires that records be kept: who did the work, against which specifications, with what agreed quality level, and what the outcome was. Traceability means that for a delivered file you can reconstruct the chain, source in, engine used, post-editor assigned, specifications applied, evaluation performed, sign-off given, file out. The records are not a nice-to-have you generate when asked. They are a requirement of the standard, which means a shop that did beautiful work but kept no records is, technically and provably, not in conformance. The audit letter at the top of this lesson is precisely a request to exercise traceability, and the panic in the PM's "Help" is the sound of a shop discovering its traceability was a fiction.
Four: Severity-Scored Evaluation
The post-editing process does not end at "the post-editor finished." Conformance with the modern standards family requires that output be evaluated, and evaluated analytically, not by vibe. This is where ISO 5060 enters: the output is scored against an MQM-aligned error typology, each error classified by dimension (accuracy, terminology, locale convention, style and fluency, and so on) and by severity (Critical, Major, Minor). The evaluation produces a structured record, not a thumbs-up. The standard's logic is that "looks fine to me" is not a quality gate, because the most dangerous error, the fluent mistranslation, is precisely the one that looks fine. Only an analytic evaluation that forces the evaluator to check each dimension catches the error the eye skips. We will walk a real scored evaluation in a moment, but for now hold the requirement: conformance means a severity-scored evaluation exists as an artifact for the file.
Five: Human Accountability
Finally, the standard nails accountability to a human. Someone signs the delivery, and that someone owns the quality. "The engine wrote it" is never an answer when a Critical error ships, because the standard makes the qualified, full-competence post-editor or evaluator the accountable professional who cleared the file. This is not only a burden, it is the source of the linguist's value: the client is not paying for the engine, which they could license themselves, but for the human who can stand behind the delivery. Conformance requires that the sign-off be real, attributed, and recorded, a named person on a named date, not an anonymous "QA passed" flag.
Five requirements, five records: a qualified post-editor (qualification on file), a defined process (written spec), traceability (the reconstructable chain), a severity-scored evaluation (the analytic report), and human accountability (the named, dated sign-off).
Mapping the Running Pipeline to Each Clause
Now we make it concrete, because the whole promise of this lesson is that conformance is a workflow you run, not a document you cite. Assume the MT-first pipeline this program has been building across every prior lesson: content comes in, gets risk-tiered at intake, the engine pre-populates every segment, a qualified post-editor revises against source and termbase, the output is severity-scored, a gate decides go or no-go, and the file ships with a record. The trick of conformance is to recognize that this pipeline, if instrumented correctly, already produces the evidence for every clause. You do not bolt conformance on at the end. You let the running process emit the proof as it goes.
Walk the map step by step, clause to step to evidence.
- Intake and risk-tiering maps to the standard's expectation that effort and process suit the content. Evidence: the intake record showing the content type, the assigned risk tier, and the routing decision (MT-ready, full PE, human-only, or MT-forbidden). This record is what proves you did not run cheap light PE on a drug label.
- Engine and configuration maps to "non-human translation output" under the revised scope. Evidence: which engine produced the draft, with what glossary and TM grounding, captured automatically by the TMS at the time of pre-population.
- Post-editor assignment maps to the qualified-post-editor clause. Evidence: the assignment record linking this file to a named linguist whose dated qualification is on file before the task opened.
- The post-editing pass maps to the defined-process clause. Evidence: the project followed the written PE specification (the brief, the termbase, the style guide, the quality target), and the CAT tool's edit log shows the segments that were changed.
- Severity-scored evaluation maps to ISO 5060 and the evaluation requirement. Evidence: the MQM/5060 error report, with each logged error's location, dimension, severity, and a note, plus the computed score against the threshold.
- The quality gate maps to the go/no-go decision the standards imply. Evidence: the gate record showing the pass/fail outcome and, crucially, that a single Critical error would have blocked delivery regardless of an otherwise clean file.
- Sign-off and delivery maps to human accountability. Evidence: the named, dated sign-off by the accountable professional, attached to the exact version delivered.
Notice what just happened. Every clause of the two standards now corresponds to a step you already run and a record that step already emits, if you instrument it. The difference between a shop that scrambles on audit-letter Tuesday and a shop that forwards the folder in ten minutes is not the quality of the post-editing. It is whether the pipeline was built to capture as it runs instead of reconstruct on demand. Conformance is an architecture decision made long before the audit, not a heroic reconstruction made after it.
The Instrumentation Principle
The single most useful mental model here is the difference between logging and remembering. A shop that "remembers" its quality work relies on people recalling what they did, hunting through email, and reconstructing a chain after the fact. A shop that "logs" its quality work has the chain fall out of the tools as a byproduct of doing the work. The TMS records the engine and the assignment automatically. The CAT tool keeps the edit log automatically. The evaluation step writes a structured report because that is how it is built. The sign-off is a recorded action, not a verbal nod. When the audit letter arrives, the logged shop does not produce new evidence. It exports evidence that already existed. That is the whole game.
What Evidence Actually Proves Conformance
Not all records are evidence. An auditor distinguishes between an artifact that merely asserts something happened and an artifact that demonstrates it, and you need to think the way the auditor thinks. The test for each record is simple: could a determined skeptic who was not present reconstruct and verify the claim from this alone? Apply that test and many comfortable "records" fail it.
Consider the qualification of the post-editor. A line in a spreadsheet that says "Anna, qualified" is an assertion. A dated qualification record, linked to Anna's credential or experience evidence, timestamped before the project opened, is proof. The difference is whether the skeptic can verify it independently or has to take your word. Consider the evaluation. "QA passed, looks good" is an assertion. A severity-scored report listing each error by segment, dimension, severity, and note, with the computed score against a stated threshold, is proof, because the skeptic can re-read the segments and check whether your severities were defensible. Consider the sign-off. A green checkmark with no name is an assertion. "Reviewed and released by M. Okafor, senior reviser, 14 March, against spec v3" is proof, because there is a human who owns it and a date that fixes it in time.
The pattern across all of these is that evidence is attributed, dated, and reconstructable. Attributed means a named human or a named system action stands behind it. Dated means it is fixed in time, and the time order matters (the qualification must predate the work; the sign-off must follow the evaluation). Reconstructable means the skeptic can rebuild the claim from the artifact without your help. A record that has all three is evidence. A record missing any one of them is a story you are asking the auditor to believe.
The Evidence Pack for a Single Delivery
For any one delivered file, conformance evidence assembles into a pack a client or certifier can read end to end. The pack is the answer to the audit letter, and a well-run shop can produce it in minutes because every piece already exists. It contains:
- The intake and risk-tier record: content type, assigned tier, routing decision, and the named person or rule that made it.
- The specification the project ran against: the client brief, quality target, termbase version, and style-guide version, all version-stamped.
- The engine and grounding record: which engine drafted the segments and against which TM and glossary.
- The post-editor qualification and assignment: the named linguist, their dated qualification, and the assignment timestamp.
- The edit evidence: the CAT edit log showing what was changed from the raw MT to the delivered text.
- The severity-scored evaluation: the MQM/ISO 5060 report with per-error detail and the score against threshold.
- The gate decision and sign-off: the go/no-go outcome and the named, dated release.
If you can hand that pack over for any delivery picked at random, you are not claiming conformance. You are demonstrating it. And the demonstration is identical whether the auditor is friendly or hostile, because evidence does not care about your reputation.
A Worked Conformance Walkthrough of a Real Delivery
Let us run the actual file from the audit letter, because abstract requirements only become real when you see them produce evidence on a concrete job. The delivery is the German (de-DE) user manual for an infusion pump, a regulated medical device. The content is high-consequence: a flipped dosage or a dropped warning is not a rework, it is a patient-safety event and a regulatory liability. Watch how each clause becomes a step and each step becomes a record.
Step One: Intake and Risk-Tiering
The file arrives. At intake it is classified, by rule, as life-safety regulated content, top risk tier. The routing decision is recorded: this content gets MT-assisted full post-editing by a qualified medical-domain linguist, with no light-PE option permitted and certain sections (the dosing tables, the contraindications) flagged for independent number-and-negation verification. Evidence produced: an intake record stamped with the tier, the routing rule that set it, and the date. This single record already answers a question the auditor will ask: "how did you decide this content warranted full treatment?" You decided by rule, at intake, and you can show it.
Step Two: Engine, Grounding, and Pre-Population
The TMS pre-populates every segment from the configured MT engine, grounded on the client's approved medical termbase and the device-family TM. The engine, the termbase version, and the TM version are captured automatically at pre-population time. Evidence produced: a configuration record naming the engine and the linguistic assets it was grounded on. Under the revised standard's expanded scope, this output is "non-human translation output" whether the engine is classic NMT or an LLM, so it is squarely in 18587 territory, and the record proves which engine you must account for.
Step Three: Qualified Post-Editor Assignment
The task is assigned to Anna, a linguist whose vendor record shows a translation degree, eight years of de-DE medical translation, and demonstrated post-editing competence, all dated and on file since before this project opened. Evidence produced: an assignment record linking the file to Anna and to her pre-existing qualification. This is the qualified-post-editor clause satisfied with attributed, dated, reconstructable proof, not a "she's good" assertion.
Step Four: The Post-Editing Pass Against the Spec
Anna works the file against the written PE specification: source segment, approved termbase, client style guide, and the full quality target. She does not trust the fluent surface. On the dosing table she catches the error that defines this whole program. The source reads "do not exceed 5 mg per hour." The MT output, grammatically perfect, reads "nicht mehr als 50 mg pro Stunde," fifty, not five, a flipped order of magnitude delivered in flawless German. A reader skimming for fluency would sail past it because the sentence is beautiful. Anna, reading the target against the source as a full-competence translator, catches it and corrects it. The CAT tool logs the change. Evidence produced: the edit log showing the raw MT, the correction, and the segment, plus the defined-process record showing she ran against spec v3, termbase v7, style guide v2.
Step Five: The Severity-Scored Evaluation
A second qualified linguist, M. Okafor, runs the ISO 5060 evaluation on a sample and on every safety-critical segment. The dosage segment, had it shipped uncorrected, would have been a Critical error under the accuracy dimension: a mistranslation that could cause patient harm. Because Anna caught it, the evaluation logs it as a corrected finding, not a shipped defect, but the evaluation still records the category for the quality history. Okafor finds, in the delivered text, two Minor style issues and zero Major and zero Critical errors. Evidence produced: the MQM/5060 report, each finding with segment, dimension, severity, and note, and the computed score against the client threshold. The report is attributed to Okafor and dated.
The dosage flip is the whole standard in one segment: a fluent, confident, grammatically perfect error that a severity-scored evaluation classes as Critical, that one Critical would have failed the file, and that only a qualified human reading target against source could catch.
Step Six: The One-Critical-Fails Gate
The delivered file carries zero Critical and zero Major errors, so it clears the gate. But the gate record explicitly states the rule it applied: a single Critical error would have blocked delivery regardless of how clean the rest looked. This matters for the audit because it proves the gate is real, not decorative. An auditor trusts a gate that can say no. Evidence produced: the gate decision record with the pass outcome and the stated fail-rule.
Step Seven: Sign-Off and Delivery
Okafor signs the release: "Reviewed and released, de-DE infusion-pump manual, against spec v3, on 14 March." The exact delivered version is attached to the sign-off so there is no ambiguity about which file the signature covers. Evidence produced: the named, dated, version-bound sign-off that satisfies the human-accountability clause.
Answering the Audit Letter
Now return to Tuesday. The auditor asked for evidence that the 14 March German infusion-pump manual was post-edited in conformance with ISO 18587 and evaluated in conformance with ISO 5060. The shop that logged as it ran does not write anything new. It exports the evidence pack: the intake and tier record, the engine and grounding record, Anna's dated qualification and assignment, the edit log with the caught dosage flip, Okafor's severity-scored evaluation, the gate decision with its fail-rule, and the version-bound sign-off. Seven artifacts, each attributed, dated, and reconstructable, assembled in the order an auditor reads them. The PM's "Help" becomes a ten-minute export. The gap between "we follow it" and "we can prove it" was closed months earlier, at the moment the pipeline was instrumented to capture instead of remember. That is conformance in practice, and it is the difference between keeping the medical-device client and losing them.
Closing the Gap on Your Own Pipeline
Awareness is useless until it changes what you build. Here is how to turn your existing MT-first pipeline into one that produces conformance evidence as a byproduct, even if you are starting from a process that currently "follows the standard" but cannot prove it.
Run a clause-to-step-to-evidence audit of your own pipeline. Take ISO 18587's five requirements (qualified post-editor, defined process, records, severity-scored evaluation, accountability) and ISO 5060's analytic evaluation, and for each one write down the step in your pipeline that satisfies it and the record that step emits. Wherever you can name the step but not the record, you have found a conformance gap: you do the work but cannot prove it. Those gaps are where the audit letter will hurt.
Convert assertions into evidence. For every record you keep, apply the three-part test: is it attributed, dated, and reconstructable? A green checkmark becomes a named, dated sign-off. A "qualified" label becomes a dated qualification linked to credential evidence. A "QA passed" becomes a severity-scored report a skeptic can re-read. You are not adding new work to the linguists. You are making the tools capture what the linguists already do.
Instrument the pipeline to log, not remember. The goal is that no one ever reconstructs evidence after the fact. The TMS captures the engine and assignment automatically, the CAT tool keeps the edit log, the evaluation step writes a structured report by construction, and the sign-off is a recorded action. When the architecture logs as it runs, the evidence pack is always an export, never a project.
Name the standard when you scope, price, and defend. Conformance is also a commercial asset. When a client pushes for commodity light-PE pricing on regulated content, you answer that ISO 18587 conformance requires full translator competence and a defined, recorded, severity-scored process, and that this content sits where consequence is high. You are not being difficult. You are being conformant, and you can prove it. The shops that get commoditized accept "you just clean up the machine." The shops that get paid hand over the evidence pack and let it speak.
Key Takeaways
- Conformance is an evidence discipline, not a quality discipline. It is the demonstrable, third-party-verifiable state of meeting a standard's requirements. "We follow ISO 18587" is a claim about intentions; "we can prove we conform to ISO 18587" is a claim about records, and from an auditor's chair, work that cannot be evidenced did not happen.
- ISO 18587 concretely requires five things: a qualified post-editor (with the revision demanding full professional-translator competence), a defined and documented process, records and traceability, severity-scored evaluation, and human accountability. Each requirement becomes a step you run and a record you keep.
- ISO 5060:2024 supplies the analytic evaluation: each error classified by dimension and severity (Critical, Major, Minor), producing a structured scored report rather than a "looks fine," because the most dangerous error, the fluent mistranslation, is exactly the one that looks fine. ISO 17100 defines the underlying human competence both standards lean on.
- A correctly instrumented MT-first pipeline already produces the evidence for every clause. Intake-and-tiering, engine-and-grounding, qualified assignment, the post-editing pass against spec, severity-scored evaluation, the one-Critical-fails gate, and the named sign-off each map to a clause and emit a record, if the pipeline logs as it runs instead of reconstructing on demand.
- Evidence must be attributed, dated, and reconstructable. A named human or system action stands behind it, it is fixed in time with the order mattering (qualification predates work, sign-off follows evaluation), and a skeptic who was not present can rebuild the claim from the artifact alone. A record missing any of the three is an assertion, not evidence.
- The gap between "we follow it" and "we can prove it" is closed by architecture, not heroics. A shop that logs its quality work exports an evidence pack in minutes; a shop that remembers it scrambles to reconstruct a chain that may not survive scrutiny. The decision to instrument is made long before the audit letter arrives.
- The worked walkthrough shows the standard living in one segment: an MT engine renders "do not exceed 5 mg per hour" as a flawless German "nicht mehr als 50 mg pro Stunde," a fluent ten-fold dosage flip that a severity-scored evaluation classes as Critical, that one Critical would fail the file, and that only a qualified human reading target against source catches. Every clause of both standards is exercised by that single catch and its record.
- Conformance is a commercial asset, not just a compliance chore. Run a clause-to-step-to-evidence audit of your own pipeline, convert assertions into attributed-dated-reconstructable evidence, instrument the tools to log rather than remember, and name the standard when you scope and price, so the evidence pack becomes the argument that moves you from commodity to indispensable.
Skill.re