Producing a Defensible Quality Record
The dispute landed in the project manager's inbox on a Friday afternoon, the worst possible time, and it was the kind that ends relationships. A medical-device client had pushed a German user guide through their own internal review and come back furious: they claimed the translation was riddled with errors, demanded a full re-do at the agency's cost, and copied their procurement lead on the email. The agency owner forwarded it to the linguist who had post-edited the file, Anika, with one line: "Did we mess this up?" Anika did not panic, and she did not write back "I'm sure it's fine, I checked it carefully." She opened a spreadsheet she had attached to the delivery three weeks earlier. It had one row per error she had found and fixed, each row carrying a segment id, the machine's raw output, her corrected target, the error category, the severity, the fix she made, her name as evaluator, and a computed score with a pass stamp on it. She had found and fixed seven errors in that file, including one Critical, before delivery. The client's "errors" turned out to be three preferential rewrites their reviewer would have phrased differently and one genuine miss that Anika's record showed had been flagged as a query, not an error, because the source itself was ambiguous. The dispute was over in an hour. Not because Anika argued well, but because she had a quality record: a structured, segment-level account of what was wrong, how serious it was, and what she did about it, written so a client and an auditor could read it without her in the room. This lesson is about building that artifact, the one that turns "trust me, I'm good" into proof.
What a Quality Record Actually Is
A quality record is a structured, persistent document that captures, error by error and segment by segment, what an evaluator found in a translation, how serious each finding was, what was done about it, and what the file scored against an agreed rule. It is the difference between a verdict and the evidence behind the verdict. The verdict is "this file passes." The record is the marked list of findings, the arithmetic, and the sign-off that makes that verdict reconstructable by someone who was not there when it was made. If the previous lesson in this program taught you the scoring model, the dimensions, the severities, the weights, the one-Critical-fails gate, this lesson teaches you to capture the output of that model in a form that survives contact with a skeptical client, a procurement audit, and your own memory six months later when you have forgotten every detail of the job.
Some vocabulary, used precisely throughout, because the whole value of a record is precision. Machine translation (MT) is any system that renders text from a source language into a target language with no human writing the words; a large language model (LLM) is a general-purpose text predictor that translates as a byproduct of its broad training. Machine-translation post-editing (MTPE), often shortened to PE, is the workflow where a human edits machine output instead of translating from scratch. A segment is the unit a translation tool works in, usually a sentence or short block, the numbered row you see in a CAT tool (a computer-assisted translation tool, the editing environment a linguist works in inside a translation-management system, or TMS). MQM is Multidimensional Quality Metrics, an analytic error-typology framework that classifies translation errors by dimension (accuracy, terminology, locale, fluency) and severity. ISO 5060:2024 is the international standard, published in 2024, that formalizes an MQM-aligned model for the human analytic evaluation of translation output. ISO 18587 is the post-editing standard, in DIS ballot with publication targeted for late 2025 into 2026, whose revision expands scope from machine translation to AI and LLM "non-human translation output" and insists the post-editor hold full professional-translator competence. And the term this lesson turns on: provenance is the documented chain of where a translation came from and what happened to it, who or what produced each segment, who changed it, why, and against what rule it was judged. A quality record is, at bottom, provenance you can hand to an auditor.
A verdict says the file passes. A record proves it. The first is an opinion; the second is an artifact a stranger can audit without you in the room.
Why a Record and Not a Memory
Working linguists resist record-keeping for an understandable reason: it feels like overhead on top of the real work, which is fixing the translation. The file is good, you fixed the errors, why write them all down? The answer is that the quality of a translation is invisible the moment it ships. A clean delivery and a sloppy-but-lucky delivery look identical from the outside; both are just a target file. The only thing that distinguishes the professional from the gambler is whether they can show their work, and human memory cannot do that. Three weeks after delivery you will not remember whether segment 88 was a terminology fix or a query, whether the date in segment 203 was wrong in the source or the target, or whether you caught the dropped negation or got lucky. The record remembers. It converts the most perishable thing in the business, your judgment in the moment, into the most durable, a written account that holds its shape under pressure. Anika won her dispute not because she was a better linguist than the client's reviewer, but because she had externalized her judgment into an artifact and the reviewer had not.
What Goes In the Record: The Anatomy of a Row
A quality record is, at its core, a table, and the table is only as defensible as its columns. Each row is one finding: one error the evaluator marked. Get the columns right and the record practically writes itself; get them wrong and you have a pile of notes nobody can audit. Here is the anatomy of a single defensible row, field by field, and why each field has to be there.
Segment ID: Where the Error Lives
Every finding must point to a precise location, and that location is the segment id, the stable identifier the CAT tool or TMS assigns to each segment. "There's a terminology error somewhere in paragraph three" is not auditable; "segment 0042, terminology error" is. The segment id is the anchor that lets a second person, the client's reviewer, a second evaluator, an auditor, navigate straight to the exact row in the exact file and see the error for themselves. Without it, a record is hearsay. With it, every finding is independently verifiable. Use the tool's native segment id, not your own ad hoc numbering, so the record aligns one-to-one with the bilingual file the client can open. If a finding spans more than one segment, record the range, but most findings live in a single segment, which is part of why segment-level evaluation is the discipline.
Source and Target: The Two Things Being Compared
The record carries the source segment and the target segment for every finding, because an accuracy judgment is meaningless without both. Accuracy, as the program has hammered, is a relationship between target meaning and source meaning, not a property of the target alone. A record that says "segment 5: mistranslation, Critical" without showing the source and the target is asking the reader to take your word for it, which is exactly what a record exists to avoid. Show both, verbatim. The reader should be able to look at the source, look at the target, and see the error you marked without any further explanation, because the two strings sitting side by side are the evidence. For a record that captures the machine-translation provenance, you may carry two target columns: the raw MT output the engine produced and the final post-edited target you delivered, so the record shows not just that an error existed but that you caught and corrected it. That pairing is the quiet proof that the human owned the quality.
Error Category: What Kind of Error It Is
Each finding is classified by error category, the MQM dimension it belongs to: accuracy (with subtypes like mistranslation, omission, addition, untranslated), terminology, locale conventions, or fluency and style. The category is not decoration; it is what makes the record diagnostic. A record whose errors cluster in terminology tells you, and the client, that the engine is drifting off the termbase, which is a fixable process problem you can quote a solution for. A record whose errors cluster in accuracy tells you the engine is hallucinating meaning, a far more dangerous signal. Without categories, a record is just a count of problems; with them, it is a map of where the problems come from, which is the difference between "this file had seven errors" and "this engine has a terminology-adherence problem that a glossary lock would fix." Use one consistent typology, and use the same category names every time, because comparability across files depends on it.
Severity: How Much It Matters
Each finding carries a severity: Critical, Major, or Minor. This is the single most consequential field in the row, because severity, not category, drives the score and the gate. A Critical here is the field that, by itself, fails the file. The record must state severity explicitly and consistently with the program's rule that severity tracks consequence, not edit size: a one-word dropped negation in a safety warning is Critical no matter how small the change, and a clumsy but harmless paragraph on a marketing page is Minor no matter how much text moved. When a client disputes a severity, and they will, the defense is the consequence question written into the record: "this error was marked Critical because a reader acting on the inverted instruction would be exposed to a hazard." A severity with a one-line consequence rationale beside it is nearly undisputable. A severity with no rationale is an invitation to argue.
Fix, Evaluator, and the Sign-Off
The fix column records what was actually done: the corrected target, or, in a pure evaluation that is not also a post-edit, the recommended correction. This closes the loop from "here is what was wrong" to "here is what is now right," and it is the field that turns a record from an accusation into a deliverable. The evaluator field names the human accountable for the finding, because accountability is the spine of the whole standards regime: "the engine wrote it" is never an answer, and the revised ISO 18587 insists a named, full-competence human owns the quality. A record with no name on it satisfies no standard and defends no one. Finally, the record carries the score and the pass/fail status for the file as a whole, the summary that sits above the rows, so a reader can see both the headline verdict and the evidence that produced it. Those last fields are where the next section lives.
Segment id, source, target, category, severity, fix, evaluator, score, pass or fail. Nine fields. Each one answers a question an auditor will ask, and the record exists so you never have to be in the room to answer it.
Computing and Presenting the Score
The rows are the evidence; the score is the headline. A record needs both, and the score has to be computed transparently enough that a reader can re-derive it from the rows, because a number nobody can reconstruct is just a vibe check wearing a decimal point. The arithmetic is the model the previous lesson built, and the record's job is to show it, not hide it.
Recall the three moves of any analytic score, constant even when the constants change: assign a severity penalty to each error, sum the penalties across the file, and normalize against the evaluated length so files of different sizes are comparable. The classic MQM weighting assigns 1 point to a Minor error, 5 points to a Major, and a large penalty to a Critical, commonly 25 or more, often set high enough to fail any realistic file on its own. Some profiles implement the Critical not as a weight at all but as an absolute gate. The record presents this in three layers, top to bottom, so a reader can stop at whatever depth they need.
- The headline: the file's quality score, its pass/fail status, and the Critical count, in one line at the top. "Score 94.2 / 100, PASS, 0 Critical." A procurement reader who trusts you reads only this line and moves on.
- The tally: the counts by severity and the total penalty, so the score is reconstructable. "1 Major (5), 2 Minor (2), total penalty 7 over 250 evaluated words." A reviewer who wants to check your math reads this and re-derives the score.
- The rows: the full segment-level list of findings, the evidence layer. "Segment 0042, terminology, Major, raw / fixed / rationale." An auditor or a disputing client reads all the way down to here and verifies every finding against the bilingual file.
Presenting the Gate, Not Just the Number
The most important thing a record presents about the score is the gate, and it must present the gate separately from the number, because the gate is not the number. A file can have a beautiful computed score of 96 and still fail, because it carries one Critical, and the record has to make that legible at a glance. The discipline is to show the Critical count as its own field next to the score, never folded silently into the arithmetic, so a reader cannot miss it. "Score 96.0, Criticals: 1, Status: FAIL" tells the true story; "Score 96.0, PASS" because the single Critical's penalty got averaged into a still-high number tells a dangerous lie. Present the gate as a hard, visible rule on the face of the record: if Critical count is greater than zero, status is FAIL regardless of the computed score. A record that hides a Critical inside a high score is worse than no record, because it launders a hazard into a pass.
Equally, a record should distinguish errors from queries, and this distinction saved Anika in the opening story. An error is a defect you marked and graded. A query is a place where the source itself was ambiguous, contradictory, or wrong, and you flagged it for the client rather than silently guessing. Queries do not score against the translation, because the defect is upstream of you, but they belong in the record so that a later dispute about that segment lands on the documented ambiguity rather than on your judgment. The record that separates "I made an error" from "the source was unclear and I told you" is the record that survives the Friday-afternoon email.
A Worked Record for a Small File
Theory settles only when you watch a real record take shape. Let us build one for a small file, the way an evaluator or post-editor actually would, so you can see every field carry its weight. The file is a 250-word excerpt of a portable oxygen-concentrator user guide, machine-translated into German and post-edited. We will walk the findings, populate the row for each, compute the score, present the gate, and stamp the verdict. For readability the source and target meanings are rendered in English; in a real record they would be the actual German strings.
The Header Block
Every record opens with a header that establishes what was evaluated, by whom, against what rule, because provenance starts before the first finding. The header for this file reads:
- File: oxygen-concentrator-userguide_DE.xliff, 250 words evaluated (full file).
- Languages: source English (en-US), target German (de-DE).
- Workflow: MT (engine: vendor NMT) + full post-edit, then analytic evaluation.
- Scoring profile: MQM-aligned, weights Minor 1 / Major 5 / Critical gate, ISO 5060-conformant typology, pass threshold 90 / 100, one-Critical-fails gate active.
- Evaluator: A. (named, qualified per ISO 18587), date of evaluation recorded.
Notice that the header already does defensive work. It names the standard and the profile, so the rules are not invented after the dispute; it names the evaluator, so accountability is fixed; it states the threshold, so "pass" means something specific and pre-agreed rather than a feeling. A client cannot later claim the bar was higher than 90 when 90 is written on the face of the record they accepted at delivery.
The Findings, Row by Row
Now the findings. The file's clean segments produce no rows, which is correct: a record lists defects, not the silent majority of correct segments, though the header's "250 words evaluated" tells the reader the whole file was assessed, not just the broken parts. Five findings were marked.
Row 1, Segment 0008. Source: "The concentrator delivers up to 5 litres of oxygen per minute." Raw MT: "The unit delivers up to 5 litres of oxygen per minute." Fixed target: "The concentrator delivers up to 5 litres of oxygen per minute." Category: terminology. Severity: Major. Rationale: the termbase specifies "concentrator" as the approved device term; the engine substituted the generic "unit," a compliance defect on regulated content where consistent terminology is a requirement. Evaluator: A. Penalty: 5.
Row 2, Segment 0019. Source: "Replace the filter every 3 to 4 months depending on usage." Raw MT: rendered the interval as the date "03/04," reading as the third of April in a day-first locale. Fixed target: "every 3 to 4 months." Category: locale conventions / accuracy. Severity: Major. Rationale: number formatting and meaning both corrupted, turning a maintenance interval into a calendar date; a reader could service the device on the wrong schedule, but the result is confusing rather than acutely hazardous. Evaluator: A. Penalty: 5.
Row 3, Segment 0026. Source: "Long-term use may occasionally cause mild dryness of the nasal passages." Raw MT: "...dryness of the nasal pathways." Fixed target: "...nasal passages." Category: fluency / style. Severity: Minor. Rationale: "pathways" is understandable but not the idiomatic term; harmless word choice. Evaluator: A. Penalty: 1.
Row 4, Segment 0026. Source/target as above. Category: fluency (punctuation). Severity: Minor. Rationale: a missing comma in the same segment, a small mechanical slip. Evaluator: A. Penalty: 1. Note that two findings in one segment get two rows: the record is one row per finding, not per segment, so the tally is honest.
Row 5, Segment 0041. Source: "Do not cover the air intake while the device is operating." Raw MT: "Keep the air intake covered while the device is operating." Fixed target: "Do not cover the air intake while the device is operating." Category: accuracy (omission of negation producing a mistranslation). Severity: Critical. Rationale: the negation collapsed and a safety prohibition inverted into a recommendation on a device whose pressure relief depends on an unobstructed intake; a reader acting on it would be exposed to a hazard. Evaluator: A. Penalty: gate.
One more line belongs in the record, and it is not a scored error. Query, Segment 0033. Source: "Store the device below 40 degrees." The source did not specify Celsius or Fahrenheit, an ambiguity in the original. The post-editor rendered it as Celsius (the document's market convention) and logged a query to the client rather than silently guessing. Category: query, source ambiguity. Severity: not scored. Rationale: the defect is upstream of the translation; flagged for client confirmation. This is the row that protects you when the client later asks why you "assumed" Celsius: the record shows you did not assume, you flagged it.
Computing and Presenting the Verdict
Now the arithmetic, shown so a reader can re-derive it. The scored findings are: one Major terminology (5), one Major locale/accuracy (5), two Minor fluency (1 + 1), for an additive penalty of 12 points from the non-Critical errors, plus one Critical accuracy error. The non-Critical penalty of 12 over 250 evaluated words is a modest penalty rate that, on its own, would land the file comfortably above the 90 threshold, a file imperfect but plausibly shippable after the terminology and locale fixes were applied, which they were. And none of that matters, because of the gate.
The headline of the record reads, in full and without softening: Critical count: 1. Status: FAIL. The one-Critical-fails gate is active, the Critical was present in the raw MT, and the gate disqualifies the file regardless of the otherwise-passing computed score. But here is the crucial subtlety that the fix column captures and that turns a failed evaluation into a clean delivery: the Critical was found and corrected in post-editing before delivery. The record therefore tells a two-state story. As the raw machine produced it, the file failed: one Critical, gate triggered. As the post-editor delivered it, the Critical was fixed, the two Majors and two Minors were corrected, and a re-evaluation of the delivered target shows zero Criticals, zero Majors, zero Minors, status PASS. That two-state record is the most valuable artifact in the whole business, because it does not merely claim the file is clean; it proves the human caught and corrected a hazard the machine produced, which is precisely the value ISO 18587 says the post-editor adds.
The raw MT failed on one Critical. The delivered file passed because a named human caught and fixed it. The record that shows both states is the proof that the post-editor, not the engine, owns the quality.
The Record as the Linguist's Protection and the LSP's Product
Everything above is mechanics. This section is why the mechanics are worth the labor, and it has two audiences: the individual linguist and the language-service provider (LSP, the vendor that delivers the localization) who sells their work.
Protection: The Record Is Your Defense
For the working linguist, the quality record is body armor. The MT era loaded a new, asymmetric risk onto the post-editor: you inherit a file the machine drafted, you put your name on the delivery, and if a silent Critical slips through, the accountability is yours and "the engine wrote it" is no defense. In that world, the record is the only thing standing between you and an unwinnable he-said-she-said with a client whose reviewer has a different opinion. When a dispute comes, and in regulated work it eventually comes, the linguist with a record says "here are the seven findings I marked, here are their severities, here is the Critical I caught and fixed, here is the source ambiguity I queried," and the conversation is over in an hour. The linguist without a record says "I'm sure I checked it carefully," which is an opinion against an opinion, and opinions against a paying client lose. The record converts your invisible diligence into visible, dated, named evidence. It is the difference between being a professional who can prove their value and a vendor who can only assert it.
There is a subtler protection too. A record disciplines you. The act of writing each finding into a row with a severity and a consequence rationale forces a rigor that vibe-checking does not. You cannot write "Critical" in a cell without articulating the hazard, and the act of articulating it catches the cases where you were about to over-grade a harmless slip or under-grade a real one. The record is not just a defense after the fact; it is a checklist that makes the evaluation better while you do it.
Product: The Record Is What the LSP Sells
For the LSP, the quality record is not overhead; it is the product differentiator that survives the race to the bottom. The market is full of vendors selling raw machine output at MTPE prices and hoping nothing critical slips. Anyone can sell cheap words. What a vendor cannot easily copy, and what a sophisticated client will pay a defensible premium for, is provable quality: "here is the throughput, here is the risk tier each content type received, here is the ISO 5060 error score with zero Criticals on delivery, here is the terminology conformance, and here is the segment-level quality record, all defensible under the revised ISO 18587." That sentence is a sales weapon a raw-MT vendor cannot say, because they have no record to point to. The quality record is the artifact that lets an LSP sell a quality tier instead of a price, and selling a tier is the only way to keep margins in a market where words themselves are nearly free.
The record is also the audit-readiness asset. A regulated client, or a certifier checking ISO conformance, does not want to hear that the LSP "has a quality process." They want to see the evidence that the process ran on this file: the marked findings, the severities, the named evaluator, the gate decision, the sign-off. An LSP that produces a quality record per file is audit-ready by construction; one that relies on its linguists' diligence and memory is one disputed delivery away from having nothing to show. The record turns a quality claim into a quality asset, and assets are what a business is built on.
Building the Habit Without Drowning in Overhead
The honest objection to all of this is time. A full segment-level record on every file would bury a linguist in clerical work and erase the MTPE economics that made the whole pipeline worth running. So the discipline is not "record everything always." It is matching the rigor of the record to the consequence of the content, the same risk-tiered logic that governs post-editing effort itself.
On high-liability content, medical, legal, financial, life-safety, the full record is non-negotiable: every finding, every severity, every fix, named and dated, because that is exactly the content where a dispute or an audit is most likely and most expensive. On lower-stakes content, a lighter record suffices: a summary tally and the score, with full rows reserved for any Major or Critical, because the cheap content rarely produces the dispute that needs the full evidence trail. The record scales with the stakes. The skill is not maximalism; it is knowing which file gets the full nine-field treatment and which file gets a one-line score stamp, and being able to defend that choice too.
Two practical habits make the record nearly free on the files that need it. First, capture as you go. The worst way to build a record is to translate the whole file, then go back and reconstruct what you fixed from memory, which is both slow and lossy. The professional marks the finding in the moment of fixing it, one row per error as it surfaces, so the record is a byproduct of the work rather than a second job after it. Most modern CAT tools and TMS platforms let you log a category and severity against a segment as you edit, and an LLM-assisted workflow can draft the raw-versus-fixed comparison and a first-pass category for you to verify, which collapses the clerical cost. Second, standardize the template. A record you reinvent every job is a record you will skip under deadline. One fixed nine-field template, the same columns, the same category names, the same severity definitions, the same score formula, reused on every file, makes the record a reflex instead of a decision. The template is the thing you build once and lean on forever.
Match the record to the stakes, capture findings as you fix them, and standardize one template you never reinvent. That is how the record stops being overhead and becomes a reflex that protects you.
Key Takeaways
- A quality record is a structured, segment-level account of what an evaluator found, how serious each finding was, what was done about it, and what the file scored, written so a client and an auditor can reconstruct the verdict without the evaluator in the room. It is provenance, the documented chain of where each segment came from and what happened to it, that you can hand to an auditor.
- The defensible row carries nine fields: segment id (the precise, tool-native location), source and target (the two things compared, ideally raw MT and final post-edited target both), error category (the MQM dimension, which makes the record diagnostic), severity (Critical/Major/Minor, which drives the gate), fix (closing the loop to what is now right), evaluator (the named, accountable human), and the file-level score and pass/fail status.
- The score is presented in three layers so a reader can stop at any depth: the headline (score, status, Critical count), the tally (counts and total penalty, so the score is reconstructable), and the rows (the segment-level evidence). A score nobody can re-derive from the rows is a vibe check wearing a decimal point.
- Present the gate separately from the number: show the Critical count as its own visible field, because a file can compute to 96 and still FAIL on one Critical. A record that folds a Critical into a high score launders a hazard into a pass and is worse than no record at all.
- Separate errors from queries. A query flags an upstream ambiguity in the source that you raised rather than silently guessing; it does not score against the translation but belongs in the record, because it is the row that protects you when a client later disputes a judgment the source itself forced.
- In the worked oxygen-concentrator record, the raw MT failed on one Critical (an inverted air-intake warning) despite an otherwise-passing 12-point penalty; the delivered file passed because the named post-editor caught and fixed the Critical, the two Majors, and the two Minors before delivery. The two-state record proving raw-failed-then-human-fixed is the artifact that shows the human, not the engine, owned the quality, exactly the value ISO 18587 requires.
- For the linguist the record is protection: it converts invisible diligence into dated, named evidence that ends a dispute in an hour instead of an unwinnable opinion-versus-opinion fight, and the act of writing each finding disciplines the evaluation while you do it. For the LSP it is the product: provable, ISO 5060-scored, ISO 18587-defensible quality is the tier a sophisticated client pays a premium for and a raw-MT vendor cannot sell.
- Match the rigor of the record to the consequence of the content (full nine-field records on high-liability work, lighter score stamps on low-stakes content), capture findings as you fix them rather than reconstructing from memory, and standardize one reusable template, so the record becomes a near-free byproduct of the work rather than a second job after it.
Skill.re