โ†
AI for Translation & Localization
Aware ยท M10 ยท lesson 10 of 19 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
ISO 5060 and MQM Error Scoring
๐Ÿ“–
now learning

ISO 5060 and MQM Error Scoring

15 min

The file came back from the evaluator with a single red flag and a one-line verdict: failed. It was a 4,200-word user manual for a portable oxygen concentrator, machine-translated into German, post-edited under a tight deadline, and it was, by any reading-for-flow standard, gorgeous. The evaluator scored it segment by segment against an error typology, an agreed list of error categories and severities, and tallied what she found: eleven Minor errors, three Major errors, and exactly one Critical error. The Critical was four words long. In a safety warning about the device's pressure-relief valve, the source said the operator must not block the air intake during operation. The German target, fluent and idiomatic and indistinguishable in tone from every other clean sentence around it, said the operator should keep the air intake covered. One inverted instruction on a medical device. The evaluator did not weigh it against the eleven Minors and three Majors and compute an average. She did not need to. The rule is older than the standard and now written into it: one Critical error fails the file, regardless of how clean everything else is. This lesson is about the model behind that verdict, ISO 5060 and the MQM error typology it formalizes, and how a real evaluation marks errors, assigns severity, weights them, computes a score, and decides whether a file ships or dies on a single sentence.

From a Vibe Check to Error-Marking

There are two fundamentally different ways to judge a translation, and the gap between them is the whole subject of this lesson. The first way is the holistic judgment: a reviewer reads the target, forms an overall impression, and renders a verdict like "this is good," "this needs work," or "publish it." It is fast, it is intuitive, and it is what most people mean when they say they "reviewed" a file. The second way is the analytic judgment: an evaluator goes through the translation systematically, finds and marks each individual error, classifies what kind of error it is and how serious it is, and arrives at a verdict by tallying those marked errors against an agreed rulebook. The holistic method produces an opinion. The analytic method produces evidence.

For most of the history of the craft, holistic judgment was the working standard, and for a skilled human translating from scratch it was often good enough, because the same expert who caught the awkward seam also caught the meaning error hiding underneath it. The machine-translation era broke that comfortable coupling. As the previous lessons in this program established, a machine produces output that is fluent first and accurate second: grammatical, idiomatic, confident prose that can mean the opposite of the source. A holistic read, an impression formed from the surface, is exactly the wrong instrument for that failure mode, because the surface is the part the machine gets right. "Looks fine to me" is calibrated to catch the errors the machine no longer makes and to miss the ones it now makes constantly. The industry needed a method that judged meaning against the source rather than prose against the reader's ear, and that method is analytic, error-marking evaluation.

A holistic review gives you an opinion about the prose. An analytic evaluation gives you marked evidence about the meaning. In the MT era, only the second one catches the error that reads perfectly.

Before going further, the vocabulary, used precisely throughout. Machine translation (MT) is any system that renders text from a source language into a target language with no human writing the words; a large language model (LLM) is a general-purpose text predictor that translates as a side effect of its broad competence. Machine-translation post-editing (MTPE), sometimes shortened to PE, is the workflow where a human edits machine output rather than translating from scratch. A segment is the unit a translation tool works in, usually a sentence or a short block, the row you see in a CAT tool (a computer-assisted translation tool, the editing environment a linguist works in). A termbase is the controlled glossary of a client's approved terms. MQM is Multidimensional Quality Metrics, an analytic error-typology framework that classifies translation errors by dimension and severity. ISO 5060:2024 is the international standard, published in 2024, that formalizes an MQM-aligned model for the human analytic evaluation of translation output. An error typology is the agreed catalogue of error categories an evaluation marks against. A dimension is the kind of error (accuracy, terminology, and so on). Severity is how much a given error matters, graded as Critical, Major, or Minor. We will build each of these into a working method, slowly, and then run a real evaluation with the numbers showing.

What Analytic Evaluation Actually Buys You

It is worth being concrete about why anyone would trade the speed of a vibe check for the labor of marking every error. Analytic evaluation buys four things a holistic verdict cannot provide. First, it is defensible: when a client disputes a delivery, "I marked three Major terminology errors in segments 14, 88, and 203, here they are" survives an audit in a way "it felt off to me" never will. Second, it is diagnostic: a tally that shows your errors cluster in terminology tells you the engine is drifting off the termbase, which is a fixable process problem, while a holistic "needs work" tells you nothing about what to change. Third, it is comparable: two evaluators applying the same typology to the same file should land close to the same score, and two engines scored the same way can be ranked honestly, which a matter of taste cannot do. Fourth, it is gateable: because it produces a structured count of errors by severity, you can write a hard pass/fail rule on top of it, which is exactly what the one-Critical-fails gate is. None of those four properties, defensibility, diagnosis, comparability, or a gate, can be built on an impression. They all require marked errors.

The MQM Dimensions: What Kind of Error Is It

The first half of an analytic evaluation is answering, for each error you find, what kind of error is this. MQM organizes errors into dimensions, and ISO 5060 aligns with the same set. You do not need the full hundred-node MQM tree to work; you need the core dimensions that a real evaluation leans on, and you need to be able to slot any error you find into the right one without hesitation. A working linguist should be able to name and recognize these on sight.

Accuracy: Does the Target Mean What the Source Means

Accuracy is the relationship between the target's meaning and the source's meaning. This is the dimension that matters most in the MT era, because it is the one fluency hides. Accuracy errors include mistranslation (the target says something the source did not), omission (something in the source is missing from the target, the dropped negation lives here), addition (the target invents content not in the source, a hallucinated number lives here), and untranslated (source text left in the target by mistake). An accuracy error is judged against the source segment, never against the reader's ear. The inverted oxygen-concentrator warning that opened this lesson is an accuracy error: a mistranslation that flipped a prohibition into a recommendation. Crucially, the prose was flawless. Accuracy and fluency are different dimensions, and an error can be perfect on one and catastrophic on the other.

Terminology: Does It Use the Approved Words

Terminology errors are failures to use the client's approved term from the termbase, or failures to use a term consistently. The engine renders the client's device as "the unit" when the approved term is "the concentrator," or translates an approved trade term with a fluent synonym the model saw more often in training. This is a distinct dimension from accuracy because a terminology error can be technically accurate (the synonym means roughly the same thing) and still be a defect, because consistency and the approved term are themselves requirements. In regulated and technical content, a wandering term is not a stylistic quibble; it can break a regulatory submission or confuse a reader who relies on one word meaning one thing across a whole document.

Locale Conventions: Does It Fit the Target Market

Locale conventions (a locale is a specific language-and-region combination, such as German for Germany or Spanish for Mexico) cover the formatting and cultural norms a target market expects: date and time formats, number and decimal separators, units of measurement, currency, address and phone formats, and formality or register. An engine that renders a US date 03/04/2026 into a market that reads dates day-first has produced a locale error that can flip a deadline by months. A figure that keeps a US decimal point in a locale that uses a comma can shift a dosage by orders of magnitude. Locale errors are subtle, they read fluently, and they are expensive precisely because the surface looks correct.

Fluency and Style: Is the Target Well-Formed

Fluency is whether the target language is itself well-formed: grammar, spelling, punctuation, register, and natural phrasing, judged on the target alone without reference to the source. Style errors are deviations from the client's style guide. This is the dimension a holistic review is actually good at, because you can assess it by reading the target by itself. It is also, in the MT era, the dimension where machines fail least, which is the whole trap: an evaluator who only scores fluency is scoring the one thing the engine reliably gets right and ignoring the dimensions, accuracy and terminology, where it fails. A file can be flawless on fluency and fail on accuracy, and that file is the dangerous one.

Fluency is judged on the target alone. Accuracy is judged on the target against the source. The machine gives you the first for free and fails the second silently, which is why an evaluation that only scores fluency scores the wrong thing.

The Severity Scale: How Much Does It Matter

Knowing what kind of error you have found is only half the evaluation. The other half, and the half that decides the file's fate, is severity: how much this particular error matters. The same dimension can span the whole severity range. A terminology error can be Minor (a slightly off synonym a reader would shrug at) or Critical (a wrong drug name). The severity scale is where ISO 5060 and MQM do their most important work, because severity, not dimension, is what feeds the pass/fail gate. There are three bands.

Minor: Noticeable but Harmless

A Minor error is a flaw that does not affect the meaning or usability of the content in any consequential way. A slightly awkward phrasing that is still perfectly understandable, a stylistic choice you would have made differently, a small inconsistency no reader would be harmed by, a missing comma. Minor errors are real, they count, and enough of them degrade a delivery, but no single Minor error is, by itself, a problem. In the worked tally below, a Minor will carry the lowest weight.

Major: Meaningfully Impairs the Content

A Major error meaningfully affects the comprehension or usability of the content. A mistranslation that distorts a non-critical part of the meaning, a wrong term a reader would notice and that undermines trust, a locale error that makes a date genuinely ambiguous, a grammatical breakdown that forces the reader to re-read to recover the sense. Major errors are serious and they count heavily. The content can, in principle, still function while someone notices and fixes the problem, but a file accumulating Majors is a file in trouble. A Major typically weighs several times a Minor.

Critical: Dangerous, Liable, or Actively Misleading

A Critical error renders the content dangerous, unusable, legally exposed, or actively misleading on a point that matters. The inverted contraindication. The flipped safety instruction. The corrupted dosage. The wrong drug name. The inverted indemnity clause that swaps which party is liable. A Critical error is not merely a worse Major; it is a different category of consequence, an error whose result is harm, liability, or the fundamental failure of the content's purpose. And it is the band that triggers the gate.

Severity Tracks Consequence, Not Edit Size

The single most common beginner mistake in scoring is to map severity onto the size of the textual change, as if a one-word error must be Minor and a rewritten sentence must be Major. Unlearn that immediately. Severity has nothing to do with how many characters changed; it is about what happens when a real reader acts on the output. The most dangerous error in this entire program, the dropped negation, is a change of one tiny word, and it is Critical, because that one word inverted a safety instruction. A sprawling, clumsy, rewritten paragraph that is awkward but still conveys the right meaning on a low-stakes marketing page might be only Minor. Always ask the consequence question, never the character-count question. A one-word flip in a drug label is Critical; a stiff paragraph in a blog post is Minor.

How Errors Are Weighted Into a Score

Here is where analytic evaluation becomes arithmetic, and where a vague "needs work" becomes a number you can put on a delivery note. In an MQM-aligned model, each marked error is assigned a penalty, a number of points, based on its severity. The exact weights are configurable per project, which matters: a client can decide their content is unusually unforgiving and crank the penalties up. But the canonical, widely used default weighting is the one to learn first, because it shows the logic, and the logic is what generalizes.

The classic MQM severity weights are 1 point for a Minor error, 5 points for a Major error, and a large penalty for a Critical, commonly 25 points or more, often set high enough to fail any realistic file on its own. Some implementations express the Critical not as a weight at all but as an absolute gate, which we will come to. The penalties are summed across every marked error to give a total penalty for the file.

A raw penalty count is not yet comparable across files, though, because a 200-word file and a 20,000-word file are not held to the same absolute error budget; the long file has more room to accumulate Minors without being worse in quality. So the model normalizes the penalty against the size of the evaluated content. The standard normalization expresses errors per unit of text, typically per word or per a fixed block such as 1,000 words. The structure of an MQM-style quality score is:

  • Total penalty = (Minor count x 1) + (Major count x 5) + (Critical count x 25-or-gate).
  • Penalty rate = total penalty divided by the evaluated word count, often multiplied out to a per-1,000-word figure so the numbers are readable.
  • Quality score = a normalized figure derived from the penalty rate, frequently expressed on a 0-to-100 scale where 100 is a flawless file and the penalty rate pulls the score down. A common formulation is roughly 100 minus the penalty rate scaled to the sample, so a file with few light errors lands in the high 90s and a file riddled with Majors falls well below a passing threshold.

The exact formula varies by implementation, and you should always read the specific scoring profile a client or LSP (a language-service provider, the vendor that delivers the localization) has defined rather than assume. But every analytic model shares the same three moves: assign a severity penalty to each error, sum them, and normalize against length. Memorize the moves, not any one client's constants.

Penalty by severity, summed across the file, normalized against length. That is the entire engine of an analytic score. The constants change per project; the three moves never do.

A Worked Evaluation: Scoring Five Segments

Theory settles only when you watch it tally. Let us evaluate a small sample by hand, the way a QE analyst (a quality-evaluation analyst, the person who runs the error-marking) would, so you can see a score emerge from marked errors. Imagine a 250-word excerpt of the oxygen-concentrator manual, machine-translated into German and post-edited. We will walk five representative segments, mark each error by dimension and severity, and then compute the file's fate. For readability the source and target meanings are given in English.

Segment by Segment

Segment 1. Source: "Place the device on a flat, stable surface at least 30 cm from any wall." Target: "Place the device on a flat, stable surface at least 30 cm from any wall." Clean. No error marked. This is the common case, and it is exactly why MT-first pipelines are economically irresistible: most segments come back correct, and trusting the clean ones is fine. The danger is never the clean majority; it is the one segment that reads just as clean and is not.

Segment 2. Source: "The concentrator delivers up to 5 litres of oxygen per minute." Target: "The unit delivers up to 5 litres of oxygen per minute." The meaning is intact and the prose is fluent, but the client's termbase specifies "concentrator" as the approved term for the device, and the engine substituted the generic "unit." This is a terminology error. Its consequence here is consistency and approved-term compliance, not harm or comprehension breakdown, so it is graded Major if the client treats termbase compliance as strict (technical and regulated clients usually do) or Minor on looser content. For this regulated manual we mark it Major. Penalty: 5.

Segment 3. Source: "Replace the filter every 03/04 months depending on usage." Target, rendered for a day-first locale: the figures "03/04" were carried over and now read as a date, the third of April, rather than the source's "every 3 to 4 months." This is a locale and accuracy tangle: the number formatting and the meaning both corrupted, turning a maintenance interval into a nonsensical calendar date. A reader could be genuinely misled about when to service a medical device, but the result is confusing rather than acutely dangerous, so we mark it Major. Penalty: 5.

Segment 4. Source: "Long-term use may occasionally cause mild dryness of the nasal passages." Target: "Long-term use may occasionally cause mild dryness of the nasal pathways." A small word choice, "pathways" instead of the idiomatic "passages." It is understandable, harmless, slightly off. This is a fluency/style error, graded Minor. Penalty: 1. We also notice a missing comma in the same segment, another Minor fluency slip. Penalty: 1.

Segment 5. Source: "Do not cover the air intake while the device is operating." Target: "Keep the air intake covered while the device is operating." The negation collapsed and the instruction inverted: a safety prohibition became a safety recommendation, on a medical device whose pressure-relief depends on an unobstructed intake. This is an accuracy error, specifically an omission of the negation producing a mistranslation, and its consequence is a hazard to the operator. It is Critical. The penalty, on a 25-point weighting, is 25; on a gate model, its mere presence fails the file outright.

Tallying the File

Now tally the five segments. We marked: one Major terminology (5), one Major locale/accuracy (5), two Minor fluency (1 + 1), and one Critical accuracy (25 or gate). On the additive weighting, the total penalty is 5 + 5 + 1 + 1 + 25 = 37 penalty points across roughly 80 evaluated words in these five segments, which is a brutal penalty rate, but the rate is almost beside the point here. Run the normalization for the full 250-word excerpt and you might land at a per-1,000-word penalty that, on the accuracy and terminology dimensions alone, already sits below most passing thresholds.

But look at what actually decided the verdict. Strip out the Critical and imagine segment 5 had been post-edited correctly. You would be left with two Majors and two Minors across the sample: a penalty of 12, a file that is imperfect but plausibly shippable after the terminology and locale fixes, a file a client would accept with light rework. Add the one Critical back in and none of that matters. The file fails. Not because the penalty arithmetic crossed a line, but because of the rule that sits above the arithmetic.

Two Majors and two Minors is a fixable file. Add one Critical and it is a failed file. The Critical does not raise the score past a threshold; it removes the file from the conversation entirely.

The Rule That One Critical Fails the File

This is the operational heart of ISO 5060-aligned evaluation, and it deserves to be stated as a law and then justified, because it feels harsh until you see why it is the only sane policy. The presence of a single Critical error fails the file, regardless of how clean every other segment is. It does not matter if the other 9,999 segments are flawless. One inverted contraindication, one flipped safety instruction, one corrupted dosage, and the deliverable does not ship.

The justification connects straight back to the fluent-error asymmetry that runs through this whole program. If errors averaged out, you could trade a Critical against a sea of perfect segments and call the file "99.99% good." But a Critical does not average. The operator who covers the air intake because the one inverted sentence told him to is not protected by the flawless translation of the other ten thousand sentences. The harm is not diluted by the surrounding quality; it is delivered in full by the single bad sentence. A scoring model that let a clean file absorb a Critical would be a model that shipped rare hazards and lawsuits as long as they were statistically uncommon. That is not a quality gate; it is a lottery with someone's safety as the stake.

So the analytic model treats the Critical as a gate, not a weight. Its presence is disqualifying on its own. Some scoring profiles implement this by setting the Critical penalty absurdly high, high enough that no realistic file can pass with one. Others implement it as an explicit hard rule layered on top of the arithmetic: if Critical count is greater than zero, status equals fail, full stop, no matter the computed score. The two implementations reach the same place. This is what people mean when they say a single Critical "fails the file," and it is why the silent critical error, the fluent one the eye skips, is the single most important thing a post-editor or evaluator hunts for. It is not one more error to count toward a total. It is the one error that, undetected, makes the entire delivery a failure no matter how good everything else is.

Why the Gate Changes How You Read

The gate has a profound effect on how you should spend your attention as an evaluator or post-editor. Because a single Critical is disqualifying and a single Minor is nearly free, your effort is not evenly distributed across the file. You do not read for an even polish. You hunt, with disproportionate intensity, for the high-consequence elements where a Critical hides: negations, dosages and numbers, drug and proper names, safety instructions, dates in regimens, and legal obligations. A model that fails on one Critical is a model that rewards finding that one error above all else. Miss ten Minors and the file still passes; miss one Critical and you shipped the hazard. The math of the gate is also a map of where to look.

Analytic vs. Holistic in Practice: Why Both Exist

It would be a misreading of this lesson to conclude that holistic judgment is worthless and analytic scoring is always the answer. The honest picture is that the two methods do different jobs, and a mature operation uses both deliberately. The analytic, error-marked ISO 5060 evaluation is the method you reach for when the verdict must be defensible: an acceptance test on a regulated delivery, an engine bake-off where two systems are ranked on the same typology, a dispute with a client, an audit, a go/no-go gate on high-liability content. It is slower and more expensive precisely because it produces marked evidence rather than an impression, and that evidence is the point.

The holistic read still has its place for speed and triage: a quick first pass to decide whether a file is even in the ballpark before you invest in a full analytic evaluation, a sense-check on low-stakes content where a marked evaluation would cost more than the content is worth, a reviewer's overall feel that flags a file for deeper analytic scrutiny. The mistake is never "using holistic judgment." The mistake is using holistic judgment as the quality gate on content where a fluent error is catastrophic, because that is exactly the place the holistic method is structurally blind. The discipline is matching the method to the consequence: vibe-check the throwaway, analytically score the file that can hurt someone, and never confuse the two.

The Score Is a Record, Not Just a Number

One last reframe that separates an evaluator from a grader. The output of an analytic evaluation is not really the score; it is the marked error report. The number, "this file scored 72 and failed on one Critical," is a summary. The asset is the structured list underneath it: segment 5, accuracy, omission of negation, Critical, here is the source, here is the target, here is the fix. That report is what makes the verdict defensible to a client, diagnostic for the process, and reproducible by a second evaluator. It is also the artifact the revised ISO 18587, the post-editing standard expanded to cover AI and LLM output and in DIS ballot with publication targeted for late 2025 into 2026, leans on when it insists that the human owns the quality. A score with no marked errors behind it is just a vibe check wearing a number. The marked report is the proof.

Key Takeaways

  • Analytic evaluation marks each error individually by dimension and severity against an agreed typology; holistic evaluation forms an overall impression of the prose. In the MT era only the analytic method catches the fluent error that reads perfectly, because a vibe check is calibrated to the surface the machine gets right and blind to the meaning it gets wrong.
  • ISO 5060:2024 formalizes an MQM-aligned analytic model. The core MQM dimensions are accuracy (target meaning against source meaning, where mistranslations and dropped negations live), terminology (approved-term and consistency compliance), locale conventions (dates, numbers, units, currency, formality), and fluency/style (the target judged on its own).
  • Severity grades how much an error matters: Minor (noticeable but harmless), Major (meaningfully impairs comprehension or usability), and Critical (dangerous, legally exposed, or actively misleading). Severity tracks consequence, not edit size: a one-word dropped negation is Critical; a stiff marketing paragraph is Minor.
  • Errors are weighted into a score by assigning a penalty per severity, classically 1 point for Minor, 5 for Major, and a large penalty (commonly 25 or more, or an outright gate) for Critical, then summing the penalties and normalizing against the evaluated word count, typically per 1,000 words, to make files of different lengths comparable.
  • The three moves of any analytic score are constant even when the constants are not: assign a severity penalty to each error, sum across the file, and normalize against length. Always read the specific scoring profile a client or LSP has defined rather than assume the default weights.
  • In the worked five-segment example, a Major terminology error, a Major locale/accuracy error, and two Minor fluency errors made a fixable file; adding one Critical accuracy error, an inverted safety instruction, failed it outright, because the Critical is a gate, not a weight.
  • One Critical error fails the file regardless of how clean every other segment is, because harm does not average out: the reader who acts on the one bad sentence is not protected by the flawless translation of the rest. The model treats the Critical as disqualifying, either by an extreme penalty or an explicit hard rule.
  • The gate is also a map: because one Critical fails and one Minor is nearly free, you spend your attention hunting the high-consequence elements where Criticals hide. The real output of an evaluation is the marked error report, not the number, because the marked report is what makes the verdict defensible, diagnostic, and reproducible, and what ISO 18587 relies on when it says the human owns the quality.