Applying MQM / ISO 5060 Error Categories
Maria opened the returned evaluation and stopped on a single cell. A junior reviewer on her team had marked a German segment in a banking app with one word in the comment column: "off." Not wrong, not a category, not a severity. Just off. The segment in question rendered the source "Your transfer will not be processed until the recipient confirms" as a fluent, idiomatic German sentence that, read aloud, sounded like every other clean line in the file. It was also a dropped negation: the target told the user the transfer would be processed, immediately. Maria knew that error cold. What she could not do with the reviewer's note was anything useful. She could not tell the engineer whether to retrain a glossary or fix a placeholder. She could not tell the client whether the file shipped or failed. She could not tell the next post-editor where to look. "Off" is a feeling. It points at nothing, it fixes nothing, and it survives no audit. The entire discipline of this lesson is the move from "off" to a row that reads: accuracy, omission, Critical, segment 142, here is the source, here is the target, here is the fix. That row is not bureaucracy. It is the difference between an evaluation that does work and a sticky note that does not. This lesson teaches you to look at any suspect segment and name, fast and defensibly, exactly what kind of error it is, because the kind of error you call it decides the fix, the severity, and whether the file ships.
Why the Category Is the Whole Job
It is tempting to treat error categorization as paperwork: you already saw the error, surely writing down which box it goes in is a clerical afterthought. That instinct is exactly backwards, and unlearning it is the point of this lesson. Naming the category is not what you do after you understand the error. Naming the category is understanding the error. The dimension you assign is a claim about what went wrong, and that claim drives three downstream decisions that "looks off" cannot touch: how the error gets fixed, who fixes it, and how severe it is.
Consider the same surface symptom, a German word that does not match the client's glossary, arriving through four completely different failure paths. If the engine rendered an approved device name with a fluent synonym, that is a terminology error, and the fix is to enforce the termbase entry and probably to tighten the glossary the engine sees. If the engine translated a word that should have stayed in English, a product name, that is an accuracy error of the untranslated-or-over-translated kind, and the fix is a do-not-translate rule. If the word is correct terminology but in the wrong grammatical case for the sentence, that is a fluency error, and the fix is a grammar correction with no glossary implication at all. And if the word is the right term but uses an American spelling in a file destined for a German-Swiss locale convention, that is a locale error, and the fix is a locale-profile setting. One symptom, four categories, four different fixes, four different people who should hear about it. Mark them all "wrong word" and you have thrown away every signal that tells the operation what to change.
The category is not a label you attach after diagnosis. The category is the diagnosis. Get it wrong and you fix the wrong thing, route it to the wrong person, and score the wrong severity.
Before we go further, the working vocabulary, defined precisely and used precisely from here on. Machine translation (MT) is any system that turns source-language text into target-language text with no human writing the words. A large language model (LLM) is a general-purpose text predictor that translates as a side effect of its broad fluency, which makes it more fluent and more confidently wrong than classic MT. Machine-translation post-editing (MTPE), often shortened to PE, is the workflow where a human edits machine output rather than translating from a blank target. A segment is the unit a translation tool works in, usually a sentence or short block, the row you see in a CAT tool (a computer-assisted translation tool, the editing environment a linguist lives in). A termbase is the controlled glossary of a client's approved terms. MQM is Multidimensional Quality Metrics, an analytic error-typology framework that classifies translation errors by dimension and severity. An error typology is the agreed catalogue of error categories an evaluation marks against. A dimension is the kind of error: accuracy, terminology, locale, fluency. An error category is a specific sub-type inside a dimension, such as mistranslation or omission inside accuracy. Severity is how much an error matters, graded Critical, Major, or Minor. ISO 5060:2024 is the international standard, published in 2024, that formalizes an MQM-aligned model for the human analytic evaluation of translation output, the rulebook that says you mark errors by dimension and severity rather than forming a vibe. We are going to take that rulebook and operate it on real segments.
The Model in One Paragraph, So We Can Apply It
A previous lesson in this program established the model; this lesson applies it, so a single paragraph of recap is enough to ground the hands-on work. Analytic evaluation marks each error individually by dimension (what kind) and severity (how much), against an agreed typology, instead of forming an overall impression of the prose. That distinction matters in the MT era specifically because machine output is fluent first and accurate second: the surface is the part the engine reliably gets right, so an impression formed from the surface is blind to the meaning errors the engine actually makes. The score is built by assigning a penalty per severity, summing across the file, and normalizing against length, and one Critical error fails the file regardless of how clean the rest reads. Everything in this lesson sits one level down from that: before you can score severity or compute a rate, you have to name the dimension correctly, because the dimension is what the rest of the machinery acts on. This lesson is the categorization layer the score is built on.
The Four Dimensions, Reframed as Questions You Ask the Segment
The MQM tree has many nodes, and ISO 5060 aligns with the same structure, but a working linguist does not categorize by walking a hundred-node taxonomy. A working linguist asks four questions, in order, and the first one that fires names the dimension. The order matters, because the dimensions are not symmetric in consequence: accuracy is checked first because it is the dimension fluency hides and the dimension that ships hazards. Train yourself to run these four questions on every suspect segment, as a reflex, until the categorization is automatic.
Accuracy: Does the Target Mean What the Source Means
Accuracy is the relationship between the target's meaning and the source's meaning, judged against the source segment and never against the reader's ear. This is the dimension that matters most in the MT era because it is the one fluency conceals, and it is the dimension a holistic read is structurally blind to. Accuracy has four working categories you must be able to name on sight:
- Mistranslation: the target says something the source did not. The classic flipped polarity, a prohibition rendered as a permission, lives here. So does a "may" rendered as "must," an "and" rendered as "or," a swapped subject and object.
- Omission: something present in the source is missing from the target. The dropped negation is the most dangerous member of this category, because a missing "not" is invisible in grammar and catastrophic in meaning. Dropped qualifiers, dropped clauses, dropped warnings all live here.
- Addition: the target invents content not in the source. A hallucinated number, an extra sentence the engine confabulated, a clarifying phrase nobody asked for that changes scope.
- Untranslated: source text left in the target by mistake, or, in the over-translated mirror case, a do-not-translate item like a product name or code identifier that got translated when it should have stayed verbatim.
The accuracy question is always the same: hold the target against the source and ask whether they assert the same thing. If they do not, you have an accuracy error, and you then name which of the four categories it is, because the category steers the fix. A mistranslation needs a meaning correction. An omission needs the missing element restored. An addition needs the invented content cut. An untranslated needs a translate-or-do-not-translate decision. Four categories, four fixes, all inside one dimension.
Terminology: Does It Use the Approved, Consistent Word
Terminology errors are failures to use the client's approved term from the termbase, or failures to use a chosen term consistently across the file. The engine renders the client's "concentrator" as "device," or translates an approved trade term with a fluent synonym it saw more often in training, or uses three different translations for one source term across a document. The defining test that separates terminology from accuracy is this: a terminology error can be technically accurate, the synonym means roughly the same thing, and still be a defect, because the approved term and consistency are themselves requirements. If the wrong word also changes the meaning, you may be looking at an accuracy error instead, or both. The reason terminology earns its own dimension is that the fix and the owner are different: terminology errors are fixed by enforcing the termbase and often by tightening what the engine is allowed to produce, and they route to the terminologist or the glossary owner, not to a general reviser. In regulated and technical content a wandering term is not a stylistic quibble; an inconsistent device name can break a regulatory submission, and a drifting financial term can mislead a reader who relies on one word meaning one thing throughout.
Locale Conventions: Does It Fit the Target Market's Rules
Locale conventions (a locale is a specific language-and-region pairing, such as German for Germany, German for Switzerland, or Spanish for Mexico) cover the formatting and cultural norms a target market expects: date and time formats, number and decimal separators, units of measurement, currency, address and phone formats, quotation-mark style, and formality or register conventions. The trap with locale errors is that they read perfectly fluently and are often technically accurate as text, which is exactly why they are expensive: the surface looks right. An engine that carries a US date 03/04/2026 into a market that reads dates day-first has flipped a deadline by months while producing a grammatical date. A figure that keeps a US decimal point in a locale that uses a comma as the decimal separator can shift a number by orders of magnitude, 1,000 meaning one thousand in one convention and one in another. A formal-register requirement, the German "Sie" versus "du," is a locale-and-style decision an engine gets wrong by defaulting to whatever it saw most. Locale errors route to a locale profile or a formatting rule, not to a meaning fix, which is why they earn their own dimension even though their consequence sometimes overlaps with accuracy.
Fluency and Style: Is the Target Itself Well-Formed
Fluency is whether the target language is itself well-formed, judged on the target alone without reference to the source: grammar, spelling, punctuation, register, and natural phrasing. Style errors are deviations from the client's style guide, a too-casual tone in a legal notice, a sentence structure the brand guide forbids. Fluency has working categories you should be able to name: grammar (agreement, case, tense, word order), spelling (including the locale-spelling overlap), punctuation, and register (formality mismatch). Here is the central irony you must internalize: fluency is the dimension a holistic read is actually good at, because you can assess it by reading the target by itself, and it is also the dimension where machines fail least. An evaluator who spends their attention on fluency is polishing the one thing the engine reliably gets right while the accuracy and terminology errors, the ones that actually fail files, slip past. Fluency errors are real and they count, but they are usually the lowest-severity dimension, and an evaluation weighted toward fluency is an evaluation looking in the wrong place.
Accuracy is judged on the target against the source. Fluency is judged on the target alone. The machine hands you the second for free and fails the first silently, so an evaluation that drifts toward fluency is scoring the wrong dimension.
A Decision Procedure: Naming the Dimension Without Hesitating
Knowing the four dimensions is not the same as being able to assign one fast and defensibly under deadline. What separates a fluent evaluator from a hesitant one is a fixed procedure, run identically on every suspect segment, so the categorization is a reflex and not a debate. Here is the procedure. Run it in this exact order, because the order encodes the consequence hierarchy.
Step One: Read the Source First, Then the Target
Read the source segment and form, in your own head, what it asserts, before you let the fluent target anchor you. This single sequencing discipline is what makes the accuracy question answerable. If you read the smooth target first, your mind accepts its claim as the meaning, and then you are comparing the source against a conclusion you have already drawn from the target. Read source first, hold its claim, then look at the target as a thing to be checked against that claim, not as the source of the claim. Every accuracy catch depends on this order.
Step Two: Run the Accuracy Gate
Ask: does the target assert the same thing the source asserts? Check the high-consequence elements specifically, the negations and polarity, the numbers and dosages and units, the named entities, the obligations and parties, the dates and scope words, the warnings and conditionals, because these are where accuracy errors hide and where they do the most damage. If the answer is no, you have an accuracy error. Name the category, mistranslation, omission, addition, or untranslated, and stop; the dimension is accuracy. Do not move on to the cheaper dimensions, because an accuracy error that you mislabel as fluency gets the wrong severity and the wrong fix. The accuracy gate fires first and fires hard.
Step Three: If the Meaning Holds, Check the Approved Terms
If the target means what the source means, ask: does it use the client's approved terms, consistently? Compare the rendered terms against the termbase. A mismatch that does not change the meaning, a fluent synonym for an approved term, is a terminology error. A mismatch that does change the meaning was already caught at the accuracy gate. This is why order matters: terminology is the dimension of the technically-correct-but-non-compliant word, and you only reach it after meaning has passed.
Step Four: Check the Locale Conventions
If the meaning and the terms hold, ask: does the formatting and register fit the target locale? Dates, decimals, units, currency, quotation style, formality. A locale convention violated, even in an otherwise meaning-correct and term-correct segment, is a locale error. This sits below terminology because a wrong term usually matters more than a wrong date format, though not always; severity, which we handle separately, can lift a locale error up when the consequence is high, as with a flipped date in a medication schedule.
Step Five: Finally, Sweep Fluency
If the segment survives the first four checks, judge the target on its own for grammar, spelling, punctuation, and register. Fluency is checked last not because it is unimportant but because it is the dimension you can assess without the source and the dimension the engine fails least, so spending your front-loaded attention there is a mis-allocation. A grammatical slip in a meaning-correct, term-correct, locale-correct segment is a fluency error, and usually a low-severity one. Run the sweep, mark what you find, move on.
Accuracy, then terminology, then locale, then fluency. The order is not arbitrary: it runs from the dimension fluency hides and that ships hazards down to the dimension the engine reliably gets right. The first question that fires names the dimension.
When a Segment Carries More Than One Error
Real segments are not tidy. A single segment can carry a terminology error in one clause and a fluency slip in another, and you mark both, as separate rows with separate dimensions and severities, because they have separate fixes. The rule is one error, one mark: do not collapse two defects into a single ambiguous note because they live in the same sentence. The harder case is when one defect could plausibly belong to two dimensions, a corrupted number that is both a locale-formatting failure and an accuracy meaning failure. Here the discipline is to mark it under the dimension whose fix and consequence dominate. If the formatting corruption changed the meaning a reader acts on, the accuracy framing governs the severity even if you note the locale mechanism. The goal is never taxonomic purity for its own sake; the goal is that the mark drives the right fix and the right severity. When in doubt, ask which categorization makes the fix and the consequence clearest, and mark that.
A Worked Batch: Categorizing Ten Segments by Hand
Theory settles only when you watch it operate on a real batch, so we will categorize ten segments the way a QE analyst (a quality-evaluation analyst, the person who runs the error-marking) would, running the decision procedure on each and producing a mark with a dimension, a category, and a severity. The sample is a mixed file: a stretch of a portable medical-device manual and a stretch of a banking app's UI strings, machine-translated into German and post-edited under deadline. For readability the source and target meanings are given in English, with the German behavior described. Watch how often the surface looks clean and the category emerges only from the procedure.
Segments One Through Five
Segment 1. Source: "Place the concentrator on a flat, stable surface." Target: same meaning, fluent, and the approved term "concentrator" is used. Run the procedure: meaning holds (accuracy gate passes), approved term present (terminology passes), no formatting (locale passes), grammar clean (fluency passes). No error. This is the common case, and it is exactly why MT-first pipelines are economically irresistible: most segments come back correct. The danger is never the clean majority; it is the one segment that reads just as clean and is not.
Segment 2. Source: "Do not operate the device near open flame." Target: fluent German that reads "Operate the device near open flame," the negation gone. Procedure: read source first, form the claim "do not operate near flame," check the target against it. The polarity is inverted. The accuracy gate fires. Category: omission of the negation producing a mistranslation. Dimension: accuracy. The consequence is a hazard to the user, so severity is Critical. Notice the prose was flawless; only the source comparison caught it.
Segment 3. Source: "The unit delivers up to 5 litres of oxygen per minute." Target: meaning intact and fluent, but the client's termbase specifies "concentrator" as the approved device term and the engine wrote "Einheit" ("unit"). Procedure: accuracy gate passes, the meaning is correct. Terminology check: the approved term was not used. Dimension: terminology. The fix is to enforce the termbase entry, and it routes to the terminologist, not a meaning reviser. On this regulated manual, where term consistency is a strict requirement, severity is Major; on looser marketing content the same error would be Minor.
Segment 4. Source: "Replace the filter every 3 to 4 months." Target: the figures were rendered "03/04," which in the day-first locale now reads as the date the third of April rather than the interval "3 to 4 months." Procedure: read source, claim is "an interval of three to four months." The target asserts a calendar date instead. This is genuinely a meaning corruption that arose through a formatting mechanism. The accuracy gate fires (the reader is misled about when to service the device), and you note the locale mechanism. Dimension: accuracy (with a locale note), category mistranslation. The consequence is confusing rather than acutely dangerous, so severity is Major.
Segment 5. Source: "Long-term use may occasionally cause mild dryness of the nasal passages." Target: fluent, meaning correct, terms correct, but renders "passages" as "Wege" ("pathways"), an idiomatic-but-slightly-off word choice, and drops a comma the German punctuation rules expect. Procedure: accuracy passes, terminology passes, locale passes, fluency sweep finds two slips. Two marks: one fluency word-choice error, Minor; one fluency punctuation error, Minor. Both real, both low-consequence, both counting toward the tally but neither threatening the file.
Segments Six Through Ten
Segment 6. Source (banking UI): "Your transfer will not be processed until the recipient confirms." Target: fluent German reading "Your transfer will be processed once the recipient confirms" reads close, but the original "will not be processed until" carries a hold-until-confirmation guarantee, and the rendering softened it into a near-opposite operational promise that the transfer proceeds. Procedure: source claim is "no processing happens before confirmation." The target weakens the prohibition into a permission. Accuracy gate fires. Category: mistranslation of the negated conditional. Dimension: accuracy. In a financial instruction that governs whether money moves, the consequence is a user acting on a false guarantee, so severity is Critical. This is Maria's "off" segment, now named.
Segment 7. Source: "Transfer fee: $2.50." Target: rendered "2,50 $" in a German locale that uses a comma decimal, which is correct, but the currency symbol position and the source's USD were carried over unchanged into a file localized for a Eurozone product where the amount should reflect the locale's currency handling per the client's spec. The meaning of the number is intact; the locale convention for currency presentation is violated. Procedure: accuracy passes (2.50 is 2.50), terminology passes, locale check fires on the currency convention. Dimension: locale. Severity is Minor here because the figure is unambiguous, though a client spec could raise it.
Segment 8. Source: "Tap Confirm to authorize the payment." Target: fluent, but the UI element label "Confirm," which the client's spec lists as a do-not-translate string matching the actual button text "Confirm," was translated to "Bestätigen" while the live button in the app still reads "Confirm." Procedure: the instruction now points the user at a label that does not exist on screen. This is an untranslated-category accuracy error in its over-translation mirror: a do-not-translate item that was translated, breaking the reference. Dimension: accuracy. Because the user cannot complete the action the instruction describes, severity is Major.
Segment 9. Source: "Contact support if the issue persists." Target: fluent German, meaning correct, approved term "support" handled per glossary, but rendered in the informal "du" register when the client's style guide for this banking product mandates the formal "Sie." Procedure: accuracy passes, terminology passes, locale-and-register check fires on the formality convention. Dimension: locale (register), bordering style. Severity is Major on a formal banking product where an inappropriately casual register damages brand and trust across every string, or Minor on a casual consumer app. We mark Major here.
Segment 10. Source: "Enter the amount you wish to send and tap Continue." Target: fluent and meaning-correct, but the German verb is in the wrong case for its object, an "Akkusativ where Dativ is required" grammar slip that a native reader notices and that mildly snags comprehension without breaking it. Procedure: accuracy passes, terminology passes, locale passes, fluency sweep fires on grammar. Dimension: fluency (grammar). Severity is Minor: noticeable, slightly jarring, not meaning-breaking.
Reading the Marked Batch as a Diagnosis
Now stand back and read the ten marks as a whole, because the categorization just produced something a vibe check never could: a diagnosis. Tally by dimension and you see two Critical accuracy errors (segments 2 and 6, both inverted polarity, both the silent fluent kind), two Major accuracy errors (segment 4's corrupted interval and segment 8's broken UI reference), one Major terminology (segment 3), two locale errors (segments 7 and 9), and three fluency Minors (segments 5 twice, and 10). Two things jump out of that distribution that "this file is off" could never tell you. First, the file fails: two Critical accuracy errors, either one of which fails it alone under the one-Critical-fails rule. Second, and this is the diagnostic payoff, the accuracy errors cluster on inverted polarity and negation handling, which is not a random scatter; it is a signature. It tells the operation that this engine, on this content, is dropping and softening negations systematically, which is a fixable process problem, a prompt or engine adjustment, not a per-segment whack-a-mole. The terminology error tells you the glossary the engine sees needs the device term locked. The locale errors tell you the locale profile needs the currency and register settings nailed down. None of that direction exists in "off." All of it falls out of correct categorization.
A marked batch is not just a verdict; it is a map. Errors that cluster in one category point at one fixable cause. "Looks off" scatters; "accuracy, omission, negation, three times" points straight at the engine's negation handling.
The Discipline of Marking by Category, Not by Feeling
Everything above is technique. This section is the professional discipline that makes the technique stick, the habits that separate an evaluator whose marks are trusted from one whose marks get argued with. The discipline has a few non-negotiable rules.
Every Mark Is a Structured Row, Never a Comment
A defensible mark is not a sentence in a comment column. It is a structured row with fixed fields: segment number, dimension, category, severity, the source text, the target text, and the recommended fix. "Off" fails on every field; it has no dimension, no category, no severity, no fix. The structured row is what makes the mark defensible to a client (here is the exact error in segment 6, here is why it is Critical), diagnostic for the process (the accuracy-negation cluster is visible only because each row carries a category), and reproducible by a second evaluator (who can find the same segment and check your dimension and severity against the same typology). When you find yourself about to write a vague note, stop and force it into the row. If you cannot name the dimension, you have not finished diagnosing the error, and the inability to fill the field is the signal that you are still at "off."
Let the Category Drive the Fix, Not Your Editing Reflex
A subtle failure mode for skilled linguists is to fix the segment first, from instinct, and categorize afterward as an afterthought, or not at all. This inverts the value. The category is the part that compounds: the fix repairs one segment, but the correctly named category, aggregated across the file, tells the operation what to change so the error stops recurring. A terminology error fixed silently is one good segment; a terminology error marked as terminology, three times, is a signal that the engine's glossary needs work and saves the next thousand segments. Mark first, or at least mark always, because the mark is the asset that outlives the individual fix. The post-editor's edit is consumed by the delivery; the evaluator's category feeds the process.
Resist the Gravity Toward Fluency
Linguists are trained to notice prose, so there is a constant gravitational pull toward marking fluency, the awkward phrasing, the inelegant word, the comma, because those are the errors a reader's ear surfaces first. Discipline means resisting that pull and spending your front-loaded attention on the accuracy gate, where the file-failing errors live and where the engine actually fails. A practical tell: if your marked batch is mostly fluency Minors with no accuracy checks recorded, you have not run the procedure; you have read the target for polish, which is the holistic vibe check wearing the costume of an evaluation. A real categorization pass shows accuracy and terminology marks, or shows explicitly that the high-consequence elements were checked and held. The fluency Minors are the easy harvest; the accuracy Criticals are the job.
Keep Severity Separate From Dimension
One discipline worth stating sharply because beginners conflate them: the dimension and the severity are two independent decisions, and you make them separately. The dimension answers "what kind of error," the severity answers "how much does it matter," and the same dimension spans the whole severity range. A terminology error can be Minor (a harmless synonym on a blog) or Major (an inconsistent device name in a regulatory file) or even Critical (a wrong drug name, which is terminology and lethal at once). An accuracy error is usually serious but not automatically Critical; a mistranslation that distorts a trivial, non-consequential detail can be Minor. Do not let the dimension imply the severity or vice versa. Name the dimension by what went wrong, then grade the severity by what happens when a real reader acts on it. Two questions, two answers, recorded in two fields. (This lesson lives one level above the severity-scoring detail, which the next lesson handles; here the rule is simply: categorize the dimension correctly, then assign severity as its own act.)
Why Correct Categorization Is Your Value in an MT-First Shop
It is worth closing on why this granular, almost clerical-seeming skill is the thing that moves a linguist up the value chain instead of out of the industry. In an MT-first pipeline, the engine already produced a fluent draft of every segment before you opened the file. Speed is no longer your scarce contribution; the engine has speed. What the engine structurally cannot do is sit above its own output and say, with marked evidence, "this segment is an accuracy omission, Critical, here is the fix, and the cluster of these tells you to fix the engine's negation handling." That sentence is judgment, and it is precisely what the revised ISO 18587, the post-editing standard expanded to cover AI and LLM non-human translation output and in DIS ballot with publication targeted for late 2025 into 2026, means when it insists the post-editor hold the same full linguistic competence as a professional translator. Categorizing an error correctly is not paperwork; it is the demonstration that you understand what went wrong well enough to fix it, route it, and stop it recurring.
The client-facing version of this is even starker. A vendor who returns "the file looks good" is selling a feeling, and a feeling loses every argument and survives no audit. A linguist who returns a marked report, dimensions named, categories assigned, severities graded, the accuracy cluster diagnosed, is selling proof, and proof is the thing a raw MT vendor cannot produce. The whole program turns on the difference between fluent and correct, and correct categorization is where that difference becomes operational: it is the moment "this is wrong" becomes "this is an accuracy mistranslation in segment 6, Critical, and here is why your money would have moved against the user's instruction." That second sentence is the credential. Maria's reviewer wrote "off." The job is to never write "off" again, and to always write the row.
Key Takeaways
- The error category is the diagnosis, not a label applied after it. The dimension you assign drives three downstream decisions a vague "looks off" cannot touch: how the error is fixed, who fixes it, and how severe it is. One symptom (a wrong word) can map to four different dimensions with four different fixes.
- The four MQM/ISO 5060 dimensions are accuracy (target meaning against source meaning, with categories mistranslation, omission, addition, and untranslated), terminology (approved-term and consistency compliance), locale conventions (dates, numbers, units, currency, register), and fluency/style (the target judged on its own for grammar, spelling, punctuation, register).
- Categorize with a fixed five-step procedure run in order on every suspect segment: read source first then target, run the accuracy gate (negations, numbers, entities, obligations, dates, warnings), then check terminology, then check locale, then sweep fluency. The first question that fires names the dimension, and the order encodes the consequence hierarchy.
- Accuracy is judged on the target against the source; fluency is judged on the target alone. Because the engine is fluent first and accurate second, an evaluation that drifts toward fluency Minors is scoring the dimension the machine reliably gets right and missing the accuracy and terminology errors that actually fail files.
- When a segment carries more than one error, mark each as a separate row with its own dimension and severity. When one defect could belong to two dimensions, mark it under the dimension whose fix and consequence dominate, so the mark drives the right action rather than chasing taxonomic purity.
- In the worked ten-segment batch, the marks revealed two Critical accuracy errors (both inverted polarity), a Major terminology error, locale and fluency errors, and, crucially, a cluster: the accuracy errors concentrated on negation handling, a fixable engine-level signature that "this file is off" could never surface.
- Dimension and severity are two independent decisions made separately: the dimension answers what kind of error, the severity answers how much it matters, and the same dimension spans the whole severity range (a terminology error can be Minor or, with a wrong drug name, Critical).
- Every mark must be a structured row (segment, dimension, category, severity, source, target, fix), never a comment, because the row is what makes the verdict defensible, diagnostic, and reproducible, and because correctly categorizing an error is the judgment an MT engine cannot perform on its own output, the value the revised ISO 18587 ties to full professional-translator competence.
Skill.re