โ†
AI for Translation & Localization
Capable ยท M18 ยท lesson 18 of 21 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Scoring Severity: Critical, Major, Minor
๐Ÿ“–
now learning

Scoring Severity: Critical, Major, Minor

15 min

The Slack message landed at 4:50 on a Friday, ten minutes before the post-editor wanted to log off. A pharmaceutical client had escalated a delivered file, a German patient-information leaflet for an injectable, and the project manager forwarded it with one line: "Client says there is a Critical. Evaluator says it is a Major. You scored it Minor. Who is right, and can you defend it by Monday?" The disputed sentence was eight words long. The source told a patient to inject the medication into the abdomen or the upper thigh. The machine-translated, post-edited German said to inject into the abdomen and the upper thigh. One word changed, "or" became "and", and three professionals had read the same eight words and assigned three different severities. That is the problem this lesson solves. Not whether you can spot an error, you can. Whether you can look at an error you have already found and assign it a severity that survives a client, an evaluator, and an auditor reading over your shoulder. Severity is the decision the whole quality gate hangs on, and it is the one most linguists make on instinct and then cannot defend. By the end of this lesson you will have a decision procedure that turns "it feels like a Major" into "it is a Major, and here is the rule that makes it one."

Why Severity Is the Whole Game

In analytic translation evaluation you do two things to every error you find. First you classify its dimension, the kind of error it is: accuracy, terminology, locale, or fluency. Second you assign its severity, how much it matters. The dimension question is usually easy. A wrong number is an accuracy error, an off-glossary term is a terminology error, a day-first date in a month-first market is a locale error, a clumsy clause is a fluency error. You can train someone to sort errors into dimensions in an afternoon. Severity is the hard one, and it is the one that decides whether the file ships, because the pass/fail gate is built on severity, not on dimension.

Before going further, the vocabulary, used precisely throughout. Machine translation (MT) is any system that renders text from a source language into a target language with no human writing the words. A large language model (LLM) is a general-purpose text predictor that translates as a side effect of its broad competence and is fluent even when wrong. Machine-translation post-editing (MTPE), often shortened to PE, is the workflow where a human edits machine output instead of translating from scratch. A segment is the unit a translation tool works in, usually a sentence, the row you see in a CAT tool (a computer-assisted translation tool, the editing environment a linguist works in). A termbase is the controlled glossary of a client's approved terms. MQM is Multidimensional Quality Metrics, the analytic error-typology framework that classifies translation errors by dimension and severity. ISO 5060:2024 is the international standard that formalizes an MQM-aligned model for the human analytic evaluation of translation output, including the Critical, Major, and Minor severity bands. Severity is how much a given error matters, and it is the entire subject of this lesson. We define each severity band in working terms below, before any of them are used to decide a file's fate.

Here is why severity carries the weight. In a severity-scored model, each error is assigned a penalty by its band, the penalties are summed, the total is normalized against the length of the file, and a pass/fail threshold is applied, with one absolute rule sitting on top: a single Critical error fails the file regardless of how clean everything else reads. That last rule is the reason severity is not just one input among several. The difference between calling an error Major and calling it Critical is not the difference between five penalty points and twenty-five. It is the difference between a file that passes with rework and a file that is dead on arrival. When the stakes are that binary, "I felt it was a Major" is not an answer. You need a reason that holds up.

Dimension tells you what kind of error you found. Severity tells you whether the file ships. The gate is built on severity, which is exactly why severity is the decision you must be able to defend.

The Three Bands, in Working Terms

The three severity bands have formal definitions in ISO 5060 and MQM, and they are worth stating cleanly before we sharpen them into a decision procedure. A Minor error is a flaw that does not meaningfully affect the meaning or usability of the content: an awkward but understandable phrasing, a stylistic choice you would have made differently, a missing comma, a slightly-off synonym a reader would shrug at. A Major error meaningfully impairs the comprehension or usability of the content: a mistranslation that distorts a non-critical part of the meaning, a wrong term that undermines trust, a locale slip that makes a date genuinely ambiguous, a grammatical breakdown that forces a re-read. A Critical error renders the content dangerous, unusable, legally exposed, or actively misleading on a point that matters: an inverted contraindication, a flipped safety instruction, a corrupted dosage, a wrong drug name, an indemnity clause that swaps which party is liable.

These definitions are correct, and they are also not yet operational. Read them again and notice the words doing the real work: "meaningfully," "genuinely ambiguous," "a point that matters." Those are judgment words, and three honest professionals can disagree about whether a given error clears them. The leaflet dispute that opened this lesson happened inside exactly that gap. Everyone agreed it was an accuracy error. They disagreed about whether "and" instead of "or" was harmless, impairing, or dangerous. The definitions alone did not settle it. What settles it is a decision procedure that converts the judgment words into questions with answers, and that procedure is the heart of this lesson.

The Three-Question Severity Test

Strip the formal language away and every severity decision reduces to a single question asked of the error: what happens when a real reader acts on this output? Severity is not a property of the text. It is a prediction about consequence. The text is just where the consequence is encoded. To turn that into something you can apply consistently and defend, run every error through three questions, in order. The first one it triggers sets the floor for its severity.

Question one: does the error mislead? Does the target lead the reader to a wrong understanding of a fact, an instruction, or an obligation? An error that merely reads awkwardly does not mislead; the reader still arrives at the right meaning, just less smoothly. An error that changes what the reader believes to be true does mislead. The dropped negation misleads, because the reader now believes the opposite of the instruction. The off-glossary synonym that still means the same thing does not mislead. Misleading is the threshold that separates a genuine content error from a cosmetic one. If the answer is no, you are almost certainly looking at a Minor.

Question two: does the error harm? If the reader acts on the misleading output, is anyone hurt, exposed to legal liability, or caused financial loss? This is the consequence question, and it is the one that separates Major from Critical. An error can mislead without harming: a mistranslated marketing tagline misleads the reader about a product benefit, which is bad, but no one is injured or sued. An error harms when the misleading content sits on a point where acting on it produces real-world damage: a flipped dosage, an inverted contraindication, a swapped party in a liability clause, a corrupted financial figure on a binding document. Harm is the threshold for Critical.

Question three: does it merely annoy? If the error neither misleads nor harms, does it still degrade the reader's experience, the brand's polish, or the content's professionalism? A stiff sentence, a missing comma, an inconsistent but understandable term, a slightly unnatural register. These are real errors. They count. They are Minor. The "merely annoy" question exists to make sure you record the error rather than dismissing it, while correctly keeping it in the lowest band.

Does it mislead, does it harm, or does it merely annoy? Annoy is Minor. Mislead without harm is Major. Mislead with harm is Critical. Three questions, asked in order, and the first one that fires sets the floor.

Mapping the Questions to the Bands

The three questions map onto the three bands cleanly, which is the point of asking them in order.

  • Merely annoys, does not mislead: the floor is Minor. The reader understands the right thing; the delivery is just less polished than it should be.
  • Misleads, but does not harm: the floor is Major. The reader is led to a wrong understanding, but acting on it does not injure, expose, or cost. The content is impaired, not dangerous.
  • Misleads and harms: the floor is Critical. The reader is led to a wrong understanding on a point where acting on it produces harm, liability, or financial loss. The content is dangerous, and the gate fires.

Two disciplines make this procedure defensible rather than just tidy. First, the questions set a floor, not a ceiling. Context can raise a severity above the floor the three questions establish, but it should rarely lower it below. If an error misleads, it is at least a Major, and no amount of "but the prose is lovely" pulls it down to Minor. Second, you answer the harm question against the worst plausible reader behavior, not the most charitable one. The question is not "would a careful, expert reader catch this and act correctly anyway?" It is "if a reasonable reader acts on exactly what the target says, what happens?" The whole danger of the fluent error is that it does not trip a careful reader's suspicion, so assuming a suspicious reader defeats the purpose of the evaluation.

Severity Tracks Consequence, Not Edit Size

The single most common and most expensive mistake in severity scoring is to map severity onto the size of the textual change. The instinct is intuitive and completely wrong: a one-word error feels like it should be Minor, and a rewritten paragraph feels like it should be Major. Unlearn this immediately, because the machine-translation era is built to punish exactly this instinct. The most dangerous error in the entire discipline, the dropped negation, is usually a change of one tiny word, and it is Critical. A sprawling, clumsy, over-translated paragraph that is awkward to read but conveys precisely the right meaning on a low-stakes page might be only Minor.

Severity has nothing to do with how many characters changed and everything to do with what happens when a reader acts on the output. The three-question test enforces this automatically, because none of its three questions ask about edit size; they all ask about consequence. Hold onto the contrast: a one-word flip in a drug label is Critical, a stiff three-sentence rewrite on a blog post is Minor. The character count points in the opposite direction from the severity in both cases, which is exactly why edit size is a trap.

This matters acutely with MT and LLM output specifically, because the errors that engines produce are often microscopic on the surface and enormous in consequence. A model that drops a "nicht," that turns "do not exceed" into "exceed," that renders "contraindicated" as "indicated," changes almost nothing in length and inverts everything in meaning. An evaluator anchored to edit size will systematically under-score the exact errors the machine is most prone to make. The discipline is to look past the size of the change to the size of the consequence, every single time.

Ask the consequence question, never the character-count question. A one-word negation flip in a safety warning is Critical. A long, ugly, correct-in-meaning paragraph on a marketing page is Minor. Edit size and severity are not the same axis.

Why the Dimension Does Not Fix the Severity

A related trap is assuming a given dimension always carries a given severity. It does not. Every dimension spans the whole severity range, and which band a particular error lands in depends entirely on consequence, not on its dimension label.

Take terminology. An off-glossary synonym for a generic UI label on a marketing microsite is a Minor terminology error: it annoys, it does not mislead. The same dimension, a terminology error, becomes Critical when the wrong term is a drug name: "renders the client's approved drug name as a different, real drug name" misleads and harms, and it is Critical, even though both errors are "terminology." Take locale. A US-style thousands separator on a casual price in a blog post is a Minor locale error. The identical dimension becomes Critical when the separator confusion shifts a dosage by a factor of a thousand on a medical label. Accuracy spans the range too: a slightly imprecise rendering of a non-essential adjective is Minor, while a flipped negation in an instruction is Critical, and both are accuracy errors. The dimension tells you what kind of error you are scoring. It tells you nothing about how much it matters. Only the consequence does that.

Worked Severity Calls on Ambiguous Cases

The clean cases score themselves. A dropped negation in a safety warning is obviously Critical; a missing comma in a blog post is obviously Minor. Nobody disputes those. The disputes, the escalations, the Friday-afternoon Slack messages, all live in the ambiguous middle, where reasonable professionals can land in different bands. Working through ambiguous calls deliberately is the only way to build a defensible instinct. Here are several, each run through the three-question test, with the reasoning shown the way you would show it to a client who challenged the score.

Case 1: The "And/Or" Flip in the Patient Leaflet

Return to the case that opened the lesson. Source: inject into the abdomen or the upper thigh. Target: inject into the abdomen and the upper thigh. Run the test. Does it mislead? Yes, unambiguously: the source offers two acceptable single sites, the target instructs the patient to use both. The reader now believes something false about the procedure. So it is at least a Major. Does it harm? Here is where it gets genuinely hard, and where the three readers diverged. The honest answer requires domain knowledge: does injecting into two sites instead of one cause harm? For most subcutaneous injections, splitting a single dose across two sites is a dosing and absorption error, potentially affecting how much drug is delivered and how fast, which on an injectable medication is a real clinical consequence, not a cosmetic one. On that reading the error harms, and it is Critical. The post-editor who scored it Minor was anchored to edit size, one word, and missed that the word changed the procedure. The evaluator who scored it Major correctly saw the misleading, but stopped short of the harm question. The client who scored it Critical asked the harm question and answered it with the domain knowledge the post-editor lacked. The defensible call is Critical, and the lesson is that the harm question often cannot be answered from the text alone; it requires knowing what the reader will do and what it costs. When you cannot answer the harm question yourself, you escalate to someone who can, rather than guessing low.

Case 2: The Approved-Term Drift

Source describes "the concentrator." The termbase mandates "concentrator" as the approved term for the device. The engine rendered it "the unit" in nine of forty segments. Does it mislead? Marginally: a reader broadly understands "the unit" refers to the same device, so the meaning is largely intact, but consistency is broken and in a regulated technical manual the approved term is itself a requirement. Does it harm? Not directly; no one is injured because the device was called "the unit." So the floor from the three questions is between Minor and Major. This is where the second discipline, context raises the floor, comes in. On a casual marketing page, term drift is Minor: it annoys. On a regulated technical manual where the client's termbase compliance is a contractual and sometimes regulatory requirement, the same drift is Major, because inconsistency in a technical document genuinely impairs usability and breaks an explicit requirement. The defensible call depends on the content tier, and the right move is to read the client's scoring profile: many regulated clients pre-declare that termbase violations are scored Major or higher by rule, which removes the judgment entirely. When the profile speaks, you follow it; when it is silent, you score by consequence in context.

Case 3: The Ambiguous Date

Source: "valid until 03/04/2026." Target carried the figures unchanged into a market that reads dates day-first, so "03/04/2026" now reads as the third of April rather than the source's intended fourth of March. Does it mislead? Yes: the reader will believe a different expiry date than the source states. At least a Major. Does it harm? This depends on what the date governs. On a coupon, a one-month error in an expiry is an annoyance and a minor commercial irritation: Major at most, arguably Minor on a throwaway promo. On a pharmaceutical product's expiry date, a reader who trusts a wrong expiry might use a product past its real shelf life or discard a still-valid one, and on a contract a wrong "valid until" date can void or extend an obligation with financial consequences: Critical. Same error, same dimension, same number of characters wrong. The severity swings from Minor to Critical entirely on what acting on the wrong date costs. This is the consequence axis doing all the work, and it is why "it's just a date format" is never a complete severity argument.

Case 4: The Confident Hallucinated Number

Source: "Tighten the bolt to the specified torque." Target: "Tighten the bolt to 45 Nm." The engine invented a specific, plausible, completely unsourced number where the source gave none. Does it mislead? Severely: the reader now believes a precise specification the source never provided, and a fabricated specification reads as more authoritative than a vague one, so the reader is more likely to act on it. At least Major. Does it harm? On furniture assembly instructions where over-torquing strips a screw, that is annoying and Major. On an aircraft maintenance manual or a medical-device assembly step where an incorrect torque value causes mechanical failure, it is squarely Critical. Note the pattern that keeps recurring: the hallucinated number is dangerous precisely because the LLM delivered it fluently and confidently, with no hedging, in prose indistinguishable from the correctly translated segments around it. The fluency is what makes the reader trust the fabrication. That is the silent critical error in its purest form, and severity scoring exists to catch it before delivery.

Case 5: The Formality Slip

Source is a formal legal notice in a language with a formal/informal distinction (the German Sie/du, the French vous/tu). The engine rendered a clause in the informal register where the formal is required. Does it mislead? No: the meaning is fully intact, the reader understands exactly what is meant. Does it harm? No: nobody is injured or exposed because a sentence addressed them informally. Does it merely annoy? Yes, but with a wrinkle: in a formal legal or institutional context, an informal register is more than an aesthetic slip; it can read as unprofessional or even disrespectful and damage the client's standing. So this is a Minor on most content and a defensible Major on high-formality institutional content where register is part of the message, never Critical, because it neither misleads nor harms. This case shows the floor working in the other direction: register errors cannot climb to Critical no matter how formal the context, because they fail the mislead and harm tests, but context can lift them from Minor to Major. The three questions cap the severity as firmly as they floor it.

Severity and Real-World Consequence: Safety, Legal, Financial

The harm question becomes much easier to answer consistently once you recognize that real-world harm in localized content falls into three recurring buckets. Almost every Critical error you will ever score lives in one of them, and naming them in advance turns the harm question from an open-ended worry into a checklist.

Safety Consequence

Safety is the most visceral bucket and usually the easiest to recognize. The error, if acted on, can cause physical harm to a person. Inverted safety instructions, flipped contraindications, wrong dosages, corrupted device-operation steps, mistranslated allergen or hazard warnings, wrong torque on a load-bearing fastener. The reader's body is at stake. Safety errors that mislead are nearly always Critical, because the harm threshold is met by definition: a misled reader acting on the content risks injury. This is the bucket that dominates medical, pharmaceutical, life-sciences, industrial, automotive, and aviation content, and it is the reason those content types are routed to full human translation or full post-editing rather than light PE.

Legal harm is less visceral but just as real: the error, if acted on or relied upon, creates legal liability, voids or alters an obligation, or misstates a right. An indemnity clause that swaps which party holds the liability. A "shall" rendered as "may," turning a binding obligation into an option. A dropped exclusion that silently expands a warranty. A mistranslated jurisdiction or governing-law clause. A consent statement that misrepresents what the reader is agreeing to. These errors do not bleed, but they can cost a company a lawsuit, a regulatory penalty, or an unenforceable contract. Legal errors that materially change an obligation or a right mislead and harm, and they are Critical. This bucket dominates contracts, terms of service, privacy notices, regulatory submissions, and consent forms.

Financial Consequence

Financial harm is the third recurring bucket: the error, if acted on, causes monetary loss or misstates a financial fact on which money moves. A corrupted price, a misplaced decimal in an interest rate, a wrong figure in a financial statement, a currency left unconverted, a discount percentage inverted, a tax rate corrupted. A misplaced decimal that turns 1,250 into 12.50 or 125,000 misstates value by orders of magnitude, and on a binding quote, an invoice, an annual report, or a financial product disclosure, money moves on the wrong number. Financial errors that misstate a material figure mislead and harm, and they are Critical. This bucket dominates financial services, e-commerce checkout flows, invoicing, and corporate reporting.

Safety, legal, financial. Almost every Critical error lives in one of these three buckets. If the error misleads and the misleading lands in any of the three, you are looking at a Critical, and the gate fires.

The Tier of the Content Sets the Stakes

The same linguistic error carries different consequences depending on the content it sits in, which is why severity scoring cannot be done blind to the content's risk tier. The error "valid until 03/04/2026 reads as the wrong date" is identical on a coffee-shop loyalty coupon and on a medication's expiry label, but the consequence, and therefore the severity, is not. This is the bridge between severity scoring and risk-tiered intake: high-liability content (medical, legal, financial, life-safety) is where Critical errors are both more likely to occur and more consequential when they do, which is exactly why that content gets the most rigorous evaluation and the strictest gate. Before you score severity, you must know what tier the content is in, because the tier tells you how to answer the harm question. Scoring a drug label as if it were a blog post is itself a severe process failure, even before a single error is marked.

How Severity Drives the Go/No-Go Gate

Everything in this lesson has been building toward one operational consequence: the severity you assign is what the delivery gate reads. The gate is the rule that converts a pile of marked errors into a single decision, ship or do not ship, and it is built almost entirely on severity bands.

The gate has two layers. The first layer is the absolute Critical rule: if the file contains even one Critical error, the file fails, full stop, regardless of how clean every other segment is, regardless of the computed score. This is not a high threshold the Critical pushes the file past; it is a separate, overriding rule. One inverted contraindication in a flawless ten-thousand-segment file fails the delivery, because the patient who acts on the one bad sentence is not protected by the ten thousand good ones. Harm does not average out. A model that let a clean file absorb a Critical would be shipping rare hazards on the theory that they are statistically uncommon, which is a lottery with someone's safety as the stake, not a quality gate.

The second layer is the weighted threshold for files with no Criticals: the Majors and Minors are penalized (classically 5 points per Major, 1 per Minor), summed, normalized against length, and compared to a passing threshold. A file with too many Majors fails even with zero Criticals, because the accumulated impairment crosses the line the client set. This is where Major-versus-Minor calls matter financially: each Major you assign carries roughly five times the penalty of a Minor, so the band you choose on the ambiguous middle cases directly moves the score and can be the difference between a pass and a fail near the threshold.

One Critical fails the file outright, no arithmetic required. With no Criticals, the weighted Majors and Minors decide it against a threshold. Every severity call you make is a vote in that decision, and the Critical call is a veto.

Why the Gate Changes Where You Look

Because a single Critical is disqualifying and a single Minor is nearly free, the gate does not just decide the file; it tells you where to spend your attention as an evaluator or post-editor. You do not read for an even polish across the file. You hunt, with disproportionate intensity, for the high-consequence elements where a Critical hides: negations, dosages and numbers, drug and proper names, safety instructions, dates in regimens and contracts, monetary figures, and legal obligations. Miss ten Minors and the file still passes. Miss one Critical and you shipped the hazard, the lawsuit, or the loss. The math of the gate is also a map of where to look, and the three-question severity test is the instrument you apply to each of those high-consequence elements in turn.

What Makes a Severity Call Defensible

A severity call is defensible when you can state, for the specific error, the three things a client or auditor will ask for. First, the dimension and the evidence: "this is an accuracy error; here is the source segment, here is the target, here is the divergence." Second, the consequence reasoning: "it misleads the reader to believe X; acting on X causes Y harm/liability/loss," or "it does not mislead, it only degrades polish." Third, the band and the rule that put it there: "misleads and harms, therefore Critical," or "client scoring profile mandates Major for termbase violations," or "annoys only, therefore Minor." A call backed by those three things survives the Friday-afternoon escalation. A call backed by "it felt like a Major" does not, and worse, it makes every other call you have ever made suspect. Defensibility is not bureaucracy; it is the entire difference between a quality professional and a grader with opinions.

One last discipline ties it together: when you genuinely cannot answer the harm question, because it requires domain knowledge you do not have, you do not guess low to keep the file passing. You flag the error, state the uncertainty, and escalate to someone who can answer it. The "and/or" injection flip that opened this lesson was scored Minor by a post-editor who could not assess the clinical consequence and resolved the uncertainty in the file's favor. The defensible move was the opposite: when the harm question is open and the content is high-tier, you escalate rather than absorb the risk. Guessing low on an unanswerable harm question is the exact mechanism by which silent Criticals ship.

Building the Severity Habit on Real Files

Knowing the procedure is not the same as running it under deadline pressure on a file the engine has already made look clean. The habit that turns the three-question test from a diagram into a reflex is built deliberately, and it is worth stating how, because this is a goldmine lesson: the skill is meant to be practiced, not just understood.

Start by separating the two decisions explicitly on every error, even when it feels mechanical. Write the dimension first, then run the three questions out loud or on paper, then write the band and the rule. The separation matters because the failure mode is collapsing the two decisions into one gut feeling, and the gut feeling is where edit-size bias and dimension bias sneak in. For the first few hundred errors you score, force the explicit steps; the reflex forms from the repetition.

Next, calibrate against a second evaluator on the ambiguous middle. The clean Criticals and clean Minors need no calibration; everyone agrees. Take ten genuinely ambiguous errors, score them independently, then compare and argue the gaps until you converge on the rule that resolves each one. The argument is the training. Two evaluators who can converge on the ambiguous cases have built a shared, defensible standard, which is exactly what ISO 5060 calls for and what a client audit tests. Disagreement on the clean cases means someone misunderstands the bands; disagreement on the ambiguous cases is normal and is resolved by the three-question test, not by seniority.

Finally, always read the client's scoring profile before you score. Many clients pre-declare severity rules that remove judgment on whole categories: termbase violations are Major, any safety-content accuracy error is Critical, locale errors on dates are Major minimum. Where the profile speaks, it overrides your judgment, and following it is itself the defensible move. Where it is silent, you apply the three-question test against the content's risk tier. The combination, declared rules where they exist plus a consistent consequence test where they do not, is what produces severity calls that hold up across evaluators, across files, and across the Friday-afternoon escalation that started this lesson.

Key Takeaways

  • Severity, not dimension, decides whether a file ships, because the pass/fail gate is built on severity bands. Classifying an error's dimension (accuracy, terminology, locale, fluency) is the easy half; assigning its severity defensibly is the half that carries the consequence and the disputes.
  • Run every error through three questions in order: does it mislead, does it harm, or does it merely annoy? Merely annoys is the floor for Minor; misleads without harm is the floor for Major; misleads and harms is the floor for Critical. The first question that fires sets the floor, and context can raise but rarely lowers it.
  • Severity tracks consequence, not edit size. A one-word dropped negation in a safety warning is Critical; a long, clumsy, correct-in-meaning paragraph on a marketing page is Minor. Edit size points the opposite way from severity on exactly the errors MT and LLMs produce most, which is why anchoring to edit size systematically under-scores the machine's worst output.
  • Every dimension spans the full severity range. A terminology error is Minor for a UI label and Critical for a wrong drug name; a locale error is Minor on a coupon and Critical when it shifts a dosage by a factor of a thousand. The dimension tells you what kind of error it is and nothing about how much it matters.
  • The harm question is answered against the worst plausible reader behavior and against the content's risk tier, not against a charitable expert reader. When you cannot answer the harm question because it needs domain knowledge you lack, you flag and escalate rather than guessing low, because guessing low on an unanswerable harm question is exactly how silent Criticals ship.
  • Almost every Critical error lives in one of three real-world consequence buckets: safety (physical harm), legal (liability, voided or altered obligations), or financial (monetary loss or a misstated material figure). Naming the buckets turns the harm question into a checklist.
  • The go/no-go gate has two layers: one Critical fails the file outright regardless of the computed score, because harm does not average across a clean file, and below that, weighted Majors and Minors are summed, normalized against length, and compared to a threshold. Every Major you assign carries roughly five times a Minor's penalty, so ambiguous-middle calls move the score directly.
  • A severity call is defensible when you can state the dimension and its evidence, the consequence reasoning, and the band with the rule that put it there. Read the client's scoring profile first, follow its declared rules where they exist, and apply the three-question test against the content's risk tier where they are silent. "It felt like a Major" survives no audit; "it misleads and harms, therefore Critical" survives every one.