โ†
AI for Translation & Localization
Aware ยท M2 ยท lesson 2 of 19 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
AI in Quality Estimation and Evaluation
๐Ÿ“–
now learning

AI in Quality Estimation and Evaluation

15 min

The dashboard glowed a calm, persuasive green. Forty-one thousand words of a German financial-services prospectus, target Spanish, sitting in the translation-management system (the TMS, the platform that routes files, applies translation memory, and runs the machine-translation pass before a human opens anything). Beside each segment was a number: the quality-estimation score, or QE score, an automatic confidence rating the system had attached to every machine-translated line. Most segments scored above 90. The project manager had already done the arithmetic in his head: a file this clean, with a QE average of 0.91, barely needs a human. Light post-editing, fast turnaround, half the margin he expected. Then the evaluator, a Spanish linguist with fifteen years in regulated finance, scrolled to segment 1,184. The score was 0.94, comfortably green, one of the highest in the file. The German said the guarantee did not cover losses arising from currency fluctuation. The Spanish, fluent and clean and scored as one of the best segments in the entire document, said the guarantee did cover losses arising from currency fluctuation. The negation was gone. The QE model had no idea. It had rated a sentence that inverted a financial liability as one of the most trustworthy lines in the project. This lesson is about what that green number actually measures, why it is a signal that tells you where to point human effort rather than a verdict that clears a segment to ship, and why the entire family of automatic translation metrics, from QE scores to BLEU, must never be mistaken for quality itself.

What a QE Score Actually Is

Quality estimation (QE) is the practice of having a machine predict how good a translation is without ever seeing a correct human reference to compare it against. That last clause is the whole trick, and it is worth slowing down on. A QE model looks at a source segment and a machine-produced target segment, and it outputs a number, often between 0 and 1, or scaled to 0 to 100, that represents its prediction of how good that target is. It does this with no answer key. There is no human translation of that segment sitting somewhere that the model checks against. The model has simply learned, from large quantities of past data, what good and bad translation pairs tend to look like, and it generalizes from that learning to the new pair in front of it.

Read the definition again with the emphasis where it belongs: QE predicts quality, it does not measure it. A confidence score is a probability, the model's estimate of the likelihood that this segment is acceptable. It is the machine's best guess about its own work, produced by the same statistical machinery, with the same blind spots, that produced the translation in the first place. When the QE model rated the inverted currency clause at 0.94, it was not lying and it was not broken. It was doing exactly what it was built to do: predicting that a fluent, grammatical, well-formed Spanish sentence that aligns reasonably with the German source is probably fine. The sentence ticked every surface box the model knows how to check. The model cannot read a contract. It does not know that one missing "no" reverses a financial obligation. It saw smooth prose and confident structure and returned a high probability, because in its training data smooth, confident prose is usually fine.

A QE score is a probability that a segment is good, produced by a machine that cannot read for meaning. It is a prediction about the work, not a verdict on it.

Reference-Based Versus Reference-Free Evaluation

To understand QE you have to see what it replaced and why it exists at all. Classic automatic translation metrics are reference-based: they need a human-produced "correct" translation, called a reference, and they score the machine output by how closely it matches that reference. The most famous of these is BLEU (Bilingual Evaluation Understudy), a metric from the early 2000s that counts how many short word sequences, called n-grams, the machine output shares with the reference. The more overlap, the higher the BLEU score. It is fast, cheap, and entirely automatic, and for two decades it was how the field reported whether one engine "beat" another.

The problem with reference-based metrics in your daily work is brutal and simple: in production, you do not have a reference. If you already had a correct human translation of the segment, you would not need the machine to translate it. References exist in research benchmarks and in controlled engine evaluations, not in the live file on your screen at 9 a.m. on a Tuesday. That is the gap QE was invented to fill. QE is reference-free. It scores the segment with no answer key, which is exactly why it can run live, on every segment, in the TMS, the moment the machine produces output. Its power is that it works where BLEU cannot. Its danger is that it buys that power by predicting quality instead of measuring it, and a prediction can be confidently, fluently wrong in precisely the same way the translation it is scoring can be.

Why a High Score Routes Effort, Not Ships Output

Here is the single most important reframing in this lesson, and it is the difference between a linguist who uses QE well and one who gets burned by it. A QE score is a triage signal, not a quality gate. Its legitimate job is to help you decide where to point your scarce human attention. Its illegitimate use, the one that ships the inverted currency clause, is to decide what does not need human attention at all.

Think about what scarce attention means in an MT-first shop. A post-editor (the linguist who edits machine output rather than translating from scratch, doing the work called MTPE, machine-translation post-editing) has a fixed number of hours and a file with thousands of segments. They cannot scrutinize every segment with equal depth, because the per-word economics of MTPE, typically 50 to 75% of full human translation rates, do not budget for it. So the question is not "is this file perfect," it is "where, in my limited time, is my attention most valuable." That is a routing question, and QE answers it beautifully when used as intended.

The Right Way to Read the Distribution

Used correctly, QE sorts the file so you spend your effort where the machine is least sure. The low-scoring segments are the ones the model itself flagged as probably weak: confusing source, unusual structure, low confidence. Those deserve a close look, and QE pointing you to them is a genuine productivity gain. You read the suspicious segments first, fix the real problems, and you have spent your hours where they bought the most quality.

But notice what that workflow does and does not claim. It claims that low scores are worth your attention. It does not claim that high scores are safe to ship unread. Those are completely different statements, and conflating them is the error that ended the prospectus story. A high QE score means the model could not see anything wrong from the outside. It does not mean nothing is wrong. The model's blindness to a dropped negation is not reduced by the negation being dropped in a fluent sentence; it is increased by it, because fluency is exactly what the model reads as a sign of health.

A low QE score earns a segment your attention. A high QE score does not earn a segment a pass. Routing is the job; clearance is the trap.

Why High Scores Are Where the Killer Error Hides

There is a cruel logic here that every QE user must internalize. The silent critical error, the fluent sentence that means the opposite of the source, lives disproportionately in the high-scoring segments, not the low ones. Why? Because a dropped negation produces a perfectly grammatical sentence. A flipped dosage produces a clean number in a clean sentence. An inverted obligation reads like any other well-formed clause. These errors do not degrade fluency, and fluency is most of what the QE model can perceive. The very property that makes these errors dangerous to a human reader, their smoothness, also makes them invisible to the QE model, and so they earn high scores. The QE distribution does not sort errors by severity. It sorts segments by surface plausibility, and the most dangerous errors are the ones that are most plausible on the surface.

This is why "trust the green, scrutinize the red" is a recipe for shipping the exact errors that cost a life or a lawsuit. The red segments are clumsy, self-announcing, and usually low-stakes garbled source. The green segments are smooth, confident, and contain the one inverted clause that fails the file. A QE score tells you where the machine was unsure. It is silent about where the machine was confidently wrong, and confidently wrong is the failure mode that matters.

What a Human MQM Evaluation Does Instead

If a QE score is a prediction, what is the thing it is predicting against? What does real quality measurement look like? The answer is an analytic human evaluation built on an error typology, and the dominant framework is MQM (Multidimensional Quality Metrics), now formalized for translation output by the international standard ISO 5060:2024 (Translation services, Evaluation of translation output). Understanding the gap between QE and MQM is understanding the gap between a guess and a judgment.

An MQM-style evaluation does not produce a single confidence number from the outside. A qualified human evaluator reads the target against the source, segment by segment, and marks each error they find. Crucially, every error gets two labels: a category (what kind of error it is) and a severity (how much it matters). The categories sort errors into dimensions the field has agreed are the ones that count. The severity decides the weight. This is called analytic evaluation, because it takes the file apart into specific, located, classified errors rather than rating it holistically with a vibe.

The Four Error Dimensions

The MQM and ISO 5060 family organizes errors into a small set of dimensions. For the working linguist, four carry most of the weight:

  • Accuracy. The relationship between the target and the meaning of the source. Mistranslation, dropped negation, added information the source did not contain, omission of a clause, a corrupted number. The pacemaker and the currency clause are accuracy errors. This dimension is where the silent critical error lives, and it is precisely the dimension a QE model is worst at, because catching it requires reading the source for meaning.
  • Terminology. Whether the approved term was used. The client recorded a specific term in the termbase (the controlled glossary of approved terms), and the engine used a common synonym instead. Fluent, natural, and wrong by the client's own rule. A QE model trained on general text often rates the wrong-but-common term higher than the right-but-unusual one.
  • Locale conventions. Whether dates, units, currency, number formatting, and formality match the target locale (the specific language-and-region combination, like Spanish for Mexico versus Spanish for Spain). A date rendered in the wrong order, a decimal comma swapped for a point, the wrong formality register. Quietly wrong, expensively wrong, invisible to fluency.
  • Fluency. Whether the target reads correctly as language: grammar, spelling, punctuation, register, awkward phrasing. This is the one dimension where machine output is usually strong and where a QE model can actually help, because fluency is largely a surface property. It is also the least dangerous dimension, because fluency errors announce themselves.

Stand back and see the shape of the problem. Three of the four dimensions that decide whether a translation is acceptable, accuracy, terminology, and locale, are dimensions a reference-free QE model perceives poorly, because all three require comparing the output against something external: the source's meaning, the client's termbase, the locale's rules. The one dimension QE handles reasonably, fluency, is the one that matters least to consequence. The metric is strong exactly where the risk is low and weak exactly where the risk is lethal.

The Severity Model the Program Builds Toward

The second label every MQM error carries, severity, is where this lesson connects to the spine of the entire program, so we will build the model carefully. Severity answers a question a confidence score never even asks: not "how likely is this segment fine" but "if it is wrong, how badly does the wrongness matter." ISO 5060 formalizes a three-level severity scale that you will use in every quality decision from here forward: Critical, Major, and Minor.

Critical

A Critical error is one that can cause real harm: physical, legal, financial, reputational, or safety harm. A flipped contraindication in a drug label. The dropped negation in the currency-fluctuation guarantee. An inverted indemnity clause. A corrupted dosage. These are not "very bad" versions of ordinary errors; they are a different kind of thing, because their consequence is not rework but damage in the real world. The defining rule of the entire quality system, the one you will meet again and again, is this: one Critical error fails the file, regardless of how clean every other segment is. A file with a QE average of 0.91 and one Critical error is a failed file. The average is irrelevant. The Critical is dispositive. This single rule is the clearest possible demonstration of why an averaged confidence score can never be a quality gate: averaging is exactly the wrong operation for a risk where one instance is catastrophic.

Major

A Major error significantly changes or obscures meaning, or breaks an important requirement, but does not rise to the level of real-world harm. A mistranslation that confuses the reader without endangering them. A wrong but non-dangerous term. A clause whose meaning is muddied but not reversed. Major errors degrade quality seriously and accumulate against the file's score, but a single Major does not automatically fail delivery the way a single Critical does. Several Majors can.

Minor

A Minor error is a small imperfection that does not impede understanding: a slightly awkward phrasing, a minor punctuation slip, a stylistic preference that is technically defensible but not ideal. Minor errors are real and they are counted, but they are the least weighted. A file with several Minors and zero Criticals and zero Majors is generally a passing file. The severity scale exists precisely so that you never treat a Minor punctuation issue and a Critical inverted obligation as the same "error," which is exactly the flattening that a single QE number performs.

Severity is the question a confidence score cannot ask. Not "how likely is this fine," but "if it is wrong, does it kill someone, lose a lawsuit, or just read a little awkwardly." One Critical fails the file.

Why Severity Breaks the Confidence Score

Now put QE and severity side by side and the incompatibility is total. A QE score is one number per segment, an averageable, sortable, surface-level probability. Severity is a judgment about consequence that cannot be averaged, because the whole point of Critical is that one of them is disqualifying. You cannot build a one-Critical-fails gate out of a metric whose only move is to average plausibility across the file. The two tools answer different questions. QE asks "where should I look." Severity asks "is this safe to ship." Confusing the first answer for the second is the structural error this lesson exists to prevent, and it is the error the prospectus dashboard committed when it implied that a 0.91 average meant a near-finished file.

Why Automatic Metrics Are Not Quality

QE is one member of a larger family of automatic metrics, and the family as a whole shares a fundamental limitation that you must be able to articulate to a client or a project manager who waves a number at you. Automatic metrics measure proxies for quality. They are not quality. Walking through the main ones makes the principle concrete.

BLEU and the N-Gram Trap

BLEU counts shared word sequences between machine output and a human reference. Consider what that rewards and what it ignores. A translation that uses different but equally correct wording from the reference scores low, because the n-grams do not match, even though it is a perfect translation. A translation that matches the reference word for word except for a single flipped "not" scores extremely high, because almost every n-gram still matches, even though that one flip is a Critical error. BLEU literally cannot tell the difference between a creative-but-correct rendering and a near-identical-but-catastrophic one, because it is counting surface overlap, not meaning. A high BLEU score and a shipped Critical error are entirely compatible. BLEU was always meant as a rough, fast research instrument for comparing engines in aggregate, never as a per-segment quality verdict, and using it as the latter is a category error.

Edit Distance Measures Effort, Not Correctness

Edit distance, often reported in localization as a post-edit distance or an "edits per segment" figure, measures how much a post-editor changed the machine output: how many insertions, deletions, and substitutions it took to turn the raw MT into the final delivered text. Vendors love it because it looks like a quality number. It is not. It measures effort, and effort is not correctness. Two failure modes prove the point. First, a segment that was completely wrong but that the post-editor under-scrutinized, fixing only the obvious surface issues, shows a low edit distance and looks "clean," when in fact a Critical error sailed through untouched: low edits, terrible quality. Second, a segment that was already correct but that a fussy post-editor rewrote to personal preference shows a high edit distance, when in fact no quality was added and budget was burned: high edits, no quality gain. Edit distance tells you how much the text moved. It is silent about whether the text is right.

The Principle Under All of Them

Across QE, BLEU, and edit distance, the same boundary holds. Every automatic metric measures something on the surface of the text, overlap, plausibility, or amount of change, because surface properties are what a machine can compute without understanding meaning. Quality, in the sense that matters to a regulated client, is a relationship between the target, the meaning of the source, the approved terminology, and the rules of the locale. That relationship requires reading for meaning, and reading for meaning is precisely what these metrics cannot do. This is not a flaw to be fixed in the next model version. It is the definitional boundary between predicting quality and judging it. A better QE model will route your effort better. It will never become a quality verdict, because a verdict on accuracy, terminology, and locale requires the one capability a reference-free metric structurally lacks: comprehension of the source.

Every automatic metric measures the surface of the text, because the surface is all a machine can compute without meaning. Quality is a relationship to the source, the termbase, and the locale, and that relationship can only be judged, not estimated.

How the Aware Linguist Uses QE Without Being Used by It

None of this is an argument to ignore QE. Ignoring a genuine triage signal in a file of ten thousand segments would be its own kind of malpractice. The argument is for using QE as exactly what it is and never as what it is not. The aware linguist holds a small set of disciplines that keep the score in its proper place.

Treat the Score as a Question, Not an Answer

A low score asks "did the machine struggle here," and the honest answer is usually yes, so you look. A high score asks nothing and answers nothing about safety; it merely failed to find a surface problem. The score never says "this is correct." It says, at most, "I did not detect a problem from the outside," and the difference between those two sentences is the difference between a shipped file and a shipped lawsuit. Train yourself to mentally rewrite every high score as "no surface problem detected," because that phrasing keeps you honest about what the green actually certifies, which is nothing about meaning.

Let Risk Tier Override the Score

The content's consequence, not its QE score, sets how hard you look. This is risk-tiered intake, the discipline of classifying content by consequence before post-editing begins. A high-scoring segment in a drug label, a contract, or a financial disclosure gets read against the source in full, every time, no matter how green the number, because the cost of a missed Critical in that content is unrecoverable. A high-scoring segment in internal documentation or a product catalog can lean on the routing more, because its errors are recoverable. The QE score informs effort within a risk tier. It never lifts content out of the scrutiny its risk tier demands. The currency clause was in a financial prospectus; its risk tier alone should have guaranteed a full source read regardless of a 0.94.

Check the Things the Score Is Blind To

Because you know QE is weakest on accuracy, terminology, and locale, you check those things deliberately on the high-scoring segments rather than relaxing on them. Negations, numbers, dosages, dates, units, approved terms, and obligations are checked by design, on every file, especially where the score is high, because high is where the fluent error hides. You are not duplicating the QE model's work. You are doing the work it cannot do: reading the target against the source for the categories of error that never show up on the surface. The QE model and the human are not redundant. They are complementary, and the human owns the half that decides whether the file ships.

Never Let the Score Be the Record

When a client or an auditor asks whether a file is good, the answer is never "the QE average was 0.91." The answer is a severity-scored evaluation: here are the errors found, here is each one's category and severity, here is the count of Criticals, which is zero, and here is the go or no-go decision that count produces. A QE distribution is an input to your process. The defensible quality record is an output of human judgment against an error typology, and the program builds you toward producing exactly that record. The score helped you allocate the hours. The evaluation is what you stand behind.

Key Takeaways

  • Quality estimation (QE) is a reference-free prediction: a machine outputs a confidence score, often 0 to 1, estimating how good a translation is with no human reference to check against. It predicts quality, it does not measure it, and the prediction comes from the same statistical machinery, with the same blind spots, that produced the translation.
  • A QE score is a triage signal that routes scarce human effort, not a quality gate that ships output. Low scores legitimately earn a segment your attention; high scores do not earn a segment a pass. Routing is the job; clearance is the trap.
  • The silent critical error hides in the high-scoring segments, because a dropped negation or a flipped dosage produces fluent, grammatical prose, and fluency is exactly what the QE model reads as health. "Trust the green, scrutinize the red" ships the errors that cost a life or a lawsuit.
  • Real quality measurement is an analytic human evaluation under MQM (Multidimensional Quality Metrics) and ISO 5060:2024: a qualified evaluator reads target against source and marks every error with a category and a severity, rather than rating the file with a single number from the outside.
  • The four error dimensions are accuracy, terminology, locale conventions, and fluency. Three of the four require comparing output against something external (the source's meaning, the termbase, the locale's rules) and are exactly where reference-free QE is weakest; the one QE handles, fluency, matters least to consequence.
  • The severity model the program builds toward is Critical, Major, Minor. One Critical error fails the file regardless of how clean everything else is, which is the clearest proof that an averaged confidence score can never be a quality gate: averaging is the wrong operation for a risk where one instance is catastrophic.
  • Automatic metrics are not quality. BLEU counts n-gram overlap with a reference and cannot distinguish a creative-but-correct rendering from a near-identical-but-catastrophic one; edit distance measures post-editing effort, not correctness. Every automatic metric measures the surface because the surface is all a machine can compute without meaning.
  • The aware linguist treats a high score as "no surface problem detected," lets risk tier override the score, deliberately checks negations, numbers, terms, dates, and obligations on the high-scoring segments, and produces a severity-scored evaluation, never a QE average, as the defensible quality record.