Reading an Automatic QE Score Skeptically
Mara had four hours and a file of 6,200 segments. The brief was a consumer medical-device app: setup screens, a troubleshooting flow, and one short section of clinical-use instructions, all machine-translated from English into Brazilian Portuguese, every segment pre-populated in the CAT tool (the computer-assisted translation editor where a linguist works segment by segment) before she opened it. Beside each line sat a quality-estimation score, the QE score, an automatic confidence number the system had attached to every machine-translated segment. The project manager had filtered the view and pasted her a cheerful summary: "QE average 0.89, only 340 segments under 0.70, should be a light pass." Mara read that sentence the way a pilot reads a weather report that is too convenient. The number was not wrong. The conclusion drawn from it was about to be catastrophically wrong. This lesson is a worked example: it walks through exactly what Mara did with that QE distribution, how she used it to route her four hours, why she refused to let a high score clear a single clinical segment, and where, predictably, the one error that could have failed the file was hiding. By the end you will be able to take a real QE-scored file and turn the score into a plan for your attention instead of a verdict on your delivery.
What the Number on Mara's Screen Actually Was
Before Mara touched a single segment, she did the one thing that separates a linguist who uses QE from one who is used by it: she said out loud what the number actually measured, in plain working terms, so she could not quietly forget it under deadline pressure. Quality estimation (QE) is a machine predicting how good a translation is without ever seeing a correct human reference to compare it against. The model looks at the English source segment and the Portuguese machine output, and it returns a number, here scaled 0 to 1, that represents its guess at the likelihood that the segment is acceptable. There is no answer key. There is no human translation of that line sitting somewhere that the model checks against. The model learned, from large quantities of past data, what acceptable and unacceptable translation pairs tend to look like, and it generalizes that pattern to the new pair in front of it.
That word, predict, is the entire lesson compressed into a verb. The QE model predicts quality. It does not measure it. A confidence score is a probability: the model's estimate of the chance that this segment is fine. It is the machine's best guess about its own output, produced by the same statistical machinery that produced the translation, carrying the same blind spots. When a segment in Mara's file scored 0.93, the system was not certifying that the Portuguese was correct. It was reporting that, from the outside, looking only at surface properties a machine can compute, this segment looked like the acceptable pairs in its training data. A 0.93 means "I did not detect a problem." It does not mean "there is no problem." Those are different sentences, and the gap between them is exactly where a file fails an audit.
A QE score is a probability that a segment is good, produced by a machine that cannot read for meaning. Read every high score as "no surface problem detected," never as "this is correct."
Reference-Free Is the Superpower and the Trap
To use QE well you have to understand the one design choice that defines it. Older automatic metrics are reference-based: they need a human "correct" translation, called a reference, and they score the machine output by how closely it matches that reference. BLEU (Bilingual Evaluation Understudy), the metric the field reported for two decades, works this way, counting how many short word sequences the machine output shares with the reference. The trouble is obvious the moment you sit in Mara's chair: in a live file there is no reference. If a correct human translation of the segment already existed, nobody would need the machine to translate it. References live in research benchmarks and controlled engine bake-offs, not in the file on your screen on a Tuesday morning.
QE was invented to fill exactly that gap. QE is reference-free. It scores the segment with no answer key, which is precisely why it can run live, on every one of Mara's 6,200 segments, the instant the machine produces output. That is a real and useful superpower, and dismissing QE entirely would be its own malpractice. But the power has a price baked into it. To score without a reference, QE must predict quality from surface signals instead of measuring it against a known-correct answer, and a prediction can be confidently, fluently wrong in exactly the same way the translation it is scoring can be. The reference-free design is what lets QE run everywhere. It is also what guarantees QE is blind to the errors that hide below the surface.
How Mara Read the Distribution
Here is where the worked example earns its keep. The PM handed Mara an average (0.89) and a count (340 segments under 0.70) and a conclusion ("light pass"). Mara discarded the conclusion and kept the data, because an average and a verdict are not the same object. The first thing she did was stop looking at the mean and start looking at the shape. An average flattens a file into one number, and flattening is precisely the operation that hides risk. She re-sorted the 6,200 segments by QE score, low to high, and looked at the distribution as a map of where the machine had told on itself.
The Low Scores Are a Gift, Take It
The 340 segments below 0.70 were, to Mara, the easy part of the decision. A low QE score is the model raising its hand: confusing source, unusual structure, a construction it had low confidence in. These segments deserve a close look, and QE pointing her straight to them was a genuine productivity gain. Without the score she would have had to find the weak segments herself by reading all 6,200; with it, the machine had pre-flagged the places it struggled. So Mara's first routing decision was simple and correct: the low-scoring band gets read first, against the source, and fixed. This is QE doing exactly the job it is good at. It sorts the file so scarce human attention lands where the machine was least sure.
Notice the economics under that decision, because they are why routing matters at all. A post-editor (the linguist who edits machine output rather than translating from scratch, doing the work called MTPE, machine-translation post-editing) works under per-word rates that typically run 50 to 75% of full human translation. The budget does not pay for scrutinizing 6,200 segments with equal depth in four hours. The honest question is never "is this file perfect," it is "where, in my limited time, is my attention worth the most." That is a routing question, and the low end of the QE distribution answers part of it well.
The High Scores Are Not a Verdict, They Are a Silence
This is the move that saved Mara's file, and it is the move the PM's "light pass" conclusion got exactly backwards. The PM's mental model was: low scores are the problem, high scores are done. Mara's model was different and correct: low scores are worth my attention, and high scores are silent about safety. Those are not two ways of saying the same thing. They are two completely different claims, and conflating them is the single error this lesson exists to prevent.
A high QE score means the model could not see anything wrong from the outside. It does not mean nothing is wrong. The model's inability to catch a dropped negation is not reduced when the negation is dropped inside a fluent sentence; it is increased by it, because fluency is the very thing the model reads as a sign of health. So Mara treated the high band not as cleared, but as unexamined. The green segments had not passed a quality check. They had failed to trip a surface alarm, which is a much weaker statement, and on medical content a much more dangerous one to act on.
A low QE score earns a segment your attention. A high QE score does not earn a segment a pass; it only means the machine did not detect a surface problem. Routing is the job. Clearance is the trap.
Where the Killer Error Was Hiding
Mara's risk-tiered instinct sent her to the clinical-use instructions before anything else, regardless of their scores. There were 70 of those segments. Their QE scores were, on average, higher than the rest of the file, because clinical instructions are written in clean, declarative, well-structured source language, and clean source produces clean, confident machine output that the QE model loves. The very segments with the gravest consequence carried the most reassuring numbers. This is not a coincidence. It is the structural cruelty at the center of QE, and it is worth stating as a law.
The silent critical error, the fluent sentence that means the opposite of the source, lives disproportionately in the high-scoring segments, not the low ones. Why? Because a dropped negation produces a perfectly grammatical sentence. A flipped number produces a clean figure in a clean clause. An inverted instruction reads like any other well-formed line. None of these errors degrade fluency, and fluency is most of what the QE model can perceive. The property that makes these errors lethal to a human reader, their smoothness, is the same property that makes them invisible to the QE model, so they earn high scores. The QE distribution does not sort errors by how badly they matter. It sorts segments by how plausible they look on the surface, and the most dangerous errors are the most plausible on the surface.
Segment 4,417
Segment 4,417 scored 0.94, one of the highest numbers in the entire file, comfortably inside the band the PM had written off as "done." The English read: Do not use the device on a patient who has an implanted pacemaker. The Portuguese, fluent, natural, grammatically immaculate, read back, when Mara translated it in her head against the source: Use the device on a patient who has an implanted pacemaker. The negation, the entire safety instruction, was gone. The machine had produced a smooth sentence that instructed a clinician to do the one thing the source explicitly forbade. The QE model had rated this inversion 0.94, near the top of the file, because the Portuguese was a perfectly well-formed sentence, and a well-formed sentence is what the model reads as quality.
If Mara had accepted the PM's routing, she would have read the 340 red segments, fixed their clumsy, self-announcing, low-stakes problems, and shipped a file that told a clinician to use a contraindicated device on a pacemaker patient, with a QE average of 0.89 and a green checkmark beside the killer line. The score would not have been lying. It would have been doing exactly what it was built to do, and exactly what it must never be trusted to do alone.
"Trust the green, scrutinize the red" is a recipe for shipping the exact errors that cost a life or a lawsuit. The red is clumsy and low-stakes. The green is smooth, confident, and where the inverted clause hides.
The Routing Plan Mara Actually Built
So what did Mara do with her four hours and her 6,200 segments? She did not read every segment with equal depth, which was impossible, and she did not trust the green, which was unsafe. She built a routing plan that used the QE score as one input among several, with the content's consequence, its risk tier, as the override. Walk through it, because this is the transferable skill.
Step One: Tier the Content Before Reading the Scores
Risk-tiered intake is the discipline of classifying content by consequence before post-editing begins, and it comes first, before the QE distribution gets a vote. Mara split the file into three tiers by what an error would cost:
- High-consequence: the 70 clinical-use instruction segments, where an error can cause physical harm. These get a full read against the source, every segment, no matter the score.
- Medium-consequence: the troubleshooting flow, perhaps 900 segments, where an error frustrates a user or causes a support call but does not harm anyone. These lean partly on the score.
- Low-consequence: the setup screens and UI labels, the remaining bulk, where errors are recoverable and cheap. These lean most heavily on the score for routing.
The tier, not the number, sets the floor on how hard she looks. A 0.94 on a clinical instruction earns no relaxation at all. A 0.94 on a UI label earns a quick confirmation and a move on. Same score, different scrutiny, because the consequence is different, and consequence is something a QE model knows nothing about.
Step Two: Route Within Each Tier Using the Score
Inside each tier, the QE score does its legitimate job of ordering attention. In the low-consequence bulk, Mara read the low-scoring band closely and sampled the high-scoring band, accepting that a missed Minor error in a setup label is a recoverable cost. In the medium tier, she read all the low scores and spot-checked the high ones at a higher rate, weighting toward anything with a number, a placeholder, or a term in it. In the high-consequence tier, the score changed nothing about coverage: she read all 70 clinical segments against the source regardless of score, and she used the score only to decide reading order, hitting the low ones first to clear the obvious problems and then reading the high ones with full attention precisely because high is where the fluent inversion hides.
Step Three: Check the Categories QE Is Blind To
On every high-consequence segment, and on the high-scoring segments she chose to examine elsewhere, Mara did not re-do the QE model's work. She did the work it cannot do. She checked the specific categories of error that never show up on the surface, the ones a reference-free model is structurally worst at:
- Accuracy against the source: negations, omissions, added information, reversed meaning. This is where segment 4,417 lived, and it is the dimension a QE model is weakest at, because catching it requires reading the source for meaning.
- Numbers, dosages, dates, units: a flipped digit or a swapped unit produces a clean, high-scoring sentence and a real-world catastrophe. Checked digit by digit, never skimmed.
- Terminology: whether the client's approved term from the termbase (the controlled glossary of approved terms) was used, or whether the engine substituted a fluent, common synonym. A QE model trained on general text often rates the wrong-but-common term higher than the right-but-unusual one.
- Locale conventions: dates, decimal separators, formality register, all things that are quietly and expensively wrong while reading perfectly fluent in the target locale (the specific language-and-region combination, like Portuguese for Brazil versus Portuguese for Portugal).
These four categories are not arbitrary. They map onto the dimensions a real human evaluation scores, and three of the four, accuracy, terminology, and locale, are exactly the dimensions a reference-free QE model perceives poorly, because all three require comparing the output against something external: the source's meaning, the client's termbase, the locale's rules. The one dimension QE handles reasonably, fluency, is the one that matters least to consequence. The metric is strongest exactly where the risk is lowest, and blindest exactly where the risk can kill.
What the Score Could Never Become: A Quality Verdict
It helps to see why this is a permanent boundary and not a temporary weakness that a better model will fix. The thing Mara produced at the end of her four hours was not a QE average. It was the beginning of a real quality evaluation, and the difference between the two is the difference between a guess and a judgment.
Real quality measurement is an analytic human evaluation built on an error typology, and the dominant framework is MQM (Multidimensional Quality Metrics), now formalized for translation output by the international standard ISO 5060:2024 (Translation services, Evaluation of translation output). Where QE outputs one confidence number per segment from the outside, an MQM evaluation has a qualified human read the target against the source and mark each error with two labels: a category (what kind of error it is, across accuracy, terminology, locale, and fluency) and a severity (how much it matters). It is called analytic evaluation because it takes the file apart into specific, located, classified errors instead of rating it holistically with a vibe or a single probability.
Severity Is the Question the Score Cannot Ask
The severity label is where QE and real evaluation become incompatible, and it is worth being precise about the three levels because they are the spine of every quality decision you will make after this lesson. ISO 5060 formalizes a three-level scale:
- Critical: an error that can cause real harm, physical, legal, financial, or safety harm. Segment 4,417's inverted pacemaker instruction is a Critical. So is a flipped dosage, an inverted indemnity clause, a corrupted contraindication. These are not "very bad" ordinary errors; they are a different kind of thing, because their consequence is damage in the real world, not rework.
- Major: an error that significantly changes or obscures meaning, or breaks an important requirement, but does not reach real-world harm. A confusing mistranslation, a wrong but non-dangerous term, a muddied clause. Several Majors can fail a file; one usually does not.
- Minor: a small imperfection that does not impede understanding, a slightly awkward phrasing, a punctuation slip. Counted, but least weighted.
The defining rule of the entire quality system, the one you will meet in every lesson that follows, is this: one Critical error fails the file, regardless of how clean every other segment is. Mara's file had a QE average of 0.89, hundreds of green clinical segments, and exactly one Critical at 0.94. That file is a failed file until the Critical is fixed. The average is irrelevant. The Critical is dispositive.
One Critical error fails the file, no matter how high the average. Averaging is the wrong operation for a risk where a single instance is catastrophic, which is the clearest possible proof that an averaged confidence score can never be a quality gate.
Why You Cannot Build a Gate Out of an Average
Put QE and severity side by side and the incompatibility is total. A QE score is one number per segment: averageable, sortable, surface-level. Severity is a judgment about consequence that cannot be averaged, because the entire point of Critical is that one of them is disqualifying. You cannot build a one-Critical-fails gate out of a metric whose only move is to average plausibility across the file. The two tools answer different questions. QE asks "where should I look." Severity asks "is this safe to ship." Mara's PM confused the first answer for the second when he read a 0.89 average as a near-finished file, and that confusion is the structural error the whole discipline exists to prevent.
The Other Metrics on the Dashboard, and Why They Mislead Too
QE is one member of a family of automatic metrics, and the PM's dashboard showed two others Mara had to be ready to argue against, because a client or a manager will eventually wave one of them at you as proof of quality. Being able to explain, calmly, what each metric actually measures is part of owning the quality.
BLEU Counts Words, Not Meaning
BLEU counts shared word sequences between machine output and a human reference. Consider what that rewards and ignores. A correct translation that happens to word things differently from the reference scores low, even though it is perfect, because the word sequences do not overlap. A translation identical to the reference except for one flipped "not" scores extremely high, even though that flip is a Critical error, because almost every word sequence still matches. BLEU literally cannot distinguish a creative-but-correct rendering from a near-identical-but-catastrophic one, because it counts surface overlap, not meaning. A high BLEU score and a shipped Critical error are entirely compatible. BLEU was built as a rough, fast research instrument for comparing engines in aggregate, never as a per-segment quality verdict.
Edit Distance Measures Effort, Not Correctness
Edit distance, often reported as "edits per segment" or post-edit distance, measures how much a post-editor changed the machine output. Vendors love it because it looks like a quality number. It is not; it measures effort, and effort is not correctness. Two failure modes prove it. A segment that was badly wrong but under-scrutinized, where the post-editor fixed only the obvious surface issues and missed the Critical, shows a low edit distance and looks clean: low edits, terrible quality. A segment that was already correct but that a fussy post-editor rewrote to personal preference shows a high edit distance: high edits, zero quality added, budget burned. Edit distance tells you how much the text moved. It is silent about whether the text is right.
The Boundary Under All of Them
Across QE, BLEU, and edit distance, the same line holds. Every automatic metric measures something on the surface of the text, plausibility, overlap, or amount of change, because surface properties are all a machine can compute without understanding meaning. Quality, in the sense that matters to a regulated client, is a relationship between the target, the meaning of the source, the approved terminology, and the rules of the locale. That relationship requires reading for meaning, which is precisely what these metrics cannot do. This is not a flaw a future model version fixes. It is the definitional boundary between predicting quality and judging it. A better QE model will route Mara's effort better. It will never become a quality verdict, because a verdict on accuracy, terminology, and locale requires the one capability a reference-free metric structurally lacks: comprehension of the source.
Every automatic metric measures the surface of the text, because the surface is all a machine can compute without meaning. Quality is a relationship to the source, the termbase, and the locale, and a relationship can only be judged, not estimated.
The Five Disciplines That Keep the Score in Its Place
Mara did not invent her workflow under pressure. She ran a small, fixed set of disciplines that keep a QE score useful without ever letting it become a verdict. Internalize these five and you can walk into any QE-scored file and route your attention correctly.
Discipline One: Rewrite Every High Score in Your Head
Train yourself to read every high QE score as "no surface problem detected," never as "this is correct." A low score asks "did the machine struggle here," and the honest answer is usually yes, so you look. A high score asks nothing and answers nothing about safety; it merely failed to find a surface problem. The score never says "this is right." It says, at most, "I did not detect a problem from the outside," and the distance between those two sentences is the distance between a shipped file and a shipped lawsuit.
Discipline Two: Tier Before You Route
Classify content by consequence before the QE distribution gets a vote. A high-scoring segment in a drug label, a contract, or a financial disclosure gets read against the source in full, every time, no matter how green, because a missed Critical there is unrecoverable. A high-scoring segment in internal documentation can lean on the routing, because its errors are recoverable. The score informs effort within a risk tier. It never lifts content out of the scrutiny its tier demands.
Discipline Three: Read the Distribution, Not the Average
An average is a flattening, and flattening hides risk by construction. Sort the file by score and read the shape: where the low band is, how heavy the high band is, and crucially, which content type the high scores cluster on. When the high scores cluster on your highest-consequence content, as they did on Mara's clinical section, treat that as a warning, not a comfort, because clean source produces confident output and confident output is where the fluent inversion lives.
Discipline Four: Check What the Score Is Blind To, by Design
Because you know QE is weakest on accuracy, terminology, and locale, you check those deliberately on the high-scoring segments rather than relaxing on them. Negations, numbers, dosages, dates, units, approved terms, and obligations are checked on every high-consequence file, especially where the score is high, because high is where the fluent error hides. You are not duplicating the model's work. You are doing the half it cannot do, and that half decides whether the file ships.
Discipline Five: Never Let the Score Be the Record
When a client or an auditor asks whether a file is good, the answer is never "the QE average was 0.89." The answer is a severity-scored evaluation: here are the errors found, here is each one's category and severity, here is the count of Criticals, which after Mara's fix is zero, and here is the go decision that count produces. A QE distribution is an input to your process. The defensible quality record is an output of human judgment against an error typology, and that record, not the score, is the thing you stand behind when your name is on the delivery.
Mara fixed segment 4,417, finished her high-consequence read, routed her remaining hours through the low and medium tiers with the score, and delivered a file with a documented evaluation and zero Criticals. The QE average was the same 0.89 the PM had celebrated. The difference between the file the PM imagined and the file Mara delivered was not the number. It was the one human reading that the number could never replace, and could only, if she had let it, have talked her out of.
Key Takeaways
- A QE (quality-estimation) score is a reference-free prediction: a machine outputs a confidence probability that a segment is acceptable, with no human reference to check against. It predicts quality, it does not measure it, and the prediction carries the same blind spots as the engine that produced the translation. Read every high score as "no surface problem detected," never as "this is correct."
- Use QE to route effort, not to clear segments. A low score legitimately earns a segment your attention; a high score is silent about safety and only means the machine did not trip a surface alarm. Routing is the job; clearance is the trap, and conflating the two is the structural error the discipline exists to prevent.
- The silent critical error hides in the high-scoring segments, because a dropped negation, a flipped dosage, or an inverted instruction produces fluent, grammatical prose, and fluency is exactly what the QE model reads as health. Clean source on high-consequence content produces high scores, so the gravest errors carry the most reassuring numbers.
- Tier content by consequence before the score gets a vote. The risk tier sets the floor on scrutiny: a 0.94 on a clinical instruction earns a full source read; the same 0.94 on a UI label earns a quick check. The score informs effort within a tier; it never lifts content out of the scrutiny its tier demands.
- Check the categories QE is structurally blind to: accuracy against the source, numbers and dosages and dates, approved terminology, and locale conventions. Three of the four error dimensions require comparing output against something external (source meaning, termbase, locale rules) and are exactly where reference-free QE is weakest; fluency, the one QE handles, matters least to consequence.
- Real quality measurement is an analytic human evaluation under MQM (Multidimensional Quality Metrics) and ISO 5060:2024, where a qualified evaluator marks every error with a category and a severity (Critical, Major, Minor). One Critical error fails the file regardless of the average, which proves an averaged confidence score can never be a quality gate: averaging is the wrong operation for a risk where one instance is catastrophic.
- Automatic metrics are not quality. BLEU counts n-gram overlap with a reference and cannot tell a creative-but-correct rendering from a near-identical-but-catastrophic one; edit distance measures post-editing effort, not correctness. Every automatic metric measures the surface because the surface is all a machine can compute without meaning.
- The defensible quality record is a severity-scored evaluation with a Critical count, never a QE average. The score helps you allocate the hours; the human evaluation against the error typology is what you stand behind when your name is on the delivery.
Skill.re