Quality-Estimation Techniques for Linguists
The file landed in Diego's queue at 9:14 on a Monday: 41,800 segments of an industrial-equipment knowledge base, English into German, every segment machine-translated and pre-populated in the CAT tool (the computer-assisted translation editor where a linguist works one segment at a time) before he opened it. Beside each line sat a quality-estimation score, the QE score, an automatic number the engine had attached to every segment as its guess at how good the translation was. The deadline was Thursday end of day. Diego could read maybe 9,000 segments in that window at the depth the content deserved, which meant he had to decide, in the first hour, which 9,000. The wrong answer was to start at segment one and read until Thursday, getting a quarter of the way through a file he had no map of. The right answer was to turn the QE distribution into a routing plan: a defensible decision about where his scarce attention would land, which segments the score could clear without him, which it could not, and which it would actively mislead him about. This lesson is that decision, worked end to end on Diego's file. By the end you will be able to take a real QE-scored file of any size, set triage thresholds you can defend, combine the score with a risk tier so it never clears content it has no business clearing, calibrate how far to trust QE on your own content, and route a week of human effort to the segments that actually need it.
From Reading the Score to Operating On It
You have already learned to read a QE score skeptically: to treat a high number as "no surface problem detected" rather than "this is correct," and to know that the silent critical error hides in the smooth, high-scoring segments where the machine and the metric both feel most confident. That is the conceptual foundation, and it is the right one. This lesson is the next move: turning that skepticism into an operating procedure you run on a real file with a real deadline. Reading the score keeps you from being fooled by it. Operating on the score is how you use it to do more work, faster, without dropping the one error that fails the file.
Let us define the operating vocabulary precisely, because the rest of the lesson stands on it. Quality estimation (QE) is a machine predicting how good a machine translation is without any human reference to compare against; it reads the source segment and the target output and returns a number, usually scaled 0 to 1 (or 0 to 100), representing its estimate of the probability the segment is acceptable. Triage is the act of sorting a large body of work into categories of urgency and effort so that limited resources go where they matter most; the word comes from emergency medicine, where a nurse decides in seconds who is seen first, and it is exactly the right metaphor for a linguist facing 41,800 segments and 30 hours. A threshold is a score boundary you set on purpose, a line in the distribution above which you do one thing and below which you do another, the mechanism that turns a continuous stream of numbers into discrete routing decisions. The whole of operational QE is this: set thresholds, triage by them, and override them with risk where the consequence demands it.
Reading a QE score keeps you from trusting it blindly. Operating on a QE score is how you turn it into a triage plan that routes your hours to the segments that need a human, without letting it clear the segments that will fail the file.
Why Routing Is the Whole Game at Scale
On a 200-segment file you can read everything, and QE is a convenience. On a 41,800-segment file you cannot, and QE stops being a convenience and becomes the thing that decides whether the file is delivered well or delivered late. The economics force the point. A post-editor (the linguist who edits machine output instead of translating from scratch, doing the work called MTPE, machine-translation post-editing) works under per-word rates that typically run 50 to 75% of full human translation, and a hybrid workflow is expected to lift throughput from roughly 2,000 words a day to 5,000 or more. That throughput is not free speed. It is bought entirely by not reading every segment with equal depth, which means it is bought by routing, which means it is bought by QE doing its one legitimate job: telling you where the machine was least sure so your attention lands there first. The moment you accept the file you have accepted a routing problem. QE is the best tool you have for solving it, and the worst tool you have for pretending it is solved.
Setting Thresholds You Can Defend
Diego's first instinct, the amateur instinct, was to ask "what is the right QE threshold." There is no right threshold in the abstract, the way there is no right speed limit in the abstract; it depends on the road, the car, and what happens if you are wrong. A threshold is a decision about where to spend attention, and that decision is yours to set against your content, your risk, and your time, not a constant you inherit from a vendor's default. What you can defend is not a magic number but the reasoning that produced it. Let us build Diego's thresholds the way a professional does.
A Threshold Is a Bet About Where Errors Live
When you draw a line at, say, 0.75 and say "below this, I read closely; above this, I sample," you are making a bet: that the density of errors worth catching is high enough below the line to justify reading every segment there, and low enough above the line that sampling will catch what matters. That bet is only as good as the relationship between the score and the actual errors on your content, and that relationship varies wildly by content type, language pair, and engine. The threshold is not a property of the QE model. It is a property of the join between the QE model and your specific file, which is why a number that is correct on your software UI strings can be dangerously wrong on your clinical instructions.
The mistake to avoid is treating the threshold as a clearance line, a number above which segments are "done." It is never that. It is a routing line: above it you spend less attention per segment, below it you spend more. Even above the highest threshold you set, content is sampled, never ignored, because, as you already know, the high band is silent about safety, not certified safe. A threshold changes how hard you look. It must never change to whether you look at all, on any content where a missed error matters.
Diego's Three-Band Scheme
Diego split the 0-to-1 score range into three bands, which is usually enough; more bands add precision you cannot act on with a single pair of hands. He did not pick the boundaries from a webinar. He picked them from a 200-segment calibration sample he scored by hand first, which the next section explains. The bands he landed on for this engine and this content were:
- Low band, below 0.60: the machine is telling on itself. Confusing source, unusual syntax, low confidence. Read every one of these against the source. This is QE doing its best work, pointing him straight at the segments most likely to be genuinely wrong in ways that show.
- Middle band, 0.60 to 0.85: the uncertain middle, where the score is least informative and the errors are mixed. Read these at a high sampling rate, weighted toward any segment carrying a number, a placeholder, a date, or a term. This is the band where most of the judgment lives.
- High band, above 0.85: the machine found no surface problem. Sample these, do not clear them, with the sampling rate set by the content's risk tier rather than by the score. On low-risk content, a light sample. On high-risk content, the high band is read in full regardless, and the score only sets reading order.
Notice what makes these thresholds defensible. Diego can say, to a PM or an auditor, exactly why each line sits where it does: it sits where his calibration sample showed the error density change, on this content, with this engine. He is not claiming 0.85 is a universal truth. He is claiming it is the right line for this file, and he has the sample to prove it. That sentence, "here is why my threshold sits here, on this content," is the difference between a professional routing decision and a guess dressed up as a number.
There is no correct QE threshold in the abstract. A threshold is a bet about where errors live on your specific content, and the only defensible threshold is one you set against a sample you scored yourself, not a vendor default you inherited.
Calibrating Trust on Your Own Content
Here is the discipline that separates a linguist who uses QE operationally from one who merely has QE turned on: before he trusted the score to route a single one of his 30 hours, Diego spent the first 90 minutes finding out whether the score deserved his trust on this content at all. He calibrated. Calibration, in this context, is the act of measuring how well the QE score actually tracks real quality on your specific content, by scoring a sample yourself and comparing your verdict to the machine's. It is the step almost everyone skips, and skipping it is how a perfectly good QE model gets trusted on content where it happens to be useless.
The 90-Minute Calibration Sample
Diego pulled a stratified sample: roughly 70 segments from the low band, 70 from the middle, and 70 from the high band, about 210 segments in total, drawn across the different content types in the file (procedures, warnings, UI labels, error messages). He read every one against the source and marked each as acceptable or not, and where it was not, he noted the error category. Then he did the only thing that matters: he laid his human verdicts next to the QE scores and looked at how often they agreed.
What he was looking for was not whether the score was "accurate" in some absolute sense. He was looking for two specific, actionable failure patterns:
- How clean is the low band, really? If a large fraction of his low-band segments turned out to be fine on close reading, the score was crying wolf, the threshold was set too high, and reading every low segment would waste hours on false alarms. If almost all of them were genuinely problematic, the low band was trustworthy and worth reading in full.
- What is leaking through the high band? This is the one that matters most. He counted how many high-band segments he, the human, judged unacceptable, and crucially, what kind of error they were. A high band leaking minor fluency slips is one thing. A high band leaking even one accuracy inversion or flipped number is a five-alarm warning that the score is blind to exactly the errors that fail the file, on this content, and that no high-band segment can be sampled lightly.
What Diego's Calibration Told Him
The low band came back about 80% genuinely problematic, which meant reading it in full was time well spent and his 0.60 line was roughly right. The middle band was a coin flip, which is exactly what a middle band should be and why it gets heavy sampling. The high band came back almost entirely clean on fluency and terminology, which is reassuring and also entirely beside the point, because among the 70 high-band segments Diego found two accuracy errors: one dropped a safety-relevant "not," and one rendered a torque specification with the wrong unit, and both had scored above 0.90. Two in seventy is not a small number when the error is the kind that fails a file. It is a measured, quantified proof, on this exact content, that the high band is silent about precisely the category that matters, and that no QE score in this file earns a clinical or safety segment a pass.
This is what calibration buys you that a vendor default never can: a number you can say out loud and defend. Diego could now tell his PM, "On a 210-segment sample of this content, the QE high band leaked roughly 3% accuracy errors including a safety negation, so I am reading all warning and procedure segments in full regardless of score." That is not a vibe. It is a measurement, and it converts "I have a bad feeling about trusting the green" into "here is the rate at which the green is wrong about the thing that matters, on this file." Calibrate once per new content type or engine, and the thresholds you set afterward are bets backed by evidence instead of hope.
Calibration is scoring a stratified sample yourself and counting what leaks through the high band, by error category. It converts "I distrust the green" into "the green leaks 3% accuracy errors on this content," which is the only sentence that lets you set a defensible threshold and route a week of work behind it.
Combining QE With Risk Tier
A QE score knows everything about the surface of a segment and nothing about its consequence. It cannot tell a software tooltip from a safety warning, because both can be fluent, well-formed German, and fluency is most of what the score perceives. The score will happily hand a 0.94 to a sentence that, if wrong, injures a technician, and a 0.94 to a sentence that, if wrong, makes a menu item slightly awkward. Identical numbers, incomparable stakes. This is why QE can never be the only input to routing. It must be combined with a risk tier: a classification of content by what an error in it would cost, assigned before the score gets a vote.
The Tier Sets the Floor, the Score Sets the Order
The clean mental model, the one that makes this combination defensible, is a division of labor. The risk tier sets the floor on how much scrutiny a segment gets, the minimum coverage it earns no matter what. The QE score sets the order and the depth above that floor, distributing attention efficiently within what the tier already guarantees. The tier answers "how little am I allowed to look at this," and the score answers "given that floor, where do I look first and hardest." Neither can do the other's job. The score has no idea about consequence; the tier has no idea about which specific segments the machine fumbled. Used together, the tier protects you from the score's blindness to stakes, and the score protects you from the impossibility of reading 41,800 segments with equal depth.
Diego tiered his file into three levels of consequence, the same triage you would run on any mixed file:
- High-consequence: safety warnings, lockout-tagout procedures, anything where an error can cause physical harm to a technician or damage to equipment. Perhaps 2,400 segments. Floor: every segment read against the source, regardless of QE score. The score is used only to set reading order within the tier, low scores first to clear the obvious, then high scores read with full attention because high is where the fluent inversion hides.
- Medium-consequence: operating procedures, configuration steps, error messages that guide a user to a fix. Perhaps 14,000 segments. Floor: all low-band read, middle band sampled heavily, high band sampled at a moderate rate weighted toward numbers, units, and terms.
- Low-consequence: UI labels, navigation, marketing-adjacent descriptions where an error is recoverable and cheap. The remaining bulk, roughly 25,000 segments. Floor: low band read, middle band sampled, high band lightly sampled. This is where QE earns its keep, carrying most of the volume on the strength of the score.
Why the Cross-Product Is the Routing Plan
The actual routing plan is the cross-product of these two axes: tier on one side, QE band on the other. A high-consequence segment scoring 0.95 gets a full source read, because the tier floor demands it and the calibration proved the high band leaks accuracy errors. A low-consequence segment scoring 0.95 gets a glance, because nothing about it can fail the file in a way that matters. Same score, opposite treatment, and the thing that flips the treatment is the tier, which is the one piece of information the QE model fundamentally does not have. Walk the grid and every cell has a rule, and every rule is defensible, because each one says, in effect, "this much scrutiny because this much consequence, ordered this way because this is what the score is good for."
The single most important cell in that grid is the high-consequence, high-score one, because it is the trap. It is the cell where the PM's instinct ("it scored 0.95, it is fine, skip it") and the correct instinct ("it scored 0.95 on safety content, which is exactly where my calibration found the leaks, read it in full") point in opposite directions. Every operational QE failure that ends in a shipped Critical lives in that one cell. The combination of QE and risk tier exists primarily to make sure that cell is never, under any deadline, routed to "skip."
The routing plan is the cross-product of risk tier and QE band. The tier sets the floor on scrutiny because the score is blind to consequence; the score sets the order and depth above that floor because the tier is blind to which segments the machine fumbled. The high-risk, high-score cell is never routed to "skip."
Where QE Is Reliable and Where It Misses
To operate on QE you need a precise, honest map of where it earns trust and where it does not, because the failures are not random. They are structural, predictable, and tied directly to what a reference-free metric can and cannot perceive. Knowing the map is what lets you lean on the score where it is strong and override it where it is blind, instead of trusting it uniformly and getting surprised.
Where QE Is Genuinely Reliable
QE is at its best at the bottom of the distribution, flagging segments where the machine genuinely struggled. A garbled rendering, a collapsed sentence, a segment where the source was so unusual the engine lost the thread: these produce real surface signals the model can read, and a low score on them is usually right. QE is also reliable as a relative sorter within a single homogeneous content type and engine: if you scored 10,000 software strings from one engine, the lowest-scoring 500 really are, on average, the most likely to need work. That relative ordering is the productivity engine of the whole workflow, and dismissing QE entirely would throw it away. The score is a good triage nurse. It is excellent at "see this one first."
Where QE Structurally Misses
QE misses, by design, the errors that do not degrade the surface. The catalog is worth memorizing, because it is the catalog of what your human read must cover on every high-consequence segment:
- A fluent Critical. This is the headline failure and the reason this whole discipline exists. A dropped negation, a flipped dosage, an inverted instruction, a reversed obligation: each produces a perfectly grammatical, natural-sounding sentence that means the opposite of the source. The QE model reads that fluency as health and scores it high. It cannot see a fluent Critical, because seeing it would require reading the source for meaning and comparing, which is exactly the capability a reference-free metric lacks. The score is blindest precisely where the consequence is gravest.
- Approved-terminology violations. When the engine uses a fluent, common synonym instead of the client's approved term from the termbase (the controlled glossary of required terms), the result reads beautifully and scores high, because a QE model trained on general text often finds the common word more "normal" than the correct specialized one. The score frequently rates the wrong-but-common term above the right-but-unusual one.
- Locale and number errors that stay fluent. A flipped digit, a swapped unit, a decimal comma where a period belongs, a date in the wrong order: these produce clean, well-formed target text and a real-world error, and the cleanliness is exactly why the score does not flag them.
- Anything requiring external knowledge. Whether a translated claim is true, whether a part number is the real one, whether a cross-reference points to the right section: the score has no access to the world outside the two segments, so it cannot judge correctness that depends on facts.
Stand back and the pattern is total: QE is reliable on fluency and on relative struggle, and structurally blind on accuracy, terminology, and locale, the three dimensions that require comparing the output against something external (the source's meaning, the termbase, the locale's rules). The one dimension QE handles well, fluency, is the one that matters least to consequence, and the three it handles poorly are the three that fail files. This is not a defect a better model fixes. It is the definitional boundary between predicting quality from the surface and judging it against meaning, and it tells you exactly which half of the work the score can do and which half is permanently yours.
The Worked Routing on Diego's 41,800 Segments
Now assemble it. Diego has calibrated thresholds, a risk-tiered file, and an honest map of where the score lies. Here is exactly how he spent his 30 hours, and why the file shipped clean.
Hours 0 to 2: Tier and Calibrate
Before reading a single segment for delivery, Diego tagged every segment into one of his three risk tiers using the file's structure (the knowledge base was organized by section type, so warnings, procedures, and UI strings were already separable) and ran his 210-segment calibration sample. This is the investment everyone wants to skip under deadline pressure, and it is the investment that makes the remaining 28 hours efficient instead of frantic. Two hours spent learning that the high band leaks 3% accuracy errors on safety content is two hours that saves him from shipping the one Critical that fails the file. The calibration is not overhead. It is the cheapest insurance in the workflow.
Hours 2 to 12: The High-Consequence Tier in Full
The 2,400 high-consequence segments got read against the source, every one, no exceptions for score. Diego ordered them low-score first to clear the self-announcing problems quickly, then read the high-score safety segments last and slowest, because his calibration had told him in numbers that high-score safety content was where the fluent inversion would be hiding. This is the inversion of the naive plan, and it is the entire point: he spent his freshest, deepest attention on the segments the score said were fine, because on this tier the score's "fine" had a measured 3% chance of being a safety error. Two segments in the high-consequence tier turned out to be exactly that: a warning that had lost its negation (scored 0.93) and a procedure step that inverted a sequence of operations (scored 0.91). Both were Criticals. Both would have shipped under "trust the green." Both were caught because the tier floor refused to let the score clear them.
Hours 12 to 24: The Medium Tier by Band
The 14,000 medium-consequence segments got the band treatment. Diego read the entire low band against the source, because calibration had shown it 80% problematic and worth the time. He sampled the middle band heavily, perhaps one in three, weighting his picks toward any segment containing a number, a unit, a placeholder, or a termbase term, because those are the carriers of the errors QE cannot see. He sampled the high band lightly but never at zero, because even a recoverable medium-tier error is worth a spot-check, and because spot-checks are how you catch a calibration that has drifted. The score did its legitimate work here: it ordered 14,000 segments so his attention landed on the ones most likely to need it, and it let him cover the whole tier in twelve hours instead of the forty it would have taken to read every line.
Hours 24 to 30: The Low Tier on the Score
The remaining 25,000 low-consequence segments, the UI labels and navigation and descriptions, leaned hardest on QE, exactly as they should. Diego read the low band, sampled the middle, and lightly spot-checked the high band, accepting openly that a missed Minor fluency slip in a menu label is a recoverable cost that does not justify a full read under deadline. This is where the throughput came from. By trusting the score most on the content where being wrong costs least, Diego cleared 25,000 segments in six hours, which is only possible because he had spent his first 24 hours making sure the content where being wrong costs most never depended on the score at all. The routing was the productivity, and the tiering was what made the productivity safe.
What Diego Delivered
He delivered the file on time, with a severity-scored evaluation attached: the errors found, each one's category and severity, a count of Criticals reduced to zero after his fixes, and a one-paragraph note on his method, including the calibration result that justified his routing. The QE average of the delivered file was whatever it was; he never reported it as the quality measure, because the average was an input to his plan, not the verdict on his work. What he stood behind was the human read on every high-consequence segment, the calibrated routing on the rest, and the zero-Critical evaluation record. A PM who had only the QE average would have called the file done after the score arrived at 9:14. Diego turned the same score into a map of where his week of attention had to go, and the difference between those two uses of one number was the difference between a clean delivery and a shipped safety error with his name on it.
The Operating Checklist for a QE-Scored File
Strip Diego's week down to the repeatable procedure and you get a checklist you can run on any QE-scored file, of any size, in any language pair. This is the transferable skill, the thing to carry out of this lesson and into Monday.
Step One: Tier Before You Touch the Score
Classify the content by consequence first, before the QE distribution gets a vote. Identify the high-consequence segments (safety, legal, medical, financial, anything where an error harms a person or creates liability) and set their floor at full source read regardless of score. Everything else gets tiered down from there. The tier is the safety net the score cannot provide, and it goes first because nothing the score says is allowed to override it.
Step Two: Calibrate on a Stratified Sample
Before trusting the score to route the bulk, pull a stratified sample across the bands and content types, score it yourself against the source, and count what leaks through the high band by error category. The number you want is the high-band accuracy-error rate, because that is the rate at which "trust the green" would ship a real error. Calibrate once per new content type or engine; the result sets your thresholds and tells you whether the high band can ever be sampled lightly on this content.
Step Three: Set Bands From the Calibration, Not From a Default
Draw your threshold lines where your sample showed the error density change, not where a vendor's setting suggested. Three bands is usually enough: a low band you read in full, a middle band you sample heavily, and a high band whose sampling rate is set by the risk tier, never by the score alone. Be able to say, for each line, why it sits where it does on this content.
Step Four: Route by the Cross-Product of Tier and Band
Build the grid: risk tier on one axis, QE band on the other, a scrutiny rule in every cell. Spend full attention on the high-consequence tier regardless of score, with the score setting only reading order. Sample the medium tier by band, weighted toward numbers, units, placeholders, and terms. Lean hardest on the score in the low tier, where errors are recoverable. Guard the high-risk, high-score cell with your life; it is where every shipped Critical hides.
Step Five: Check the Categories the Score Cannot See
On every segment you do read closely, do the work the score cannot: check accuracy against the source (negations, omissions, reversed meaning), numbers and dosages and units digit by digit, approved terminology against the termbase, and locale conventions. You are not redoing the model's work. You are doing the half it is structurally incapable of, the half that decides whether the file ships.
Step Six: Record the Evaluation, Not the Average
What you deliver and stand behind is a severity-scored evaluation under MQM (Multidimensional Quality Metrics, the error typology formalized for translation output by ISO 5060:2024) with a Critical count, not a QE average. The score allocated your hours; the human evaluation against the error typology is the record. When a client or an auditor asks whether the file is good, the answer is the evaluation and its zero Criticals, never the number the engine printed beside the segments.
Tier, calibrate, set bands from the calibration, route by the cross-product of tier and band, check what the score is blind to, and record the evaluation, not the average. That sequence turns a QE score from a verdict you cannot trust into a routing tool you can defend.
Key Takeaways
- Operating on QE, not just reading it, is the L3 skill: turning the score into a triage plan that routes scarce human hours across a large file, so you post-edit more, faster, without dropping the one error that fails the file. On a file too big to read in full, routing is the whole game, and QE is the best tool for it and the worst tool for pretending the file is done.
- There is no correct QE threshold in the abstract. A threshold is a bet about where errors live on your specific content, engine, and language pair, and a defensible threshold is one set against a sample you scored yourself, never a vendor default. A threshold is a routing line that changes how hard you look, never a clearance line that changes whether you look.
- Calibrate before you trust: pull a stratified sample across the bands and content types, score it against the source yourself, and count what leaks through the high band by error category. The number that matters is the high-band accuracy-error rate, because that is the rate at which "trust the green" ships a real error. Diego's high band leaked roughly 3% accuracy errors including a safety negation, which is the measured proof that no score on that content earns a safety segment a pass.
- Combine QE with a risk tier, because the score is blind to consequence: it hands an identical high number to a safety warning and a menu label. The risk tier sets the floor on scrutiny (high-consequence content is read in full regardless of score); the QE score sets the order and depth above that floor. The routing plan is the cross-product of tier and band.
- The high-risk, high-score cell is the trap. It is where the naive instinct ("it scored 0.95, skip it") and the correct instinct ("it scored 0.95 on safety content, where calibration found the leaks, read it in full") diverge, and where every shipped Critical hides. The whole point of combining QE with risk tier is to make sure that cell is never routed to "skip."
- QE is reliable at flagging genuine struggle at the bottom of the distribution and at ordering segments within one homogeneous content type and engine. It is structurally blind to a fluent Critical (dropped negation, flipped dosage, inverted instruction), to approved-terminology violations, and to fluent locale and number errors, because all three require comparing the output against something external that a reference-free metric cannot read.
- The one dimension QE handles well, fluency, is the one that matters least to consequence; the three it handles poorly, accuracy, terminology, and locale, are the three that fail files. This is the definitional boundary between predicting quality from the surface and judging it against meaning, not a defect a better model will fix.
- The deliverable is a severity-scored MQM / ISO 5060:2024 evaluation with a Critical count of zero, plus the calibration note that justifies the routing, never a QE average. Tier, calibrate, set bands from the calibration, route by the cross-product, check what the score is blind to, and record the evaluation: that sequence is the defensible operating procedure for any QE-scored file.
Skill.re