โ†
AI for Translation & Localization
Strategic ยท M18 ยท lesson 18 of 20 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Standing Up an Evaluation Program
๐Ÿ“–
now learning

Standing Up an Evaluation Program

15 min

The evaluation program collapsed on a Tuesday, and it collapsed because Marta took her daughter to the pediatrician. A localization operation running six languages for a medical-device account had exactly one person the leadership trusted to score quality: Marta, a senior reviewer with fourteen years in regulated content, who could look at a post-edited German file and tell you in ninety seconds whether it was safe to ship. For two years she had been the quality function. Not a member of it, the whole of it. Files routed through her, she scored them in her head against a rubric that also lived mostly in her head, and the delivery record read "reviewed and approved, M.K." Then her daughter got sick, Marta took three weeks, and the operation discovered that it did not have an evaluation program. It had Marta. In her absence, two other linguists tried to score the same files and produced numbers eleven points apart on a per-thousand-word scale where the client's ceiling was fifteen. One passed a file the other failed. A batch of insulin-pump instructions shipped without a real evaluation because nobody else knew what "real" meant, and the client's own scoring came back with a Critical the operation had never caught. The head of localization sat in the post-incident review and said the sentence that names this entire lesson: our quality is a person, and that person is a single point of failure. This lesson is about the opposite of Marta: building an evaluation program, meaning the organizational function that makes quality scoring a reproducible system run by qualified, calibrated, interchangeable people, so that the number a file receives does not depend on who happened to open it.

From a Hero to a Function

Almost every localization operation starts with a Marta. Quality begins as a person, because a person is the cheapest way to get quality when you have one client and one skilled reviewer. It works right up until it does not, and the moment it stops working is always the same moment: when the operation scales past the one person, or when the one person is unavailable, or when a client asks a question the person's private judgment cannot answer with evidence. The failure is not that the hero was bad at the job. Marta was excellent. The failure is that excellence lodged in one head is not a program, it is a liability wearing a competence costume, and it fails silently until the day it fails catastrophically.

Let me fix the vocabulary, because this lesson assembles an organizational machine and the parts need exact names. An evaluation program is the organizational function that produces reproducible quality scores across people, files, and time: the qualified evaluators, the rubric they score against, the calibration that keeps them aligned, the sampling strategy that decides what gets scored, the tooling that computes the numbers, the reporting that turns numbers into decisions, and the continuous-improvement loop that keeps the whole thing honest. It is a system, not a step. Machine translation (MT) is any system that renders text from one language to another with no human writing the words; a large language model (LLM) is a general text predictor that translates as a side effect and is fluent even when wrong; machine-translation post-editing (MTPE), or PE, is the workflow where a human edits machine output. MQM is Multidimensional Quality Metrics, the analytic error-typology framework that classifies each error by dimension (what kind of error) and severity (how much it matters). ISO 5060:2024 is the international standard that formalizes the MQM-aligned human analytic evaluation of translation output: the Critical, Major, Minor severity bands and the normalized error score. ISO 18587 is the post-editing standard, in revision toward publication in 2026, that expands to cover AI and LLM output and insists the post-editor hold full professional-translator competence. Inter-rater agreement, also called inter-rater reliability, is the degree to which independent evaluators scoring the same content arrive at the same result: the measurable property that tells you whether your program is a system or a collection of individuals. Calibration is the recurring exercise that produces and maintains that agreement. A rubric is the written rule set that fixes every scoring decision the program does not leave to per-error judgment. Use these precisely from here on.

The distinction between the previous lesson and this one is the distinction between a skill and an institution. You already know the six-step scoring workflow: sample, mark, assign severity, weight, normalize, output. That is what one qualified evaluator does to one file. This lesson builds the organization around that workflow: the function that ensures the operation has many qualified evaluators instead of one, that they all score the same way, that their agreement is measured and defended, and that the program improves rather than decays. A workflow is what a person runs. A program is what survives that person leaving.

Quality that lives in one skilled person is not a program, it is a single point of failure wearing a competence costume. An evaluation program is the function that makes the score independent of who opened the file.

Why One Skilled Evaluator Is a Single Point of Failure

It is worth staying with the failure mode, because leadership routinely underrates it. A single expert evaluator fails in four distinct ways, and the operation running on Marta was exposed to all four. The first is availability: the person gets sick, takes leave, resigns, or is simply busy on another account, and quality scoring stops or degrades to whoever is nearest, exactly the eleven-point divergence the story opens on. The second is non-reproducibility: even when the expert is present, their score is a private judgment no one else can replicate, so when a client produces a different number there is no shared basis to reconcile the dispute, only two assertions of expertise. The third is unfalsifiable drift: a lone evaluator's standard shifts slowly over months, tightening or loosening with mood, fatigue, and the last file they scored, and because there is no second evaluator and no calibration, nobody can see the drift or correct it. The fourth is concentration risk in the audit: when a certifier or a regulated client asks "show me that your evaluation is systematic and repeatable," an operation whose evaluation is one person's judgment has nothing to show, because a system that produces different numbers in different hands, or works in only one pair of hands, certifies nothing. The revised ISO 18587 and ISO 5060 both assume a documented, repeatable procedure precisely because a standard met only by a specific individual is not a standard. Building the program is how you convert a hero into an institution, and the conversion is the strategist's job, not the evaluator's.

The Six Organs of an Evaluation Program

An evaluation program has six components, and they are organs rather than steps because they run continuously and depend on each other, not in a line. Naming them fixes the shape of the thing you are building, so that "stand up an evaluation program" stops being a vague mandate and becomes six concrete pieces you can staff, fund, and audit. The rest of the lesson takes each in turn, but here is the whole body first.

  • Evaluator qualification and training: how a person becomes an authorized evaluator, what competence they must prove, and how new evaluators are onboarded, so that "qualified" is a demonstrated status and not a job title.
  • The rubric and the scoring profile: the written, versioned rule set (dimensions, severity definitions, weights, normalization, thresholds, client-specific overrides) that every evaluator scores against, so the rules live in a document and not in a head.
  • Calibration and inter-rater agreement: the recurring exercise that keeps evaluators aligned and the metric that proves they are, so the program can measure and defend its own reproducibility.
  • Sampling strategy: the program-level policy for what gets scored, how much, and how selected across content, clients, and risk tiers, so evaluation effort is spent where consequence lives.
  • Tooling: the environment that presents the file, captures the marked errors, computes the arithmetic, and stores the record, so the human spends judgment on errors and the machine handles everything mechanical.
  • Reporting and continuous improvement: the layer that turns individual scores into trends, diagnoses, and decisions, and feeds what it learns back into the rubric and the training, so the program gets better instead of merely running.

The order matters only loosely. A real program stands them up roughly together, because each is weak without the others: a perfect rubric with unqualified evaluators produces confident wrong scores; qualified evaluators without calibration drift apart; calibration without reporting cannot prove its value to the leadership that funds it. The strategist's work is to build all six as a system, and to resist the temptation, which every operation feels, to build only the rubric and declare victory, because the rubric is the easiest organ and the least sufficient one.

Evaluator Qualification and Training

The first organ answers the question the Marta operation never asked: what makes someone an evaluator? Not a translator, not a post-editor, but a person authorized to put a defensible number on a file. Treating "evaluator" as an informal role that any senior linguist drifts into is the root cause of the whole hero problem, because it means the operation never defined the competence it depends on, never verified it, and therefore cannot reproduce it. A program makes evaluation a qualified status: a demonstrated, revocable, documented authorization, earned by proof, not seniority.

Qualification rests on three kinds of competence, and a candidate must show all three. The first is linguistic and domain competence: the evaluator must have the full professional-translator competence in the relevant language pair that the revised ISO 18587 now demands of the post-editor too, plus enough command of the content domain to recognize a subtle terminology or accuracy error a generalist would miss. A brilliant generalist who cannot tell that a rendered drug name is the wrong salt form is not qualified to evaluate medical content, however fluent their scoring elsewhere. The second is rubric competence: the evaluator must know the program's dimensions, severity definitions, weights, normalization, and thresholds cold, and apply them as written rather than as felt. The third, and the one most often missing, is calibration competence: the evaluator must be able to score in agreement with the rest of the program, which is a different and harder skill than scoring correctly in isolation. An evaluator can be linguistically excellent and rubric-literate and still be miscalibrated, systematically stricter or more lenient than their peers, and until they are calibrated they are a source of variance, not a source of quality.

The Qualification Gate, Not the Qualification Assumption

The program authorizes an evaluator through a qualification gate: a concrete, passable, documented step, not an assumption based on tenure. The standard shape is a gold-standard test set: a set of files that experienced evaluators have already scored to an agreed reference, with the errors pre-marked and the severities pre-agreed. The candidate scores the same files cold, and their scoring is compared to the reference. Passing is not "found roughly the same problems"; it is a measured threshold: the candidate's error identifications, severities, and final normalized scores must fall within a defined tolerance of the reference, for example agreeing on every Critical, on at least a set fraction of Majors, and landing the file score within a small band of the reference score. A candidate who misses a Critical fails outright, because a missed Critical is the exact failure the program exists to prevent, and an evaluator who does not reliably catch Criticals cannot be authorized regardless of how elegantly they score the rest.

Two disciplines make the gate real rather than ceremonial. First, qualification is scoped: a person is qualified for specific language pairs, content domains, and risk tiers, not for "evaluation" in the abstract. The evaluator authorized for French marketing is not thereby authorized for French pharmaceutical labels; the gate is passed per scope, because the competence is per scope. Second, qualification is revocable and renewed: it is not a permanent badge but a status maintained by continued calibration performance, and an evaluator who drifts out of agreement on calibration rounds loses authorization until they re-qualify. This is what turns the Marta situation into its opposite: instead of one irreplaceable expert, the program deliberately qualifies a pool, so that any file in a given scope can be scored by any qualified evaluator in that scope, and the loss of any one evaluator degrades capacity rather than eliminating the function. The strategic instruction is blunt: never let the number of qualified evaluators for a live scope fall to one. One qualified evaluator is a Marta waiting to happen.

Qualification is a demonstrated, scoped, revocable status proven against a gold-standard test set, not a seniority badge. Missing a single Critical fails the gate, and no live scope should ever run on one qualified evaluator.

Training as the On-Ramp to Qualification

The gate presumes a way in, and that is training. New evaluators are not thrown at the test set and expected to pass; the program runs an onboarding path that teaches the rubric, walks the dimensions and severities with worked examples, has the candidate score practice files under supervision with immediate feedback against a mentor's scoring, and only then presents the qualification test. The most valuable training material is the program's own history of resolved calibration disputes: the real ambiguous cases the team argued to a rule, which teach the candidate not the clean cases they will get right anyway but the hard middle where evaluators diverge. Training is also where the program transmits its culture of evidence: that an error is not scored until it is marked with its span and note, that the profile overrides instinct where it speaks, that a Critical is a verdict and not a tally entry. An operation that skips training and qualifies people by reputation is rebuilding the hero problem with more heroes, each carrying a slightly different private standard, which is not progress but the same disease at larger scale.

The Rubric, Calibration, and Inter-Rater Agreement

The next two organs work as a pair, and it helps to see them together. The rubric is the written rule set that makes many evaluators score the same way; calibration is the recurring exercise that keeps them scoring that way over time, and inter-rater agreement is the metric that proves it. A rubric without calibration is a shared document that evaluators slowly drift away from; calibration without a rubric has nothing to converge on. Together they are what turn a pool of qualified individuals into one reproducible function.

The Rubric and the Scoring Profile as Shared Infrastructure

The second organ is the written rule set, and at program level its purpose shifts from the previous lesson's. There, the rubric made one evaluator's scoring reproducible over time. Here, the rubric is shared infrastructure: the single artifact that makes many evaluators score the same way, the thing that converts a pool of individuals into a function. Marta's fatal characteristic was that her rubric lived in her head, so it could not be shared, checked, taught, or defended. The program's rubric lives in a document, is versioned, and is the source of truth every evaluator reads before scoring and refers to during it.

The program rubric carries the same core the workflow lesson defined: the four MQM dimensions (accuracy, terminology, locale conventions, fluency), the three severity bands with the three-question test (does it merely annoy, does it mislead, does it mislead and harm), the weights per band (classically 1 for Minor, 5 for Major, a large penalty commonly 25 for Critical), the normalization to an error score per thousand words, and the pass threshold. What the program adds is the machinery that lets one rubric serve many clients and many evaluators without fragmenting into private variants.

Client Scoring Profiles as Overlays on a Common Core

Different clients need different rules, and the naive response is to give each client its own rubric, which multiplies the number of rule sets the program must maintain and the evaluators must remember until consistency is impossible. The disciplined response is a common core plus client profiles: one base rubric that defines the dimensions, the severity logic, the arithmetic, and the evidence discipline, and a thin per-client scoring profile that overrides only what that client needs, its weights, its threshold, and its declared severity rules ("any termbase violation is Major minimum," "any accuracy error in safety content is Critical"). An evaluator moving between accounts learns one core and a small overlay per client, not a dozen incompatible rubrics. The profile is versioned and referenced in every score output, so a score can always be traced to the exact rules that produced it, which is what lets an evaluator score a file the same way a colleague did last month and the same way the program will defend it next year.

The profile does more organizational work than it appears to. Every rule a client pre-declares in the profile is one fewer judgment call that can vary between evaluators, which means the profile is not just a client-service convenience but a reproducibility instrument: it converts the most-disputed categories, the ones where evaluators diverge most, into lookups the whole pool applies identically. The strategist's move is to push as much as the client will agree to into the profile, because every declared rule shrinks the surface of irreducible judgment and every shrinkage of that surface raises inter-rater agreement. The rubric core plus the client profiles is the program's constitution: written down, versioned, shared, and binding, the opposite of a standard that lived in one head and died when that head was on leave.

Calibration: The Alignment Engine

The third organ is the one that most directly kills the hero problem, and it is the one operations most often skip because it costs evaluator time and produces no billable deliverable. Calibration is the recurring exercise where the program's evaluators independently score the same sample, then compare and reconcile their scores. Its purpose is not to catch bad evaluators, though it does that; its purpose is to keep good evaluators aligned, because even excellent evaluators drift apart without it, each one's standard hardening in a slightly different direction until two of them score the same file ten points apart and neither can say who is right. Calibration is the maintenance that keeps a pool of qualified people scoring as one function instead of as several individuals who happen to share a rubric.

The exercise is mechanical in structure and rich in what it surfaces. The program selects a calibration sample, ideally one carrying the ambiguous cases where divergence lives, not the clean cases everyone gets right. Each evaluator scores it independently and blind, without seeing the others' scores. The program then computes inter-rater agreement: how closely the independent scores match, at the level of error identification (did they find the same errors), severity assignment (did they band them the same), and final normalized score (did they land the file in the same place). Then the team convenes and reconciles: the clean agreements pass without comment, and the disagreements are argued to a shared resolution that is written back into the rubric. The clean cases teach nothing; the value is entirely in the ambiguous middle, the term drift that might be Minor or Major, the date error whose harm depends on what the date governs, and calibration is where those get argued to a rule the whole pool then applies.

Measuring Agreement as a Program Metric

Inter-rater agreement is not just a diagnostic the team looks at once; it is a program-level metric the operation tracks over time and reports, and this is where calibration earns its keep with the leadership that funds it. A program that measures agreement can make a claim no hero operation can make: "our qualified evaluators agree within two points on calibration samples, and agree on 100% of Criticals." That claim is exactly the evidence an ISO 5060 conformance audit looks for and exactly the reassurance a regulated client wants, because it demonstrates that the score a file receives does not depend on which evaluator opened it. Tracking agreement over time also reveals the health of the program: rising agreement means the rubric and calibration are working; falling agreement means the rules have gotten ambiguous, a new evaluator is miscalibrated, or the content has shifted into territory the rubric does not cover, each of which is a signal to act.

Two failure patterns in the agreement data mean different things and demand different responses. Disagreement on clean cases, where evaluators split on an obvious Critical or an obvious Minor, is an alarm: it means someone misunderstands the severity bands and needs retraining or re-qualification, because the clean cases are the ones the program cannot afford to get wrong. Disagreement on ambiguous cases is normal and healthy: it is the program discovering a rule it has not yet written, and the correct response is to argue the case to a resolution and update the rubric, not to defer to the most senior evaluator in the room. The strategic point is that calibration converts disagreement from a hidden liability into visible, resolvable information. The Marta operation had maximal disagreement, an eleven-point spread, and no way to see it, because it never calibrated; the moment two evaluators were forced to score the same file, the disagreement that had always existed simply became undeniable. A program calibrates on purpose and continuously, so the disagreement surfaces in a low-stakes exercise rather than in a client's audit.

Calibration is the maintenance that keeps qualified evaluators scoring as one function. Inter-rater agreement is the metric that proves it, and tracking it over time is the evidence a conformance audit and a regulated client both demand.

Sampling, Tooling, and the Reporting Loop

The final three organs decide what the program spends its scarce evaluation effort on, what it automates so evaluators do not diverge on mechanics, and what it does with the numbers once they exist. Sampling allocates the effort, tooling removes the drift, and reporting closes the loop that keeps the whole program improving instead of decaying. Treated together, they are the difference between a program that scores files and a program that manages quality.

Sampling Strategy at the Program Level

The fourth organ is sampling, and at the program level it is a strategy rather than a per-file rule. One evaluator scoring one file needs a sampling rule; a program scoring thousands of files across many clients, languages, content types, and risk tiers needs a sampling strategy: a policy that allocates the operation's finite evaluation capacity across everything competing for it, so that scoring effort concentrates where consequence lives and does not evaporate on low-stakes content. Sampling is where the economics of the program are decided, because evaluation costs real linguist time, and a program that tries to score everything at 100% destroys the MT economics that justified the pipeline, while a program that samples too thin misses the errors it exists to catch.

The strategy fixes several things as policy, once, so that no individual file's sampling is improvised. It sets the baseline sampling rate for ordinary content, a defensible shape such as "the greater of 1,000 words or 10% of the delivery, random stratified." It sets the risk-tier overrides: high-liability content (dosages, contraindications, legal obligations, financial figures) is scored at 100% and never sampled, because there is no acceptable amount of effort to save on a drug label, while low-stakes content may be sampled lightly or scored on a rotating basis. It sets the stratification dimensions: the sample must spread across the linguists who worked the file, the content types present, and the risk tiers, so no slice escapes evaluation by luck. And it sets the escalation triggers: conditions under which sampling intensifies, a new engine or a new post-editor on an account, a recent failed batch, a client complaint, each of which should raise the sampling rate until confidence is re-established.

Sampling as Risk Allocation, Not a Fixed Quota

The mistake a program makes here is treating sampling as a fixed quota, the same percentage everywhere, which is either wasteful on safe content or negligent on dangerous content. Sampling is a risk-allocation decision: it deliberately spends more evaluation on content where a missed error costs more and less where it costs less, and it adjusts dynamically as evidence accumulates. An account with a long history of clean batches from a proven post-editor and a stable engine can be sampled at the baseline; the same account after an engine change reverts to intensified sampling until the new engine has proven itself, because the risk profile changed even though the content did not. This is the same risk-tiered thinking that governs intake, applied to evaluation: match the intensity of the check to the consequence of missing, and let the intensity move with the evidence. A program that samples by a flat rule is leaving both money and safety on the table; a program that samples by risk spends its evaluation budget where it buys the most protection.

Tooling and the Instrumented Program

The fifth organ is tooling, and its job is to remove from the evaluator every task that is not irreducible human judgment. The workflow lesson established that automating the weighting and normalization eliminates the denominator mistake and the multiplication slip; at program level, tooling does more, because it is the layer that enforces the rubric, captures the record, and produces the data the reporting organ needs. The strategic principle is that repeatability is highest when the program asks the evaluator for the smallest possible number of irreducible judgments, the genuine mislead-and-harm calls, and handles everything else, sampling selection, profile-rule application, weighting, normalization, record capture, by mechanism. Human judgment is the scarce, valuable, variable input; the tooling exists to spend it only where it is irreplaceable.

Good evaluation tooling does specific things that directly raise inter-rater agreement and cut the cost of the program. It presents the sample already drawn by the sampling strategy, so the evaluator does not choose what to score and cannot steer toward flattering passages. It enforces the marking discipline, requiring a dimension, a source span, a target span, and a note for every error before it can be recorded, so no error enters the record as a bare assertion. It applies the client profile automatically, flagging where a declared severity rule governs so the evaluator applies the rule rather than their instinct. It computes the arithmetic, so the penalty total, the normalization against the recorded scored word count, and the Critical-rule verdict are never typed by a human. And it stores the complete record, every error with its spans and notes, the sample definition, the arithmetic, the profile version, so the output is an auditable artifact by construction rather than by a diligent evaluator's memory. Tooling that does these things does not just speed the program up; it raises agreement by removing the many small places where evaluators would otherwise diverge, and it produces the record the reporting layer and the audit both depend on.

Reporting and Continuous Improvement

The sixth organ is what turns a scoring operation into a program that learns, and it is the organ most operations neglect entirely, running the scoring and letting the numbers pile up unread. Reporting is the layer that aggregates individual scores into the information leadership and clients actually need, and continuous improvement is the loop that feeds what the reporting reveals back into the other five organs. Without this organ the program produces verdicts and nothing else; with it, the program produces management information, audit evidence, and its own steady improvement.

Reporting operates at three altitudes, and a mature program produces all three. At the file level it produces the score output the workflow lesson defined: the verdict, the Critical count, the error list, the arithmetic, the profile version, the artifact that defends a single file in a single dispute. At the account and program level it produces trends and diagnoses: the per-thousand-word score over time per client and per engine, the dimension breakdown that shows whether failures cluster in accuracy, terminology, locale, or fluency, the Critical-error rate as its own tracked line because Criticals are the metric that matters most, and the inter-rater agreement that proves the program's own reproducibility. At the leadership level it produces the dual-axis story the strategist owes the executive who funds the program: throughput and cost on one axis, quality-risk (the Critical rate, the trend, the conformance posture) on the other, so that quality is visible as a managed risk and not an invisible assumption. The reporting layer is what lets a strategist walk into a leadership review and say the program is working, with evidence, rather than asserting it, which is the difference between a program that gets funded and one that gets cut in the next budget round.

The Improvement Loop That Closes the Program

Reporting that no one acts on is decoration. The continuous-improvement loop is what makes the reporting load-bearing: it reads the trends and the diagnoses and feeds them back into the other organs. A dimension breakdown showing terminology errors running at ten times the accuracy rate routes a fix to the glossary and the engine grounding, not to post-editor retraining, because the data localized the problem. A drop in inter-rater agreement routes a fix to calibration and the rubric, surfacing the ambiguous cases that need a new written rule. A pattern of a specific evaluator diverging routes a fix to training and re-qualification. A recurring Critical type, dropped negations in safety warnings, say, routes a fix upstream to the post-editing checklist and the intake routing, so the error is prevented rather than merely caught. The loop is what distinguishes a program that improves from one that merely runs: the same scores that gate this week's files become, in aggregate, the evidence that reshapes next quarter's rubric, training, sampling, and even the engines the operation chooses. A program without the loop decays, because rubrics go stale, evaluators drift, and content shifts; a program with the loop compounds, because every file it scores makes the next batch of scoring a little more accurate and a little more defensible.

Reporting turns scores into decisions; the improvement loop turns decisions into a better program. Without the loop a program decays as rubrics stale and evaluators drift. With it, every file scored makes the next score more defensible.

A Worked Program Setup: From Marta to a Function in Ninety Days

Abstract organs become a program only when you stand them up on a concrete operation, so take the Marta operation and build its evaluation function across a first quarter, one organ at a time but all six by the end. The operation runs six languages for a medical-device account, roughly forty thousand words a week of instructions, labeling-adjacent content, and UI strings, machine-translated first and post-edited by a pool of linguists, under a client contract that specifies a per-thousand-word error score not exceeding fifteen with zero Criticals, scored to ISO 5060.

Weeks 1 to 3: rubric and profile. The strategist writes the base rubric first, because it is the artifact everything else references: the four MQM dimensions with examples drawn from the account's own past errors, the three severity bands with the three-question test, the classical weights (Minor 1, Major 5, Critical 25), normalization to an error score per thousand words, and the pass threshold of 15 from the contract. On top of it sits the client profile: the declared rules the medical account needs, "any accuracy error in safety content is Critical," "any termbase violation is Major minimum," "any dropped or added negation in an instruction is Critical," plus the absolute zero-Criticals gate. Marta's private standard, extracted from her head in a week of structured interviews before she is allowed to be the only one who holds it, becomes the seed of the written rubric. The rubric is versioned as v1.0 and every future score references it.

Weeks 3 to 6: qualification and training. The strategist builds a gold-standard test set: twenty files from the account's history, scored to an agreed reference by Marta and one external senior reviewer scoring independently and reconciling, so the reference itself is calibrated and not just Marta's opinion. The training path is written from the rubric and the account's resolved-dispute history. Then the pool is put through it: the linguists who tried and failed to score during Marta's leave, plus two new hires, train on the rubric and the worked examples, score supervised practice files, and finally take the qualification test cold. The gate is explicit: agree on 100% of Criticals in the test set, at least 80% of Majors, and land each file's normalized score within three points of the reference. Of six candidates, four pass on the first attempt; two miss Criticals and are sent back to training before re-testing. The operation now has, for the first time, a scoped, qualified pool instead of a single hero, and no live scope runs on one evaluator.

Weeks 6 to 8: sampling strategy and tooling. The sampling strategy is fixed as policy: baseline of the greater of 1,000 words or 10%, random stratified across the six languages and the post-editors, with 100% coverage of every safety-tier segment, and an escalation trigger that intensifies sampling whenever a new post-editor or a new engine version enters an account. The tooling is configured to draw the sample automatically, enforce the marking discipline (dimension, source span, target span, note per error), apply the profile's declared severity rules with a flag, compute the penalty, normalize against the recorded scored word count, fire the Critical rule, and store the complete record. The denominator mistake and the improvised weight become structurally impossible.

Weeks 8 onward: calibration and reporting stand up as the recurring engine. Monthly calibration begins: all qualified evaluators score a shared sample carrying the account's ambiguous cases, blind, then reconcile, and the program computes inter-rater agreement and writes the resolutions back into the rubric as v1.1, v1.2, and so on. The first calibration round exposes exactly the disagreement that ambushed the operation during Marta's leave, an eleven-point spread on a term-drift case, but now it surfaces in a low-stakes exercise, gets argued to a written rule, and never recurs. Reporting stands up at all three altitudes: file-level score outputs on every batch, account-level trends and dimension breakdowns and a tracked Critical rate and inter-rater agreement, and a leadership dashboard pairing throughput and cost against the quality-risk line. The improvement loop closes when the first month's dimension data shows terminology errors dominating, which routes a glossary-and-grounding fix to the engine rather than a futile round of post-editor scolding, and the terminology rate falls the following month.

What the Operation Can Now Say That It Could Not Before

The measure of the program is the sentence the operation can now say to the client and to an auditor, the sentence it could not say when quality was Marta. Before: "Marta reviewed it." After: "This batch was scored by a qualified evaluator against rubric v1.3 and the account's profile v2.1; the sample was a 1,000-word stratified draw with 100% safety coverage; the normalized score was 6.0 against a threshold of 15 with zero Criticals; our evaluators agree within two points on calibration and on 100% of Criticals; here is the error record, the arithmetic, and the trend over the last quarter." That sentence does not depend on any one person. Marta can take three weeks, resign, or move to another account, and the number a file receives is unchanged, because quality stopped being a person and became a function. The head of localization who said "our quality is a single point of failure" in the post-incident review can now say the opposite, and can prove it, which is the entire purpose of standing up an evaluation program: not to replace the hero, but to make the operation no longer need one.

Key Takeaways

  • A single skilled evaluator is a single point of failure in four ways: availability (they leave or fall sick and scoring stops), non-reproducibility (their private judgment cannot be replicated when a client disputes it), unfalsifiable drift (their standard shifts and no one can see it), and audit concentration risk (a system that works in only one pair of hands certifies nothing). An evaluation program converts the hero into a function so the score does not depend on who opened the file.
  • An evaluation program has six interdependent organs, not steps: evaluator qualification and training, the rubric and scoring profile, calibration and inter-rater agreement, sampling strategy, tooling, and reporting with continuous improvement. Building only the rubric, the easiest organ, and declaring victory rebuilds the hero problem at larger scale.
  • Evaluator qualification is a demonstrated, scoped, revocable status proven against a gold-standard test set, not a seniority badge. Competence has three parts (linguistic/domain, rubric, and calibration), a missed Critical fails the gate outright, and no live scope should ever run on one qualified evaluator, because that is a Marta waiting to happen.
  • The rubric at program level is shared infrastructure: a common core (MQM dimensions, severity bands, weights, normalization, threshold) plus thin versioned client profiles that override only what each client needs. Every rule pushed into the profile is a reproducibility instrument, converting the most-disputed judgment calls into lookups the whole pool applies identically.
  • Calibration is the recurring exercise that keeps qualified evaluators aligned; inter-rater agreement is the tracked metric that proves it. The value lives in the ambiguous middle, not the clean cases: disagreement on clean cases signals someone needs retraining, disagreement on ambiguous cases is the program discovering a rule to write and version back into the rubric.
  • Sampling at program level is a risk-allocation strategy, not a flat quota: a defensible baseline (the greater of 1,000 words or 10%, random stratified), 100% coverage of high-liability content, stratification across linguists and content types and risk tiers, and escalation triggers that intensify sampling when an engine or post-editor changes or a batch fails.
  • Tooling exists to ask the evaluator for the smallest possible number of irreducible judgments and handle everything else by mechanism: it presents the drawn sample, enforces the marking discipline, applies profile rules, computes the arithmetic and the Critical verdict, and stores the complete auditable record, raising inter-rater agreement by removing the small places where evaluators would otherwise diverge.
  • Reporting turns scores into decisions at three altitudes (file, account/program, leadership), and the continuous-improvement loop feeds diagnoses back into the right organ: a terminology cluster to the glossary and engine, an agreement drop to calibration and the rubric, a recurring Critical type upstream to post-editing and intake. A program with the loop compounds; a program without it decays as rubrics stale and evaluators drift.