AI Terminology Every Linguist Should Know
The kickoff call was twelve minutes old and Nina had already lost the thread. She had translated medical and technical content for fifteen years, and she was very good at it, but the program manager on the screen was speaking a language she only half recognized. "So the source comes in through the TMS, we run NMT pre-translation, the QE flags route anything under the threshold to full MTPE, the rest gets light PE, the termbase is locked so watch for term drift, the placeholders are non-translatable so leave the tags alone, and we score the deliverable against 5060 before sign-off. Locale is fr-CA, not fr-FR, and the i18n team already extracted the strings so you should not see any concatenation." Someone typed "+1 on the MQM gate" in the chat. Nina wrote down four acronyms, realized she was not sure what two of them meant, and felt the specific, quiet panic of a skilled professional who suddenly cannot tell whether she is being asked to do her job or a different one. None of these words are difficult. Every one of them describes something she already does or already knows. The problem is that nobody had ever defined them for her in the language of the file, in working terms, with a plain answer to the only question that matters to a person on a deadline: what does this mean for me, right now, in this segment. This lesson is that glossary. Not the textbook definitions, the working ones, grouped the way the work is actually grouped, so that the next time the acronyms fly you are not writing them down in fear, you are nodding along because you own them.
How to Read This Glossary
A glossary that just lists words alphabetically is a dictionary, and a dictionary is useless on a deadline. The terms in a localization pipeline are not independent. They cluster, because the work clusters. There is a group of words about the workspace you sit inside, a group about editing a machine's output, a group about the meaning units you move through, a group about terminology, a group about quality and scoring, and a group about the technical container the words live in. We are going to take them one cluster at a time, in roughly the order content moves through a shop, and for each term you get three things: a plain definition in the language of someone who post-edits for a living, a concrete example from a real file, and a single line that begins "Why it matters to you" so you never have to guess at the stakes.
Two ground rules before we start, because they prevent the two most common confusions. First, an acronym is not a concept; it is a label stuck on a concept, and the label is only as useful as your grip on the thing underneath. We will always define the thing, then the label. Second, many of these words have a textbook meaning and a shop meaning, and they are not always the same. When a program manager says "QE," they do not mean the academic field of quality estimation in the abstract. They mean the number on your screen that decides whether you read this segment carefully or skim it. We will always give you the shop meaning, because that is the one that is true on the call.
An acronym is a label stuck on a concept. Learn the concept first, then the label will never frighten you again.
The Workspace: Where the Words Live
Before any of the verbs, you need the nouns for the rooms you work in. Three tools form the physical and digital workspace of modern localization, and almost every other term is something that happens inside one of them. Get these three solid and the rest has somewhere to attach.
CAT Tool
Plain definition. CAT stands for computer-assisted translation. A CAT tool is the editing environment a linguist actually works in: a screen split into source segments on one side and target segments on the other, with a translation memory, a termbase, and now machine-translation output all wired in to help fill the target side. It is the cockpit. If you translate or post-edit professionally, you spend your day inside one. Crucially, "computer-assisted" does not mean "the computer translates." It means the computer organizes the work, remembers your past translations, surfaces approved terms, and increasingly pre-fills suggestions, while the human makes the calls.
Example. You open a file and see two columns. Left column, German: "Das Gerat darf nicht bei Patienten mit aktivem Schrittmacher verwendet werden." Right column, an editable French field where a suggestion already sits. The grid around you, the match percentages down the margin, the term highlights, the keyboard shortcut that confirms a segment: that whole apparatus is the CAT tool.
Why it matters to you: the CAT tool is where every other term in this lesson becomes a thing you click, so knowing its parts is knowing where to look when something goes wrong.
TMS
Plain definition. TMS stands for translation-management system. If the CAT tool is the cockpit one linguist sits in, the TMS is the air-traffic control above the whole airport. It is the platform that receives the client's files, breaks projects into jobs, assigns them to linguists, tracks deadlines and status, applies the translation memory and termbase, runs machine-translation pre-translation, and pushes the finished work back out. The pre-filled segments that now greet you on every file were put there by the TMS before you ever opened the job.
Example. The program manager's sentence "the source comes in through the TMS, we run NMT pre-translation" describes the TMS doing its job: ingesting the German manual, running it through a machine-translation engine so every segment arrives populated, and routing the result to you. You never touched the TMS directly, but it shaped the file before you saw it.
Why it matters to you: the TMS decides what state your file is in when it lands, so when a file arrives pre-translated, term-locked, or split oddly, the explanation is almost always something the TMS did.
TM (Translation Memory)
Plain definition. TM stands for translation memory. It is a database of your own (and your team's) past translations, stored as pairs: this source segment was translated as this target segment, approved, and saved. The next time an identical or similar source segment appears, the CAT tool offers you the stored translation so you do not redo work you have already done. A TM is a record of human decisions. This is the single most important distinction to hold onto in this entire lesson: a TM is not machine translation. A TM replays something a human already approved. Machine translation generates something new on the spot.
Example. Three months ago you translated "Charge the battery before first use" into French and approved it. Today the same English string appears in a new file. The TM surfaces your old French translation instantly, marked as a 100% match. You did not retranslate; you reused a verified human decision.
Why it matters to you: a clean TM compounds your quality and your speed across years, while a polluted TM propagates a single past error into every future file that matches it, so what you save into the TM matters as much as what you ship.
Segments and Matches: The Units You Move Through
Localization does not happen in documents or paragraphs. It happens in segments, and the relationship between a new segment and the segments already in your translation memory is described by a small, precise vocabulary of "matches." This is the cluster that, once it clicks, makes the whole margin of your CAT tool suddenly legible.
Segment
Plain definition. A segment is the unit of translation, usually a sentence, sometimes a heading, a list item, a cell, or a UI string. When a file enters the CAT tool, it is automatically broken into segments by a process called segmentation, and you work through them one at a time, confirming each before moving on. The segment, not the document, is the atom of the work. Everything else, the matches, the scores, the errors, attaches to a segment.
Example. A 40,000-word manual is not "a translation job" to your CAT tool. It is roughly 2,500 segments. The contraindication that nearly shipped in Nina's story was not "a paragraph." It was segment number 4, a single sentence, scored and confirmed on its own.
Why it matters to you: because every quality decision, every match, and every error is scoped to a single segment, "the file looks fine" is never a real judgment; the only real judgments happen segment by segment.
Exact Match (and 100% / ICE)
Plain definition. An exact match means the source segment in front of you is character-for-character identical to a segment already stored in the translation memory, so the TM offers its stored translation as a perfect, 100% match. A stricter version, sometimes called an ICE match (in-context exact), means it is identical and the surrounding segments are identical too, so the context is the same and the stored translation is even safer to trust.
Example. Your TM holds "Click Save to continue." The new file contains "Click Save to continue," spelled and punctuated identically. That is an exact 100% match, and the CAT tool pre-fills the stored French without you typing a word.
Why it matters to you: exact matches are usually charged at the lowest rate and trusted the most, but "identical source" does not guarantee "correct in this context," so a 100% match still deserves a glance when meaning depends on surroundings.
Fuzzy Match
Plain definition. A fuzzy match is a partial match: the new source segment is similar but not identical to one in the translation memory, and the CAT tool reports how similar as a percentage, typically anywhere from 50% to 99%. A 95% fuzzy match means almost everything is the same and one small thing changed; a 75% fuzzy match means a meaningful chunk is different. The tool shows you the stored translation with the differences highlighted so you can adapt rather than retranslate.
Example. Your TM holds "Charge the battery before first use." The new segment is "Charge the battery fully before first use." That is perhaps a 90% fuzzy match. The tool offers your old French and highlights that "fully" is the new piece, so you insert one adverb instead of building the sentence again.
Why it matters to you: fuzzy matches are where money and danger hide, because a high fuzzy percentage tempts you to accept the stored translation while the one changed word, a "not," a number, a negation, is exactly the part that flips the meaning.
The Editing Verbs: Post-Editing and Its Cousins
Now the verbs. This cluster describes what you actually do to machine output, and it is the cluster clients argue about most, because each level of effort carries a different price. Getting these precise is not pedantry; it is the difference between being paid fairly and being asked to do full translation at a light-editing rate.
MT, NMT, LLM (a Quick Recap)
Plain definition. MT is machine translation, the umbrella term for any system that turns source text into target text without a human writing the words. NMT is neural machine translation, the narrow, translation-only neural engine that runs most high-volume pipelines: you give it a source segment, it returns a target segment, and that is all it does. LLM is a large language model, a general-purpose text predictor trained on broad human writing that can translate as a side effect of predicting plausible next words. NMT is a specialist that drifts off the source in predictable ways; an LLM is a generalist that is often more fluent and more confidently, creatively wrong.
Example. "Run NMT pre-translation" on the call meant a dedicated translation engine filled every segment. If the same shop said "we used an LLM to draft the marketing taglines," they meant a general model improvised creative copy, a different machine doing a different job.
Why it matters to you: the engine that produced your first draft predicts how it will fail, so knowing whether you are editing NMT output or LLM output tells you what kind of error to hunt for.
Generative Drafting
Plain definition. Generative drafting is using an LLM to produce content rather than to translate a source: drafting copy from a brief, expanding a bullet into a paragraph, inventing a tagline, summarizing. There is no source segment to be faithful to, so the measure of quality is not accuracy but appropriateness. It is creative writing assistance wearing the same interface as translation, which is exactly why it gets dangerously confused with translation.
Example. "Translate this product description into Spanish" is translation: there is a right answer the output must match. "Write three punchy Spanish taglines for this product" is generative drafting: there is no source, only a goal. An engine asked to "translate this and make it catchier" is being asked to do both at once, and it will resolve the conflict by inventing.
Why it matters to you: the question you ask of the output changes completely between the two, "is this faithful to the source?" versus "is this good content?", and answering the wrong question is how invented claims ship as if they were translations.
Post-Editing (PE)
Plain definition. Post-editing, abbreviated PE, is the act of editing machine-translation output into an acceptable final translation, rather than translating from a blank target. It is the defining verb of the MT-first era. When every segment arrives pre-filled, you are no longer primarily producing text; you are primarily correcting and verifying machine-produced text, which is a different and in some ways harder skill, because you must catch what is wrong in something that someone, or something, else wrote fluently.
Example. The NMT engine filled your French target. Your job is not to ignore it and start over; it is to read the German source, read the French suggestion, decide whether the French faithfully renders the German, and fix it where it does not. That act is post-editing.
Why it matters to you: your role has inverted from author to verifier, and the pay, the pace, and the failure modes all changed with it, so naming the work accurately is the first step to doing it safely.
MTPE
Plain definition. MTPE stands for machine-translation post-editing. It is the named, productized workflow in which a human post-edits MT output as the standard process, with its own rates, its own service tier, and its own quality expectations. In 2026 it is not a niche line item; it is the shape of mainstream localization work. Industry research from Nimdzi shows MTPE adoption rose from roughly 26% in 2022 to about 46% in 2024, and 81.1% of language-service providers now offer it.
Example. A client says "we want MTPE pricing on this manual." They mean: run it through MT first, then have a human post-edit, and charge accordingly, typically 50 to 75% of a full human-translation rate. The whole commercial arrangement of the file is contained in that one acronym.
Why it matters to you: MTPE is the workflow your livelihood now runs through, and the term carries a price expectation, so when a client says "MTPE" they are also saying "at a discount," which makes the next two terms, light versus full, the ones you must be able to argue.
Light PE vs. Full PE
Plain definition. These are two effort levels of post-editing. Light post-editing aims for output that is accurate and understandable but not necessarily polished: you fix meaning errors, mistranslations, and serious terminology problems, and you leave stylistic imperfections alone. Full post-editing aims for output indistinguishable from quality human translation: accurate, terminologically correct, fluent, and on-style. Light is faster and cheaper, sometimes priced as low as $0.02 per word; full costs more because it demands more. The revised ISO 18587 post-editing standard is moving away from this rigid two-way split toward an effort spectrum matched to content, but the words light and full are still everywhere on real projects.
Example. An internal knowledge-base article for employees might warrant light PE: make it correct and clear, do not polish the prose. A customer-facing legal warranty warrants full PE: every term, every nuance, every comma, at human-translation quality. Applying the wrong level is expensive in both directions, over-editing the throwaway burns budget, under-editing the warranty ships risk.
Why it matters to you: the effort level is the contract, so if a file is priced for light PE but its content demands full PE, that gap is a conversation you must have before you start, not after a Critical error ships.
Match the post-editing effort to the consequence of the content, never to the price someone wishes the content cost.
Terminology: The Words That Must Not Drift
Some words in a project are not yours to choose. The client has approved exactly one rendering for their device, their feature, their drug, their brand, and your job is to use that one and no other. This cluster is about the machinery that captures and enforces those approved words, and about the specific way machines undermine it.
Termbase / Glossary
Plain definition. A termbase (often called a glossary in casual use, though a termbase is the richer, structured database version) is the controlled list of approved terms for a project or client: the source term, its single approved target translation, and often a definition, a part of speech, a usage note, and a list of forbidden alternatives. Where the translation memory stores whole segments, the termbase stores individual terms. In the CAT tool, when a source term that lives in the termbase appears, it is highlighted and its approved translation is shown, so you do not have to remember or guess.
Example. The client makes a device they insist on calling a "cardiac monitor," never a "heart monitor," and in French "moniteur cardiaque," never "moniteur de coeur." That rule lives in the termbase. Every time "cardiac monitor" appears in the source, the CAT tool flags it and shows the one allowed French rendering.
Why it matters to you: the termbase turns the client's preferences into enforceable rules, so using the approved term is not a stylistic choice you can override, it is a requirement you can be scored against.
Term Drift (and Why the Engine Causes It)
Plain definition. Term drift is what happens when the approved term silently gets replaced by a synonym across a file, usually because a machine prefers a more common word. An NMT engine or an LLM was trained on general text where "heart monitor" is far more common than "cardiac monitor," so left to itself it produces the common word, segment after segment, drifting off the approved term without any signal that it has done so.
Example. The termbase says "moniteur cardiaque." The NMT pre-translation, confident and fluent, writes "moniteur de coeur" in 30 of the 40 segments where the term appears. Nothing is grammatically wrong. Every one of those 30 segments violates the client's approved terminology, and if you trust the smooth output, you ship 30 terminology errors.
Why it matters to you: term drift is invisible to a reader checking for fluency and obvious to a reader checking against the termbase, so terminology is a thing you verify deliberately, never something you assume the engine got right.
Term Extraction
Plain definition. Term extraction is the process of pulling candidate terms out of a source text so they can be defined, translated once, approved, and added to the termbase before bulk translation begins. It can be manual, tool-assisted, or now AI-assisted, with an engine proposing candidate terms for a human to accept, reject, or refine. Extraction is the front-loaded work that prevents term drift later: settle the words once, up front, so they hold across every segment.
Example. Before translating a 200-page software manual, a terminologist runs extraction to surface the 300 product-specific terms ("dashboard," "widget," "tenant," "webhook"), decides the one approved translation for each, and locks them into the termbase. Now every linguist on the project, and the engine, has one source of truth.
Why it matters to you: the terms you settle before the project starts are the terms that hold during it, so investing in extraction up front is what makes consistency cheap instead of impossible to retrofit.
Quality and Scoring: How a File Passes or Fails
This is the cluster that decides whether your work ships, and it is the one most often waved at vaguely on a call ("we'll run it through QE," "score it against 5060") and least often defined. These are the words that turn "looks fine to me" into a defensible verdict. Learn them and you stop being scored by a process you do not understand.
QE (Quality Estimation)
Plain definition. QE stands for quality estimation. It is an automatic, machine-produced guess at how good a translation is, segment by segment, expressed as a confidence score, often without ever comparing the output to a human reference. A high QE score means the machine thinks the segment is probably fine; a low score means the machine is uncertain. The critical word is guess. QE is a signal that routes human attention, not a verdict that clears a segment for delivery. A confident QE score on a fluent hallucination is exactly the trap, because the same fluency that fools your eye fools the estimator.
Example. "QE flags route anything under the threshold to full MTPE" on the call meant: the system scores every segment, and segments the machine is unsure about get sent for careful human editing, while high-scoring segments get a lighter pass. The threshold decides where human effort is spent.
Why it matters to you: QE tells you where to look harder, not what to trust, so a high QE score is permission to focus your scrutiny elsewhere, never permission to skip reading the source.
MQM
Plain definition. MQM stands for Multidimensional Quality Metrics. It is a standardized framework for evaluating translation quality by classifying each error along defined dimensions, accuracy, terminology, locale conventions, fluency, style, and so on, and assigning each error a severity. Where QE is an automatic guess at a score, MQM is a human, analytic evaluation: a person reads the output against the source, marks each actual error, categorizes it, and rates how bad it is. MQM is the vocabulary a real quality evaluation speaks.
Example. An evaluator reviewing your post-edited file does not write "pretty good, 8 out of 10." They mark segment 4 as an accuracy error, Critical severity (dropped negation in a contraindication); segment 17 as a terminology error, Major severity (term drift off "moniteur cardiaque"); segment 31 as a locale error, Minor severity (wrong date format). That structured record is MQM in action.
Why it matters to you: MQM is how your work is judged when it is judged seriously, so understanding its categories lets you self-check against the exact dimensions an evaluator will use, before the file leaves your hands.
ISO 5060
Plain definition. ISO 5060:2024 is the international standard that formalizes MQM-aligned human evaluation of translation output. It defines an analytic model: classify each error by type, assign a severity of Critical, Major, or Minor, weight and tally the errors, and reach a pass-or-fail decision on the file. When someone on a call says "score it against 5060," they mean run this structured, severity-based evaluation rather than offering a vague impression. ISO 5060 is the rulebook that makes "this file passes" a defensible statement instead of an opinion.
Example. The client's contract requires an ISO 5060 evaluation on every delivery. Your post-edited file is scored: error types tallied, severities assigned, the result checked against the agreed threshold. Because one segment carried a Critical error, the file fails regardless of how clean the other 2,499 segments are.
Why it matters to you: ISO 5060 is increasingly the standard your deliverables are measured against, so its severity scale is not academic, it is the literal gate between "shipped" and "returned."
Critical, Major, Minor (Error Severity)
Plain definition. These are the three severity levels MQM and ISO 5060 use to rate an error's seriousness, independent of its category. A Minor error is a small imperfection that does not impede understanding or use: an awkward phrasing, a slightly off preposition. A Major error meaningfully degrades the translation: a real mistranslation, a wrong but non-dangerous term, a confusing sentence. A Critical error is one that can cause real harm, safety, legal, financial, or reputational: a flipped dosage, a dropped negation in a warning, an inverted obligation in a contract. The defining rule of the whole quality system: one Critical error fails the file, no matter how clean everything else is.
Example. In Nina's manual, the awkward phrasing in segment 200 is Minor, the term drift in segment 17 is Major, and the dropped negation in segment 4, the one that turns "must not be used on pacemaker patients" into "should be used," is Critical. Segment 4 alone fails the entire 40,000-word file.
Why it matters to you: severity, not error count, is what decides whether you pass, so a file with twenty Minor errors can ship while a file with a single Critical cannot, which is why hunting the one dangerous error matters more than polishing twenty small ones.
Error Typology
Plain definition. An error typology is the structured menu of error categories an evaluation uses: accuracy (mistranslation, omission, addition), terminology, locale conventions, fluency (grammar, spelling, punctuation), style, and so on, each crossed with a severity. It is the difference between "there's something wrong here" and "this is an accuracy error, mistranslation, Critical severity, in segment 4." The typology gives every error a name and a place, which is what makes a quality record reproducible and arguable rather than a matter of taste.
Example. Two evaluators using the same typology on the same file should classify the dropped negation the same way: accuracy / mistranslation / Critical. Without a shared typology, one calls it "a big mistake" and the other calls it "a slip," and the disagreement is unresolvable because there is no common vocabulary.
Why it matters to you: the typology is the shared language that makes quality objective, so framing your own checks in its categories means your defense of a translation speaks the same language as the evaluation that judges it.
The Technical Container: i18n, l10n, Locale, and Tags
Words do not float in space; they live inside files, code, and interfaces with rules of their own. This cluster is the engineering vocabulary a linguist needs, not to become an engineer, but to recognize the container the words sit in and the specific ways a machine breaks it.
i18n and l10n
Plain definition. These two odd abbreviations are everywhere, and they are simply numeronyms: the first and last letters of a long word with the letter count in between. i18n is internationalization ("i," then 18 letters, then "n"): the engineering work of building a product so it can be adapted to other languages and regions, separating text from code, allowing for text expansion, handling different scripts and date formats. l10n is localization ("l," 10 letters, "n"): the work of actually adapting the product for a specific language and region, including translation but also dates, units, currency, images, and cultural fit. i18n is the preparation; l10n is the adaptation. You cannot localize well what was never internationalized.
Example. "The i18n team already extracted the strings so you should not see any concatenation" meant: the engineers pulled the translatable text out of the code into separate string files, and they built sentences as whole units rather than gluing fragments together, so you will not be handed half a sentence to translate blind. That is good internationalization making your localization possible.
Why it matters to you: when internationalization is done badly, your localization job becomes nearly impossible, broken sentences, no room for text expansion, so recognizing an i18n failure lets you raise it instead of silently absorbing the damage.
Locale
Plain definition. A locale is a specific language-and-region combination with all its conventions: not just "French" but "French as used in Canada" versus "French as used in France," each with its own spelling preferences, terminology, date and number formats, currency, and formality norms. A locale is written as a code like fr-CA (French, Canada) or fr-FR (French, France). The same language can have multiple locales that differ in ways that matter to a reader and that an engine, asked vaguely for "French," will get subtly and expensively wrong.
Example. "Locale is fr-CA, not fr-FR" on the call was a precise instruction. A date the France locale writes as "12/04/2026" and certain terms differ; Canadian French has its own conventions and a stronger preference for translated rather than borrowed technical terms. Translating into generic "French" instead of fr-CA produces output that is grammatical and wrong for the audience.
Why it matters to you: locale errors are fluent by definition (the wrong date format is still a valid date), so they sail past a fluency check and only get caught by someone who knows the target locale's actual conventions, which is you.
Placeholder / Tag
Plain definition. A placeholder is a piece of code embedded in a string that gets replaced with real data at run time: {username}, %d, {{count}} become "Nina," "5," "12" when the software runs. A tag is embedded formatting or markup, often shown in the CAT tool as a little numbered token, that carries bold, a link, a line break, or other structure. Both are non-translatable: they must survive translation untouched and in the right position, because they are not words, they are machinery. Translate or delete a placeholder and you break the code; move a tag and you mangle the formatting or the link.
Example. The source string is "You have {count} new messages." The placeholder {count} must appear, unaltered, in your French: "Vous avez {count} nouveaux messages." An engine that helpfully "translates" {count} into {nombre}, or drops it, produces a string that crashes or displays "{count}" literally to the user. "The placeholders are non-translatable so leave the tags alone" was a warning about exactly this.
Why it matters to you: a broken placeholder is a Critical-class failure that is invisible in the linguistic text and only surfaces when the software runs, so protecting tags and placeholders is a verification step you do deliberately on every string, not a thing you assume the engine respected.
The Words That Frame the Work: Transcreation, LSP, Provenance
A last cluster of framing words: terms that describe whole categories of work, the businesses that sell it, and the records that prove it was done well. These are the words that come up when the conversation rises above the segment to the project, the contract, and the audit.
Transcreation
Plain definition. Transcreation is the creative adaptation of a message so it has the intended effect in the target culture, even if that requires departing substantially from the literal source. It is what a slogan, a pun, a brand promise, or an emotional campaign needs: not a faithful rendering of the words, but a re-creation of the intent, the feeling, and the persuasion in a form that lands for the new audience. It is the opposite end of the spectrum from a contraindication, where literal faithfulness is everything. Transcreation is the work a machine most conspicuously cannot do, because it requires understanding what the message is for.
Example. An English tagline plays on a rhyme that does not exist in French. A literal translation is accurate and dead. Transcreation throws out the literal words and invents a new French line that carries the same playful energy and brand promise, perhaps with a completely different image. The source was a springboard, not a target to match.
Why it matters to you: transcreation is high-value, human-only work that the engine flags as easy and gets fluently wrong, so recognizing when a file needs transcreation rather than translation is both a quality safeguard and a way to be paid for judgment the machine cannot supply.
LSP
Plain definition. LSP stands for language-service provider: a company that sells translation, localization, and related services, sitting between the client who needs content in many languages and the linguists who produce it. LSPs range from giant multi-thousand-person operations to two-person boutiques. They run the TMS, hold the client relationship, set the workflows and quality standards, and assign work to in-house or freelance linguists. When statistics say "81.1% of LSPs now offer MTPE," they are describing what these provider companies sell.
Example. A pharmaceutical company needs its device manual in fifteen languages. It does not hire fifteen freelancers directly; it contracts an LSP, which scopes the project, runs MT pre-translation through its TMS, assigns post-editing to qualified linguists, scores the result against ISO 5060, and delivers. The LSP is the orchestrator.
Why it matters to you: whether you work in-house, freelance, or run your own shop, the LSP is the structure that sets your rates, workflows, and quality bar, so understanding its role clarifies who is deciding the standards you are held to.
Provenance / Quality Record
Plain definition. Provenance is the documented history of how a translation was produced: which engine pre-translated it, who post-edited it, what term decisions were made, what errors were found and at what severity, and who signed off. The artifact that captures this is the quality record: a segment-level account a client or an auditor can reconstruct to verify that the work was done to standard. In an MT-first world where "the engine wrote it" is never an acceptable excuse for a Critical error, the quality record is how human accountability is proven rather than just asserted.
Example. A regulated client audits your delivery. You produce a record showing: this content was risk-tiered as high-liability, routed to full post-editing by a qualified linguist, scored against ISO 5060 with zero Critical errors, with terminology conformance to the locked termbase and a named human sign-off. That record turns "trust me" into proof.
Why it matters to you: the quality record is the difference between a price you raced to the bottom on and a quality tier you can defend, so the linguist who can produce one is selling something a raw MT vendor structurally cannot.
In the MT-first era, the deliverable is not just the translation. It is the translation plus the proof of how its quality was owned.
Putting the Vocabulary Back on the Call
Return now to Nina's kickoff call, and read the program manager's sentence again with everything this lesson gave you. "The source comes in through the TMS" (the management platform ingests the file), "we run NMT pre-translation" (a narrow translation engine pre-fills every segment), "the QE flags route anything under the threshold to full MTPE" (automatic confidence scores send the uncertain segments to careful human post-editing), "the rest gets light PE" (the confident segments get a faster, correctness-only pass), "the termbase is locked so watch for term drift" (the approved terms are fixed, and the engine will silently swap synonyms if you let it), "the placeholders are non-translatable so leave the tags alone" (the code tokens must survive untouched), "score the deliverable against 5060" (run a structured, severity-based human evaluation before sign-off), "locale is fr-CA, not fr-FR" (target Canadian French conventions specifically), and "the i18n team already extracted the strings so you should not see any concatenation" (the engineers prepared the content properly, so you will get whole sentences, not fragments).
Not one of those phrases is mysterious anymore. Every one describes something concrete, something you can verify, something with a clear stake. That is the whole point of owning the vocabulary: not to impress anyone on the call, but to convert a wall of acronyms into a precise map of what you are responsible for and what to watch. The fear Nina felt was never about the words being hard. It was about the words being undefined. Define them in working terms, once, and the panic becomes fluency, the good kind, the kind that knows the difference between sounding right and being right.
Key Takeaways
- The workspace nouns anchor everything: the CAT tool (computer-assisted translation) is the cockpit you edit in, the TMS (translation-management system) is the platform that routes and pre-translates files, and the TM (translation memory) is your database of past human translations, which is not machine translation and must be kept clean because it propagates whatever you save into it.
- Work happens in segments, not documents. An exact match is identical to a stored segment (100%, or ICE in-context), and a fuzzy match is partial with a similarity percentage; high fuzzy matches are where the one changed word, a "not" or a number, hides the danger.
- The editing verbs carry price tags: post-editing (PE) is editing machine output, MTPE is the productized machine-translation post-editing workflow (about 46% adoption in 2024, offered by 81.1% of LSPs), and light versus full PE is the effort level that must match the content's consequence, not the price someone wishes it cost.
- Terminology must not drift. A termbase (glossary) holds the client's single approved term per concept, term drift is the engine silently swapping in a common synonym, and term extraction settles the words up front so consistency is cheap instead of impossible to retrofit.
- Quality is scored, not vibed: QE (quality estimation) is an automatic guess that routes attention, MQM (Multidimensional Quality Metrics) is the human analytic evaluation framework, and ISO 5060 formalizes it, with severities Critical, Major, Minor where one Critical error fails the entire file regardless of how clean the rest is.
- The technical container has its own vocabulary: i18n (internationalization, the engineering preparation) enables l10n (localization, the adaptation); a locale like fr-CA is a specific language-and-region with its own conventions; and placeholders/tags are non-translatable code and formatting that must survive untouched.
- Transcreation is creative, human-only adaptation of intent (the opposite pole from a contraindication), an LSP is the provider company that orchestrates the work, and the quality record / provenance is the documented proof of how quality was owned, the artifact that turns "trust me" into something an auditor can verify.
- An acronym is a label on a concept; learn the concept first and the call stops being frightening. The goal of the vocabulary is not to impress anyone but to convert a wall of acronyms into a precise map of what you are responsible for and what to watch in every segment.
Skill.re