Linguistic-Asset and Data Governance
The breach notification did not come from the localization team. It came from the legal department, forwarded at 21:40 on a Thursday to the head of localization with a single line above it: "Is any of this ours?" Attached was a security researcher's disclosure. A public large-language-model service, one that a handful of the enterprise's contract linguists had quietly been using to speed up their post-editing, had been found to retain and, in a narrow set of conditions, surface fragments of the prompts users had pasted into it. Among the surfaced fragments were product code names that had not yet been announced, a paragraph of a drug's not-yet-approved indication, and, most alarming to legal, a customer's full name and home address embedded in a support ticket that a linguist had pasted in verbatim to translate. None of it had been authorized to leave the building. All of it had been fed, one convenient paste at a time, into an engine whose data-handling terms nobody on the localization team had ever read, because nobody on the localization team had ever thought of a translation memory or a support ticket as the kind of asset a data-protection officer loses sleep over. The head of localization spent that night learning a lesson this entire chapter has been building toward: everything defensible about an enterprise localization operation, every quality tier, every ISO conformance claim, every risk-tiered pipeline, rests on a foundation most operations never governed at all. That foundation is the linguistic assets themselves, and the data that flows through them. This lesson is about governing the crown jewels: the translation memories and termbases that are the operation's compounding institutional value, and the confidential data that moves through third-party engines every time a segment is machine-translated.
The Assets Are the Operation, Not a Byproduct of It
Start with a reframe that most localization budgets get backwards. An enterprise localization operation spends its visible money on people, tools, and vendors: linguists, project managers, the translation-management system, the machine-translation engine, the language-service providers. Those are the line items leadership sees. But the thing that actually compounds in value, the thing that would be most expensive and slowest to rebuild if it vanished overnight, is almost never on the budget as an asset at all. It is the accumulated linguistic assets: the translation memory and the termbase.
Define these precisely, because at the enterprise level the definitions carry weight they do not carry at the desk. A translation memory (TM) is a database of previously translated, approved source-target sentence pairs that the computer-assisted translation tool reuses before the machine-translation engine runs, so the same sentence is never solved twice. A termbase is the controlled vocabulary of the operation: the approved translation of every product name, feature, legal term, and unit of art, with its definition, its forbidden alternatives, and the rules for when to use it. Together, a mature TM and termbase are the encoded memory of every linguistic decision the enterprise has ever made across every language, product, and regulatory filing. They are, in the most literal sense, the operation's institutional knowledge rendered into a form a machine can reuse.
Now weigh what they are worth. A ten-year-old, multi-million-segment TM across forty languages represents tens of thousands of hours of expert human work, much of it reviewed twice, some of it defended in a regulatory submission. Rebuilding it from scratch is not a matter of money alone; it is a matter of years, because much of the source content that generated those segments no longer exists to be re-translated. A termbase that encodes how a regulated device is named in every market, and which forbidden synonyms will fail a regulatory review, is not a spreadsheet. It is the difference between a market submission that clears and one that bounces. These assets are appreciating: every clean segment added makes every future project cheaper and every future translation more consistent. They are the one thing in the operation that gets more valuable with age, and they are almost always the one thing nobody has assigned an owner, a valuation, or a protection policy.
The linguists, the tools, and the vendors are replaceable on a timeline of weeks to months. The translation memory and the termbase are not. They are the crown jewels: the appreciating, hard-to-rebuild, institutional asset the whole operation compounds on top of, and the first thing an enterprise governance program has to name, value, and protect.
Why the Enterprise View Differs From the Desk View
Earlier in this program, translation-memory governance was taught as an operational discipline: provenance on every segment, penalties on machine-origin leverage, a working memory separated from a trusted master, promotion reviews, periodic audits. All of that remains true and load-bearing, and this lesson assumes it. But the enterprise leader's questions sit one level up from the operational controls. The operational question is "how do I keep the master TM clean." The enterprise questions are different and harder: Who owns this asset, the enterprise or the vendor who happens to store it? What is it worth, and what is our exposure if we lose it or cannot move it? What confidential and regulated data is inside it or flows through the engines that populate it, and who is legally accountable when that data leaks? Can we prove, to a regulator or an acquirer, that the asset is what we say it is? These are not tool-configuration questions. They are ownership, confidentiality, portability, and provenance questions, and they are the substance of asset-and-data governance at the enterprise level.
Ownership and Stewardship: Who Actually Holds the Crown Jewels
The first governance question is the one that sounds too basic to ask and is almost never answered cleanly: who owns the translation memory and the termbase? Not who uses them. Who owns them, as property, with the right to copy, move, delete, and refuse to hand them over. The answer decides everything downstream, and in a shocking number of enterprises the honest answer is "we are not entirely sure."
The ambiguity arises because the assets are usually built, stored, and touched by parties other than the enterprise itself. Consider the common arrangement. The enterprise's content is translated by a language-service provider (LSP), whose linguists produce the segments. The segments accumulate in a translation memory that physically lives inside the LSP's translation-management system, or inside a cloud platform the LSP administers, on infrastructure the enterprise never sees. The termbase may have been seeded by the LSP's terminologist. Over years, the asset grows inside the vendor's walls, built by the vendor's people, stored on the vendor's systems. When the relationship is good, nobody asks who owns it. When the relationship ends, or when the enterprise wants to bring a second vendor in, or when an acquirer performs due diligence, the question becomes urgent, and the enterprise discovers that ownership was never contractually established.
The governance answer is a clear, written separation between ownership and stewardship, two roles that must never be allowed to blur:
- Ownership belongs to the enterprise, unambiguously and contractually. The TM and termbase are derivative works of the enterprise's own source content and encode the enterprise's own approved decisions. The enterprise must own them as intellectual property, with the contractual right to receive a complete, clean copy in a standard, portable format at any time and on termination, without fee, delay, or degradation. This is not a courtesy to negotiate later. It is a clause that must exist before the first segment is committed, because an asset you cannot compel the return of is an asset you do not own.
- Stewardship is the operational responsibility for the asset, and it can be delegated. A steward runs the day-to-day: commits, promotions, cleanup, the operational governance taught earlier. Stewardship can sit with an internal TM manager, with an LSP under contract, or split between them. Delegating stewardship is normal and often wise. Delegating ownership is a category error that enterprises make by accident, through silence in a contract, and pay for at the worst possible moment.
Inside the enterprise, stewardship needs a named human owner, not a diffuse "the localization team." Somebody, typically a terminology lead or a head of localization, must be accountable for the crown-jewel assets the way a database administrator is accountable for a production database and a records manager is accountable for the corporate archive. That person owns the answer to "what state are the assets in, who has copies, and can we prove their integrity." Without a named steward, the assets belong to everyone, which means they belong to no one, and no one is accountable when they decay, leak, or turn out to be irretrievable from a departed vendor.
Own the asset; delegate the stewardship. The enterprise must hold contractual ownership and the right to a clean, portable copy on demand, while the day-to-day stewardship may sit with an internal lead or a vendor. Blur the two, and you will one day ask for your own crown jewels back and be told they belong to someone else.
The Acquisition and Vendor-Transition Test
There is a simple test that reveals whether ownership is truly settled: imagine you must transition every asset to a new vendor, or hand it to an acquirer's due-diligence team, in thirty days. Can you produce a complete, current, clean, portable copy of every TM and termbase, with its provenance intact, without depending on the goodwill of the outgoing vendor? If the answer is yes, ownership is real. If the answer involves phrases like "we would have to ask them to export it" or "most of it is in their platform and we are not sure what format we would get," then what you have is not ownership; it is access, granted by a party whose interests diverge from yours precisely at the moment you most need the asset. Enterprise governance turns that access into ownership before the transition, not during it.
Confidentiality: What Leaves the Building Every Time a Segment Is Machine-Translated
Now the harder half of the lesson, and the one that produced the Thursday-night breach notification. Ownership governs the asset at rest. Confidentiality governs the data in motion, and in an MT-first operation, data is in motion constantly, in a direction most people never think about. Every time a segment is sent to a third-party machine-translation or LLM engine, the source text leaves the enterprise's control and enters someone else's system to be processed. The comforting mental model is that translation is a closed loop inside the building. In an MT-first pipeline, it is not. It is a continuous outbound flow of the enterprise's own content to external processors, and the content flowing out is frequently the most sensitive the enterprise holds.
Define the terms this section turns on. Confidentiality is the obligation to prevent unauthorized disclosure of information, an obligation that in an enterprise is simultaneously contractual (owed to customers and partners), legal (owed under data-protection and sector regulation), and competitive (owed to the enterprise itself). Data governance is the framework of policies, roles, and controls that determines how data is classified, who may access it, where it may flow, how long it is kept, and how its handling is proven. And PII (personal identifiable information) is any data that identifies a specific person, such as a name, address, account number, health record, or the combination of fields that together single someone out. These are not abstractions in a localization context. They are the literal contents of the segments the pipeline machine-translates every day.
The Three Classes of Content That Must Not Leak
An enterprise localization pipeline routinely carries three classes of high-sensitivity content, each with a different reason it must be governed before it touches an external engine:
- Personal identifiable information (PII) and regulated personal data. Support tickets, customer communications, user-generated content, and app strings that carry user data all flow into localization. A support ticket pasted verbatim into a public engine can contain a named customer's contact details, account number, or health complaint. Sending that to an external processor without a lawful basis and a data-processing agreement is a data-protection violation before it is a localization decision. Under regimes like the European GDPR, the enterprise is the data controller and remains legally accountable for that personal data no matter which engine or linguist actually mishandled it. "A contractor pasted it into a chatbot" is not a defense; it is an admission.
- Trade secrets and unreleased competitive information. Unannounced product names, roadmap details, pricing, unreleased marketing, and pre-filing technical content all pass through localization, often earlier than they pass through any other function, because content must be localized before a global launch. This is the class the breach notification surfaced: code names and an unannounced indication. A trade secret loses its legal protection the moment it is disclosed without appropriate confidentiality controls. Feeding it to an engine that may retain or train on the input can, depending on the engine's terms, constitute exactly that disclosure.
- Regulated content under sector rules. Medical, pharmaceutical, financial, and legal content carries confidentiality and handling obligations specific to its sector, on top of general data protection. Clinical data, patient information, financial records, and privileged legal material each come with rules about where they may be processed, by whom, and under what contractual protections. Routing them through an engine whose data residency, retention, and confidentiality posture has never been assessed against those sector rules is a compliance failure waiting for an auditor to find it.
The unifying danger across all three is the same, and it is subtle. The linguist pasting content into an engine is not being reckless in their own frame. They are being efficient, using a tool that produces a fast, fluent first draft, exactly as the whole MT-first program taught them to. The content looks, to them, like text to be translated, not like a data-protection event. The gap between "text to be translated" and "regulated personal data leaving the enterprise's control" is invisible at the desk and enormous in the eyes of a regulator. Enterprise data governance exists to close that gap with policy and architecture, because it cannot be closed by asking every linguist to become a privacy lawyer.
In an MT-first operation, every machine-translated segment is an outbound transfer of the enterprise's own content to an external processor. The most sensitive content the enterprise holds, personal data, trade secrets, and regulated material, flows through that pipe daily, disguised as ordinary text to be translated. Governing what may enter the pipe, and under what protections, is the core of localization data governance.
The Engine-Tier Distinction That Decides Exposure
Not all engines carry the same exposure, and the single most important architectural control is knowing which kind of engine each piece of content is allowed to reach. There is a spectrum, and the enterprise must place every engine on it explicitly:
- Public consumer engines and chatbots. The free or consumer-tier web interfaces of general MT and LLM services. Their terms frequently permit retention of inputs and, in some cases, use of inputs to improve the service, which can mean the enterprise's content becomes training data or is stored on systems the enterprise has no contract with. This tier is where the breach originated. For an enterprise, unsanctioned use of this tier for any non-public content is the single highest-probability data-governance failure, precisely because it is so easy and so invisible.
- Enterprise or contracted engine services with data protections. The same underlying technology, procured under an enterprise agreement that contractually commits the vendor to no retention of inputs beyond processing, no use of inputs for training, defined data residency, and confidentiality terms. This is the minimum acceptable tier for confidential content, and the difference between it and the public tier is entirely in the contract, not the technology. Identical models, opposite exposures.
- Private, self-hosted, or single-tenant deployment. The engine runs inside the enterprise's own environment or a dedicated instance where content never leaves a controlled boundary. This is the tier for the most sensitive regulated and trade-secret content, and it is the only tier where "the content never left our control" is literally true.
The governance move is to map content classes to engine tiers as an enforced rule, not a suggestion. Public content may use any tier. Confidential and personal data may use only the contracted or private tiers. The highest-sensitivity regulated and trade-secret content may be restricted to the private tier or to full human translation with no external engine at all. This mapping is the localization-specific expression of the "some content is MT-forbidden" principle the whole program insists on, extended from a quality rule ("the machine must not translate this") to a data rule ("this content must not leave our control"). Both rules run off the same risk-tiering discipline, and at the enterprise level they must be governed together.
Contractual Data Protections With Engine and Vendor Partners
Because the difference between a safe engine and a dangerous one is the contract, the contract is a governance instrument, not a procurement formality. An enterprise localization leader who cannot read the data terms of an engine agreement is governing the operation with the most important control invisible to them. The specific protections that must be present, in writing, before confidential content flows through any external engine, are a short and non-negotiable list:
- No retention beyond processing. The vendor commits that inputs are used only to produce the translation and are not stored afterward, or are stored only for a defined, minimal period under stated controls. Retention is where content that has left your building becomes content that persists in someone else's, and it is the property that turned a paste into a breach.
- No use for training. The vendor commits that the enterprise's inputs will not be used to train, tune, or improve the vendor's models. Without this clause, the enterprise's confidential content can become embedded, irretrievably, in a model that other customers use. This is the clause that separates an enterprise agreement from a consumer service.
- A data-processing agreement (DPA) with defined roles. A contract that establishes the enterprise as the data controller and the vendor as the processor, specifying what data is processed, for what purpose, with what security measures, and under which regulatory regime. Under data-protection law, this document is not optional for personal data; its absence is itself a violation.
- Defined data residency and sub-processor transparency. Where, geographically, the content is processed and stored, and which sub-processors the vendor uses, because a chain of undisclosed sub-processors is a chain of undisclosed exposures. Regulated content frequently carries residency requirements that a vendor's default routing will silently violate.
- Confidentiality, breach notification, and audit rights. Binding confidentiality obligations, a contractual duty to notify the enterprise promptly of any breach, and the right to audit or receive evidence of the vendor's controls. Without breach notification, the enterprise learns about its own exposure from a security researcher on a Thursday night, which is precisely the failure mode this lesson opened on.
These protections apply not only to the engine vendor but to every LSP and platform in the chain, because an LSP that sub-contracts to a linguist who uses a public engine has reopened the exact hole the enterprise's own engine contract closed. The governance requirement flows down: the enterprise's data-handling standard must be contractually imposed on every vendor, who must impose it on every sub-contractor, so that the protection is unbroken from the enterprise to the last human or machine that touches a segment. A confidentiality regime that stops at the first vendor and does not flow down to the linguist pasting into a chatbot is a regime that protects the paperwork and not the data.
The difference between a safe engine and a breach is not the technology; it is the contract. No retention beyond processing, no training on your inputs, a data-processing agreement with the enterprise as controller, defined residency, and breach notification, imposed on every vendor and flowed down to every sub-contractor. Read the data terms, or govern blind.
Retention, Deletion, and Portability of the Assets
Two more properties of the crown-jewel assets need enterprise governance, and they pull in opposite directions, which is exactly why they need a policy to reconcile them: how long data is kept, and how freely it can move.
Retention, Deletion, and the Segment That Cannot Be Forgotten
A translation memory is, by design, a permanent record. It keeps segments forever, because their whole value is that they can be reused years later. But some of the content inside those segments is legally required not to be permanent. Under data-protection regimes, personal data must be deleted when the purpose for holding it ends, and a data subject may have the right to have their personal data erased. Now consider the collision: a support ticket containing a customer's name and health complaint was translated, and the source-target pair, with the personal data inside it, was committed to the TM. Years later, that customer exercises their right to erasure. The enterprise deletes them from the support system, the CRM, and the data warehouse, and forgets entirely that a copy of their personal data is sitting inside a translation memory, reused into who-knows-what since. The TM has quietly become an ungoverned shadow copy of personal data the enterprise is legally obligated to be able to delete.
The governance response has two parts. First, prevent the collision upstream: personal data should be redacted or pseudonymized before content enters the localization pipeline, so that the TM captures the reusable linguistic structure without capturing the identifying data. A well-governed pipeline translates "Dear [CUSTOMER_NAME], regarding your [CONDITION]" and never lets the real name and condition into the memory at all. Second, where personal data has entered the assets, the enterprise must be able to find and remove it, which is another reason provenance and segmentation matter: a memory whose segments carry client, domain, and origin metadata can be queried for the class of content that might contain personal data, in a way an undifferentiated mass of segments never can. Retention policy for the linguistic assets must therefore state, explicitly, how long segments are kept, how personal data is kept out of them or removed from them, and how a deletion obligation is honored in a store that was designed to forget nothing.
Portability and the Open Interchange Standards
Portability is the opposite pressure: the asset must be able to move, cleanly and completely, whenever the enterprise decides. Portability is the operational expression of ownership. An asset you own but cannot extract in a usable form is an asset owned on paper only. The enterprise's protection here is to insist that its assets live in, or can be exported without loss to, the open interchange standards the industry maintains precisely so that linguistic assets are not hostages to a single vendor's format:
- TMX (Translation Memory eXchange). The open standard for exchanging translation memories between tools. A TM the enterprise can export to complete, well-formed TMX, with its metadata intact, is a portable asset. A TM that only exists inside one vendor's proprietary database is a hostage.
- TBX (TermBase eXchange). The open standard for termbases. A termbase that exports to clean TBX carries its terms, definitions, forbidden variants, and approval states to any tool. One that exports to a flat list, dropping the definitions and the approval states, has lost exactly the governance metadata that made it an enterprise asset rather than a glossary.
- XLIFF (XML Localization Interchange File Format). The open standard for the bilingual work files that move through the pipeline, so that in-progress work is portable too, not just the finished memory.
The governance discipline, taught in the previous lesson as a portability audit, is to test the export rather than trust it. Periodically, and always before any vendor transition, export each asset to its open standard and verify that the export is complete and clean: that the TMX carries the provenance metadata, that the TBX preserves the approval states and forbidden variants, that nothing silently drops. An export that loses the governance metadata has technically moved the data while destroying the thing that made it defensible. Portability is not the ability to get a file; it is the ability to get the asset, whole, with everything that makes it trustworthy still attached.
Retention and portability pull opposite ways: the TM is built to keep everything forever, while the law requires some of it to be deletable, and ownership requires all of it to be movable. Govern both: keep personal data out of the assets by redaction upstream, and keep the assets in open standards (TMX, TBX, XLIFF) so ownership is portable in fact, not just on paper.
Provenance and Quality as the Strategic Value of the Asset
The final property that enterprise governance must protect is the one that determines whether the crown jewels are actually valuable or merely large: the provenance and quality of what is inside them. Provenance is the recorded origin and history of every segment: where it came from, who committed it, whether it was reviewed, and whether it began as human translation, post-edited machine output, or a raw machine guess. Quality is the trustworthiness of the segments themselves. At the enterprise level these two properties stop being operational hygiene and become the difference between an asset that appreciates and an asset that silently rots.
Consider why provenance is strategic and not merely operational. A translation memory with clean provenance is an appreciating asset because every segment in it can be trusted to the exact degree its provenance records, and its trust can be demonstrated. A translation memory without provenance, one that has been accumulating raw machine output confirmed at speed by rushed post-editors under an MT-first pipeline with no controls, is a depreciating asset wearing the disguise of an appreciating one. It is growing in size while shrinking in trustworthiness, because the share of unverified machine-origin segments is rising and there is no way to tell them from the reviewed human work. The enterprise that leverages heavily from such a memory is compounding on a foundation that is quietly turning to sand, and it will not discover this until a regulator or a client asks the question the whole chapter turns on: which of these segments can you actually vouch for?
This is why the MT-first era makes provenance an enterprise-strategic concern and not just a desk practice. The very throughput that makes MT-first attractive, the lift from around two thousand words a day to five thousand or more, is also the mechanism by which unverified machine output floods into the linguistic assets faster than human review can keep pace. Without governance, the crown jewel most exposed to MT-first contamination is the very asset the enterprise most depends on. Provenance is the control that lets the enterprise keep the speed while preserving the asset's integrity, by ensuring that fast machine-origin leverage is always labeled as what it is and never silently promoted into the trusted, client-facing, regulation-defensible master. The revised ISO 18587, expanding its scope to cover AI and LLM non-human translation output and insisting the post-editor hold full professional-translator competence, and ISO 5060:2024, formalizing the Critical, Major, and Minor error scoring that decides whether output ships, both point the same direction: an enterprise increasingly has to prove not just that its output is good but that it can demonstrate who was accountable for it. A provenance-rich asset is the artifact that furnishes that proof. A provenance-blind one cannot, no matter how good its translations happen to be.
An asset's size is not its value. A translation memory that grows while losing provenance is a depreciating asset disguised as an appreciating one: larger every quarter, less trustworthy every quarter, and impossible to defend when a regulator asks which segments you can vouch for. Provenance is what keeps the crown jewels appreciating instead of quietly rotting under the throughput the MT-first era pours into them.
A Worked Asset-and-Data Governance Framework
The controls become a program only when they are written as a framework that names the assets, assigns ownership, classifies the data, maps content to engines, imposes the contracts, and sets the cadence. Here is a worked linguistic-asset-and-data governance framework, the kind of document that would have made the Thursday-night breach notification a policy citation instead of an emergency. It is deliberately concrete; a real one carries the enterprise's specific regulatory regimes, engine names, and vendor list.
1. Asset register and ownership
Every translation memory and termbase in the operation is entered in an asset register that records the asset, the language pairs and domains it covers, its business owner (a named internal steward, typically the terminology or localization lead), where it is physically stored, which vendors hold copies, and its last verified portable export. The enterprise asserts contractual ownership of every asset, with a written right to a complete, clean, standard-format copy on demand and on termination. No asset may be created or stored with a vendor without an ownership and return clause in place first.
2. Data classification
All content entering localization is classified before it is routed: public, confidential (including trade secrets and unreleased material), personal data (PII and regulated personal data), and regulated-sector (medical, financial, legal). Classification is assigned at intake, travels with the content, and determines which engine tier and workflow the content is permitted to use. Content whose class is unknown is treated as confidential until classified, never as public by default.
3. Engine tiering and content-to-engine mapping
Every machine-translation and LLM engine used in the operation is placed on a tier: public consumer (permitted for public content only), contracted enterprise with data protections (minimum tier for confidential and personal data), and private or self-hosted (required tier for the highest-sensitivity regulated and trade-secret content, alongside full human translation where the machine is forbidden). Use of any public consumer engine for non-public content is prohibited and enforced through the pipeline, not left to individual judgment. The mapping is written, and it is the localization expression of the "some content the engine must never touch" rule extended from quality to data.
4. Vendor and engine data contracts
No confidential or personal-data content flows through any external engine or vendor without contractual protections in place: no retention beyond processing, no use of inputs for training, a data-processing agreement establishing the enterprise as controller, defined data residency and sub-processor transparency, and confidentiality with breach notification and audit rights. These terms are imposed on every LSP and platform and flow down contractually to every sub-contractor, so no linguist may lawfully route confidential content through an uncontracted engine.
5. Confidentiality and PII handling
Personal data is redacted or pseudonymized before content enters the pipeline wherever the linguistic work does not require the real values, so that personal identifiers do not accumulate inside the translation memory. Where personal data must be processed, it is confined to contracted or private engine tiers under a data-processing agreement. Linguists are trained, and bound by policy, never to paste enterprise content into any unsanctioned public engine, and the pipeline is architected so the sanctioned path is the easy path.
6. Retention, portability, and provenance
Each asset has a stated retention policy, including how personal data is kept out of it or removed from it to honor deletion obligations. Each asset lives in, or exports without loss to, the open interchange standards (TMX for memories, TBX for termbases, XLIFF for bilingual files), and the export is tested, not trusted, on a fixed cadence and before any vendor transition, with the export verified to carry provenance and governance metadata intact. Every segment carries provenance (origin type, committer, review status, timestamp, client, domain), and the trusted master is separated from working memory and populated only by reviewed promotion, so machine-origin leverage never silently becomes client-facing institutional memory.
7. Accountability and audit-readiness
The named asset stewards own the register, the classification, the contracts, and the audit schedule. At any time the operation must be able to answer, by query and by document, four questions: who owns each asset and can we retrieve it whole; what data class flows through which engine under what contract; what personal data lives in or passes through the assets and can we delete it on obligation; and which segments can we vouch for and on what provenance. The asset register, the data-classification records, the vendor contracts, and the provenance metadata together constitute the conformance and data-protection evidence, assembled as a standing capability rather than a scramble triggered by a breach notice or an audit.
A governance framework is the asset that makes the other assets defensible. Its real test is a single Thursday-night question: when legal forwards a breach disclosure and asks "is any of this ours, and how did it get out," is the answer a documented policy citation, or a night spent discovering that nobody had ever governed the crown jewels or the pipe they flow through?
Reading the Framework as One System
The seven sections interlock the way the operational TM policy did, but at a higher altitude. The asset register and ownership (section 1) establish what the crown jewels are and that the enterprise holds them. Data classification (section 2) is the layer everything downstream routes off, exactly as provenance was the layer the operational controls queried. Engine tiering (section 3) and the vendor contracts (section 4) are the two halves of confidentiality in motion: the architecture that decides where content may flow and the contract that makes that flow safe. Confidentiality and PII handling (section 5) closes the human gap, the paste into a chatbot, that no contract reaches on its own. Retention, portability, and provenance (section 6) protect the asset at rest, so it can be deleted where the law requires, moved where ownership requires, and trusted where quality requires. Accountability (section 7) is the claim the whole structure exists to make defensible. Pull out classification and the engine mapping has nothing to route on. Pull out the contracts and the sanctioned engines are no safer than the public ones. Pull out provenance and the asset rots invisibly under MT-first throughput. Pull out ownership and portability and the enterprise governs an asset it will one day be unable to retrieve. The framework is not a checklist of independent good ideas; it is a single system in which each control closes a gap the others leave open, and every gap it closes is a specific way that a localization operation, running fast and confident on top of machines and vendors, quietly turns its most valuable asset and its most sensitive data into the liability that arrives on a Thursday night.
Key Takeaways
- The translation memory (TM, the database of approved source-target pairs the CAT tool reuses) and the termbase (the operation's controlled, approved vocabulary) are the crown jewels: the appreciating, hard-to-rebuild, institutional asset the whole operation compounds on. Enterprise governance must name them, value them, and protect them as property, not treat them as a byproduct of the vendors and tools that happen to build and store them.
- Own the asset; delegate the stewardship. The enterprise must hold contractual ownership with a written right to a complete, clean, portable copy of every TM and termbase on demand and on termination, while day-to-day stewardship may sit with an internal lead or an LSP. The acquisition-and-transition test reveals the truth: if you cannot produce every asset whole in thirty days without the outgoing vendor's goodwill, you have access, not ownership.
- In an MT-first operation, every machine-translated segment is an outbound transfer of enterprise content to an external processor. Three high-sensitivity classes flow through that pipe daily, disguised as ordinary text: personal identifiable information (PII, any data identifying a specific person), trade secrets and unreleased competitive material, and regulated-sector content. Under regimes like GDPR the enterprise is the accountable data controller regardless of which linguist or engine mishandled the data.
- Exposure is decided by engine tier, not technology: public consumer engines may retain and train on inputs and are the highest-probability failure point; contracted enterprise engines carry data protections; private or self-hosted engines keep content inside a controlled boundary. Map content classes to engine tiers as an enforced rule, extending the program's "some content the machine must never touch" principle from a quality rule to a data rule.
- The difference between a safe engine and a breach is the contract, so the contract is a governance instrument: no retention beyond processing, no use of inputs for training, a data-processing agreement with the enterprise as controller, defined residency and sub-processor transparency, and confidentiality with breach notification and audit rights, imposed on every vendor and flowed down to every sub-contractor so protection is unbroken to the last human or machine that touches a segment.
- Retention and portability pull opposite ways and both need policy: a TM is built to keep everything forever, but data-protection law requires personal data to be deletable, so keep personal data out of the assets by redacting or pseudonymizing upstream and be able to find and remove it via provenance. Meanwhile keep the assets in open interchange standards (TMX for memories, TBX for termbases, XLIFF for bilingual files) and test the export, so ownership is portable in fact, with governance metadata intact, not just on paper.
- Provenance and quality are the strategic value of the asset, not operational hygiene. A memory that grows while losing provenance is a depreciating asset disguised as an appreciating one, flooded with unverified machine output by the very MT-first throughput that makes it attractive. Provenance keeps the crown jewels appreciating and furnishes the accountability proof that the revised ISO 18587 and ISO 5060:2024 increasingly demand: not just that output is good, but that you can show who was accountable for it.
- A worked seven-part framework (asset register and ownership, data classification, engine tiering, vendor data contracts, confidentiality and PII handling, retention-portability-provenance, and accountability) turns the controls into a program in which each section closes a gap the others leave open. Its real test is one Thursday-night question: when legal forwards a breach disclosure asking "is any of this ours and how did it get out," is the answer a documented policy citation or a scramble to discover nobody ever governed the crown jewels or the pipe they flow through?
Skill.re