Identifying Novel, Defensible Use Cases
The pitch arrived in a slide deck with a countdown clock on it. A vice president of digital experience, freshly back from a conference, had seen a demo of an AI agent that translated a live customer-support conversation in real time, both directions, no human in the loop, and he wanted it live across all thirty markets by the end of the quarter. The deck was titled "AI-First Global CX," and it listed, on one bright optimistic slide, every content type the company produced as a candidate for the same treatment: support chat, help articles, product pages, the mobile app, the terms of service, the pharmacovigilance safety notices for the medical-device line, and the patient-facing instructions for use. All of it, the slide said, was now a real-time AI translation opportunity. The head of localization sat in that room with a decision to make that had nothing to do with whether the demo was impressive. It was impressive. The decision was which of those content types belonged in that future and which of them, if put there, would eventually put a mistranslated dose in front of a patient or an inverted liability term in front of a court, delivered in prose so fluent that no one would catch it until the harm was done. This lesson is about how a localization leader makes that decision defensibly: how to tell a genuinely novel, valuable, safe use case from a hyped one that will hold up in a demo and collapse in an audit, and how to say no to a powerful executive with a countdown clock without saying no to the future.
What "Defensible" Means, and Why It Is the Whole Job
Start with the three words that carry this lesson, because at the enterprise level the imprecise use of them is how millions of dollars and one shipped critical error both happen. A use case is a specific pairing of a content type and a workflow: not "AI translation" in the abstract, but "machine-translation-first drafting of tier-one help-centre articles into fourteen locales with light post-editing," or "real-time large-language-model translation of live support chat with no human review." The unit of decision is never "should we use AI." It is always "should we apply this workflow to this content type at this level of human oversight." Executives think in the first frame. Your job is to force every conversation into the second, because that is the only frame in which a defensible answer exists.
Machine translation (MT), for our purposes, is any system that produces target text with no human writing the words: the neural engine or the large language model (LLM) that drafts a segment before a human sees it. MT-first is the operating posture where the machine produces the default draft for a content type and humans edit, review, or sample it rather than translating from scratch. And a consequence is the worst realistic outcome of a single undetected fluent error in a piece of content, measured not by how likely it is but by how bad it is and how hard it is to reverse. Consequence is the axis this entire lesson turns on, because in localization the danger of content is never visible in the output. MT and LLM output is fluent first and accurate second: the engine is structurally excellent at producing grammatical, confident prose and structurally unable to verify that the prose means what the source meant. A sentence that says "take 25 mg" reads exactly as smoothly as one that says "take 2.5 mg," and only one of them is safe.
Now the word that organizes the leader's job. A use case is defensible when you can stand in front of three specific audiences and justify it: a certifying auditor asking why this content was allowed on this workflow under ISO 18587 and ISO 5060, a regulator or a court asking who was accountable when it went wrong, and your own linguists asking whether you have quietly transferred the risk of the machine's error onto their signatures without transferring the authority to refuse it. If you can answer all three in advance, the use case is defensible. If your only answer to any of them is "the demo looked great" or "the competitor is doing it" or "the engine is really good now," it is not defensible, and the fact that it has not failed yet is not evidence that it will not. Defensibility is a property you establish before deployment, not a story you assemble after an incident.
A use case is defensible only if you can justify it in advance to an auditor, a regulator, and your own linguists. "It works in the demo" answers none of those three, which is exactly why the demo is where the trouble starts.
The Executive Mistranslation of the Word "Opportunity"
The slide in that room made a category error that is worth naming precisely, because you will see it in every hyped pitch you ever receive. It treated "content the company produces" and "content that is an AI translation opportunity" as the same set. They are not. Every content type is an opportunity to move faster. Only some are opportunities to move faster safely. The executive's optimism is not wrong about the technology; the real-time agent genuinely is impressive and genuinely does translate fluently. The optimism is wrong about the boundary. It assumes that because the capability is uniform across content types, the wisdom of applying it is uniform too, and it is precisely the opposite: the capability is uniform and the consequence is wildly not. The same engine that turns a help article around in seconds with trivial risk will, applied to a dosing table, produce a fluent, confident, wrong number with catastrophic risk, and it will do both with identical polish. The leader's entire contribution to that room is to separate the two, because the technology cannot separate them and the executive did not know they needed separating.
The Five-Test Framework for Evaluating a New Use Case
A framework is only worth having if it produces a decision that a smart, motivated person cannot argue their way around under deadline pressure. The one below is built to be run in order and to fail fast: if a use case fails an early test, you stop, because a later strength cannot buy back an earlier disqualification. Run these five tests, in this sequence, on every proposed new AI localization use case, and require the answers in writing before a pilot is funded.
Test One: Value. Is the prize real, and is it worth the exposure?
Begin with value, because a use case that is not worth doing does not deserve the scrutiny of the tests that follow, and half of hyped pitches die here honestly. State the prize in numbers the operation can measure: languages added, turnaround reduced, cost-per-word displaced, coverage extended to content that today gets no translation at all. The last of those is where the genuinely novel value usually lives, and it is worth dwelling on. The most defensible new use cases are frequently not "do the existing work cheaper" but "translate content that was previously untranslated because human translation was economically impossible": the long tail of user-generated content, the back catalogue of support articles no one had budget to localize, the internal knowledge base that was always English-only. That is additive value, and additive value is easier to defend than substitutive value because it does not displace a human deliverable that was meeting a quality bar; it reaches content that was reaching no one. If the value is real, proceed. If the value is "we should use AI because everyone is," stop here: that is not a value, it is a fear, and fear is not a use case.
Test Two: Feasibility. Can the pipeline actually do this, grounded on your assets?
Feasibility asks whether the workflow can be built on your real linguistic assets and infrastructure, not on a vendor's staged demo. The demo ran on the vendor's data. Your production will run on your translation memory (TM), your termbase, your locale rules, your content-management plumbing, and your engines, against your language pairs, some of which are low-resource and behave nothing like the English-to-Spanish the demo used. Ask concretely: can the engine be grounded on our approved terminology so it does not drift to a common synonym for a device name? Can it be constrained to our locale conventions and length budgets so it does not blow a mobile UI's pixel limit or invert a date? Do we have the integration to route content, capture a quality record, and reconcile against the TM? A use case can be valuable and simply not feasible on your stack this year, and calling that out early is a service, not an obstruction. Feasibility failures are recoverable: they become a roadmap item, not a rejection. But shipping a use case that is not actually feasible on your assets is how you discover, in production, that the demo was the only place it ever worked.
Test Three: Quality Risk. What is the worst realistic error, and how would you know it happened?
Now the test that most pitches never run at all, and the one your discipline exists to run. For this specific content on this specific workflow, name the worst realistic undetected error and trace its path into the world. Not the average error: the worst plausible one, because a fluent critical error is low-frequency and high-consequence, and you manage it by consequence, not by expected value. For real-time support chat, the worst error is a fluently mistranslated instruction that leads a customer to take a harmful action, or a confidently invented answer to a question about a paid feature. For high-volume user-generated content, the worst error is usually reputational or a moderation miss, unpleasant but bounded and reversible. For a drug label, the worst error is a flipped dose that reaches a patient and cannot be recalled. The second half of the test is the one that separates governance from theatre: if that worst error occurred, how would you know? What signal, review, sampling, or gate would surface it before it caused harm, and how fast? A use case where the worst error is severe and the detection is "a customer would eventually complain" is not a use case with a quality plan; it is a use case with a hope. Recall the medical-content evidence the whole discipline is built around: LLM output shows error rates around fifty-nine percent on drug names, sixty percent on dates and times, and sixty-six percent on adverse events, every one delivered in grammatically perfect prose. On content where those categories carry consequence, "the output looks fluent" is not reassurance. It is the warning.
Test Four: The MT-Forbidden Test. Is this content the machine must never draft?
This test is a hard gate, not a slider, and it overrides the first three entirely. Some content is MT-forbidden: content where the machine must never produce the trusted draft, because one fluent error in it is catastrophic and irreversible, or because a law, a regulation, or a client contract simply prohibits machine processing. If a proposed use case applies an MT-first or, worse, an unreviewed real-time workflow to MT-forbidden content, it is disqualified regardless of how strong its value, feasibility, and detection story are. There is no value large enough to buy back an irreversible harm to a patient or a court, and no feasibility clever enough to make prohibited processing permitted. The test is a scan for unmistakable markers, the same ones your intake discipline already knows: the audience acts on the text physically (a human will swallow, inject, operate, or evacuate based on it); the text allocates legal rights or obligations (indemnity, warranty, liability, "shall" and "shall not"); a regulator will read and act on it; a signature must certify it before a court or authority; or the content is contractually or legally ring-fenced from machine processing. Any one marker pulls the content out of the AI-first candidate pool for the trusted draft. This is the test that the "AI-First Global CX" slide never ran, and it is the one that quietly deletes the pharmacovigilance notices and the patient instructions from the list before anything else is discussed.
Test Five: Reversibility. If it goes wrong, can you take it back?
The final test asks a question the other four do not: not how bad the error is, but whether you can undo it once it has occurred. Reversibility is the difference between an error you can patch next sprint and an error that has already reached a body, a court, or a regulator and cannot be recalled. A wrong help article is highly reversible: you correct it, you push the fix, the damage is a few confused readers and a support ticket. A wrong internal knowledge-base note is reversible: the audience is your own staff, the feedback loop is fast, and nothing left the building. A user-generated-content translation is usually reversible: you can re-translate, hide, or correct it. A real-time support answer that told a customer to do something harmful is much less reversible, because it was consumed live and acted upon before any review existed. A shipped drug label is effectively irreversible once it is in a pharmacy. Reversibility is what lets you deploy an imperfect workflow safely: if you can detect an error and undo it before it causes harm, an occasional error is a manageable operating cost. If you cannot take it back, every error is permanent, and permanence is what pushes content toward the MT-forbidden gate. High reversibility does not make a use case defensible on its own, but low reversibility is a strong signal that even a low error rate is unacceptable, because you will be living with every mistake forever.
Value and feasibility tell you whether a use case is worth doing and buildable. Quality risk, the MT-forbidden gate, and reversibility tell you whether it is safe to do at all, and any one of the last three can disqualify a use case the first two loved.
Defensible Extensions: Where MT-First Genuinely Belongs
The framework is not a machine for saying no. Run correctly, it says yes to a set of genuinely novel, high-value extensions with clear consciences, and the leader who only ever says no is as useless as the one who says yes to everything. Here are the extensions that pass the five tests convincingly, and why each one passes, because the reasoning is the transferable skill, not the list.
Real-Time Chat and Support, Bounded Correctly
Real-time translation of live customer chat is a defensible extension when it is scoped by consequence rather than deployed as a blanket. It passes on value: it lets a company support customers in languages it could never staff human translators for, in the moment, which is additive coverage of the best kind. It can pass on feasibility if the engine is grounded on the product's terminology and constrained to the domain. It passes on quality risk and reversibility for the majority of conversational content, where a small error produces a re-clarifying message and no harm. The discipline is in the boundary: the moment a conversation turns to a medical symptom, a dosing question, a legal obligation, a financial instruction, or anything a customer will physically act on, the workflow must route to a human, because that content just tripped the MT-forbidden markers even though it arrived inside a "support chat" container. The defensible version of the executive's dream is not "real-time AI across all support." It is "real-time AI for the conversational bulk, with a hard, tested handoff to a human the instant the conversation crosses into consequence." That handoff is the entire difference between a defensible extension and a liability, and building it is the work the demo skipped.
High-Volume User-Generated Content
User-generated content, the reviews, forum posts, community answers, and marketplace listings that arrive in enormous volume and were never translated at all because human translation was economically impossible, is one of the most defensible novel use cases in the entire enterprise. It passes value overwhelmingly, because it is pure additive coverage: content that reached no one now reaches a global audience. It passes reversibility, because a wrong translation can be corrected, hidden, or re-run. Its quality-risk profile is generally bounded, because the consequence of an imperfect review translation is a slightly awkward reading experience, not a harmed body. The genuine risk here is narrower and manageable: moderation and safety filtering, so the engine does not fluently translate content that is abusive, fraudulent, or dangerous, and rare cases where user content itself carries consequence (a user posting medical advice, for instance). Grounded MT with automated safety screening and human sampling is a defensible workflow for this content precisely because the consequence is low, the reversibility is high, and the volume makes the additive value enormous. This is the use case that most rewards the leader who can see additive value where the incumbent process saw only untranslatable overflow.
Internal Content and Enablement
Internal content, the knowledge base, enablement material, internal communications, and process documentation that companies historically left English-only, is a quietly excellent extension. It passes value as additive coverage: employees in every market gain access to content they never had. It passes reversibility, because the audience is internal, the feedback loop is fast, and nothing leaves the company. Its quality risk is usually bounded, because an internal reader who hits a confusing passage can ask, escalate, or check the source, and the accountability chain is inside the organization rather than facing a customer, a regulator, or a court. The exceptions are real and worth naming so the extension stays honest: internal content that governs safety-critical procedures, that carries legal or compliance obligations, or that contains confidential material a machine may not process, all of which route out of the MT-first lane on the same markers as any other content. But the ordinary bulk of internal enablement is a defensible, high-value place to extend MT-first, and it has the additional virtue of being a low-stakes environment in which to mature the pipeline before pointing it at anything customer-facing.
Undefensible Use Cases: Where MT-First Must Not Go
The other half of the leader's judgment is the set of use cases the framework refuses, and refusing them is not caution for its own sake; it is the recognition that consequence and reversibility have hard floors below which no value and no feasibility can rescue a use case. These are the ones judged by consequence rather than hype, and they fail the fourth and fifth tests categorically.
Regulated life-safety content is the clearest. Drug labels, patient information leaflets, instructions for use of a medical device, dosing tables, clinical-trial and informed-consent documents: on all of these, a single fluent error can reach a patient and cannot be recalled, and the medical-content error rates cited above make the probability of such an error unacceptably high on unreviewed machine output. An informed-consent form trips three MT-forbidden markers at once: a regulator reads it, it allocates legal obligations, and a human acts on it physically. No real-time or MT-first workflow for the trusted draft is defensible on this content, and the "AI-First Global CX" slide's inclusion of pharmacovigilance notices and patient instructions was not ambition; it was the single most dangerous line on the page.
Binding legal instruments are the second. Indemnity and liability clauses, warranties, contracts, sworn and certified documents: content whose entire value is the precise, verified allocation of rights and obligations, where a flipped negation or an inverted clause is a lawsuit, and where a signature may be legally required. A machine may sit beside a legal translator as a reference at most; it may not produce the trusted draft, and it certainly may not produce it in real time with no review.
Regulatory submissions are the third. Any content a government agency reviews, approves, or files against: marketing-authorization language, disclosures, prospectuses, patent claims. The consequence is a regulatory finding or a rejected submission, the reversibility is poor, and the accountability faces an authority that will not accept "the engine wrote it."
The pattern across all three is identical and worth internalizing as a single rule: where the consequence of a fluent error is a harmed body, a lost lawsuit, or a regulatory sanction, and where that outcome cannot be reversed once it occurs, the content is MT-forbidden for the trusted draft no matter how impressive the technology or how loud the pressure to extend it. The hype does not change the consequence, and the consequence is what governs.
The undefensible use cases are not undefensible because the technology is weak. They are undefensible because the consequence is irreversible, and no demo, no competitor, and no countdown clock changes what a flipped dose or an inverted indemnity does once it ships.
How to Say No to a Hyped Use Case
Knowing a use case is undefensible is half the leader's job. Saying so, to a vice president with a countdown clock and a board expectation, without becoming the person who "blocks innovation," is the other half, and it is a skill as learnable as the framework itself. The failure modes are two: the leader who says a flat "no" and is routed around as an obstacle, and the leader who says "yes" against their judgment and owns the incident. Neither is the move. The move is to say yes to the ambition and no to the specific unsafe pairing, with the framework as the reason, so the refusal reads as rigor rather than resistance.
Separate the Vision From the Scope
Never argue with the vision, because the vision is usually right and arguing it makes you the obstacle. "AI-first global customer experience" is a good goal. Agree with it loudly and immediately. Then move the disagreement to where it belongs, the scope: "The goal is right, and here is the scope of it that we can ship this quarter safely, and here is the specific slice that will hurt us if we ship it the same way." You have now reframed the conversation from "localization is blocking AI" to "localization is telling us which parts of this are safe to ship first," which is the story an executive can carry to the board and which is also true. The countdown clock now runs on the defensible scope, not the whole slide, and you are the person who made the deadline achievable rather than the person who fought it.
Translate the Refusal Into Consequence and Liability
An executive does not feel a quality abstraction, but every executive feels a consequence with a name and an owner. Do not say "the quality risk is high." Say "if the real-time engine mistranslates a dosing question in a support chat, a patient can take a wrong dose, we cannot recall it, and the accountability sits with this company, not the vendor, because we deployed it." Do not say "this content is out of scope." Say "this is the content a regulator reads, and 'the AI wrote it' is not a defence in front of that regulator." Name the worst realistic error, name who it harms, name that it cannot be reversed, and name who is accountable when it happens. You are not being negative; you are being the only person in the room doing the risk assessment the demo skipped, and you are doing it in the language of consequence and liability that the executive is actually paid to manage.
Offer the Defensible Path, Not Just the Refusal
The refusal must always come attached to a path, because a "no" with no "instead" is heard as obstruction and a "no, and here is the yes" is heard as leadership. Bring the defensible scope, the pilot that would generate evidence, the handoff design that makes the risky slice safe, the risk tiers that route the forbidden content out. "We are not shipping real-time AI on all support this quarter. We are shipping it on the conversational bulk with a tested human handoff on consequence, we are running the additive user-generated-content and internal use cases in parallel because those are pure upside, and we are keeping the regulated content on the human workflow it legally requires. Here is the evidence plan that tells us in ninety days whether to widen it." That is not a smaller version of the executive's dream. It is the only version of it that survives contact with an auditor, and delivering it is how the localization leader becomes the person the executive brings into the room early rather than the person they route around.
You do not say no to the vision. You say yes to the vision and no to the unsafe scope, in the language of consequence and liability, with a defensible path attached. A refusal without a path is obstruction; a refusal with a path is leadership.
A Worked Use-Case Evaluation
Take the executive's own proposal through the framework end to end, because the value of the five tests is entirely in watching them force a decision the slide could not. The proposed use case, stated properly in the content-plus-workflow frame the leader insisted on: real-time LLM translation of live customer-support chat, no human in the loop, across all thirty markets and all support topics. Run it.
Value. Real and large. The company cannot staff human support translators for thirty markets in real time; the engine can extend live support to languages that today get none. This is additive coverage, the strongest kind of value. The use case passes Test One convincingly, which is exactly why it is dangerous: a use case with weak value dies early and harmlessly, but a use case with strong value tempts everyone to wave the later tests through. Proceed, with the guard up.
Feasibility. Partial. The engine can be grounded on product terminology and constrained to the support domain for the common conversational cases, and the real-time integration is buildable. But "all thirty markets" includes low-resource language pairs where the engine's behaviour is materially worse than the demo's high-resource pair, and "all support topics" is not a domain, it is every domain, which no grounding can constrain uniformly. Feasibility passes for the conversational bulk in the higher-resource markets and is shakier at the edges. This is a scope signal, not yet a disqualification: it tells us the real feasible use case is narrower than the proposed one.
Quality risk. Here the proposal breaks. The worst realistic undetected error on "all support topics with no human in the loop" is a fluent, confident mistranslation of a consequential instruction: a symptom described, a dose asked about, a financial action instructed, a legal right stated, delivered live and acted upon before any human sees it. And the detection story for the no-human-in-the-loop design is essentially nonexistent: by construction, no one reviews the output before the customer acts on it, so the worst error is discovered only after the harm. Recall the error rates on exactly the categories that carry consequence. On this design, the quality risk is severe and the detection is absent, which is the definition of a use case with a hope instead of a plan. The proposal as stated fails Test Three.
The MT-forbidden test. "All support topics" necessarily includes topics that trip the forbidden markers: a customer asking about a medication dose (physical action), about a contractual right (legal obligation), about a regulated financial product (a regulator reads the answer). The blanket scope drags MT-forbidden content into an unreviewed real-time workflow, which is a categorical disqualification for those topics. The proposal as stated fails Test Four for a definable slice of its own scope.
Reversibility. Low, and that is decisive. A real-time answer is consumed and acted upon live; there is no window in which to catch and undo the error before the customer acts. Low reversibility on top of severe, undetected quality risk means every error on the consequential slice is a permanent one. The proposal as stated fails Test Five for the same slice.
The decision. The blanket proposal is undefensible and is not shipped. But the framework does not stop at "no"; it has, in the course of failing the proposal, described the defensible use case hiding inside it. Re-scope to: real-time LLM translation of live support chat for conversational, non-consequential topics in the higher-resource markets, with a tested, hard handoff to a human the instant the conversation crosses into medical, legal, financial, or safety-consequential territory, and human sampling of a live audit stream. That version passes all five tests: the value survives, the feasibility matches the narrowed scope, the quality risk is bounded by the handoff, the MT-forbidden content is routed out by the handoff triggers, and the reversibility problem is contained because the only content left in the unreviewed lane is content where an error is low-consequence and correctable. The additive user-generated-content and internal use cases run in parallel as pure-upside pilots. The regulated content stays on its required human workflow. The executive gets a real-time AI support capability shipped this quarter, the auditor gets a defensible record of which content got which workflow and why, and the head of localization gets to be the person who made the ambition safe rather than the person who fought it. The countdown clock still runs; it just runs on the version that will not end in a courtroom.
The framework's job is not to kill the use case. It is to find the defensible one hiding inside the hyped one, by letting each failed test carve away the unsafe scope until what remains is a use case you can ship and defend.
Key Takeaways
- The unit of decision is never "should we use AI" but "should we apply this workflow to this content type at this level of human oversight." A use case is a content-plus-workflow pairing, and a use case is defensible only when you can justify it in advance to an auditor, a regulator, and your own linguists, not when it merely works in a demo.
- Evaluate every proposed use case with five tests in order: value (is the prize real, and is additive coverage of previously untranslated content the strongest form of it), feasibility (can the pipeline do this grounded on your real assets, not the vendor's demo), quality risk (what is the worst realistic undetected error and how would you know it happened), the MT-forbidden test (is this content the machine must never draft), and reversibility (if it goes wrong, can you take it back).
- Value and feasibility ask whether a use case is worth doing and buildable; quality risk, the MT-forbidden gate, and reversibility ask whether it is safe to do at all, and any one of the last three can disqualify a use case the first two loved. Judge by consequence, not by hype.
- The defensible extensions are real and high-value: real-time chat and support scoped with a hard human handoff on consequence, high-volume user-generated content as pure additive coverage with safety screening, and internal content and enablement as a low-stakes place to mature the pipeline. Each passes because the consequence is bounded and the reversibility is high.
- The undefensible use cases fail categorically: regulated life-safety content (drug labels, dosing tables, patient instructions, informed consent), binding legal instruments (indemnity, warranties, sworn documents), and regulatory submissions. They fail not because the technology is weak but because the consequence of a fluent error is a harmed body, a lost lawsuit, or a regulatory sanction that cannot be reversed.
- The MT-forbidden test is a hard gate that overrides value and feasibility, triggered by any one marker: the audience acts on the text physically, the text allocates legal rights or obligations, a regulator reads it, a signature must certify it, or the content is contractually or legally ring-fenced from machine processing. Any one marker pulls the content out of the AI-first candidate pool for the trusted draft.
- Saying no to a hyped use case is a skill: agree with the vision and disagree only with the unsafe scope, translate the refusal into a named consequence with an owner and an accountability chain rather than a quality abstraction, and always attach a defensible path. A refusal without an "instead" is heard as obstruction; a refusal with one is heard as leadership.
- A worked evaluation kills the blanket proposal and, in the process of failing each test, carves out the defensible use case hiding inside it. The framework's purpose is not to say no but to find the version of the ambition that survives an audit, ships on time, and does not end in a courtroom.
Skill.re