Standing Up a Localization-AI Governance Group
The VP of localization at a fast-growing software company thought she had the AI question handled, right up until three decisions landed on her desk in the same week and none of them had an owner. On Monday, an engineering manager had quietly swapped the default MT engine for the whole German pipeline to a cheaper LLM API, because it benchmarked two points higher on an internal test and shaved a line off the cloud bill. On Wednesday, a project manager under a brutal deadline had routed a batch of pharmaceutical package-insert copy through machine-translation post-editing at the light tier, because the client had asked for "MTPE pricing" and nobody had told the PM that this particular content was supposed to be human-only. On Friday, a linguist forwarded her a screenshot: the new engine had rendered a financial disclosure with an inverted negation, fluent and confident and completely wrong, and it had shipped to the client three days earlier. Three people, three reasonable-looking calls, three different corners of the org, and every one of them was a governance decision that no single role was authorized to make and no forum existed to catch. The engine choice belonged to nobody. The MT-forbidden policy lived in a linguist's head. The incident had no review path. The VP realized she did not have a quality problem or an engineering problem or a PM problem. She had a governance vacuum, and it was filling itself with whoever happened to be nearest to the decision when it needed making. This lesson is about how you fill that vacuum on purpose, with a body that has the seats and the authority to set policy and make it stick.
Why Governance Cannot Be a Side-of-Desk Job
Let us define the word before we build the thing, because "governance" gets used loosely and the looseness is exactly what kills it. Governance, in the localization-AI sense, is the standing authority to decide policy about how machine translation and large language models are used across the operation, and the mechanism to enforce those decisions so they hold everywhere rather than in the corners where a conscientious person happens to apply them. It is not a document. It is not a one-time policy PDF that a director wrote and emailed once. It is a living function: a body that meets, decides, records, and enforces, and that keeps doing so as engines change, clients change, and the people running the pipeline turn over. The document a governance group produces is downstream of the group. The group is the thing that keeps the document true.
The reason governance cannot live inside any one existing role is that a localization-AI decision is almost never confined to one discipline. Consider the engine swap the engineering manager made. Purely on the engineering axis it was defensible: cheaper, marginally better on a generic benchmark, a smaller bill. But the choice of engine is simultaneously a quality decision (does the new engine make more Critical errors on regulated content, and how would anyone know), a terminology decision (does it honor the client's approved termbase or drift to common synonyms), and a commercial decision (has the operation now taken on concentration risk by standardizing on a single vendor's API). No one of those four people, the engineer, the quality lead, the terminologist, the account owner, could have made that call correctly alone, because each of them can only see one face of a four-sided decision. The engineer optimized the face he could see and was blind to the other three, which is not a failing of the engineer. It is a structural inevitability of asking a single discipline to make a cross-disciplinary call.
This is the load-bearing insight of the whole lesson. A localization-AI decision is cross-functional by nature, so the authority to make it must be cross-functional by design. If it is not, the decision does not disappear. It defaults to whoever is nearest, optimizing the one axis they can see, and the org discovers the other three axes only when they fail, usually in the form of a shipped Critical error, a blown budget, or a client who read the disclosure before the linguist caught it.
Governance is not a document you write once. It is a standing, cross-functional body with the authority to set policy and the mechanism to enforce it, because a localization-AI decision has four faces and no single role can see all of them.
The Difference Between a Policy and a Governance Group
It is tempting to think you can skip the body and just write the policy. Write down "regulated content is human-only, these three engines are approved, one Critical fails the file," circulate it, and be done. This fails, and it fails predictably, for a reason worth sitting with. A policy is a static artifact, and the world it governs is not static. New engines appear monthly. A client brings a content type your tiering never anticipated. An engine that passed evaluation last quarter regresses after a silent model update. A linguist finds an edge case the policy never addressed. Every one of these is a question the written policy cannot answer on its own, because a document cannot reason, cannot weigh a new trade-off, and cannot revise itself. Someone has to. The governance group is the someone. The policy is the current state of the group's decisions; the group is the engine that keeps the policy alive and correct as reality shifts underneath it. An operation with a policy and no group has a snapshot of good intentions that goes stale the day it is written. An operation with a group has a policy that stays true because there is a body whose standing job is to keep it true.
Who Sits at the Table and Why Each Seat Exists
A governance group is defined first by its seats, because the seats are what make the decisions correct. The whole premise is that a localization-AI decision has multiple faces, so the table must have someone who can see each face and is accountable for it. The failure mode of an immature governance attempt is a table that is really one discipline in a trench coat: a group nominally cross-functional but where engineering makes the calls and quality is a courtesy invitee, or where quality dominates and engineering is treated as plumbing. A real governance group gives each seat genuine standing. We will take the core seats one at a time, because who is missing from the table is exactly where the org will be blindsided.
The Quality Seat
The quality lead owns the question the entire program exists to answer: is the output actually correct, measured against a real error typology rather than a feeling. This is the person who runs or owns the severity-scored evaluation, who thinks in MQM (Multidimensional Quality Metrics, the error typology that classifies each mistake by dimension and severity) and ISO 5060 (the 2024 standard that formalizes the MQM-aligned Critical, Major, and Minor scoring that decides whether output ships). Without a quality seat with real authority, every other seat will, under pressure, optimize its own axis at quality's expense: engineering will pick the cheaper engine, PM will route to the faster tier, and the account owner will promise the aggressive price, and quality will find out when the Critical ships. The quality seat is the one that says "the German engine you swapped in makes more accuracy errors on regulated content, here is the score, and until that is fixed it is not approved for that tier." It is the seat with the authority to fail a thing that everyone else finds convenient.
The Terminology Seat
The terminologist owns whether the client's approved language actually survives the pipeline. This seat is often folded into quality, and folding it is a mistake, because terminology is a distinct axis with its own failure mode: an engine that is accurate in the ordinary sense can still systematically drift off an approved term because it "prefers" a more common synonym, and that drift is invisible to an accuracy check that is only asking "does this mean the right thing." The terminology seat governs the termbase itself as an asset (its version control, its authority, its grounding into the engines) and holds the line that an engine which will not honor the approved term for a medical device is not fit for that client regardless of how fluent it is. When the terminology seat is empty, terminology becomes everyone's job and therefore no one's, and it degrades one convenient synonym at a time until a client notices their brand name has been quietly re-translated across ten thousand strings.
The Engineering Seat
The localization engineer owns the plumbing that the linguistic quality actually rides on: engine integrations and API access, grounding configuration (how the termbase and translation memory are fed to the engine), placeholder and tag safety, length and encoding constraints, and the CI/CD (continuous integration and delivery) pipeline that builds and ships localized content. This seat exists because a governance decision that is linguistically perfect but technically un-implementable is a wish, not a policy. When the group decides "all engines must be grounded on the client's termbase," it is the engineering seat that says whether that is a config change shipping Friday or a two-sprint integration project, and that owns the reality that the swap the engineering manager made unilaterally was, in fact, an engineering decision that should have come to this table. Give engineering a real seat and the impulse to make the change alone converts into the discipline of bringing it to the group. Deny engineering a real seat and the changes keep happening in the shadows because the plumbing has to move whether governance authorizes it or not.
The PM and Account Seat
The project or program manager owns the point where policy meets the deadline, which is precisely where policy tends to die. This is the seat that lives closest to the client's pressure to go faster and cheaper, and that under that pressure made the call to run pharma copy at the light tier. Putting PM at the governance table does two things. It gives the group a true picture of the operational reality (what deadlines actually look like, where clients push, where the friction between policy and delivery is real rather than theoretical), and it gives the PM the authority and the cover to say no to the client, because "our governance group forbids machine translation on this content type" is a defensible institutional line in a way that "I personally think we shouldn't" never is. The PM seat is also the group's early-warning system: the PM is the one who sees the new content type, the unusual request, the edge case the policy did not anticipate, and who can bring it to the table before it becomes an incident rather than after.
Who Chairs, and Who Else Belongs
The group needs a chair with genuine executive authority, typically the head of localization or a quality director, because a body that can only recommend is not a governance group, it is a discussion club, and discussion clubs lose to deadlines every time. The chair is the person whose sign-off makes a decision binding on the operation. Beyond the four core seats, the table flexes to the org: a legal or compliance representative is essential the moment the operation touches regulated content, because the MT-forbidden policy is partly a legal-risk judgment; a vendor or procurement voice belongs there when engine and tool contracts are on the agenda; and a data or security representative belongs there whenever confidentiality of the translation memory and termbase, or the data-handling terms of an engine vendor, are in question. The principle is constant: every axis a localization-AI decision can fail on needs a seat that can see that axis and is accountable for it. The art is keeping the group small enough to decide and broad enough to be right.
The seats are the safeguard. Each represents one face a localization-AI decision can fail on: correctness (quality), approved language (terminology), implementability and integrity (engineering), and the collision with deadline and client (PM). An empty seat is exactly where the org gets blindsided.
What Authority the Group Actually Holds
A governance group is only as real as the authority it holds, and this is where most attempts quietly fail. A body that meets, discusses, and produces recommendations that any director can override on any Tuesday under deadline pressure is not governing anything. It is generating advice that gets ignored precisely when it matters most, which is under pressure, which is the only time governance is ever tested. The difference between a governance group and a working group is the difference between decide and suggest. The charter has to say, in plain words, that on the matters within its scope the group's decisions are binding on the operation, and that changing them happens through the group, not around it.
Concretely, a localization-AI governance group must hold three kinds of authority, and it is worth being precise about each because a group that holds only the first is toothless.
- The authority to set policy. The group decides the standing rules: which engines and tools are approved for which content, what the quality tiers are and what each guarantees, which content is MT-forbidden, what the delivery gate is, and what metrics the operation reports. These are not suggestions the group offers to line managers. They are the operating rules of the pipeline.
- The authority to enforce policy. Setting a rule that anyone can quietly ignore is theater. The group must own the enforcement mechanism, which in practice means the controls are built into the tools and the process so that violating the policy is hard rather than merely discouraged: the TMS refuses to assign MT-forbidden content to a machine workflow, an unapproved engine is not available to select, the gate blocks delivery on a Critical. Enforcement lives in the system, and the group owns the system's rules.
- The authority to review and revise. The group has the standing right to review incidents, to audit whether policy is actually being followed, and to change the policy when reality demands it. This is the authority that keeps the policy alive rather than stale, and it is the authority that turns a shipped Critical error from a blame exercise into a governance input.
Here TMS is the translation-management system, the platform that routes and tracks jobs; MT is machine translation, LLM is a large language model, MTPE is machine-translation post-editing, and CAT is the computer-assisted translation tool the linguist works in. MT-forbidden, a term this program uses deliberately, means content that the machine must never touch at all, not even as a draft a human then post-edits, because the consequence of a fluent error in that content (a flipped dosage, an inverted indemnity, a mis-stated contraindication) is a harm no post-editing efficiency can justify risking. Deciding what falls into that category, and enforcing that it never reaches an engine, is one of the group's most consequential powers.
The Escalation and Override Path
Real authority does not mean unaccountable authority, and a mature charter says so. The group's decisions are binding, but there must be a defined escalation path for the case where a decision is genuinely wrong or where a business emergency requires a documented exception. The point of naming this path is not to create a loophole. It is to ensure that when someone does override the group (a senior executive who accepts a documented risk to win a strategic account, say), the override is explicit, attributed, dated, and recorded as an exception rather than executed silently by whoever was nearest the decision. A governance group that can be silently bypassed is not a governance group. A governance group whose overrides are visible, owned, and logged is a mature one, because the exception becomes a governed event rather than the very vacuum the group was created to eliminate.
What the Group Governs: The Standing Agenda
The scope of a localization-AI governance group is not open-ended. It governs a specific and knowable set of things, and defining that set precisely is what keeps the group from either overreaching into everyone's daily work or shrinking into irrelevance. Five domains form the standing agenda, and each maps directly to one of the vacuum-filled decisions from our opening story.
Engine and Tool Approval
The group owns which MT and LLM engines, and which CAT and TMS tooling, are approved for use, and for which content. This is the domain that would have caught the Monday engine swap. The rule is not that engineering may never change an engine; it is that an engine reaches the approved list only after it has been evaluated on the axes that matter (accuracy and critical-error rate on the operation's real content, terminology adherence, locale correctness, data-handling terms, and cost), and that approval is scoped to content types rather than granted globally. An engine can be approved for high-volume marketing content and forbidden on regulated content in the same decision. Crucially, the group evaluates engines on domain-relevant quality and critical-error rate, not on a generic benchmark score, because an engine that wins on a broad academic metric can still be the worse choice on the specific content that can hurt you. Approval is also revocable: an engine that regresses after a model update comes back to the table.
Quality Tiers and the Delivery Gate
The group owns the definition of the quality tiers (raw MT, light post-editing, full post-editing, full human) and, critically, owns the delivery gate that decides whether output ships. The gate is where governance becomes concrete: a shared scoring model against the ISO 5060 typology, a numeric threshold, and the non-negotiable absolute rule that a single Critical error fails the file regardless of how clean the average looks. The group owns this gate because a gate that means different things on different accounts is not a gate, and only a cross-functional body with real authority can hold the line that the gate is identical everywhere and that the Critical rule is never averaged away under deadline pressure. This is the domain that connects the group's policy to the individual file: the tier and the gate are how the abstract governance decision reaches the linguist's screen.
The MT-Forbidden Policy
The group owns the list of content that the machine must never touch, and the enforcement that keeps machines away from it. This is the domain that would have caught the Wednesday routing of pharma copy to light MTPE. The MT-forbidden policy is a risk judgment that genuinely requires the whole table: quality knows the error rates, the terminologist knows where the language is unforgiving, engineering knows how to build the block into the TMS so the content structurally cannot be assigned to a machine workflow, PM knows where client pressure will push against the policy, and legal knows the liability. No single seat can draw that line correctly, which is exactly why it belongs to the group and not to a linguist's memory. And because the policy is only as good as its enforcement, the group's job is not merely to publish the list but to ensure the system refuses to violate it, so that a PM under deadline cannot route forbidden content to an engine even if they wanted to, because the tool will not let them.
Incident Review
The group owns the review of shipped Critical errors and near-misses, and this is the domain that would have caught the Friday incident. Incident review is not a witch hunt for the person who shipped the error; treated that way, it teaches people to hide incidents, which is the opposite of what governance needs. It is a standing, blameless examination of what in the system allowed the error through: was the engine unapproved for that content, was the tier wrong, did the gate miss it, was the content mis-classified at intake. The output of incident review is almost never "that linguist should be more careful." It is a change to policy or enforcement: an engine de-approved, a content type moved to MT-forbidden, a gap in the gate closed, an intake rule tightened. This is the feedback loop that makes the whole program learn, and it only works if the group has both the authority to review incidents and the authority to change policy in response.
Metrics and Reporting
The group owns which metrics define the health of the localization-AI operation and what gets reported to leadership. This matters more than it sounds, because the wrong metrics actively hide risk. A dashboard that shows only throughput and edit-distance and cost-per-word will look magnificent right up until the moment a Critical error ships, because none of those numbers can see the silent mistranslation. The group's job is to insist that the metrics include the risk axis: critical-error rate, terminology conformance, the proportion of content correctly risk-tiered, gate pass and fail rates. Governing the metrics is how the group ensures leadership sees the true dual-axis story (speed and provable quality together) rather than a flattering speed number that conceals the liability building underneath it.
Five domains, one standing agenda: engine and tool approval, quality tiers and the delivery gate, the MT-forbidden policy, incident review, and metrics. Each maps to a decision that, ungoverned, defaults to whoever is nearest and optimizes the one axis they can see.
The Charter and the Cadence
Two artifacts turn a good intention into a running body: a charter that says what the group is, and a cadence that makes it actually meet and decide. Neglect either and the group evaporates. A group with a beautiful charter that never meets is a document; a group that meets without a charter is a recurring argument with no authority to settle anything.
What the Charter Must Say
The charter is the founding document that establishes the group's existence, scope, membership, and authority. It is short, it is signed by an executive who can make it binding, and it says, without ambiguity, the handful of things that determine whether the group can actually govern. A working charter states the group's purpose in one or two sentences; its scope, the five domains it governs and, just as importantly, what it does not govern so it neither overreaches nor leaves gaps; its membership, the named seats and who fills them, with the chair identified; its decision authority, the plain statement that its decisions are binding on the operation within scope and that changes go through the group; its decision rule, how the group actually decides (consensus where possible, chair's call where consensus fails, so a single dissenter cannot deadlock a needed decision); its escalation path, how a decision is overridden and how that override is recorded; and its cadence, how often it meets and how urgent matters are handled between meetings. That is the whole charter. Its power is not in its length but in its precision on authority: a charter that hedges on whether the group's decisions are binding has already lost, because the first deadline will expose the hedge.
The Cadence That Keeps It Alive
A governance group needs two rhythms, because it does two kinds of work. The first is the standing meeting, typically monthly, where the group works its recurring agenda: reviewing any engine or tool approval requests, reviewing incidents since the last meeting, checking the metrics, and revisiting any policy that reality has strained. The monthly rhythm is deliberate. Too infrequent and decisions pile up behind the meeting, so the pipeline stalls waiting for governance or, worse, routes around it; too frequent and the group burns senior time on a cadence that has nothing to decide. Monthly is usually the sweet spot for the steady-state work, with the standing agenda ensuring the meeting is never a blank-page discussion but a disciplined pass through the five domains.
The second rhythm is the exception path, the mechanism for the decision that cannot wait a month. An engine regresses badly and needs de-approving today; a client brings an urgent content type the tiering does not cover; a Critical error ships and containment cannot wait for the next monthly. The charter names how these are handled: usually a fast-track decision by the chair plus the relevant seats, made on the record and ratified at the next standing meeting. Without an exception path, one of two bad things happens: either urgent decisions wait a month and the pipeline suffers, or people make the urgent decision unilaterally and the vacuum reopens exactly where the group was supposed to close it. The exception path is what lets the group hold authority over fast-moving decisions without becoming a bottleneck that people learn to bypass.
The Record Is Not Optional
Every decision the group makes, and every exception granted, is recorded: what was decided, by whom, on what date, and why. This is not bureaucracy for its own sake. It is the difference between a group whose policy is knowable and enforceable and one whose decisions live in the memory of whoever was in the room. The decision record is what a new PM consults to know whether an engine is approved for their content; it is what an auditor reads to see that the MT-forbidden policy was set deliberately and dated; it is what incident review examines to see whether a shipped error violated a standing decision or exposed a gap in one. A governance group that decides but does not record is only marginally better than the vacuum, because its decisions are as fragile as the memory of the people who made them, which is precisely the hero problem the group was built to solve, relocated from the linguist to the committee.
A Worked Governance-Group Setup
Let us make it concrete by standing up the group for the software company from our opening, the one with the three ungoverned decisions. The VP has decided she will not solve those three incidents one at a time; she will build the body that would have caught all three and every future one like them. Here is how she actually does it, in the order that works.
Step One: Name the Seats and the Chair
She starts with people, because the seats are the safeguard. She chairs the group herself, as head of localization, so that its decisions carry executive weight from day one and cannot be dismissed as a working group's musings. She names four core seats: the quality lead who owns the ISO 5060 evaluation, the terminologist who owns the termbases, the senior localization engineer who owns the engine integrations and the TMS controls (pointedly, the same function whose manager made the unilateral swap, now given a real seat so the impulse to change engines flows to the table instead of around it), and a lead PM who represents the delivery reality and the client pressure. Because her company handles pharmaceutical and financial content, she adds a compliance representative as a standing seat rather than an occasional guest, since the MT-forbidden policy is partly their call. She keeps it to six, small enough to decide, broad enough to see every face of the decisions they will make.
Step Two: Write the One-Page Charter
She writes a charter that fits on a page. Purpose: to set and enforce policy on the use of MT and LLMs across the localization operation so that speed never ships an ungoverned risk. Scope: the five domains, engine and tool approval, quality tiers and the delivery gate, the MT-forbidden policy, incident review, and metrics, and an explicit note that the group does not manage individual projects or staffing, so it neither overreaches nor blurs into line management. Authority: the plain sentence that within scope the group's decisions are binding on the operation and change through the group. Decision rule: consensus where the seats agree, chair's call where they do not, so one dissenter cannot deadlock a needed decision. Escalation: a documented, executive-signed exception path for overrides. Cadence: monthly standing meeting plus a chair-led fast track for urgent matters. She gets it signed by her own VP so the "binding" clause has teeth above her as well as below.
Step Three: Populate the Standing Agenda from the Real Backlog
Rather than starting from a blank page, she seeds the first meetings with the actual mess. The engine the manager swapped in goes to the engine-approval domain: the quality lead scores it on the company's real regulated content, and the group discovers it does make more accuracy errors on financial disclosures, so it is approved for marketing content and forbidden on regulated, and the German pipeline is reverted for regulated jobs. The pharma-copy routing goes to the MT-forbidden domain: the group formally classifies pharmaceutical package inserts and financial disclosures as MT-forbidden, and tasks the engineering seat with building the TMS block so that content type structurally cannot be assigned to a machine workflow again. The shipped disclosure error goes to incident review as the group's first case, examined not to punish the linguist but to find the systemic gap, which turns out to be all three failures at once, an unapproved engine on forbidden content that the gate did not catch, and each gap becomes a policy change.
Step Four: Build Enforcement into the Tools, Not the Reminders
This is the step that separates a governance group from a memo. The VP insists that every policy the group sets is enforced by the system rather than by people remembering it. The approved-engine list becomes the only set of engines selectable in the TMS for each content type. The MT-forbidden classification becomes a hard block: forbidden content cannot be routed to an engine, and the tool refuses rather than warns. The delivery gate, with its absolute Critical rule, is wired into the pipeline so a Critical cannot ship. The metrics dashboard is rebuilt to show the risk axis, critical-error rate and terminology conformance alongside throughput and cost, so leadership sees the dual-axis truth. The principle is that enforcement lives in the system and the group owns the system's rules, because a policy that depends on a tired PM under deadline remembering to do the right thing is a policy that will be violated on exactly the day it matters most.
Step Five: Run the Cadence and Let the Group Learn
She runs the first monthly meeting and discovers the agenda is full, because the backlog of ungoverned decisions was large. By the third month the standing meeting has found its rhythm: a pass through engine requests, incidents, metrics, and strained policies, most months short because most months are quiet, with the occasional fast-track exception handled between meetings and ratified at the next. The decision record fills up, and it starts paying off in unexpected places: a new PM joins and reads the record to learn what is approved for what, rather than asking around; a client asks how the operation governs AI quality, and the VP answers with the charter and the decision log instead of a shrug; an auditor asks who decided that pharma is human-only, and the answer is a dated, attributed decision rather than a linguist's recollection. The vacuum the VP found on that three-decision week is gone, not because those three problems were solved, but because the body that would have caught them now exists and keeps catching the next ones.
What the Group Lets the Operation Say
The payoff is a set of sentences the operation could not honestly say before. To a client: "Every engine we use on your content was approved by a cross-functional governance group for that specific content type, and here is the dated decision." To an auditor: "Our MT-forbidden policy was set deliberately by a body with quality, engineering, and compliance at the table, it is enforced by a hard block in our TMS, and here is the record." To leadership: "Here is the dual-axis dashboard the governance group maintains, throughput and provable quality together, so you are never looking at a speed number that hides a liability." And internally, to the engineer, the PM, and the linguist who each made a lonely call under pressure: "You are not the one who has to carry this decision alone anymore. There is a table for it, it has your seat on it, and its answer will hold." That last sentence is the quiet one, and it may be the most important, because a governance vacuum does not only ship bad output. It puts impossible cross-functional decisions on individuals who never had the authority or the visibility to make them well, and standing up the group is as much a relief to them as it is a control on the risk.
Key Takeaways
- A governance vacuum fills itself with whoever is nearest. An ungoverned localization-AI decision (which engine, which tier, is this MT-forbidden, why did this ship) does not disappear when no body owns it. It defaults to the person closest to it, who optimizes the one axis they can see and is structurally blind to the other three, and the org learns about those axes only when they fail.
- Governance is a standing cross-functional body, not a document. A policy is a static snapshot that goes stale the day it is written; the governance group is the living engine that keeps the policy true as engines, clients, and people change. An operation with a policy and no group has good intentions with an expiry date.
- The seats are the safeguard, and each represents a face a decision can fail on. Quality (is it correct against ISO 5060), terminology (does the approved language survive), engineering (is it implementable and tag-safe), and PM (where does policy collide with the deadline and the client). An empty seat is exactly where the org gets blindsided; add legal, vendor, and data seats as the content demands.
- The group must hold three kinds of authority: to set policy, to enforce it, and to review and revise it. A body that can only recommend loses to deadlines every time. Enforcement must live in the tools (the TMS refuses MT-forbidden content, unapproved engines are unselectable, the gate blocks a Critical) so that violating policy is hard rather than merely discouraged.
- The group governs five domains: engine and tool approval (scoped to content type, judged on domain-relevant critical-error rate not a generic benchmark), quality tiers and the delivery gate (with the absolute one-Critical-fails rule), the MT-forbidden policy (a whole-table risk judgment, enforced as a hard block), incident review (blameless, aimed at the systemic gap not the person), and metrics (insisting the risk axis is visible so speed cannot hide liability).
- The charter makes the group real, and its precision on authority is everything. A one-page charter names purpose, scope (including what the group does not govern), seats and chair, a plain binding-authority clause, a decision rule so one dissenter cannot deadlock, an escalation path for documented overrides, and a cadence. A charter that hedges on whether decisions are binding has already lost.
- Two rhythms keep it alive: a monthly standing meeting and a fast-track exception path. Monthly is frequent enough that decisions do not pile up and rare enough not to waste senior time; the exception path handles the decision that cannot wait so the group never becomes a bottleneck people learn to bypass. Every decision and exception is recorded, attributed, and dated, or the group recreates the hero problem inside a committee.
- The worked build has an order: name the seats and an executive chair, write the one-page charter, seed the agenda from the real backlog of ungoverned decisions, build enforcement into the tools rather than reminders, then run the cadence and let incident review teach the program. The payoff is a set of sentences the operation can finally say to a client, an auditor, leadership, and its own people about how AI quality is actually governed.
Skill.re