Enterprise AI Policy for a Regulated Discloser
The assurance partner slides a single page across the table. It is your enterprise AI policy, printed from the intranet, and she has a yellow highlighter on three lines. "This says your team may use generative AI to draft ESRS narrative. Show me the control that stopped it from inventing the 2027 target on page 41." The room goes quiet. The target is real, it turns out. But nobody in the room can point to the sentence in the policy that would have caught it if it were not. That gap, not the target, is the finding.
Why an AI Policy Is a Control, Not a Memo
Most enterprise AI policies read like a staff memo: be careful, do not paste confidential data into public tools, use good judgment. That is fine for a marketing team drafting social posts. It is not fine for a company whose sustainability disclosure is read, line by line, by an external assurer under a limited-assurance engagement, and can be reopened by a regulator years later. For a regulated discloser, the AI policy is a control. It is part of the same control environment the assurer tests when they ask how you know your Scope 3 number is right, and it will be requested, read, and challenged like any other control.
The distinction matters because a control has to do two things a memo does not. It has to be specific enough that a person can follow it without interpreting your intentions, and it has to leave evidence that it operated. "Use good judgment" fails both tests. An assurer cannot test judgment. They can test whether every AI-assisted figure in the inventory carries a provenance tag, whether the model was configured to refuse when it could not cite a source, and whether a named human signed the datapoint. A policy written as a control tells your people exactly what to do and, just as importantly, tells the assurer exactly what to look for.
The 2026 backdrop makes this non-optional. CSRD survived the Omnibus: Directive (EU) 2026/470, in force 18 March 2026, kept the largest undertakings in scope, more than 1,000 employees and more than EUR 450M turnover, with member-state transposition due 19 March 2027. Roughly 73% of large global companies now obtain external assurance on at least some sustainability disclosures, up from 51% in 2019 (a number worth verifying against the current IFAC study, but directionally the trend is one way). The companies left in scope are the biggest ones, where a failed disclosure is a board-level event. When your board says "use AI for everything," the AI policy is the document that lets you say yes without inheriting a restatement.
There is a second reason the policy has to be a control rather than a memo, and it is about behavior under pressure. A memo describes the ideal; a control governs what actually happens at 11pm the night before a filing deadline, when an analyst is missing one supplier's number and a fluent model is offering a plausible-looking substitute. In that moment, "use good judgment" is an invitation to fill the gap and move on. A control, by contrast, has already decided: the tool refuses to generate an ungrounded figure, the provenance gate blocks the datapoint, and the analyst is routed to a documented escalation instead of a quiet fabrication. The whole point of writing the policy as a control is to make the right behavior the path of least resistance precisely when the temptation to cut a corner is highest. Policies that only work when everyone is calm and unhurried are not controls; they are wishes.
Allowed and Forbidden Uses, Named by Task
A policy that says "AI may be used to assist with sustainability reporting" is useless, because it treats a single word, AI, as if it named one activity. It does not. Extraction, classification, estimation, and generative drafting are four different jobs with four different failure modes, and your policy has to draw the lines at the level of the task, not the technology. The reader who has come up through this program already knows this distinction cold; the policy is where it becomes an enforceable rule for the whole enterprise.
Write the allowed and forbidden lists against concrete reporting tasks. The point is that a carbon accountant, an analyst, and a disclosure lead should each be able to read the policy and know, for their specific task, whether AI is permitted, permitted with conditions, or forbidden.
| Reporting task | AI status | Binding condition |
|---|---|---|
| Extracting activity data from invoices, bills, supplier files | Allowed | Source location preserved on every field; human reconciles to prior period |
| Drafting supplier questionnaires and triaging responses | Allowed | No response reclassified from estimate to primary by the model |
| Clustering stakeholder and impact inputs for materiality | Allowed with conditions | Clustering is input to a human materiality decision, never the decision; basis documented |
| Suggesting an emission factor | Allowed with conditions | Factor must resolve to a named, dated, authoritative library entry or be refused |
| Generating an estimate where primary data is unavailable | Allowed with conditions | Labeled as an estimate with method and uncertainty; never presented as measured |
| Drafting ESRS or ISSB narrative datapoints | Allowed with conditions | Every claim, figure, and target verified against evidence before it ships |
| Inventing, inferring, or "reasonably assuming" a target, factor, or activity figure not in the evidence base | Forbidden | None. This is fabrication, whatever it is called |
| Autonomous filing or sign-off with no human decision on the record | Forbidden | None. Accountability cannot be delegated to a model |
Notice that almost nothing is flatly allowed and almost nothing is flatly forbidden. The interesting cases live in the middle, "allowed with conditions," because that is where the real work of the policy happens. The condition is the control. An assurer does not care that you permit AI factor lookup; they care that your policy forces every suggested factor back to a named source or forces a refusal. Write the conditions as if the assurer will read them, because they will.
One more discipline separates a task-based table that works from one that merely looks thorough: the conditions have to be observable. "Human reconciles to prior period" is a good condition because an internal auditor can pull the reconciliation and see it happened. "Analyst uses appropriate care" is not, because there is nothing to inspect. When you draft each condition, ask what artifact its operation would leave behind, and if the answer is "nothing," rewrite it until it produces evidence. A control that leaves no trace cannot be tested, and a control that cannot be tested is, in an assurance engagement, indistinguishable from no control at all. This is the same standard the assurer applies to your financial controls, brought over to the reporting-AI world where most organizations have never applied it before.
Provenance Requirements as a Policy Clause
Provenance is the spine of the whole program, so it deserves its own clause with teeth. The iron rule is simple to state and expensive to violate: every figure you publish must trace to evidence, and "the AI estimated it" is not evidence. The policy has to turn that rule into a requirement a system can enforce and an assurer can test.
A workable provenance clause specifies, for every datapoint that reaches the inventory or the disclosure, a minimum set of attributes that must exist before the datapoint is allowed to move forward. Name them explicitly:
- Source. The named document, database entry, or supplier response the figure came from, with enough locator detail that someone else can find it.
- Data type. Primary (measured or supplier-reported) versus secondary (estimated, averaged, spend-based). A primary figure and an estimate must never look identical in the file.
- Method. For any estimate, the calculation method and the emission factor used, each factor resolving to a named, dated library entry.
- Human decision. The named person who accepted, adjusted, or overrode the value, and when.
- AI involvement. Whether AI touched the datapoint and how (extracted, suggested, drafted), so the file is honest about where the machine was in the loop.
A number without provenance is not a fast number. It is an unfiled restatement waiting for an assurer to pull the thread.
The critical design choice is that provenance is a precondition, not an afterthought. The policy should state that a datapoint lacking any required attribute cannot be included in a disclosed figure. That single sentence converts provenance from a nice-to-have into a gate, and it is the sentence the assurance partner in our opening scene was looking for and did not find.
The Human-Accountability Boundary
The cardinal rule of the program is that disclosure accountability stays human. "The model recommended it" is never a defense to an assurer or a regulator. The enterprise policy has to draw that boundary as a bright line and then, crucially, name who stands on the human side of it for each artifact.
Drawing the boundary well means being precise about what a human must personally decide versus what a human may accept from AI after review. A useful way to structure it is by the weight of the judgment. Boundary decisions, what is material, where the organizational and operational boundary sits, which Scope 3 categories are in scope, must be made by a named, accountable person and cannot be produced by a model. Verification decisions, whether a specific extracted figure matches its source, whether a factor is the right one, can be AI-assisted but must be signed by a human who takes responsibility for the specific value. Drafting, the actual prose of a narrative datapoint, can be AI-generated, but the person who publishes it owns every claim in it as if they had written it themselves.
Make the accountability concrete by role. The policy should map artifacts to accountable owners: the materiality conclusion to the person who chairs the materiality assessment, the GHG inventory to the carbon accounting lead, the CBAM declaration to the authorised declarant, the full statement to the CSO and controller who sign it. An assurer's first question about any AI-assisted number is "who is responsible for this?" The policy should answer before they ask.
Model-Change Control and the Reproducibility Problem
Here is a failure mode that catches organizations who otherwise do everything right. An assurer, mid-engagement, asks you to reproduce how a figure was derived. You go back to the AI-assisted step, run it again, and get a different answer, because the underlying model was silently upgraded three weeks ago. Nothing was fabricated. But you can no longer reconstruct your own number, and reconstructability is exactly what limited and reasonable assurance test. The version that produced your disclosed figure no longer exists.
This is why a regulated discloser's AI policy needs a model-change control clause that a general-purpose corporate AI policy would never think to include. The clause should require, at minimum, that the model and configuration used for any reporting task be recorded with the output, that model or version changes affecting reporting workflows go through a documented change process before they take effect, and that a superseded model's behavior can be described well enough to explain a prior-year figure during a restatement or re-assurance.
You do not need to freeze models forever, which is neither realistic nor desirable. You need to know which model produced which number, and you need changes to be deliberate rather than invisible. The distinction between a controlled change and a silent one is the difference between a clean walkthrough and an assurer writing "unable to verify" next to a figure that is, in fact, correct.
The "Cite or Refuse" Default
The most powerful single line in a regulated discloser's AI policy is a default behavior: when the model cannot cite a source from the approved evidence base, it must refuse rather than generate. This flips the model's natural tendency. Left alone, a generative model will always produce a fluent, confident answer, and a plausible invented emission factor looks exactly like a real one. The whole danger of AI in disclosure is that fabrication and measurement are indistinguishable in the output. "Cite or refuse" makes refusal the safe default and generation the exception that requires grounding.
In practice this becomes a policy requirement that reporting AI tools be configured, through system prompts, retrieval grounding, and tool design, so that the model answers from your factor library and source documents, not the open web or its own training memory, and returns "I cannot find a source for this" rather than a manufactured number. The policy states the default; the L2 and L3 techniques in this program are how it is implemented; the assurer sees the result as a workflow that does not produce unsupported figures.
The behavioral payoff is enormous. A team operating under "cite or refuse" stops treating AI output as an answer to be checked and starts treating it as a grounded draft that already carries its sources. The verification burden drops, and the failure mode that ends careers, the hallucinated factor that sails into a public, assured number, is designed out rather than caught after the fact.
It also changes what a refusal means culturally. In most organizations, a model saying "I cannot answer that" feels like a failure of the tool. Under a well-written policy, a refusal is the control working. It is the moment the system declined to manufacture a number and handed the decision to a human, exactly as designed. Leaders who understand this stop measuring their reporting AI by how often it produces an answer and start measuring it by whether it refuses when it should. A tool that never refuses on a disclosure task is not a confident tool; it is an ungrounded one, and it is quietly producing figures that will not survive the assurer. The policy should say this out loud, so that a refusal is read as evidence of a working control rather than a defect to be tuned away.
The Minimum Contents Checklist
Pulling the threads together, a regulated discloser's AI policy is not defensible unless it contains, at a minimum, each of the following. Treat this as the checklist you hand to whoever drafts or reviews the document, and treat any missing item as a gap the assurer will find before you do.
- Scope and definitions. Which tools, which reporting tasks, and which data are governed, named specifically enough that no one has to guess whether the policy applies to them.
- Task-based allowed and forbidden uses. The distinctions between extraction, estimation, and drafting, each with observable conditions, and the bright-line prohibitions on fabrication and autonomous sign-off.
- Provenance requirements. The attributes every datapoint must carry and the statement that missing any of them blocks inclusion in a disclosed figure.
- The human-accountability boundary. What a human must personally decide, and the artifact-to-owner map that names who is accountable for each output.
- Model-change control. Recording model and configuration with outputs, the change process for reporting workflows, and retention of superseded behavior for restatement.
- The cite-or-refuse default. The requirement that reporting AI answer from the approved evidence base and refuse rather than generate ungrounded figures or claims.
- Escalation and exception handling. What happens when the tool refuses, when a condition cannot be met, and who may grant a documented exception.
- Review and evidence. How the policy's operation is tested, how often it is reviewed, and where the evidence of its operation lives for the assurer to inspect.
Notice that this checklist is not long, and that is deliberate. A policy that runs to forty pages of aspiration is harder to follow and harder to test than a tight one built entirely of controls. The measure of a good reporting-AI policy is not its length; it is whether every clause names a behavior someone can perform and an artifact an assurer can inspect.
Worked Example: Turning a Weak Policy Into a Control
Watch a real clause evolve. A large manufacturer in CSRD scope had this line in its enterprise AI policy: "Employees may use approved AI tools to improve efficiency in sustainability reporting, provided outputs are reviewed for accuracy." It sounds responsible. It is nearly worthless as a control, and here is why the assurer tore it apart.
"Approved AI tools" was undefined, so nobody could say which tools were in scope. "Improve efficiency" named no task, so the policy applied equally to invoice extraction and to inventing targets. "Reviewed for accuracy" named no reviewer, no standard, and left no evidence, so the assurer could not test whether review happened. And there was no mention of provenance, no data-type labeling, no accountability owner, no model versioning, and no refusal default. The policy would have permitted the exact failure, an AI-drafted target the company never set, that it was supposed to prevent.
Now the rewritten version, built as a control. The company replaced the single sentence with a task-based allowed/forbidden table, a provenance clause making the five attributes a precondition for inclusion, a role-mapped accountability boundary, a model-change control requirement, and a "cite or refuse" default configured into the reporting tools. The efficiency the original clause promised did not go away; it got faster, because grounded output needed less rework. But now, when the assurer asked "show me the control that would have caught an invented target," the disclosure lead pointed to the provenance precondition (no source, no inclusion), the accountability owner (the datapoint was signed), and the refusal default (the tool was configured not to generate a target it could not ground). The finding closed. The board got its AI, the assurer got a testable control, and the same policy did both jobs.
The lesson is the one that runs through the whole program: the discipline that makes an AI output traceable is the same discipline that makes it assurable. A policy written as a control captures speed and tightens the audit trail, because they are the same move.
Key Takeaways
- For a regulated discloser, the enterprise AI policy is a control the assurer tests, not a staff memo, so it must be specific enough to follow and must leave evidence that it operated.
- Draw allowed and forbidden uses at the level of the task (extraction, estimation, drafting), not the technology; the interesting rules live in "allowed with conditions," where the condition is the control.
- Make provenance a precondition for inclusion: source, data type (primary vs. secondary), method and factor, human decision, and AI involvement must all exist before a datapoint reaches a disclosed figure.
- Keep accountability human and named: boundary and materiality decisions cannot be produced by a model, and every artifact maps to an accountable owner who answers "who is responsible for this?"
- Add model-change control that general corporate AI policies lack, so you can always say which model produced which number and can explain a prior-year figure during restatement or re-assurance.
- Set "cite or refuse" as the default so the model returns "no source found" instead of a fluent invented factor, designing out the hallucination that ends careers rather than catching it later.
- The 2026 stakes are real: CSRD (Directive (EU) 2026/470, in force 18 March 2026) keeps the largest undertakings in scope, and roughly 73% of large companies now obtain external assurance, a figure worth verifying but pointing one direction.
- The policy lets you say yes to the board's "use AI for everything" without inheriting a restatement, because the discipline that makes output traceable is the same discipline that makes it assurable.
Skill.re