โ†
AI for ESG & Sustainability Reporting
Visionary ยท M10 ยท lesson 10 of 16 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Running Reporting AI Pilots to Evidence
๐Ÿ“–
now learning

Running Reporting AI Pilots to Evidence

15 min

The pilot review deck lands on your desk and it is glorious. Ninety percent time saved on the first draft. A materiality matrix generated in an afternoon that used to take three weeks. Screenshots of a model producing clean ESRS narrative on demand. Everyone in the room is nodding. Then your assurance partner, who you invited on purpose, asks one question: "For this figure the pilot produced, can you show me where it came from?" The screen goes to the next slide. There is no next slide for that. The demo was real. The evidence trail was never built. And in that silence you learn the thing this lesson exists to teach: a pilot that produces a demo has proven nothing. A pilot that produces an evidence trail has proven the only thing that matters.

The Demo Trap

Most reporting AI pilots are designed, without anyone deciding it on purpose, to answer the wrong question. They are built to answer "can the model do the task?" The answer is almost always yes. Modern models can draft, cluster, extract, summarize, and reconcile with startling fluency, so a pilot aimed at that question is a foregone conclusion dressed up as an experiment. It generates enthusiasm, a favorable slide deck, and a dangerous false confidence, because it has tested the half of the problem that was never in doubt.

The question a pilot in a regulated, externally-assured function must answer is different: "can the model do the task in a way that survives assurance?" That is not a foregone conclusion at all. It is the entire risk. And it is invisible in a demo, because a demo shows you the output and hides everything an assurer actually cares about: where the number came from, how it was derived, whether it can be reconstructed, and who signed off. A pilot that only shows outputs is measuring the wrong variable with great precision.

The correction is a redefinition, and it should be written into the pilot charter before any work begins. A successful pilot is one that produced an evidence trail, not just a demo. Speed is a finding, not the finding. The pilot succeeds if, at the end, you can hand your assurance partner a file that reconstructs every figure the pilot produced, from raw data to output, without the analyst in the room. If you cannot, the pilot failed, no matter how fast it was, because it proved the model can generate numbers your organization cannot defend, which is the worst possible thing to prove.

It helps to notice why the demo trap is not a mistake anyone makes on purpose. It is a gravitational pull. A demo is exciting to build, easy to present, and flattering to everyone involved: the vendor whose model performed, the analyst who ran it, the sponsor who funded it. An evidence trail is none of those things. It is unglamorous, it exposes the workflow's weak points, and it is capable of embarrassing the very people who championed the pilot. So the natural momentum of any pilot, left ungoverned, bends toward the demo and away from the evidence, which is exactly why the redefinition has to be imposed deliberately and in writing at the start, rather than hoped for at the end. The charter is not bureaucracy. It is the counterweight to a bias every pilot has built into it.

There is also a subtler cost to the demo pilot that is worth naming, because it is the one that compounds. A demo pilot that succeeds creates organizational momentum to scale the very workflow that was never tested for assurability. Enthusiasm becomes a roadmap, the roadmap becomes a budget, and the budget becomes a deployment, all riding on a pilot that measured the wrong variable. By the time an assurer finally asks where a number came from, the untraceable workflow is no longer a pilot you can quietly stop. It is embedded across the report. The evidence redefinition, applied at pilot stage, is the cheapest possible place to catch a problem that becomes exponentially more expensive at every later stage.

The output of a pilot is not the artifact it generated. It is the evidence file that proves the artifact can be trusted.

Designing the Pilot Backwards From the Assurance File

Because the goal is an evidence trail, you design the pilot backwards from the file the assurer will read, not forwards from the model's capability. Start by writing down, before you build anything, what the assurance file for this workflow would need to contain if it were live in a real disclosure. That list becomes the pilot's actual specification.

For an emission-factor workflow the file must show, for every factor the pilot used, a named and dated source database and the category match. For an activity-data extraction workflow it must show, for every field, the source document, page, and location it came from. For an estimation workflow it must show the method, the assumptions, and the disclosed uncertainty on every estimated figure, and it must keep those figures visibly distinct from measured ones. For a narrative workflow it must show a per-claim link from each quantitative statement back to its supporting evidence. Whatever the workflow, write the file requirements first.

Now the pilot has a real specification, and it is a far more demanding one than "make the model produce the output." The pilot is not testing whether the model can draft a narrative. It is testing whether the model, embedded in a workflow you designed, can draft a narrative and simultaneously emit the evidence that makes each claim defensible. That is the thing that is genuinely uncertain, genuinely worth piloting, and genuinely predictive of whether the workflow can go live. Designing backwards from the file also has a quiet benefit: it forces you to confront, in the safe environment of a pilot, exactly the questions the assurer will ask in the dangerous environment of a live engagement.

The Gates a Pilot Must Pass

A pilot without gates is a demo with a longer runtime. Gates are the predefined checkpoints where the pilot must prove a specific property or stop, and they are what turn a pilot from a source of enthusiasm into a source of evidence. Set them before the pilot starts, in writing, so no one is tempted to move the goalposts once the outputs look impressive.

Gate one: the provenance gate

Take a sample of the figures the pilot produced and try to trace each one to its source using only what the workflow captured. If any figure cannot be traced without asking the analyst what the model "was thinking," the workflow fails the gate. This is the cheapest and most revealing gate, and many pilots die here, which is exactly what a pilot is for: to kill an indefensible workflow before it becomes a live one.

Gate two: the reconstruction gate

Hand the pilot's file to someone who was not involved and ask them to rebuild a number from raw data to output. If they can, the workflow has demonstrated reconstructability, the property an assurer tests directly under a limited or reasonable assurance engagement. If they cannot, you have learned that the workflow depends on knowledge that lives in a person's head rather than in the file, which is precisely the dependency that fails assurance.

Gate three: the labeling gate

Check that primary data, secondary data, and estimates are distinguishable in the output, each carrying its method. A pilot that produces a stream of numbers where a supplier-reported figure and an industry-average estimate look identical has failed, even if every number is individually plausible, because the assurer cannot rely on a file that blurs measured and estimated data.

Gate four: the accountability gate

Confirm the workflow includes and records a named human review and sign-off at the points where judgment is exercised. A pilot that quietly removed the human to demonstrate end-to-end automation has not demonstrated a viable workflow. It has demonstrated an indefensible one, because "the model recommended it" is never a defense.

Pass all four gates and you have something rare: evidence that the workflow is both faster and assurable. Fail one, and you have learned something specific and valuable about what must change before it goes live. Either way the pilot has produced knowledge, not just an artifact. A pilot that cannot fail is not a pilot.

Two disciplines keep the gates honest. The first is that the pass criteria are written and fixed before anyone sees an output. The single most common way a pilot's gates get quietly defeated is not fraud, it is drift: the outputs come back impressive, everyone wants the pilot to succeed, and the definition of "traced well enough" softens by degrees until it means "we are confident it is roughly right." Fixing the criteria in advance removes that degree of freedom. The second is that at least one gate, ideally the reconstruction gate, is run by someone with no stake in the pilot's success. A pilot team reconstructing its own numbers will unconsciously supply the missing context from memory and conclude the file is sufficient when it is not. An uninvolved reviewer, working only from the file, tests the thing that actually matters, which is whether the file stands alone. If it needs the pilot team in the room to make sense, it will need them in the room during the assurance engagement too, and they will not be allowed there.

A Worked Example: Two Pilots, One Task

Two teams at a large undertaking in CSRD scope pilot the same use case: AI-assisted drafting of an ESRS narrative datapoint on climate transition. Same model, same task, opposite outcomes, and the difference is entirely in the discipline.

Team One runs the demo pilot. They prompt the model with a short brief, it produces two pages of fluent, confident narrative about the company's transition progress, and they screenshot it for the review. The deck says: draft time cut from two days to twenty minutes, ninety percent saved. In the review, the assurance partner reads the draft and asks about a sentence stating the company reduced a category of emissions by a specific percentage against its target. Where did that number come from? Nobody knows. The model produced it. It is plausible. It may even be roughly right. But there is no source in the file, and worse, on inspection the narrative has softened a negative impact the company is obligated to disclose, because the model optimized for a smooth story. Team One's pilot succeeded at generating a draft and failed as a pilot, because what it actually proved is that this workflow can put an unsupported figure and a greenwashing risk into a disclosure at ninety percent lower cost. That is not a win. That is a documented liability.

Team Two runs the evidence pilot. Before building anything they write the file requirement: every quantitative claim in the draft must carry a footnote pointing to the source figure in the evidence pack, and the workflow must draft only from that pack, never from open generation. They build the workflow that way. In the review, the same draft appears, slightly less silky, and every number in it is footnoted to a specific figure in the inventory or the prior disclosure. The assurance partner asks the same question about the reduction percentage. Team Two clicks the footnote. There is the source. Then they walk the four gates: provenance, yes, every claim linked; reconstruction, they hand the file to an uninvolved colleague who rebuilds the reduction figure; labeling, the one estimated figure in the draft is marked as an estimate with its method; accountability, the reviewer's sign-off is recorded. The time saving is smaller than Team One's headline, maybe seventy percent rather than ninety. And it is real, because it is a saving on work the organization can actually keep.

Here is the lesson the two pilots teach together. Team One's pilot looked more successful and was a failure. Team Two's pilot looked more modest and was the only real success, because it produced the one thing a pilot in a regulated function exists to produce: proof that the workflow can go live without importing a restatement. When you take these two results to a governance body, Team Two's slightly-slower, fully-traced result is the one you scale. Team One's dazzling, untraceable result is the one you stop, and the stopping is the pilot working exactly as intended.

Making the Lessons Transfer

A pilot that produces evidence for itself is good. A pilot whose lessons transfer to the next workflow, the next framework, and the next business unit is what a transformer is actually after, because at enterprise scale you cannot pilot everything from scratch. Transferable lessons come from documenting the pilot in a form that outlives it.

Capture three things from every pilot, win or lose. First, the file specification you wrote backwards from the assurer, because that specification is reusable: the next narrative workflow needs the same per-claim footnoting, the next estimation workflow needs the same method labeling. Second, the gate results, including exactly which gate a failed pilot failed and why, because that is the most valuable knowledge a pilot produces and it is the first thing that gets lost when only the successes are written up. Third, the controls that made it assurable, stated as reusable requirements rather than as one-off fixes, so they become the default starting point for the next pilot instead of being rediscovered painfully each time.

A word on the failed pilots, because a mature program treats them as its most valuable output rather than an embarrassment to bury. A pilot that fails a gate has told you, cheaply and safely, exactly where a class of workflow breaks. If your narrative pilot failed the provenance gate because the model could not reliably attach a source to each claim, that is not a disappointment. It is a discovery that saves you from failing the same way in the CBAM narrative, the ISSB narrative, and the materiality rationale, all of which share the structure. The failed pilot is a scout that mapped the minefield with a single controlled explosion instead of a series of live ones across your report. A program that only writes up its successes throws that map away and re-steps on every mine at scale. So the honest documentation of failure is not a compliance nicety. It is how a transformer converts one pilot's pain into every future pilot's head start.

Documented this way, your pilots stop being isolated demos and become an accumulating body of evidence about which reporting AI workflows survive assurance and why. That body of evidence is what lets you scale deliberately rather than hopefully. It is also, not incidentally, exactly what an assurer wants to see when deciding how much to rely on your AI-assisted workflows across the whole report: a track record of pilots that were gated on assurability, documented honestly, and scaled only when they passed. The pilot discipline, in the end, is not a hurdle before the real work. It is the mechanism by which a regulated discloser earns the right to move fast.

Key Takeaways

  • Redefine success before the pilot starts: a successful pilot produced an evidence trail, not just a demo. Speed is a finding, not the finding.
  • Most pilots answer the wrong question, "can the model do the task," which is a foregone conclusion. The real question is "can it do the task in a way that survives assurance."
  • Design the pilot backwards from the assurance file. Write down what the file would need to contain if the workflow were live, and make that the pilot's specification.
  • Gate the pilot on four properties, in writing, before it starts: provenance, reconstruction, labeling, and accountability. A pilot that cannot fail is not a pilot.
  • The reconstruction gate is the assurer's actual test: can an uninvolved person rebuild a number from the file alone. If the workflow depends on knowledge in someone's head, it fails.
  • A dazzling, untraceable pilot is a documented liability, not a win. Stopping it is the pilot working as intended; scaling the slower, fully-traced pilot is the correct outcome.
  • Make lessons transfer by capturing the file specification, the gate results including failures, and the assurability controls as reusable requirements for the next pilot.
  • An accumulating record of honestly-gated pilots is what earns an assurer's reliance on your AI workflows and lets a regulated discloser scale deliberately rather than hopefully.