โ†
AI for ESG & Sustainability Reporting
Aware ยท M16 ยท lesson 16 of 19 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
What AI Is and Isn't for Disclosure
๐Ÿ“–
now learning

What AI Is and Isn't for Disclosure

15 min

It is March, the assurance partner is in the room, and she points at one number in the draft sustainability statement: a Scope 3 figure, 412,000 tonnes of CO2 equivalent, sitting in the purchased-goods line. "Walk me from this number back to where it came from," she says. The disclosure lead opens the file. The number was produced by an AI tool the company licensed last year, the one the vendor said would "do your CSRD." There is no source behind it. There is only a confident sentence the model wrote. That silence in the room is the whole subject of this lesson.

The Marketing Story That Gets People Fired

Walk any sustainability conference floor in 2026 and you will hear the same promise in a dozen accents: our AI does your CSRD. It is a seductive sentence because it is almost true and completely wrong at the same time. AI can touch nearly every step of a sustainability disclosure. It cannot own a single number in it. The gap between "touch" and "own" is where careers end and restatements begin.

To see why, you have to drop the marketing frame entirely and replace it with a colder, more useful question. Not "can AI do my reporting" but "which specific, separable job inside my reporting is AI actually doing, and what happens to that job when an external assurer reads the output." A sustainability disclosure is not one task. It is a chain of very different tasks: deciding what is material, gathering activity data, choosing emission factors, calculating an inventory, drafting narrative, tagging datapoints, and assembling an evidence file. The phrase "AI does your CSRD" smears all of those into one undifferentiated blob, and that blob is exactly the thing no assurer will accept.

Here is the term that anchors the whole program. Assurance is the independent, external examination of your sustainability disclosures by an accredited third party, who issues a formal opinion on whether the information is fairly stated. As of 2026, 73% of large global companies obtain external assurance on at least some sustainability data, up from 51% in 2019, and greenhouse gas emissions are the most-assured category. Why you care: an assured number is an audited number. If you cannot show the assurer where it came from, it is not a number, it is a liability.

AI can draft your disclosure. It cannot defend it. The defending is still your job, and the assurer is the examiner.

Four Jobs, Not One Machine

The single most useful thing a reporting professional can learn about AI is that the word "AI" hides at least four different jobs, each with a different relationship to evidence. Conflate them and you will trust a tool to do a job it was never doing. Separate them and you can place each one precisely on your reporting cycle, with the right verification attached. The four jobs are classification, extraction, estimation, and generative drafting.

Classification: Sorting and Labeling

Classification means the model puts an input into a category. Is this supplier line item a "purchased good" or a "capital good"? Does this stakeholder comment relate to climate, water, or labor? Is this invoice a fuel purchase or a maintenance contract? The model is not inventing a number and not pulling a fact out of a document. It is sorting. Classification is genuinely useful and relatively low-risk, because a human can spot-check a sort quickly: pull twenty rows, see if the labels are right. The failure mode is a mislabel, and a mislabel is usually visible. This is AI at its most load-bearing and least dangerous, the quiet workhorse of a reporting cycle.

Extraction: Pulling a Fact From a Source

Extraction means the model reads a document you gave it and pulls out a specific value that is actually in that document: the kilowatt-hours on a utility bill, the litres of diesel on a fuel invoice, the tonnage on a supplier's questionnaire. The crucial word is in. A correct extraction is a value that exists in the source and is faithfully copied. The job has a source by definition, which is what makes it the friendliest AI job for an assurer: you can point to the bill and the number on it. The failure mode is misreading: the model reports 14,200 when the bill says 142,000, or grabs the wrong line. That is catchable, because there is a document to check against.

Estimation: Computing a Proxy

Estimation means the model computes a stand-in for a number you do not have. You cannot get a small supplier's actual emissions, so a tool estimates them from how much you spent with that supplier and an industry average. This is where the danger sharpens. An estimate is not measured data. It is a defensible guess if, and only if, it is labeled as an estimate, its method is disclosed, and its uncertainty is acknowledged. The catastrophic failure mode is laundering: the estimate gets written into the inventory looking exactly like a measured figure, and nobody downstream can tell the difference. That is the move that fails assurance.

Generative Drafting: Writing Prose

Generative drafting means the model writes language: an ESRS narrative paragraph, a policy summary, a description of your transition plan. This is the job the marketing demos love, because the output looks finished and impressive. It is also the job with the most spectacular failure mode: fabrication. A drafting model will, with total fluency, write that the company has "committed to a 2030 net-zero target" when no such target exists, or insert an emission factor that sounds authoritative and was never in any database. The prose is plausible. Plausibility is not truth, and an assurer does not grade on fluency.

Why Separating the Jobs Is the Whole Skill

Notice what happens when you stop saying "AI" and start naming the job. "The AI did my Scope 3" becomes four sharper, answerable questions. Did it classify spend lines into the right categories? Did it extract tonnages from supplier responses that actually contain them? Did it estimate the missing suppliers, and is that estimate labeled? Did it draft the methodology narrative, and does every claim in that narrative trace to something real? Each question has a different verification and a different risk. The blob has none.

This is also how you immediately spot a tool that is doing a riskier job than it admits. A vendor says their platform "calculates your Scope 3." Calculation sounds like arithmetic, safe and mechanical. But if the platform is silently estimating the 60% of your suppliers who never responded, it is doing estimation, and if those estimates are not labeled, it is laundering. The word on the box ("calculates") hid the job that matters ("estimates the gap"). Naming the four jobs is your X-ray.

When someone says "the AI did it," your only safe response is: which of the four jobs, and where is the source.

Where AI Is Actually Load-Bearing in 2026

Drop the "automate everything" fantasy and a far more useful picture appears: a one-page map of the reporting cycle with the real, defensible AI use case in each box. This is the artifact worth keeping on your wall.

Reporting stageThe real AI jobWho still owns the answer
Double-materiality assessmentClassification: cluster hundreds of stakeholder and impact inputs into themesThe human decides what is material and documents the basis
Activity-data collectionExtraction: pull kWh, litres, tonnes from bills, invoices, supplier filesThe analyst verifies each value against the source document
Emission-factor selectionExtraction with provenance: retrieve a factor from a named, dated databaseThe carbon accountant confirms the factor and its source
Filling genuine data gapsEstimation: compute a labeled proxy for missing suppliersThe human labels it an estimate, discloses method and uncertainty
Disclosure narrativeGenerative drafting: a first draft of ESRS or ISSB narrative textThe discloser checks every claim and figure against evidence
Datapoint taggingClassification: map content to the right ESRS or ISSB datapointThe reporting lead confirms the tag is correct and complete
Assurance file assemblyExtraction and classification: gather and organize evidenceThe human signs the basis-of-preparation

Read the right-hand column down the page. The answer is always owned by a person. AI moves load on the left; accountability never moves on the right. That is not a limitation to apologize for. It is the design of a regulated disclosure, and it is exactly why "AI does your CSRD" is a category error: CSRD is a system of human accountability into which AI feeds drafts and pulls, never a thing a machine can "do."

Notice too that the same vendor tool might sit in three different boxes at once. A carbon-accounting platform can extract activity data, estimate the gaps, and draft a methodology note in a single run. The danger is that it presents all three as one seamless "calculation." Your job as the discloser is to mentally un-blend the run back into its jobs, because the assurer will. Vendor categories exist (reporting platforms, carbon-footprinting tools, document-extraction engines), but naming a vendor never tells you which of the four jobs produced a given number, and a number is never trustworthy just because a named tool produced it. The obligation does not transfer to the platform.

A Worked Example: Before and After

Return to the Scope 3 figure from the opening, 412,000 tonnes in purchased goods, and watch two versions of the same workflow.

Before (the blob). The team uploads a year of procurement spend to the licensed tool and clicks "calculate Scope 3." The tool returns a clean number: 412,000 tCO2e. It looks authoritative. It goes into the draft statement. When the assurer asks "walk me back to this number," the disclosure lead discovers that the tool classified spend lines, applied spend-based factors to everything, and silently estimated every supplier with no primary data, which was most of them. None of that was labeled. The 412,000 is a wall of unlabeled estimation wearing the costume of a measured figure. The assurer cannot get comfortable. The number is pulled, the timeline slips, and a quiet question about the company's controls is now in the file.

After (the four jobs, named). The same team runs the same tool but treats it as four jobs, not one. Classification: the tool sorts spend into categories, and an analyst spot-checks twenty lines. Extraction: for the suppliers who returned primary data, the tool pulls reported tonnages, and each is traced to the supplier file. Estimation: for non-responding suppliers, the tool produces spend-based estimates, and every one is tagged "secondary, spend-based, estimated" with the method noted. Generative drafting: the tool drafts the methodology paragraph, and the analyst checks that it honestly says, in plain words, how much of the 412,000 is primary versus estimated. Now when the assurer asks the same question, the lead answers in one breath: here is the primary share with sources, here is the estimated share with method and uncertainty, here is the basis-of-preparation. Same tool, same number, completely different fate, because the jobs were separated and each was verified for what it actually was.

The lesson is not that AI is dangerous. It is that an undifferentiated "AI did it" is dangerous, and a precisely named "here is which job AI did, and here is the source" is defensible. The number did not change. The accountability did.

How to Interrogate Any AI Claim in a Meeting

The four-jobs frame is most useful as a set of questions you can fire in real time, when a vendor is presenting, when a colleague forwards an AI output, when the CFO asks why the report is not done yet. You do not need to understand the model's internals. You need to force the conversation from the foggy noun "AI" down to the specific job, and from the specific job to the source. Here is the sequence, in order.

First: which job is this? Make the speaker name it. Is the tool classifying, extracting, estimating, or drafting? If the answer is a vague "it does everything," that is your signal that several jobs are blended and at least one is hiding. Second: where is the source? For an extracted value, demand the document. For a drafted claim, demand the evidence behind the claim, not the eloquence of the sentence. Third: for anything estimated, is it labeled? An estimate that cannot tell you it is an estimate, with a method and an uncertainty, is not yet a defensible number. Fourth: who signs? The answer is always a person, and if no one can name that person, the workflow has a hole an assurer will find.

These four questions are deliberately boring. They are not about whether the AI is impressive. They are about whether its output can survive being read by someone whose job is to disbelieve it until shown evidence. A reporting professional who asks them habitually becomes the person in the room who cannot be sold a number, which is exactly the person a CFO and an assurance partner both want owning the disclosure.

The Vendor Conversation, Rewritten

Picture a demo. The vendor says, "Our platform automatically generates your full Scope 3 inventory and the disclosure narrative." In the old frame you would nod at the speed. In the new frame you translate instantly: "generates your inventory" means it is extracting where suppliers responded, estimating where they did not, and drafting the narrative, three jobs in one sentence, and the dangerous one is the silent estimation. So you ask: for the categories where we have no supplier data, what method does the platform use, and does it label those figures as secondary estimates with their method and uncertainty in the output the assurer sees? If the honest answer is that everything comes out as one clean number, you have just learned that the platform launders estimates by default, and that no demo polish can fix. The same translation works on every tool you will ever evaluate.

The Iron Rule, Stated Once

Everything in this program reduces to one sentence you should be able to recite cold. Every figure you publish must trace to evidence, and "the AI estimated it" is not evidence. AI can accelerate the work of getting to evidence. It cannot be the evidence. An estimate can be defensible, but only when a human labels it, discloses the method, and owns the call. Generative fluency can produce a beautiful paragraph, but a paragraph is not a fact until the fact behind it is shown. Classification can sort your world, but the sort is checked by a person who signs. The assurer reads all of it, and the regulator can reopen the filing. In that world, the most valuable skill is not generating the report faster. It is knowing exactly which of the four jobs you just asked a machine to do, and refusing to let any of them reach the disclosure without a source a human stands behind.

So when the partner points at the next number and says "walk me back," you do not reach for the tool's confidence. You reach for the job. You say: this came from extraction, here is the bill. This came from estimation, here is the label and the method. This sentence came from drafting, here is the evidence behind the claim. That answer, spoken without hesitation, is what this entire program is built to give you.

Key Takeaways

  • "AI does your CSRD" is a category error: a disclosure is a chain of distinct tasks bound by human accountability, not one thing a machine can own.
  • The word "AI" hides four different jobs: classification (sorting), extraction (pulling a fact from a source), estimation (computing a proxy), and generative drafting (writing prose).
  • Each job has its own failure mode: mislabel, misread, laundered guess, and fabrication. Knowing which job you asked for tells you which failure to look for.
  • Extraction is the assurer's friend because a source exists by definition; estimation is dangerous only when it is unlabeled; generative drafting is the most fluent and the most fabrication-prone.
  • The X-ray skill is replacing "the AI did it" with "which of the four jobs, and where is the source," which instantly exposes a tool doing a riskier job than its label admits.
  • On the reporting cycle, AI is load-bearing on the left (moving work) but accountability never moves on the right (a human signs every answer).
  • An assured number is an audited number: 73% of large global companies now obtain external assurance, and the assurer will walk any figure back to its source.
  • The iron rule of the whole program: every published figure must trace to evidence, and "the AI estimated it" is never evidence.