Where AI Genuinely Helps in Sustainability Reporting
A carbon accountant opens a folder of 1,400 supplier emails, a stack of utility bills in four currencies, and a survey that 300 vendors half-answered. The deadline is six weeks out. An AI tool offers to read all of it tonight. The relief is real, and so is the trap: the question is never whether AI can help, it is which jobs it helps with, and what each of those jobs forces you to verify before a single number reaches the disclosure an assurer will read.
The Real Question Is Not Whether AI Helps, It Is Where
If you work in sustainability disclosure in 2026, you have heard both extremes. One side says AI will "automate your CSRD" and hands you a demo of a glossy report generated in minutes. The other side, often the assurance partner, says "not until you can show me the audit trail." Both are reacting to the same fact and missing the same nuance. AI genuinely helps with sustainability reporting. It helps in specific, namable places, and in each of those places it creates a verification obligation that did not exist before. The skill is not believing the hype or rejecting it. The skill is knowing exactly where the help is real and exactly what each kind of help demands back from you.
This matters because the stakes per filer have risen. Assurance is independent external checking of your sustainability numbers, the same idea as a financial audit applied to your emissions and impact data. As of 2026, roughly 73% of large global companies obtain external assurance on at least some sustainability disclosures, up from 51% in 2019 (IFAC, AICPA and CIMA). That is a number worth verifying against the source rather than repeating blindly, but the direction is unmistakable: every disclosed figure is now an audited figure. So when AI helps you produce a number faster, it has also handed you a number you will have to defend in front of someone whose job is to pull the thread.
Let us walk the places where the help is genuine, evidenced, and worth building into your workflow, and pair each one with the verification it demands.
Materiality Input Clustering: AI Themes, You Decide
Double materiality is the CSRD requirement to assess both how sustainability issues affect your company (financial materiality) and how your company affects people and the environment (impact materiality). To do it, you gather inputs: stakeholder survey responses, interview notes, peer benchmarks, risk registers, sometimes hundreds of separate documents. Then you have to find the themes inside that pile and decide which issues clear your materiality threshold, the line above which an issue is significant enough to report.
This is where AI earns its place. A language model can read 400 stakeholder comments and cluster them into themes such as water stress, labour conditions in the supply chain, or product safety, in an afternoon rather than a fortnight. It can map a raw comment to the relevant ESRS standard (the European Sustainability Reporting Standards, the detailed rulebook that says what a CSRD report must contain). It can draft a first-pass rationale for why a theme might be material.
Here is the hard boundary. The cluster is a hypothesis, not a verdict. The model has grouped comments by surface similarity; it has not applied your threshold, it has not weighed severity against likelihood, and it has not made the materiality call. That call stays human, and it stays documented. The verification the clustering demands is this: every theme must trace back to the specific source comments that produced it, the threshold you applied must be written down, and any issue the AI grouped away or dropped must be a decision you can defend, not a silent omission. An assurer who asks "show me the basis for this materiality conclusion" is asking to reconstruct your reasoning from the inputs. If the trail runs cold at "the AI clustered it," you do not have a basis, you have a guess.
There is a subtler trap inside clustering that is worth naming, because it catches careful people. A model clusters by what is frequently said, and frequency is not the same as severity. Suppose only three stakeholders raised a serious human-rights concern in your supply chain, while two hundred mentioned office recycling. A naive use of clustering would bury the human-rights theme as a small cluster and elevate recycling as a large one, exactly inverting their materiality. The human reading the clusters has to remember that a rare but severe impact can be highly material and a common but trivial one need not be. AI gives you the map of what was said; you still have to apply judgment about what matters. That judgment, and the record of it, is the part an assurer is actually testing.
Document and Activity-Data Extraction: The Tedious Win
The single most reliable place AI helps is reading documents and pulling structured numbers out of them. In GHG accounting, activity data is the raw measure of something you did: litres of diesel, kilowatt-hours of electricity, tonne-kilometres of freight. You multiply activity data by an emission factor (a published conversion rate, for example kilograms of CO2 per litre of diesel) to get emissions. Before any of that math, somebody has to find the activity data, and it lives in invoices, utility bills, fuel receipts, travel records, and supplier spreadsheets that arrive in no consistent format.
An AI extraction tool reads a messy PDF utility bill and returns a structured field: 48,200 kWh, billing period, meter ID, site. It does this across hundreds of documents without fatigue. This is a genuine, large, boring win, and boring wins are the best kind in disclosure because they are easy to verify.
And verify you must. Extraction has a specific failure mode: the model can transcribe a number wrong, read the wrong line, or quietly "helpfully" fill a missing value. The verification that extraction demands is provenance plus a spot check. Provenance means each extracted figure carries a pointer back to where it came from: this document, this page, this line. With provenance, an assurer can sample any number and trace it to source in seconds, and you can too. The discipline is simple to state and non-negotiable in practice: an extracted number with no source pointer is not data you can use, it is a number you would have to defend without evidence.
The reason extraction is such a good fit for AI is worth understanding, because it tells you where to lean on the tool and where to stay alert. Extraction is a closed problem: the answer is already in the document, and the model's job is to find and transcribe it, not to invent anything. When the task is "tell me what this bill says," the model has a ground truth to be checked against, namely the bill itself. Contrast that with a question like "estimate the emissions of a supplier who sent no data," which is an open problem with no document to check against. The closed problems are where AI is safest and most valuable, precisely because verification is cheap: you hold the output next to the source and look. Build your AI use to favor closed problems, and reserve your deepest skepticism for the open ones, where there is no source to hold the output against.
Supplier-Survey Drafting and Triage: Cutting the Scope 3 Bottleneck
Scope 3 is the GHG Protocol's category for emissions in your value chain that you do not own or directly control: purchased goods, business travel, the use of your sold products, and twelve other categories, fifteen in total. For a typical company, Scope 3 is around 75% of the total footprint. It is also the hardest data to get. In the Sphera 2025 Scope 3 Report, 62% of reporters cited internal data quality and 79% cited supplier-data availability as their top barriers. Those are numbers to verify against the source, and they describe a real, expensive bottleneck: most of your footprint depends on data sitting inside companies you have to ask nicely.
AI helps here in two practical ways. First, drafting. It can generate a clear supplier questionnaire tailored to a category, in the supplier's language, with the exact fields you need. Second, triage. When 300 suppliers return responses in inconsistent formats, AI can parse them, sort the complete from the partial, flag the contradictory, and route the gaps to your follow-up list. This collapses weeks of inbox archaeology into a structured queue.
The verification the survey work demands is about labelling and gaps. Every datapoint that comes back must be tagged as primary data (the supplier actually measured and reported it) or secondary data (an estimate or industry average standing in for the real thing), because an assurer treats those two very differently and they cannot look identical in your file. And the suppliers who did not respond must show up as visible gaps, not get quietly replaced by an average that makes the table look complete. The temptation, when 79% of your peers cannot get supplier data, is to let AI "fill the gap." Filling a gap with a labelled, disclosed estimate is defensible. Filling it with a plausible number that reads like measured data is the move that fails assurance.
ESRS and ISSB Narrative Drafting: A Fast First Draft, Not a Final Word
A large part of a sustainability report is prose: the narrative datapoints that describe your policies, your transition plan, your governance, your impacts. Both ESRS and the ISSB standards (IFRS S1 and S2, the global baseline now adopted or planned across 30-plus jurisdictions) require a great deal of structured narrative. Writing it from a blank page, in the required structure, against the required disclosure requirements, is slow.
AI drafts this well. Give it your underlying facts, the relevant ESRS disclosure requirement, and your prior-year narrative, and it returns a structured first draft in minutes. For a disclosure team under deadline, that is a real acceleration of a real task.
AI can write the sentence. It cannot decide whether the sentence is true. That decision, and the evidence behind it, stays with you.
The verification narrative drafting demands is the most subtle of all, because the failure is quiet. A model drafting from your bullet points can soften a negative impact, round a target into something more flattering, or assert a commitment the company never actually made. None of that looks wrong on the page; fluent prose is convincing precisely when it should not be. So every claim in an AI-drafted narrative must be checked against the evidence file, every figure must trace to a source, and every target must match what your governance actually approved. The draft is scaffolding. The truth check is the job.
A practical way to keep narrative drafting safe is to feed the model only what you can already support and to instruct it not to add. If you give the model your approved targets, your verified figures, and the relevant disclosure requirement, and you tell it to draft strictly from those inputs and flag anything it lacks rather than fill it, you have turned an open-ended generator into something much closer to a closed problem. It is still not free of risk, because a model can drift even from good instructions, but you have stacked the odds in your favor. The professionals who get the most from narrative drafting are not the ones who ask for the most; they are the ones who constrain the model to the evidence and then check that it stayed there.
The Adoption Reality, as Numbers to Verify
It helps to hold the landscape in your head, with the caveat that every figure below is something you should confirm against its source before you cite it in a meeting. 73% of large global companies now obtain external assurance on at least some sustainability data; that is why the assurance lens runs through everything. Scope 3 averages roughly 75% of a company's footprint across the fifteen GHG Protocol categories; that is why so much AI effort goes into the value chain. And 79% of reporters cite supplier-data availability as a top barrier, with 62% citing internal data quality; that is the exact bottleneck AI survey tooling is sold against.
Notice what these numbers do and do not tell you. They tell you the help is aimed at real pain. They do not tell you that any specific AI output is correct. A tool that closes the 79% gap by inventing supplier data has not solved the problem, it has converted a visible data gap into an invisible misstatement. The adoption story is a reason to use AI carefully, never a reason to trust its output on sight.
Worked Example: An AI Supplier Summary, Before and After
Watch a realistic AI output sail toward a disclosure, then watch an informed professional turn it into something assurable.
The Before: A Tidy, Dangerous Summary
An analyst feeds 280 supplier survey responses into an AI tool and asks for a Scope 3 Category 1 (purchased goods and services) summary. The tool returns a clean paragraph: "Based on supplier responses, purchased goods and services emissions total 142,000 tCO2e. Coverage is strong across all major suppliers. Data quality is high." It looks finished. It looks like something you could paste into the inventory. It is a liability.
Why? The 142,000 figure blends primary supplier-reported numbers with AI-estimated fill-ins for the 90 suppliers who never responded, and nothing in the summary says which is which. "Coverage is strong" is the model's editorial flourish, not a measured coverage rate. "Data quality is high" is unsupported. If this reaches the disclosure and an assurer samples it, the thread unravels in one question: "Which of these suppliers actually reported, and which did you estimate?"
The After: The Same Speed, Now Defensible
The informed analyst keeps the speed and adds the discipline. They re-run the task with instructions that force labelling: tag every supplier figure as primary or secondary, never blend them into one total, list non-responders as explicit gaps, and cite the source document for each primary figure. The output now reads: "Primary data received from 190 of 280 suppliers, covering 68% of Category 1 spend, total 96,400 tCO2e (each traced to a supplier submission). 90 suppliers (32% of spend) did not respond; these are flagged as gaps pending a labelled spend-based estimate, disclosed separately with its method and uncertainty." That is longer, less tidy, and infinitely stronger. It survives the assurer's first question because it answers it before being asked.
The lesson of the contrast: AI produced both summaries in the same few seconds. The difference in defensibility came entirely from the human framing the task around what the file would have to prove.
Key Takeaways
- AI genuinely helps in four evidenced places in sustainability reporting: clustering materiality inputs, extracting activity data from documents, drafting and triaging supplier surveys, and drafting ESRS and ISSB narrative. Each is a real, expensive task, and each creates a verification obligation.
- Materiality clustering produces a hypothesis, not a verdict. The theme must trace to its source comments, the threshold must be documented, and the materiality call stays human.
- Document extraction is the most reliable win because it is the easiest to verify: every extracted number must carry provenance, a pointer back to the exact document, page, and line.
- Supplier-survey work cuts the Scope 3 bottleneck, but every datapoint must be labelled primary or secondary, and non-responders must appear as visible gaps, never as a quietly inserted average.
- Narrative drafting accelerates a slow task, but the failure is quiet: a softened impact, an invented target, an unsupported claim. Check every claim against the evidence before it ships.
- Treat the adoption numbers (73% assured, Scope 3 around 75% of the footprint, 79% citing supplier-data barriers) as figures to verify against their sources, not slogans to repeat.
- The bright line across all four wins: a labelled, disclosed estimate is defensible; a plausible number dressed up as measured data is the move that fails assurance.
- AI provided the speed in every example here. The defensibility came from the human who framed the task around what the assurance file would have to prove.
Skill.re