AI for ESG & Sustainability Reporting
Capable · M4 · lesson 4 of 24 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
AI-Assisted ESRS and ISSB Narrative Datapoints
📖
now learning

AI-Assisted ESRS and ISSB Narrative Datapoints

15 min

It is the last week before the sustainability statement goes to the board, and you have eighty narrative datapoints to draft. You paste the climate transition section into a model, ask for a clean ESRS narrative, and ninety seconds later you are reading a paragraph so fluent it could already be in the annual report. It mentions a 2030 target. The problem is that your company never set a 2030 target.

What an ESRS Narrative Datapoint Actually Is

Before you can let a model draft disclosure, you have to be precise about what it is drafting. Under the European Sustainability Reporting Standards, a datapoint is a single, specified piece of information the standard requires you to report. The ESRS define well over a thousand of them across the environmental, social, and governance topics. A datapoint is not a vibe or a section heading. It is a discrete obligation: report this, in this form, with this content. An assurer works through them one at a time, and a digital-tagging process maps each one to a machine-readable tag. The "why you care" is blunt: a missing or mis-stated datapoint is not a stylistic weakness, it is a gap in a regulated filing.

Datapoints come in two flavours, and confusing them is the first way an AI draft goes wrong. A quantitative datapoint is a number with a unit: tonnes of CO2 equivalent, cubic metres of water withdrawn, the percentage of the workforce covered by collective bargaining. A narrative datapoint is prose: a description of your transition plan, an account of how a policy is implemented, an explanation of the process you use to identify material impacts. ESRS deliberately mixes the two. A single disclosure requirement might ask for a number and then ask you to narrate the policy, the action, and the target that sit behind it. The narrative is not decoration. It is the disclosure.

This matters for AI because the two flavours fail differently. A model that fabricates a quantitative datapoint hands you a wrong number, and a wrong number is at least checkable against a source. A model that fabricates a narrative datapoint hands you a confident, well-formed sentence that asserts something about your company that may be entirely untrue, and a fluent false sentence is far harder to catch than a wrong number, because nothing about it looks broken.

Why ISSB Rides Along With ESRS

You are rarely drafting for one framework. A large undertaking still in CSRD scope after Directive (EU) 2026/470 is filing ESRS, and the same company adopting IFRS S1 and S2 under the ISSB is producing climate and general sustainability disclosure for investors across the thirty-plus jurisdictions now converging on those standards. The fact base is shared: the same transition plan, the same governance, the same emissions. The narrative obligations are similar but not identical. ESRS leans into double materiality and impact; ISSB centres on financial materiality and the information a reasonable investor needs. When you draft narrative datapoints, you are usually mapping one underlying truth into two slightly different disclosure shapes, and the model is very willing to blur the two if you let it.

That blurring is not harmless. An ESRS narrative about a labour-rights impact in your supply chain is framed around the impact on people; the equivalent ISSB-adjacent disclosure, where one applies, is framed around the financial consequences to the company. Ask a model to "write the sustainability narrative" without naming the framework and it will produce something that sounds like both and satisfies neither, often importing impact language into an investor disclosure or financial-risk language into an impact disclosure. The cost is not stylistic. A datapoint written to the wrong materiality lens can be technically present and still fail to discharge the obligation, because it answers a question the standard did not ask. Naming the framework and the specific disclosure requirement at the top of the prompt is the cheapest way to keep the two shapes distinct, and verifying against the right framework's obligations is how you confirm the model held the line.

What the Model Does Brilliantly, and Where It Turns

The case for using AI here is real and you should not pretend otherwise. Narrative datapoints are repetitive, structured, and tedious. The model is genuinely excellent at taking your raw inputs, a policy document, a set of bullet points from the climate team, last year's disclosure, and producing a first draft that is grammatical, on-structure, and roughly the right length. It can rephrase a clumsy internal memo into disclosure-grade prose. It can take a German-language policy and produce an English narrative. It can hold the ESRS structure in mind and slot your facts into the expected shape. For a team drafting eighty narratives in a week, that is not a toy. It is hours saved per datapoint.

The turn comes from the same machinery that makes it useful. A language model produces the most probable next words given everything it has seen. When your inputs are thin, when the climate team gave you three bullets and the standard wants a paragraph, the model does not stop and tell you it is short of facts. It fills the space with the most plausible-sounding continuation. Plausible, here, means "the kind of thing a company like yours would say." So it writes that you are "committed to reducing absolute Scope 1 and 2 emissions by 2030," because companies like yours say that, even though you never committed to it. It is not lying in any intentional sense. It is completing a pattern. The danger is that the completion reads exactly like a disclosure you would be proud of.

Hold onto why this is specifically a disclosure problem and not just a writing problem. In most uses of AI, a confident-but-wrong sentence is a nuisance you fix on the next read. In disclosure, that sentence is going into a regulated filing that an external assurer will test under a limited- or, increasingly, a reasonable-assurance engagement, and that a regulator can reopen later. The assurer's whole method is to pull claims and ask for their basis. So a fluent fabrication is not a draft imperfection; it is a latent assurance finding sitting in your file with your company's name on it. The model that saved you twenty minutes can hand you a sentence that costs you a restatement, and the two outcomes are indistinguishable at the moment you read the draft. That asymmetry, cheap to produce, expensive to defend, is exactly why the speed must be paired with a verification discipline rather than trusted on its own.

The model writes the sentence you wish were true. Your job is to check whether it is.

The Three Narrative Failure Modes That End in a Finding

Across thousands of AI-drafted narratives, three failures recur, and each maps to a specific assurance risk.

The invented target. The model asserts a commitment, a percentage, or a date the company never set. This is the most dangerous because a target is a forward-looking claim that an assurer, a regulator, and an NGO can all hold you to. A 2030 net-zero pledge that exists only in your AI draft is a greenwashing exposure with your signature under it.

The softened impact. ESRS requires you to describe your actual and potential negative impacts honestly. The model, trained on a world of corporate communications, has a strong gravitational pull toward the reassuring. Ask it to narrate a human-rights risk in your supply chain and it will tend to wrap the risk in mitigation language, shrinking the negative impact until it reads like a managed issue rather than a real one. Under double materiality, understating a negative impact is itself a misstatement.

The unsourced figure. The model drops a number into the narrative, "our renewable electricity share rose to 42%," that did not come from your data. Sometimes it is a half-remembered number from training. Sometimes it is interpolated from a trend. Either way it is a quantitative claim wearing narrative clothing, and it has no line in your evidence file.

The Verification Pass: Every Claim to Its Evidence

The discipline that makes AI-assisted narrative safe is not a better prompt. It is a verification pass you run on every draft before it moves an inch toward the disclosure. The pass has one organising idea: a narrative is a bundle of claims, and every claim needs evidence. You decompose the fluent paragraph back into the individual assertions it makes, and you check each one against a source.

Walk a sentence. "In 2025 we reduced absolute Scope 1 and 2 emissions by 12% against our 2021 baseline, in line with our science-based target validated in 2023." That single sentence carries four separate claims: a 12% reduction, a 2021 baseline, a science-based target, and a 2023 validation date. Each one is a hook an assurer can pull. The verification pass asks four questions, not one. Is the 12% in the inventory? Is 2021 the documented baseline? Is there a validated target on file? Is 2023 the right date? A narrative that reads as one smooth thought is, for verification, four obligations.

Classify Each Claim, Then Check It

Sort every claim in the draft into one of four buckets, because each is verified differently.

  • Quantitative claims (a number, a percentage, a date) trace to the inventory, the source data, or the dated record. The number in the narrative must equal the number in the file, to the figure.
  • Target and commitment claims (a pledge, a goal, a deadline) trace to a board-approved or formally adopted target document. If there is no document, the claim does not exist, full stop.
  • Policy and process claims ("we have a supplier code of conduct," "our board reviews climate risk quarterly") trace to the actual policy, the charter, the minutes. The claim must match what the document says, not what the model assumes a good company does.
  • Impact and outcome claims (the description of a negative impact, the effect of an action) trace to your materiality assessment and your actual records. Here you check not only that the claim is supported but that it is not softened: does the narrative describe the impact as honestly as your own assessment does?

Anything that does not trace to a source is not "probably fine." It is removed or rewritten until it does. The model does not get the benefit of the doubt, because the model has no way to earn it.

One subtlety separates a good verifier from a careless one: the direction of the check. It is not enough to read the source and confirm the narrative is consistent with it, because the narrative can contain extra claims the source never made, and a consistency read glides right past them. You verify from the claim outward, taking each assertion in the prose and demanding the source that produced it, rather than reading the source and nodding along to the prose. The invented target survives a consistency read precisely because nothing in your inputs contradicts it; there was simply nothing there to contradict. It only dies when you point at the sentence and ask, where did this come from, and find that the answer is nowhere. Make the claim prove itself; do not let the absence of contradiction stand in for the presence of evidence.

The Softening Test for Impact Claims

Impact claims deserve a check the other three buckets do not need, because they can fail while being technically supported. A quantitative claim is right or wrong; a target exists or it does not. But an impact claim can trace to your assessment and still misrepresent it by tone. The model's pull toward reassurance shows up as adjectives and framing: "well-managed," "substantially mitigated," "limited," "being addressed." Each softens a negative impact a notch, and a paragraph of notches can turn a serious unresolved risk into a managed footnote. The test is comparative, not binary: read the narrative and your underlying assessment side by side and ask whether a reasonable reader would come away with the same sense of severity from both. If the narrative reads calmer than your own assessment, it is softened, and under double materiality a softened negative impact is a misstatement even though every individual word might be defensible. You are not just checking that the claim is supported; you are checking that it is supported at the right intensity.

Worked Example: One AI Narrative, Before and After

Here is a raw model output for an ESRS climate transition narrative datapoint. The inputs the analyst gave the model were modest: a one-page internal summary saying the company has an emissions inventory, is "working on" a transition plan, and reduced emissions last year.

Before (the AI draft): "The Group is firmly committed to achieving net-zero emissions across its operations by 2040, with an interim target to reduce Scope 1 and 2 emissions by 50% by 2030. In 2025, the Group reduced its absolute greenhouse gas emissions by 15%, driven by an accelerated shift to renewable energy, which now accounts for 60% of electricity consumption. The Group's transition plan is fully aligned with a 1.5C pathway and is overseen by the Board's Sustainability Committee."

It is beautiful. It is also a minefield. Run the pass. The 2040 net-zero commitment: no target document exists, the company is only "working on" a plan. Fabricated. The 50%-by-2030 interim target: same, no source. Fabricated. The 15% reduction: the inventory shows 6%. Wrong number. The 60% renewable electricity: nothing in the source data supports it. Unsourced. The "fully aligned with a 1.5C pathway": that is an alignment assertion no one has validated. Unsupported claim. The Board Sustainability Committee oversight: the company has no such committee. Fabricated governance. Six claims, six failures. Every single one would be an assurance finding, and two or three would be greenwashing exposures a regulator could act on.

After (the verified draft): "In 2025, the Group reduced its absolute Scope 1 and 2 greenhouse gas emissions by 6% compared to 2024 (see GHG inventory, basis of preparation section 3.2). The Group is developing a climate transition plan; as of the reporting date, no quantified emissions-reduction target has been formally adopted, and this is disclosed as a known gap. Renewable electricity accounted for 28% of total electricity consumption, sourced from the energy-procurement records. Climate-related matters are reviewed by the Group Risk Committee; the Group has not established a dedicated sustainability committee."

The after-version is shorter, plainer, and far less impressive. It is also true, sourced line by line, and honest about what does not yet exist. That last quality matters: ESRS expects you to disclose that a target is not yet set rather than invent one. The verified draft turns the absence of a target from a temptation to fabricate into a disclosed fact. That is the move. The model gave you fluent fiction; the analyst turned it into a traceable, assurable narrative. Notice the after-version is not slower to write than the before-version was to fix. The speed is real. It just lives in the drafting, not in the shipping.

The Honest Gap Is a Disclosure, Not a Weakness

The single most useful habit a narrative analyst can build is treating "we do not have this yet" as a legitimate, often required, disclosure. The model's instinct is to paper over gaps with plausible prose. The standard's instinct, and the assurer's, is the opposite: tell us what you do, tell us what you do not yet do, and do not dress the second up as the first. A disclosed gap is defensible. A fabricated strength is a liability with a fuse on it.

Prompting to Reduce the Damage, Not Eliminate It

You can shape the model's behaviour so the verification pass has less to catch, but you can never delete the pass. Ground the model: give it the actual policy, the actual inventory extract, last year's disclosure, and instruct it to draft only from the material you provided and to write "[NOT IN SOURCE]" wherever the standard wants content you did not supply. That single instruction converts the model from a confident fabricator into an honest flagger. Instead of inventing a 2030 target, it writes "[NOT IN SOURCE: no emissions-reduction target provided]," which is exactly the signpost you want.

Tell it the framework explicitly: drafting an ESRS narrative datapoint is different from drafting an ISSB one, and naming the standard and the specific disclosure requirement keeps the model in the right shape. Ask it to keep negative-impact language at the same intensity as your inputs, which pushes back against the softening instinct. None of this makes the output trustworthy on its own. A well-grounded model still drifts, still rounds, still occasionally invents. The prompt lowers the failure rate; the verification pass is what makes the output shippable. Treat prompting as harm reduction, not as a substitute for checking.

Key Takeaways

  • An ESRS datapoint is a single specified disclosure obligation; a quantitative datapoint is a number with a unit, a narrative datapoint is required prose, and the two fail differently when a model drafts them.
  • A fluent false sentence is harder to catch than a wrong number, because nothing about it looks broken; that is what makes AI-drafted narrative dangerous.
  • The three narrative failure modes are the invented target, the softened negative impact, and the unsourced figure, and each maps to a specific assurance or greenwashing risk.
  • Verify by decomposing every paragraph back into its individual claims, then checking each claim against a source; one smooth sentence can carry four separate obligations.
  • Sort claims into quantitative, target, policy/process, and impact buckets, because each is verified against a different kind of evidence.
  • You are usually mapping one shared fact base into both ESRS and ISSB shapes; name the framework and the specific disclosure requirement so the model does not blur them.
  • A disclosed gap is a legitimate, often required, disclosure; let the model flag "[NOT IN SOURCE]" rather than paper over what does not yet exist.
  • Grounding and prompting lower the failure rate but never replace the verification pass; the model writes the sentence you wish were true, and your job is to check whether it is.