โ†
AI for ESG & Sustainability Reporting
Aware ยท M2 ยท lesson 2 of 19 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
AI in Disclosure Drafting and Tagging
๐Ÿ“–
now learning

AI in Disclosure Drafting and Tagging

15 min

The disclosure lead has the sustainability statement open in one window and an AI assistant in the other. The narrative it just drafted is genuinely good: fluent, on-message, structured exactly the way the ESRS expect. It describes the company's climate transition plan, cites a 30% reduction target, and explains progress against it. There is one problem. The company never set a 30% target. The model wrote a plausible number into a public, assured disclosure, and it reads so cleanly that three reviewers might sign it before anyone checks. This is the quiet danger of AI in disclosure drafting: the better the prose, the easier it is to ship a claim that the evidence file cannot support.

A Disclosure Is Two Outputs, Not One

Modern sustainability disclosure has two faces, and AI touches both. The first is the narrative: the prose and the quantitative datapoints a human reads, the climate transition plan, the policy descriptions, the targets and metrics, the explanations of impacts, risks, and opportunities. The second, newer and easier to forget, is the machine-readable digital tag: the structured, coded version of the same disclosure that software reads. Under the ESRS digital reporting requirements and the ISSB's push toward structured data, your disclosure is not just published as text; it is tagged so that each datapoint is marked with a standardised code identifying what it is, letting regulators, investors, and analytics tools extract and compare it automatically.

Here is the point most people miss: the tag is as auditable as the prose. A datapoint tagged with the wrong code, or a number that says one thing in the narrative and another in the tagged data, is an error in the disclosure exactly like a wrong sentence is. An assurer and a regulator can read the tagged layer, and increasingly they do, with software that checks the structured data against the narrative and against the standard. So when you use AI to help draft the narrative and to help tag the datapoints, you have not one verification job but two: the words must be true and supported, and the tags must be correct and consistent with the words. Treating the tag as a formatting afterthought is how a clean report fails a digital validation it never saw coming.

It helps to understand why the digital tag exists at all, because the reason is exactly what makes it auditable. The whole point of structured, machine-readable disclosure is that a regulator or an investor should not have to read your 200-page statement to extract your Scope 1 figure or your transition target; their software can pull it directly from the tagged data and line it up against every other company's. That comparability is the feature. But it cuts both ways. The same machine that makes your number instantly comparable also makes your number instantly checkable, against the narrative, against last year's filing, against the standard's validation rules, and against your peers. A figure that looks fine buried in prose stands exposed the moment it is a clean, coded datapoint a tool can query. So the digital layer is not a lower-stakes formatting step that happens after the real work. It is a second, more literal version of your disclosure that is in some ways easier to audit than the prose, because a machine can check it in seconds and never gets tired or charitable.

What an ESRS Datapoint Is

An ESRS datapoint is a specific, defined piece of information the standard requires you to disclose: a particular metric, a particular narrative element, a particular policy disclosure. The ESRS define hundreds of them, each with an identity. Some are quantitative (a tonnage, a percentage, a monetary figure) and some are narrative (a description of a policy, a transition plan, a governance arrangement). Why you care: the datapoint is the unit the whole machine operates on. It is what you draft, what you tag, what the assurer tests, and what the digital validation checks. When AI drafts "the narrative datapoints," it is drafting these defined units, and each one carries both a claim that must be true and a tag that must be right. The ISSB's IFRS S1 and S2 work the same way conceptually, with defined disclosures that map to structured data, which is why a single fact base can feed multiple frameworks if your datapoints are clean.

This framing matters because it reframes what disclosure drafting actually is. It is tempting to think of the report as a document you write, and AI as a faster writer. But under ESRS and ISSB the report is closer to a structured database with a narrative wrapped around it. Each datapoint has an identity, a defined meaning, a place in the standard, and increasingly a code. When you draft, you are not just writing prose; you are populating defined fields, each of which carries a truth obligation. That is why the same number can appear in your inventory, your narrative, your tagged data, and a CBAM or ISSB filing, and must be identical and traceable in all of them. AI that helps you populate those fields faster is valuable. AI that populates them with plausible-but-wrong values is populating a database that regulators query, which is a far more exposed mistake than a loose sentence in a brochure.

Where AI Genuinely Helps in Drafting

AI is a strong drafting assistant for disclosure, and the gains are real. Narrative datapoints are repetitive and structured: the ESRS tell you what each one must cover, and AI is good at producing a first draft that hits the required elements in the required order. Point it at your evidence (your policies, your inventory, your materiality basis) and ask it to draft the climate transition plan narrative, and it will give you a structured, complete-looking first pass in minutes instead of hours. It is also good at consistency: keeping terminology uniform across a long statement, ensuring the same metric is described the same way in every section, and adapting one underlying fact into the slightly different phrasings that ESRS and ISSB each require. And it can help with the tagging itself, suggesting which standardised code applies to a given datapoint, which is genuinely useful when there are hundreds of possible tags and the mapping is fiddly.

Every one of these gains is real and worth having. And every one comes with the same condition that runs through this whole program: the draft is a draft, and nothing ships until a human has verified every claim and every figure against the evidence, and confirmed every tag against the standard. AI accelerates the production of the disclosure. It does not, and cannot, take responsibility for the disclosure being true.

There is a particular trap worth naming in narrative drafting, because it is subtler than an invented number. A generative model is built to produce text that is coherent and complete. When the evidence is thin, the model does not leave a gap and tell you it is missing; it fills the gap with the most plausible-sounding continuation, because that is what produces smooth prose. In ordinary writing that instinct is helpful. In disclosure it is exactly wrong, because the places where your evidence is thin are precisely the places a careful disclosure should be cautious, hedged, or silent, and the model's instinct is to be confident there instead. So the danger is not random; it is concentrated at your weakest evidence, which is the worst possible place for it. This is why you cannot verify an AI disclosure draft by reading it for plausibility. Plausibility is what the model is optimising for, and it will be most plausible exactly where it is least supported.

The better the AI prose reads, the more dangerous an unsupported claim inside it becomes, because fluent confidence is exactly what slips an invented target past three tired reviewers.

The Failure Modes That End in a Restatement

Three drafting failures recur, and all three are made worse by how good the prose looks. The first is the invented figure or target: the model writes a specific number, a 30% reduction target, a 2030 net-zero date, a percentage of renewable energy, that the company never set or that does not match the evidence. It is the hallucinated emission factor's cousin, and it is just as fatal, because a public, assured disclosure stating a target the company never adopted is a misstatement and a greenwashing exposure. The second is the softened impact: asked to describe a negative impact, a generative model trained to be helpful and positive can quietly downgrade it, turning a serious harm into neutral language, which both misleads the reader and can drop a material impact below where the standard requires it to sit. The third is the unsupported claim: a fluent sentence asserting progress, leadership, or compliance that sounds right but has nothing in the evidence file behind it. Each of these reads beautifully. That is the trap. In disclosure, fluency is not a sign of truth; it is a reason to check harder.

And the Tagging Failures Hiding Underneath

The tag has its own failure modes, quieter because almost no one reads the structured layer by eye. The mistagged datapoint: AI applies the wrong standardised code, so a number that is correct in the narrative is filed under the wrong identity, and the digital validation flags it or, worse, an analytics tool compares it against the wrong thing. The narrative-tag mismatch: the prose says 42,000 tonnes and the tagged datapoint says 24,000, a transposition or a stale figure, so the two layers of your own disclosure contradict each other. The missing tag: a required datapoint is drafted in the narrative but never tagged, so the digital disclosure is incomplete against the standard. None of these are visible if you only proofread the words. All of them are caught only if you treat the tag as a first-class part of the disclosure that must be verified, every figure in the tagged layer reconciled to the same figure in the narrative and to its source.

Before and After: A Transition-Plan Datapoint, Verified

Watch one narrative datapoint, the climate transition plan disclosure, move from a clean draft to a shippable one.

Before (the AI draft, what almost shipped): The disclosure lead asks the AI to "draft our climate transition plan narrative for the ESRS." It returns a polished three-paragraph passage. It states that the company "has committed to a 30% reduction in absolute Scope 1 and 2 emissions by 2030 against a 2020 baseline," describes "significant progress" in the first year, and notes the plan is "fully aligned with a 1.5 degree pathway." It then suggests tags for each datapoint. The passage is fluent, structured, and ESRS-shaped. It would pass a quick read. And it is a minefield: the company's actual board-approved target is 25% by 2030 against a 2019 baseline, "significant progress" is unsupported because year-one data shows a small increase, and "fully aligned with a 1.5 degree pathway" is a claim no internal analysis backs. Three invented or unsupported assertions, each in clean prose, each tagged and ready to file.

After (the verified, shippable datapoint): The disclosure lead treats every claim and figure as something to prove against the evidence. The target is corrected to the board-approved 25% by 2030 against a 2019 baseline, traced to the minuted board decision, and the tag for the target datapoint is set to the correct code and the tagged value reconciled to the narrative figure. "Significant progress" is replaced with the actual position: a small increase in year one, with the reasons disclosed, because the standard wants the truth, not a flattering story, and an assurer will compare the claim to the inventory. The "1.5 degree pathway" claim is removed entirely because no analysis supports it, and putting it in a public assured disclosure would be greenwashing. Each remaining datapoint is then checked twice: the narrative claim against its source, and the tag against the standard, with the tagged figure reconciled to the narrative figure. What ships is shorter, less flattering, and completely defensible: every sentence supported, every number traced, every tag correct and consistent. The AI still saved the lead an hour of drafting. It just could not be trusted with the truth, which is exactly what the human is for.

The lesson repeats one more time in a new place. The model produced the draft fast and produced three landmines doing it. The human kept the speed and defused the landmines by refusing to let any claim, figure, or tag ship without tracing it to evidence. Speed from the model, truth from the human, and a disclosure, in both its layers, that the file can support.

Notice one detail in the after version that is easy to skip past: the verified disclosure was shorter and less flattering than the AI draft, and that is not a coincidence. The unsupported sentences were the impressive ones, the 1.5 degree alignment, the significant progress, the rounder target. Verification stripped them out precisely because they were the claims with nothing behind them. This is a pattern you will see constantly. An AI draft tends to be more confident and more glowing than your evidence justifies, and the act of verifying it is largely the act of bringing the prose back down to what you can actually prove. A disclosure lead who feels their verified version is disappointingly modest compared to the AI draft is usually not weakening the report; they are removing exactly the greenwashing the assurer and the regulator are watching for. Modest and true beats impressive and unsupported every time a thread gets pulled.

Working Rules for Disclosure Drafting and Tagging

The rules close the loop on the whole reporting cycle. Use AI to draft narrative datapoints and to suggest tags, because both save real time. Verify every claim and every figure in the narrative against its source before it ships, treating fluent prose as a reason to check harder, not a reason to relax. Watch specifically for invented figures and targets, softened negative impacts, and unsupported claims of progress or alignment, because these are the drafting failures that become restatements. Treat the machine-readable tag as a first-class part of the disclosure: confirm every tag against the standard, and reconcile every figure in the tagged layer to the same figure in the narrative and to its source, so the two layers never contradict each other and no required datapoint goes untagged. Remember that one fact base can feed ESRS, ISSB, and other frameworks, but only if your datapoints are clean and verified at the source, so the verification you do once protects every framework it flows into. Do all of this and AI gives you a faster disclosure that survives both the assurer reading the prose and the software reading the tags. The accountability never moved. The model drafted; you verified; the file proves it.

Key Takeaways

  • A modern disclosure has two outputs: the narrative datapoints a human reads, and the machine-readable digital tags that software reads, and AI touches both.
  • The tag is as auditable as the prose: a mistagged datapoint or a narrative-tag mismatch is an error in the disclosure exactly like a wrong sentence, and assurers and regulators increasingly read the tagged layer.
  • An ESRS datapoint is a specific defined unit of disclosure, quantitative or narrative; it is what you draft, tag, and what the assurer and the digital validation test, and the ISSB frameworks work the same way.
  • AI genuinely helps by drafting structured narrative datapoints fast, keeping terminology consistent, adapting one fact to multiple frameworks, and suggesting which tag applies.
  • The drafting failures that end in restatements are the invented figure or target, the softened negative impact, and the unsupported claim, all made more dangerous by how fluent the prose looks.
  • The tagging failures are quieter: the mistagged datapoint, the narrative-tag mismatch, and the missing tag, none visible if you only proofread the words.
  • Verify twice: every narrative claim and figure against its source, and every tag against the standard with the tagged figure reconciled to the narrative figure, so both layers are true and consistent.
  • One clean, verified fact base can feed ESRS, ISSB, and other frameworks; the verification you do once at the source protects every framework it flows into, but the accountability for truth always stays human.