AI for ESG & Sustainability Reporting
Proficient · M21 · lesson 21 of 24 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
The Datapoint-to-Disclosure Workflow
📖
now learning

The Datapoint-to-Disclosure Workflow

15 min

A disclosure lead opens the draft ESRS statement at 6 a.m. and finds a single line: "In FY2025 the undertaking emitted 412,000 tonnes CO2e in Scope 1 and 2 (market-based)." It reads clean. It is also a loaded gun. Eleven months from now an assurer will put a finger on that number and ask one question: where did it come from, and who said it was right? If the answer lives only in the head of an analyst who has since left, the disclosure is not assured, it is asserted. This lesson is about the path that turns the second case into the first.

Why a Workflow, Not a Task

Most reporting teams treat a disclosure figure as a task: get the number, paste it in, move on. That works right up until someone independent reads it. An ESRS datapoint is a specific, defined unit of disclosure required by the European Sustainability Reporting Standards: a single quantitative or narrative element, such as gross Scope 1 emissions in tonnes CO2e, or the description of a transition plan, that the standard names and expects you to report. Why you care: the standard does not ask for a tidy paragraph, it asks for thousands of named datapoints, each of which an external assurer can isolate and test in its own right. The day a disclosure becomes an assured disclosure, the unit of trust is not the report, it is the datapoint.

So the right mental model is not "write the report." It is "run each datapoint through a repeatable path that ends in a number an independent reader will accept on sight." The path has four stages: draft, verify, tag, sign off. AI is genuinely useful in two of them and dangerous in the other two if you let it lead. The whole art of the practitioner is knowing which is which, and capturing the right evidence at every stage so the figure can be rebuilt without you in the room.

The reason this matters in 2026 specifically: 73% of large global companies now obtain external assurance on at least some sustainability disclosures, up from 51% in 2019, and the companies still inside the post-Omnibus CSRD net are the largest undertakings, where a failed disclosure is a board-level event. Treat that 73% as a number to verify against your own sector, not a slogan, but the direction is unmistakable. Every figure you publish is now an audited figure, and the iron rule of the program applies: every figure must trace to evidence, and "the AI estimated it" is not evidence.

There is a second reason the workflow framing matters, quieter but just as important: continuity. Reporting teams turn over. The analyst who knows in their bones where the 412,000 came from leaves, gets promoted, goes on parental leave. If the only record of how a figure was built is that person's memory and a folder of loose spreadsheets, then the day they walk out the door, every figure they touched becomes unreconstructable. A workflow that captures evidence at each stage is, among other things, an insurance policy against your own team's turnover. The test an assurer applies, can someone rebuild this number without you in the room, is also the test of whether your reporting function can survive a resignation. The two are the same test, and a stage-by-stage workflow is how you pass it.

Stage One: Draft (AI Leads, Lightly)

Drafting is where AI earns its keep, and where it is least dangerous, because nothing it produces here is yet a published claim. Give the model the datapoint definition from the standard, the source data you have already gathered, and a hard instruction: draft only from the material provided, and where a figure or a fact is missing, say so explicitly rather than filling the gap. The output you want is a candidate disclosure plus a list of every number it used and where that number supposedly came from.

Watch the move that separates a useful draft from a poisoned one. Ask the model to draft the Scope 1 and 2 disclosure and it will happily produce fluent prose with a market-based and a location-based figure, a year-on-year comparison, and a sentence about methodology. Every one of those is a claim. Some trace to your activity data and emission factors; some the model may have smoothed, rounded, or quietly invented. The draft is not the disclosure. It is a hypothesis about the disclosure, written fast, that now has to be proven true line by line.

The evidence to capture at this stage is small but real: the prompt and source set you gave the model, the raw draft it returned, and the model and version that produced it. You are not capturing this because the draft is trustworthy. You are capturing it so that if a verified figure later diverges sharply from the draft, you can see exactly what the model did and whether it introduced an error you nearly inherited.

Grounding the Draft

Drafting quality collapses the moment the model reaches past your evidence into its training data. A model asked for "the standard emission factor for natural gas" will produce a confident number that may be plausible, outdated, or simply wrong. The discipline is to ground the model on your own factor database and source files and forbid it from supplying any figure not present in them. If the factor is not in the grounding set, the correct output is a flag, not a guess. A flagged gap is a task for a human. A guessed number is a misstatement waiting to be assured.

Stage Two: Verify (Human Leads, Every Number to Source)

Verification is the stage the whole workflow exists to protect. Here a human ties every number in the draft back to its source, and the rule is absolute: no figure survives to the next stage without a verified source. Not most figures. Every figure.

Concretely, the verifier takes the draft's list of numbers and, for each one, opens the evidence and confirms three things. First, the number matches the source: the 412,000 tonnes ties to the inventory calculation, which ties to the activity data and the emission factor, which ties to a named, dated factor database. Second, the source is the right one: a current factor, the correct boundary, the period that matches the disclosure. Third, the path is recorded: a verifier should be able to point at the cell, the document, the database entry that produced the figure, not gesture at a folder.

A number with a confident tone and no traceable source is not a disclosure. It is a misstatement that has not been caught yet.

This is also where you catch the failure modes that end careers. A hallucinated emission factor is a plausible but invented number with no source behind it; verification kills it because there is nothing to tie it to. Fabricated activity data, a target the company never set, a softened negative impact in AI-drafted narrative, an estimate dressed up as measured data: every one of these surfaces the instant a human demands the source and finds none. Verification is not bureaucracy. It is the single control that stands between an AI draft and a public, assured, legally exposed figure.

Notice what makes AI-drafted disclosure uniquely dangerous at this stage, compared with a junior analyst's draft. A junior analyst who does not know a figure will usually leave it blank, or flag it, or guess in a way that looks like a guess. A capable model does the opposite: where it lacks the fact, it produces something fluent, confident, and correctly formatted, a number with the right number of digits and a methodology sentence that reads exactly like the real thing. The fluency is the hazard. A blank cell announces that work remains; a plausible fabrication announces that the work is done. Verification is the discipline that refuses to be reassured by fluency and insists on the source underneath it, which is why it cannot be delegated back to the model that wrote the draft. You do not ask the author to confirm its own confidence; you ask the evidence.

A practical way to run verification is to make the draft hand you the checklist. If the draft stage forced the model to list every figure with its claimed source, then verification is a structured pass down that list: open each claimed source, confirm the figure is there and current and correctly bounded, mark it verified with the location and your initials, or send it back as a gap. Figures that the model flagged as missing are already tasks; figures the model claimed a source for are claims to test. Worked this way, verification is not an open-ended re-investigation of the whole disclosure, it is a finite, auditable sweep, and the artifact it produces, the per-number source record, is the most valuable thing in the entire file.

The evidence to capture here is the heart of the assurance file: for each datapoint, the verified figure, the named source, the location within that source, the verifier's identity, and the date verified. This is what an assurer means by an audit trail. A figure that arrives with this packet is reconstructable. A figure that arrives bare is a finding.

Primary Versus Secondary, Made Visible

Verification also forces a label that assurers live by. A supplier-reported, measured figure is primary data; an estimate, an industry average, or a spend-based proxy is secondary data. The two cannot look identical in the file. If your draft blended a measured Scope 1 figure with an estimated Scope 3 line and the verified record does not distinguish them, the assurer cannot judge the quality of your inventory, and a defensible estimate looks exactly like a fabricated one. Verification is where the label gets attached, permanently, to each number.

Stage Three: Tag (Machine-Readable, Same Audit Weight)

A verified figure is not yet a filed datapoint. Under the ESRS digital reporting regime, the disclosure must be machine-readable: each datapoint is marked up with a tag that identifies which element of the taxonomy it represents, so a regulator's or investor's software can read your gross Scope 1 emissions without parsing your prose. Digital tagging is the act of attaching that machine-readable label to a value, mapping your 412,000 tonnes to the exact taxonomy element the standard defines. Why you care: a mistagged datapoint is a disclosure error in its own right, even if the underlying number is perfect, because the machine-readable layer is what downstream systems actually consume.

AI assists here, the human verifies. A model can propose the taxonomy element for a given value far faster than a person hunting through the tag list, especially across thousands of datapoints. But the proposal is exactly that. If the model maps a market-based figure to the location-based element, or tags a current-year value as restated, the number is right and the disclosure is still wrong. The verifier confirms each tag against the taxonomy definition the same way they confirmed each number against its source. The tag carries the same audit weight as the narrative, so it earns the same verification.

The evidence to capture: the tag applied to each datapoint, the taxonomy element it maps to, who confirmed the mapping, and any tag the model proposed that the human overrode. That override log is quietly valuable. It is the proof that a human reviewed the machine's tagging rather than rubber-stamping it.

Stage Four: Sign Off (Accountable Owner)

Sign-off is the stage with no AI in it at all. A named, accountable owner reviews the verified, tagged datapoint and accepts it into the filing. This is a governance act, not a generative one, and it is the point where accountability becomes a person rather than a process. "The model recommended it" is never a defense to an assurer or a regulator; the file has to show a human who looked at the evidence and decided.

The owner is not re-doing verification. They are confirming that verification happened, that the source packet is complete, that primary and secondary data are labeled, that the tag is confirmed, and that they personally stand behind the figure entering the public statement. For a material datapoint, that owner is senior enough to carry the consequence: a controller, a disclosure lead, ultimately the people who sign the statement. The evidence captured is the simplest and most important of all: who signed, what version they signed, and when. That record is what converts a draft into a disclosure.

The version detail in that record does real work. Disclosures change between draft and filing, a figure gets corrected, a narrative gets sharpened, a tag gets fixed. If sign-off names a version, then the file shows exactly which state of the datapoint the owner accepted, and a later change after sign-off is visible as a later change, not silently folded into what was approved. Without versioned sign-off, you cannot tell whether the figure that was filed is the figure that was signed, and that gap is precisely where a post-sign-off edit, well intentioned or not, can slip an unapproved number into the public statement. Tie the signature to a version, and the chain from approval to filing stays unbroken.

It is also worth being clear about why sign-off has to be a person and not a control the system performs automatically. A rule that says "advance any datapoint whose packet is complete" can be satisfied by a packet that is complete and wrong: every field populated, every source attached, and a boundary judgment that an experienced owner would have questioned. The owner's value is the judgment a checklist cannot encode, the moment of an accountable human looking at a material figure and deciding it is right to publish. That is what an assurer and a regulator are ultimately testing for when they ask who decided. The answer has to be a name, and behind the name, evidence that the person actually looked.

A Worked Example: One Datapoint, Two Endings

Take the datapoint: gross Scope 1 GHG emissions, FY2025, in tonnes CO2e. Watch it travel the path twice.

The ending you do not want. An analyst prompts a general model: "Draft our Scope 1 emissions disclosure, we burned natural gas and diesel across four sites." The model returns a clean paragraph with a figure of 118,000 tonnes, a confident emission factor for natural gas, and a sentence comparing it favorably to last year. It looks done. It goes into the draft statement. Eleven months later the assurer asks for the basis of the 118,000. The factor the model used cannot be found in any database the company subscribes to; it was generated. The four-site total never reconciled to the utility bills because the model estimated the diesel volume the analyst never gave it. The figure unwinds, the disclosure is restated, and the restatement becomes the story. The number was never wrong on purpose. It was simply never tied to anything.

The ending you want. The same analyst grounds the model on the activity-data file and the company's factor database, and instructs it to draft only from those and flag anything missing. The model drafts the same paragraph but lists every input and flags that diesel volume for one site is absent. That flag becomes a task; the analyst pulls the bill. In verification, each figure ties to source: gas volumes to meter reads, the factor to a named, dated database entry, the site total to the reconciled bills. The Scope 1 figure resolves to 121,400 tonnes, not 118,000, because the model's draft had quietly rounded and dropped the late diesel line. The verified figure is labeled primary, tagged to the correct taxonomy element by the analyst over a model proposal that had picked the wrong element, and signed off by the disclosure lead. When the assurer points at the 121,400, the analyst opens one packet: source, factor provenance, verifier, tag, signer, date. The thread is on the table. The assurer moves on.

Same model, same datapoint. The difference was not intelligence. It was the workflow: grounded draft, human verification of every number, confirmed tag, accountable sign-off, evidence captured at each stage.

Running the Path at Scale

One datapoint is easy to imagine running cleanly. A full ESRS statement is thousands of them, and that is where teams either keep the discipline or quietly abandon it. The way to hold the line is to let each stage's natural owner work at the stage's natural speed. Drafting is fast and broad: the model can draft and propose tags across the whole statement in a fraction of the time a person would take, so you let it. Verification is slow and deep, but it is finite if the draft handed you a clean list of figures and claimed sources, so you resource it as the real work it is rather than treating it as a formality squeezed into the final days. Sign-off is concentrated on the people accountable for material datapoints. The mistake is to let the speed of drafting set the tempo for verification; the draft being done in an afternoon does not mean the disclosure is, and a team that confuses the two ships unverified numbers.

Risk-based effort is how you keep verification both rigorous and feasible. Every figure ties to a source, that floor does not move, but the depth of senior review concentrates on the datapoints that matter most: the material ones, the ones with the highest uncertainty, the ones built on estimates rather than measured data. A clearly measured Scope 1 figure tied to a meter read needs confirmation; a Scope 3 estimate for a category with 79% supplier non-response needs scrutiny of its method, its uncertainty, and its label. Documenting the basis for how you allocated review effort is itself part of the assurance posture, because it shows the assurer that your attention went where the disclosure risk was, deliberately, rather than spread evenly and thin.

And the whole path feeds the thing this program keeps returning to: a file that survives the engagement. Each datapoint that travels draft, verify, tag, sign off, capturing evidence at each stage, arrives at the assurer's desk as a small, complete, reconstructable package. When the assurer samples ten datapoints and asks you to rebuild each from raw data, you are not assembling anything new under pressure; you are opening packages you built as you went. That is the difference the workflow buys: the assurance file is not a document you write after the report, it is the residue of having done the report properly, one datapoint at a time.

Key Takeaways

  • The unit of trust in an assured disclosure is the datapoint, not the report. Run each ESRS datapoint through a repeatable path: draft, verify, tag, sign off.
  • AI leads the draft and assists the tag; humans lead verification and own sign-off. Conflating those roles is how an unsupported number reaches a filed statement.
  • Drafting must be grounded on your own source data and factor database, with a hard instruction to flag gaps rather than fill them. A flagged gap is a task; a guessed number is a misstatement.
  • Verification is non-negotiable and absolute: every number ties to a named, dated, locatable source, or it does not advance. This is the single control that stops hallucinated factors, fabricated activity data, and laundered estimates.
  • Primary versus secondary data must be labeled at verification, because a measured figure and an estimate cannot look identical in the file.
  • Digital tagging carries the same audit weight as the narrative. A mistagged datapoint is a disclosure error even when the number is correct, so AI proposes the tag and a human confirms it against the taxonomy.
  • Sign-off is a human governance act with a named, accountable owner. "The model recommended it" is never a defense; the file must show who decided.
  • The evidence captured at each stage, draft inputs, verified source, confirmed tag, signer and date, is the assurance file. A datapoint that arrives with its packet is reconstructable; one that arrives bare is a finding.