AI for ESG & Sustainability Reporting
Capable · M7 · lesson 7 of 24 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Building Verification Checklists for ESG AI
📖
now learning

Building Verification Checklists for ESG AI

15 min

A new analyst joins the sustainability team in October, three weeks before the inventory closes. She is good, careful, and skeptical by instinct. But her skepticism lives in her head, as a set of reflexes she developed over years: check the factor, sniff at a number that jumped, never trust a claim with no source. When she leaves the room, her reflexes leave with her. The carbon accountant beside her has different reflexes, sharper on factors, looser on narrative. A contractor brought in for the crunch has almost none. Three people are screening AI output before it reaches the inventory, and they are screening it three different ways, with three different blind spots, none of it written down. This lesson is about turning that private skeptic's reflex into a public, repeatable control: a verification checklist anyone on the team can run, that catches the invented factor, the mislabeled estimate, and the unsupported claim before any of it reaches the inventory or the disclosure, and that produces a sign-off an assurer can read.

Why a Reflex Is Not a Control

An experienced reviewer's instinct is genuinely valuable, and it is also genuinely fragile, for three reasons that a checklist exists to fix. It does not transfer. The skeptic's reflex lives in one person's experience and cannot be handed to the new analyst or the contractor, so the quality of the screen depends entirely on who happens to be doing it. It is not consistent. The same reviewer is sharper at 9 a.m. than at 7 p.m. the night before a deadline, sharper on the figures they care about than the narrative they find tedious, and sharper on a normal day than the day the CFO is asking why the report is late. It leaves no evidence. When an assurer asks how AI output was checked before it entered the inventory, "our analysts are careful" is not an answer the engagement accepts. A reflex performed and forgotten produces nothing the file can show.

A checklist fixes all three at once. It transfers, because anyone can run a written list whether or not they have the years of instinct behind it. It is consistent, because the list does not get tired or bored or rushed in the way a person does; it asks the same questions in the same order every time. And it leaves evidence, because a completed, signed checklist is an artifact the assurance file can hold, the documented control that proves the screen happened. The checklist does not replace the skeptic; it captures what the best skeptic does and makes it available to everyone, every time, on the record. It is the difference between hoping the right person looked carefully and being able to show that a defined check was run and passed.

There is a deeper point here about how assurance actually thinks, and it is worth absorbing because it reframes the whole exercise. An assurer does not primarily ask whether your numbers are right; the assurer asks whether you have a control that makes them reliably right, and then tests whether that control operated. A brilliant analyst catching errors by instinct is not a control in this sense, because it cannot be described, cannot be tested, and cannot be shown to have operated on any particular output. It is a talented individual, which is admirable and unassurable. A written checklist with a defined stop and a sign-off is a control, because it can be described ("we run these checks on every AI output"), tested ("re-perform check A2 on this sampled figure"), and evidenced ("here is the signed checklist for that output"). This is why turning the reflex into a checklist is not bureaucratic box-ticking that gets in the way of the real work. In an assured disclosure, the checklist is part of the real work, because a figure that was screened by an undocumented instinct is, to the engagement, a figure that was not screened at all.

An instinct that is not written down protects only the prompts the right person happened to review on a good day. A checklist protects every output, run by anyone, and leaves proof it ran.

The Three Things the Checklist Must Catch

A verification checklist for ESG AI can grow long, but it earns its keep by reliably catching the three failure modes that fail assurance and end careers. Everything else on the list is in service of these three.

The Invented Factor or Figure

The most dangerous AI failure in carbon accounting is the hallucinated emission factor: a plausible-looking number with no real source, which silently corrupts every result it touches. The checklist's countermeasure is a question with a binary, checkable answer: for this factor, can I name the database, the table, the version, and the year, and can I open that source and find this exact number? If yes, it passes. If no, it fails and the figure does not move, regardless of how reasonable it looks. The same applies to any activity figure: can I trace this number to a named file, page, and line in the source document? The check is not "does this seem right." Seeming right is exactly how a fabricated factor passes a tired reviewer. The check is "can I reopen the source and confirm it," which a fabrication cannot survive, because there is no source to open.

The Mislabeled Estimate

The second failure is subtler and just as fatal: an estimate dressed up as measured data. A spend-based or average-data figure is legitimate when it is labeled as an estimate with its method and uncertainty, and it is a misstatement when it sits in the inventory looking identical to a metered, supplier-reported number. The checklist asks: is this figure primary or secondary, is it tagged as such, and does the tag match how it was actually produced? An estimate with no label, or worse, an estimate wearing a primary label, is the failure here. The check catches the moment where AI helpfully "filled a gap" and the gap-filler quietly inherited the appearance of measured data. The assurer lives by the primary-versus-secondary distinction, so a mislabeled estimate is not a small tagging error; it is a figure claiming an evidentiary basis it does not have.

The Unsupported Claim

The third failure lives in narrative rather than numbers: an unsupported claim, a sentence asserting a target, a commitment, an achievement, or a characterization of an impact that the evidence does not support. AI drafting drifts here because fluent prose wants to sound positive and complete, so it softens a negative impact, rounds a partial achievement into a full one, or states a target the company never set. The checklist asks of every claim in AI-drafted narrative: what evidence supports this exact sentence, and does that evidence actually say this? A claim that cannot be tied to a specific piece of supporting evidence does not ship. This is the greenwashing-by-fluency check, and it is the one reviewers skip most, because reading numbers feels like verification and reading prose feels like editing. The checklist forces the prose to be verified as rigorously as the figures.

The reason this check is so easy to skip and so dangerous to skip deserves spelling out. A fabricated number announces itself if you look: it has no source, it fails to reconcile, it sits oddly against last year. A fabricated or inflated claim hides inside well-formed language that reads exactly like the supported sentences around it. "The company is on track to meet its 2030 reduction target" is grammatically and tonally identical whether the company set such a target or not, whether it is on track or badly behind. The fluency that makes AI drafting useful is precisely what makes its overstatements invisible, because there is no surface signal distinguishing a sentence the evidence supports from one it does not. The only defense is to refuse to read the paragraph as prose and instead interrogate each claim as an assertion that must be backed: pull the target sentence out, ask what document sets that target, open it, and confirm the target exists and the trajectory is as described. An assurer trained on greenwashing does exactly this, sentence by sentence, and a regulator empowered against misleading sustainability claims can act on a single overstated one. The narrative check is not a courtesy to good writing; it is the control that keeps a fluent draft from quietly becoming a misleading disclosure.

A Worked Example: The Actual Checklist

Here is a verification checklist a team could run on any AI-assisted output before it enters the inventory or the disclosure. It is deliberately plain, phrased as yes-or-no questions, because a checkable question is one that does not depend on the reviewer's mood or experience to answer.

Section A, every figure.

A1. Source named? For each figure, is the source named specifically (file, page, line for activity data; database, table, version, year for a factor)? Yes / No / N.A.

A2. Source opened and confirmed? Did I open the named source and find this exact number there? Yes / No.

A3. Recomputed? Does the result recompute from its activity value and factor? Yes / No / N.A.

A4. Unit correct? Is the unit explicit and consistent through the calculation? Yes / No.

A5. Prior-period sane? Is the figure within a plausible range of last year, or is any jump explained? Yes / No.

Section B, every factor.

B1. Real and current? Is the factor from a named, authoritative database, and is the version or year the right one for this reporting period? Yes / No.

B2. Right factor for the activity? Does the factor actually match this activity, region, and unit, not just a near neighbour? Yes / No.

Section C, every figure's data type.

C1. Primary or secondary tagged? Is the figure tagged primary or secondary? Yes / No.

C2. Tag honest? Does the tag match how the figure was actually produced (no estimate wearing a primary label)? Yes / No.

C3. Estimate labeled fully? If secondary, are the method and uncertainty stated? Yes / No / N.A.

Section D, every narrative claim.

D1. Evidence tied? Does each claim link to a specific piece of supporting evidence? Yes / No.

D2. Evidence actually says it? Does that evidence actually support this exact sentence, with no softening, no invented target, no rounded-up achievement? Yes / No.

Section E, sign-off.

E1. Any No unresolved? Are there any No answers that have not been fixed or escalated? If yes, the output does not move.

E2. Reviewer and date. Name and date of the person who ran this checklist.

How to Read the Checklist's Design

Notice three deliberate choices. First, every answer is yes, no, or not-applicable, never "looks fine." A binary question cannot be passed by a vague impression, which is precisely the impression a fabrication exploits. Second, A2 is separate from A1, because naming a source and confirming a source are different acts, and the dangerous failure is a figure with a named source that, when you open it, does not contain the number. Splitting them forces the reviewer to actually open the source rather than nod at the citation. Third, Section E makes the checklist a control with teeth: a single unresolved No stops the output, and the reviewer's name and date turn the act of checking into a signed, dated artifact the assurance file can hold. Without Section E a checklist is advice; with it, it is a documented control with a sign-off.

Making It a Real Control, Not a Ritual

A checklist can decay into theatre, ticked without being run, and a few disciplines keep it real. Make the structured record do half the work. If the output already arrives as a structured record with source, factor provenance, primary-or-secondary, and method in their own fields, then most checklist questions become fast field-level confirmations rather than prose hunts. The structured-output discipline from the previous lesson and the checklist are partners: the schema lays out the answers in legible fields, and the checklist confirms each one is true. Require evidence of A2, not just the tick. The strongest version of the control records not only that the source was confirmed but, for sampled figures, a note of what was found, so the sign-off is auditable rather than self-asserted. Tie unresolved Nos to a real stop. The checklist only works if a No actually blocks the figure; a team that checks the box and ships anyway has a ritual, not a control. Right-size it. A checklist long enough to be ignored protects nothing, so keep it to the questions that catch the three failure modes plus the few that catch common errors, and resist the urge to add a question for every theoretical risk.

It is also worth being honest about the failure mode of the checklist itself, because a control you cannot criticize is a control you cannot trust. The danger is normalization of the tick: a reviewer who has run the list a hundred times and never found a problem begins to answer Yes by reflex, and the binary question that was supposed to defeat impression-based reviewing quietly becomes an impression again. Three things hold this off. Rotate the evidence requirement, so that on a sample of figures the reviewer must record what they actually found when they opened the source, which cannot be done from memory and forces a genuine look. Track what the checklist catches, because a checklist that has never returned a single No across a whole cycle is either screening perfect inputs, which is implausible, or being run without being performed, which is likely, and the catch rate is the signal that tells you which. Keep the stakes visible, by reminding the team that the No they are tempted to skip is exactly the one the assurer will find, and that a catch made internally is a quiet fix while the same miss found externally is a finding, a restatement, and a headline. A checklist stays a real control only when the team treats a Yes as a claim they would defend, not a box they clear.

One more point that elevates the checklist from a chore to a credential. The completed, signed checklist is precisely the kind of evidence an external assurer wants to see, because it demonstrates a control: a defined screen, applied consistently, with a stop on failure and a named reviewer. When the assurer asks how the team kept AI output from corrupting the inventory, the answer is not a description of careful people; it is a stack of signed checklists showing the screen was run on every output. That is the move that turns the skeptic's private reflex into an institutional control, and turns "trust us, we checked" into "here is the documented evidence that we checked, output by output." A team that can hand the assurer that stack is a team whose AI-assisted inventory the assurer can trust faster, which is the whole point.

Before and After: The Same Output, Two Screens

Return to the October crunch and the three people with three sets of reflexes. Before the checklist, an AI-drafted Scope 3 section reaches the inventory. The careful analyst would have caught the spend-based figure wearing a primary label, but she was on a different task, and the contractor who reviewed it has no reflex for the primary-versus-secondary distinction, so the mislabeled estimate sails through. There is no record that anyone checked it, and no way to know what was and was not screened. The failure is invisible until the assurer finds it.

After the checklist, the same section reaches the same contractor, who runs Section C. C2 asks whether the tag matches how the figure was produced, and the contractor, who needs no special instinct because the question is explicit, finds that a spend-based estimate is tagged primary. That is a No. Section E stops the output, the figure goes back to be retagged and relabeled with its method and uncertainty, and the corrected record moves on with a signed checklist attached. The contractor with no reflexes just performed the catch the best analyst would have made, because the checklist carried the analyst's instinct to him. And when the assurer arrives, the signed checklist proves the screen ran. The reflex became a control, the control caught the failure, and the catch is on the record.

Key Takeaways

  • A skeptic's reflex is valuable but fragile: it does not transfer to new staff, it is not consistent across people, time, and pressure, and it leaves no evidence; a checklist fixes all three by being transferable, consistent, and documented.
  • The checklist exists to catch three failure modes that fail assurance: the invented factor or figure, the mislabeled estimate, and the unsupported narrative claim.
  • The invented-factor check is binary and unforgiving: can I name the database, table, version, and year and open the source to find this exact number, not does this seem right, because seeming right is how a fabrication passes a tired reviewer.
  • The mislabeled-estimate check tests whether the primary-or-secondary tag matches how the figure was actually produced, catching an estimate that quietly inherited the appearance of measured data.
  • The unsupported-claim check forces narrative to be verified as rigorously as numbers: every claim must tie to specific evidence that actually says it, catching the softened impact, the invented target, and the rounded-up achievement.
  • Design choices make the checklist a real control: yes-or-no answers that no vague impression can pass, a separate step to open and confirm the source rather than nod at the citation, and a sign-off section where a single unresolved No stops the output.
  • Structured output and the checklist are partners: when source, factor provenance, and tags arrive in their own fields, most checks become fast field-level confirmations instead of prose hunts.
  • A completed, signed checklist is exactly the evidence an external assurer wants, turning the private reflex into a documented institutional control and trust us, we checked into here is the proof we checked.