AI for ESG & Sustainability Reporting
Capable · M18 · lesson 18 of 24 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Recognizing Bad AI Output in Reporting Work
📖
now learning

Recognizing Bad AI Output in Reporting Work

15 min

An analyst is reviewing an AI-drafted section of a GHG inventory before it goes to the carbon accountant. The draft reads cleanly: a tidy table of Scope 3 activity data, a few emission factors, a short narrative on the company's progress. It looks finished. Then she slows down and starts reading it not as a reader but as a skeptic. The first factor has no database name. The purchased-goods figure is exactly 10,000 tonnes, suspiciously round. One factor is in kgCO2e per kilogram while the activity data is in litres. A sentence claims "emissions fell 12% year on year" with no figure behind it. And a paragraph on a known pollution incident has been smoothed into "the company continued to monitor environmental performance." None of it is flagged. All of it is wrong, or unprovable, in a way that would fail assurance. Catching it took her four minutes, because she had a checklist and a reflex. This lesson builds that reflex.

The Skeptic's Reflex, Before Anything Reaches the Inventory

The single most valuable habit a disclosure professional can develop around AI is not prompting and not tooling. It is a reflex: to read every AI output as a skeptic looking for the reason to reject it, before any of it touches the inventory or the disclosure. Fluent text invites trust. A clean table looks authoritative. A confident sentence reads as true. The skeptic's reflex is the deliberate refusal to grant that trust on the basis of fluency, because in disclosure fluency and truth are unrelated, and the things that fail assurance are precisely the things that look fine.

This is not paranoia; it is the job. Every figure and claim in your report is something an assurer can pull and test. Better that you find the weak ones first, at your desk, in minutes, than that the assurer finds them later, in front of your CFO, in a finding. The checklist that follows is a set of fast, mechanical tells, each one a question you ask of every AI output, each one cheap to run and expensive to skip. Run them as a reflex and most bad output stops at your desk. Skip them and the bad output flows downstream toward a disclosure, gathering credibility at every step it is not challenged.

Treat the checklist as covering three kinds of content, because each fails differently: emission factors, activity data, and narrative claims. A factor fails by lacking provenance. A number fails by being implausible or mismatched. A claim fails by being unsupported or quietly softened. The tells below are organized so that whichever kind of content you are looking at, you know what to hunt for.

The Checklist for Emission Factors and Activity Data

Numbers carry the disclosure, and numbers are where AI does its quietest damage, because a wrong number looks exactly like a right one. Five tells catch most of it.

No Source, Reject

The first and most important tell is the simplest: a figure or factor with no source is rejected, full stop. Not investigated, not assumed correct until proven wrong, rejected until a source is produced. This is the cardinal rule of the program in operational form. An emission factor without a named, dated database behind it is not a weak factor; it is a non-factor, regardless of how reasonable its value looks. The reflex is to let your eye go straight to the provenance, and if there is none, the number is out until one appears. Do not let a plausible value buy a number a pass it has not earned.

The Round-Number Tell

Real activity data is messy. A utility bill says 47,318 kWh, not 47,000. A supplier reports 4,210 tonnes, not 4,000. So when an AI output presents a suspiciously round number, exactly 10,000 tonnes, precisely 5%, a clean 1,000,000 litres, treat the roundness itself as a warning. A round number is often a tell that the figure was estimated, assumed, or invented rather than measured, because measured reality rarely lands on a round figure. It is not proof of error; sometimes things genuinely round. But it is a flag that says "ask where this came from," and in AI output a round number with no source behind it is very often a number the model produced from thin air to fill a slot.

The Unit Mismatch

This one catches errors that are catastrophic and invisible at once. Check that the factor's units match the activity data's units. If your activity data is in litres of diesel and the emission factor is expressed per kilogram, multiplying them produces a number that is not just wrong but wrong by whatever the density conversion would have been, and the result looks perfectly normal. Unit mismatch is the silent killer of inventories because nothing about the final figure announces the error; it is only visible if you check that litres met a per-litre factor and kilograms met a per-kilogram factor. AI is entirely capable of pairing a real activity figure with a real factor that is in the wrong unit, and the product will sail straight into the inventory unless someone checks the dimensions.

The Factor With No Date or Database

A factor is not fully sourced just because a number is attached to a name; it needs a date and a specific database edition. "An IPCC factor" is not enough. Which database, which year, which edition? Emission factors are revised annually, and using last year's edition or a different database's value is a real and common error. The tell is a factor that names no edition or no year. A properly sourced factor reads like "DEFRA 2024 conversion factors, road diesel, kgCO2e per litre, [row]." A factor that reads "standard diesel factor" has not told you what you need to know to defend it, and an assurer will notice the gap immediately.

For any consequential number, ask whether there is a path from it to a specific piece of evidence: a bill, a supplier response, a factor database row, a documented calculation. A number with no evidence link is unprovable, and in an assured disclosure unprovable is functionally the same as wrong. The reflex here is to mentally pull the thread on each number and see whether it leads anywhere. If pulling the thread leads to "the AI calculated it" and nothing further, the number is not yet ready to be in the report, no matter how reasonable it appears.

The Checklist for Narrative Claims

Narrative is where the subtlest and most reputationally dangerous failures live, because a softened or invented claim does not look like an error. It looks like good writing. Two tells matter most.

The Claim With No Evidence Behind It

AI-drafted narrative loves to assert progress: "emissions fell 12%," "the company is on track for its target," "supplier engagement improved." Each of these is a factual claim that needs a figure and a source behind it, and AI will produce them whether or not the support exists, because they are the kind of sentence that belongs in a sustainability report. The reflex is to treat every quantitative or evaluative claim in narrative as a figure in disguise, demanding the same source as a number in a table. "Emissions fell 12%" is not prose; it is a calculation that must trace to two reported figures. If the draft asserts it without them, it is an unsupported claim, and an unsupported claim in a public disclosure is exactly the kind of thing a greenwashing complaint is built from.

The Softened Negative Impact

This is the most insidious tell of all, and the one a skeptic must hunt for deliberately because the model will not flag it. AI-drafted narrative has a tendency to smooth, to make things sound better, to round the sharp edges off a bad fact. A real pollution incident becomes "the company continued to monitor environmental performance." A missed target becomes "the company remains committed to its ambitions." A material negative impact, which CSRD's double-materiality logic specifically requires you to disclose, gets quietly minimized into something palatable. The reflex is to ask, of any narrative on a sensitive topic, what did this soften, and compare the draft against the underlying facts. A softened negative is not just a quality problem; it is a disclosure failure and a greenwashing risk, because the standard requires the negative to be reported plainly, and the assurer and the regulator are specifically looking for the impact you made disappear.

No source, reject. A round number, a unit mismatch, an undated factor, a claim with no evidence link, or a softened negative each fails assurance. Build the reflex before any output reaches the inventory.

Worked Example: Running the Checklist on a Clean-Looking Draft

Return to the analyst's AI-drafted inventory section. It looked finished. Watch the checklist take it apart in four minutes, and watch each tell map to a specific assurance failure she has just prevented.

The draft's first line: "Purchased goods and services: 10,000 tCO2e, using a standard spend-based emission factor." She runs the reflex. Round-number tell: exactly 10,000 is suspicious for real activity data. No-source and undated-factor tells: "a standard spend-based emission factor" names no database, no year, no edition. Verdict: rejected pending a named, dated factor and the underlying spend the calculation rests on. Two tells, one line, one prevented finding.

The next line: "Mobile combustion: 14,200 litres of diesel multiplied by 2.1 kgCO2e per kilogram." She runs the reflex. Unit mismatch: litres of activity data multiplied by a per-kilogram factor. The product is dimensionally wrong, and the result, around 29,820, would have entered the inventory looking entirely plausible. Verdict: rejected; the factor must be per litre, or the diesel must be converted to kilograms by density first, and either way the basis must say which. This is the silent killer caught before it killed anything.

The narrative paragraph: "Emissions fell 12% year on year as the company continued to monitor environmental performance at its coastal facility." She runs the reflex. Unsupported claim: "fell 12%" has no two figures behind it; it is a calculation asserted as prose. Softened negative: "continued to monitor environmental performance" at a facility she knows had a reportable discharge event this year. The draft has converted a material negative impact into a reassuring nothing. Verdict: the 12% is held until the two figures are produced and the calculation traced, and the coastal sentence is rewritten to disclose the actual incident plainly, because the standard requires it and softening it is a greenwashing exposure.

In four minutes, with no special tools, the analyst caught a fabricated round number, an undated factor, a catastrophic unit mismatch, an unsupported quantitative claim, and a softened negative impact. Each one was invisible to a reader and obvious to a skeptic. The difference was not intelligence or effort. It was a checklist run as a reflex, before any of it reached the inventory.

Why the Reflex Must Come Before the Output Moves

The timing in that last phrase is the whole point. These tells are only cheap if you run them before the output flows downstream. A round number caught at your desk is a four-minute fix. The same round number caught after it has been entered into the inventory, rolled up into a total, drafted into the narrative, and reviewed by three people who each assumed the person before them had checked it, is a restatement. Every step an unchallenged figure travels, it gains credibility it has not earned, and the cost of removing it rises. The skeptic's reflex is cheapest at the first moment the output appears, which is exactly when fluency is working hardest to make you skip it.

So build the reflex into the moment of receipt. The instant an AI output lands, before you read it for content, read it for the tells: source on every factor, sanity on every number, units that match, dates on every factor, an evidence path for each figure, and a hunt for the claim that has no support and the negative that has been softened. It takes minutes. It runs on outputs you produced and outputs a teammate produced and outputs a vendor tool produced, because the tells do not care where the number came from, only whether it can be defended. A team that runs this reflex on everything has a simple and powerful property: bad output stops at the desk it lands on, and never gathers the false authority that turns a small catch into an expensive restatement.

One closing reframe. The checklist is not about distrusting AI specifically. It is the same skepticism a good disclosure professional has always applied to any number from any source, now pointed at a tool that produces plausible numbers faster than anything before it. AI did not create the need for the reflex; it raised the stakes on having one, by making it trivially easy to generate fluent, confident, unsupported content at volume. The professionals who thrive are not the ones who trust AI more or less. They are the ones whose skeptic's reflex is fast enough to keep up with how fast the tool can produce the very things that fail assurance.

Turning the Checklist Into a Team Habit

A reflex in one analyst's head protects only what that analyst touches. The real prize is a team where the checklist is a shared, named, repeatable step that every output passes through, regardless of who produced it or how trusted the source feels. Three moves turn a personal reflex into a team habit.

Name the Step and Make It Mandatory

Give the check a name (call it the receipt-time review, the skeptic pass, whatever sticks) and make it a required gate before any AI output enters the inventory, the calculation, or the draft. A named, mandatory step resists the erosion that informal good intentions suffer under deadline. "Did this pass the skeptic pass?" is a question a reviewer can ask and a colleague can answer, where "did you check it?" is too vague to enforce. The naming converts a private habit into a public standard the whole team can hold each other to.

Write the Verdict Down

For each flagged item, record a one-line verdict: the tell that fired, the missing element, and the disposition (rejected pending a source, held pending evidence, accepted with provenance noted). This costs seconds and pays twice. First, it forces the reviewer to be specific rather than vaguely uneasy. Second, and crucially, it becomes part of the assurance trail: when the assurer asks how AI-assisted figures were controlled, the team can show a record of outputs reviewed, tells caught, and figures rejected or fixed. The checklist, documented, is not just a quality gate; it is evidence that a quality gate exists, which is itself something an assurer wants to see.

Apply It Regardless of Source

The tells do not care whether a number came from your own prompt, a teammate's draft, or a vendor tool's dashboard, and neither should the review. The most dangerous outputs are often the ones that arrive with built-in authority: a figure from an expensive platform, a draft from a senior colleague, a number you yourself produced and feel confident about. Authority of source is exactly the bias that lets an unsourced or mismatched figure through, so the discipline is to run the same checklist on everything, with the same indifference the assurer will show. A figure from a named platform and a figure from a chat window are judged by one question: can it be defended.

A team that does these three things has converted a fragile individual habit into a durable institutional control. Bad output stops at the first desk it reaches, the catch is documented as evidence, and no source is trusted enough to skip the gate. That is the difference between a team that happens to catch some bad AI output and a team that systematically prevents it from reaching a disclosure, and only the second kind can tell an assurer, with a straight face, that its AI-assisted figures are under control.

Key Takeaways

  • The most valuable habit around AI is a reflex: read every output as a skeptic looking for the reason to reject it, before any of it touches the inventory or disclosure. Fluency and truth are unrelated, and the things that fail assurance look fine.
  • No source, reject. A figure or factor with no named, dated source is a non-factor until provenance appears, no matter how reasonable its value looks.
  • The round-number tell: measured reality rarely lands on a round figure, so an exact 10,000 or a clean 5% is a flag that the number may have been estimated, assumed, or invented.
  • The unit mismatch is the silent killer: a per-kilogram factor against litres of activity data produces a dimensionally wrong number that looks normal, visible only if you check that the units match.
  • A factor needs a date and a specific database edition, not just a name. "A standard diesel factor" has not told you what you need to defend it; "DEFRA 2024, road diesel, kgCO2e per litre" has.
  • Treat every quantitative or evaluative narrative claim as a figure in disguise: "emissions fell 12%" is a calculation that must trace to two reported figures, and an unsupported claim is greenwashing exposure.
  • Hunt deliberately for the softened negative impact: AI smooths bad facts into palatable nothing, and CSRD double materiality requires the negative to be disclosed plainly, so softening it is a disclosure failure the assurer and regulator look for.
  • Run the reflex first, at the moment of receipt, because every step an unchallenged figure travels it gains unearned credibility, turning a four-minute desk catch into a restatement.