โ†
AI for ESG & Sustainability Reporting
Aware ยท M10 ยท lesson 10 of 19 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Extraction, Estimation, Generation: Three Different AIs
๐Ÿ“–
now learning

Extraction, Estimation, Generation: Three Different AIs

15 min

Three numbers sit side by side in a Scope 3 inventory, each produced with AI, each looking equally clean. The first, 1,240 MWh of purchased electricity, was read off a utility bill. The second, 8,900 tonnes of CO2 equivalent for a non-responding supplier, was computed from spend and an industry average. The third, "an estimated 12% reduction versus the prior year," was a sentence the model wrote into the narrative. To the eye they are identical: three tidy figures in three tidy cells. To an assurer they are three completely different risks, and the team that cannot tell them apart is the team that ships an unsupported number into a public, assured disclosure. This lesson is about telling them apart.

Three Jobs, Three Failure Modes

In the framing lesson we separated four AI jobs. Three of them put a value or a claim into your disclosure, and those three are where unsupported numbers are born: extraction, estimation, and generation. Each does something genuinely different to the relationship between your number and the truth, and each fails in its own distinct way. Memorize the three pairings and you have a diagnostic you can run on any AI output in seconds.

Extraction pulls a value that already exists in a source you provided. Its failure mode is misreading. Estimation computes a proxy for a value you do not have. Its failure mode is laundering, passing a guess off as measured data. Generation writes prose, including sentences that contain numbers and claims. Its failure mode is fabrication, inventing a value or a commitment from nothing. Same surface, three different machines, three different ways to fail. Conflating them is the single most common path by which a number that cannot survive assurance reaches a filing.

Extraction has a source by definition. Estimation has a method, or it has nothing. Generation has neither unless you supplied it. Know which one you are looking at before you trust the number.

Extraction: The Honest Job That Still Misreads

Extraction is the most assurance-friendly of the three because it is structurally honest: the value it produces exists in a document you can point to. When the model extracts 1,240 MWh from a utility bill, the bill is right there, and verification is fast and conclusive. You open the bill, you find the line, you confirm the number. This is why extraction is the workhorse of AI-assisted reporting and the job you should push as much work toward as possible. The closer an AI task sits to extraction, the easier it is to defend.

But honest does not mean infallible. The failure mode of extraction is misreading: the model grabs the wrong line, transposes digits, misreads a unit, or picks up a subtotal instead of a total. It reports 14,200 when the bill says 142,000. It reads litres as gallons. It pulls last quarter's figure because it appeared higher on the page. The error is real and can be large, but it has a saving grace: because a source exists, the error is catchable by anyone willing to open the source and look. Extraction never asks you to trust the model. It asks you to check the model against a document that is sitting right there.

How Extraction Actually Fails in Practice

The most dangerous extraction errors are the plausible ones. A model that reports a wildly wrong figure gets caught by a sanity check. A model that reports 1,420 MWh instead of 1,240, a transposition, sits quietly in the inventory because it is the right order of magnitude. The defense is not cleverness; it is discipline. Trace a sample of extracted values back to their source documents every time, and reconcile totals to something independent, like the prior period or a control total. Extraction's gift is that the truth is always one document away. Your job is to actually walk that one document away.

Estimation: The Guess That Must Confess

Estimation is where the real peril lives, because estimation produces a number for something you genuinely cannot measure. The 8,900 tonnes for the non-responding supplier was never measured. It was computed: take the spend with that supplier, multiply by an industry-average emission factor per unit of spend, and you get a plausible figure. There is nothing wrong with doing this. In Scope 3, where ~75% of a typical footprint lives and where 79% of reporters cannot get reliable supplier data, estimation is unavoidable and entirely legitimate. The GHG Protocol expects it.

The peril is not estimating. The peril is laundering: letting the estimate enter the inventory looking exactly like a measured figure, with no label, no method, no uncertainty. Once it is laundered, nobody downstream can tell that 8,900 is a spend-based guess and 1,240 is a metered fact. They sit in adjacent cells wearing the same costume. When the assurer asks "which of these are primary and which are estimated," the team that laundered cannot answer, and the entire inventory's credibility wobbles.

The distinction the assurer lives by is primary versus secondary data. Primary data is directly measured or supplier-reported activity data, the metered MWh, the supplier's own reported tonnage. Secondary data is estimated or averaged, the spend-based proxy. Why you care: a primary figure and a secondary estimate carry completely different reliability, and an assurer needs to see, on the face of the file, which is which. A laundered estimate erases that distinction, and erasing it is the failure.

A defensible estimate confesses. It says, on its face: I am estimated, here is my method, here is my uncertainty. A laundered estimate hides, and a hidden estimate is a misstatement.

The Bright Line of Estimation

Hold the line in one image. "Estimated and disclosed" is on the right side of the law: the 8,900 carries a label (secondary), a method (spend-based), and an acknowledgment that it is uncertain. "Made up" is on the wrong side: the 8,900 sits unlabeled next to the metered figures, indistinguishable. The number can be byte-for-byte identical. What changes its fate is whether it confesses what it is. The whole skill of defensible estimation is making the guess wear a sign that says "I am a guess, and here is how I was built."

Generation: The Sentence That Invents

Generation is the job people least suspect of producing numbers, which is exactly why it is dangerous. We think of generation as writing prose, and it is, but prose contains numbers and claims. "An estimated 12% reduction versus the prior year" is a sentence the model wrote, and inside that sentence is a quantitative claim that may correspond to nothing. The failure mode of generation is fabrication: the model invents a value, a percentage, a target, or a commitment because that is the kind of content the sentence wanted, not because any data supports it.

Generation fabrication is worse than estimation laundering in one respect: an estimate at least started from real inputs (the spend, the average). A fabricated figure started from nothing but the model's sense of what sounds right. The 12% might be invented whole cloth. The "net zero by 2040" in a drafted narrative might be a commitment the board never made. And because generation is fluent, the fabrication arrives polished and confident, embedded in a paragraph that otherwise reads beautifully. The number hides inside good writing, which is the perfect camouflage.

The defense against generation fabrication is to treat every number and every claim inside generated prose as a separate object that must trace to evidence, independent of how good the surrounding sentence is. Good writing is not a source. A fluent paragraph earns no trust for the figures it contains. You extract the claims, you check each against the evidence file, and you cut or correct any that do not trace.

The Diagnostic: Running It on Real Output

Here is the diagnostic as a single habit. For any AI-produced value heading toward your disclosure, ask: which job produced this, and therefore which failure am I hunting?

The jobWhat it doesIts failure modeYour verification
ExtractionPulls a value that exists in a provided sourceMisreading (wrong line, transposed digits, wrong unit)Open the source, find the line, confirm the value
EstimationComputes a proxy for an unmeasured valueLaundering (a guess passed off as measured data)Confirm it is labeled secondary, with method and uncertainty
GenerationWrites prose containing numbers and claimsFabrication (an invented value, target, or commitment)Trace each claim and figure to the evidence file independently

Run this on the three opening numbers. The 1,240 MWh is extraction: open the bill, confirm. The 8,900 tonnes is estimation: confirm it is labeled secondary, spend-based, with uncertainty noted. The 12% reduction is generation: find the underlying figures and recompute, or cut the claim. Three numbers that looked identical now have three different verifications, and each is answerable. The team that runs this diagnostic never hands an assurer a number it cannot place.

Why Conflation Is the Default, and the Danger

If telling the three jobs apart is so important, why do teams conflate them constantly? Because the tools are built to hide the seams, and the disclosure surface flattens everything. A modern carbon-accounting platform takes your procurement file and returns a tidy inventory in which the metered electricity, the spend-based supplier estimates, and the auto-drafted methodology note all appear as elements of one finished product. The interface does not say "this cell is extraction, this cell is estimation, this paragraph is generation." It says "here is your Scope 3." The blending is a feature from a usability standpoint and a hazard from an assurance standpoint, and you, the discloser, are the only one who can un-blend it.

The disclosure surface compounds the problem. A published table of emissions presents every figure in the same font, the same column, the same number of decimal places. Visually, a metered fact and a laundered guess are indistinguishable, which is exactly why the primary-versus-secondary label has to be carried as a separate, explicit field rather than left to the appearance of the number. The eye cannot see the difference. Only the label can. A team that relies on the figures looking different will never catch a laundered estimate, because laundering is precisely the act of making a guess look like a fact, and it succeeds by default unless someone actively labels.

The three jobs do not announce themselves. The tool blends them and the table flattens them, so naming the job is work you must do deliberately, every time, on every value.

The Cost of Getting It Wrong

Consider what each failure costs once it reaches a filing. A misread (extraction) is the cheapest to fix: you re-open the source, find the right number, correct it, and the provenance is restored. A laundered estimate (estimation) is more expensive: you must go back, relabel every affected figure as secondary, document the method and uncertainty you should have captured the first time, and possibly re-collect primary data to replace estimates the assurer will not accept at the current coverage. A fabrication (generation) is the most expensive of all: the invented target or percentage may already be public, so correcting it can mean a restatement, an awkward conversation with the assurer about how an unsupported claim got into an assured disclosure, and exposure to a greenwashing finding. The order of cost (misread, then laundering, then fabrication) is also the order in which the jobs move away from having a source, which is the deep reason the diagnostic matters: it sorts your verification effort by exactly how dangerous each failure is.

A Worked Example: Before and After

Before (conflated). A reporting analyst assembles a Category 1 section using an AI tool and treats every output the same way: glance, accept, paste. The metered electricity, the spend-based supplier estimates, and a narrative line claiming "a 12% year-on-year improvement in supplier emissions intensity" all land in the draft looking equally solid. The section ships to the assurer as a uniform block of confident figures. The assurer asks three questions: where is the bill behind the electricity, which suppliers are primary versus estimated, and what is the basis for the 12%. The analyst can answer the first, fumbles the second because nothing is labeled, and cannot answer the third at all because the 12% was a generated sentence with no calculation behind it. The section is sent back. The 12% becomes a finding.

After (separated). The same analyst runs the diagnostic on each value. The metered electricity is tagged as extraction with the bill referenced. The supplier figures are tagged secondary, spend-based, with an uncertainty note, clearly distinct from the primary suppliers who reported actuals. The "12% improvement" claim is traced: the analyst tries to reconstruct it from the underlying data, finds it does not hold, and removes it, replacing it with the actual change the evidence supports and a note on how it was computed. Now the section reaches the assurer pre-answered: extraction with sources, estimation with labels and method, generation with every claim traced or cut. The three questions are already answered in the file. The section passes.

Nothing about the analyst's tools changed. What changed was that three jobs stopped being one blur. The metered fact, the labeled guess, and the generated claim were each handled as the distinct thing it was, and that is the entire difference between a section that fails and a section that holds.

Make this the habit you carry out of the lesson. When any AI output arrives on its way to a disclosure, do not ask whether it looks right, because all three jobs produce output that looks right, and looking right is exactly the trap. Ask instead which job produced it. The moment you name the job, the failure to hunt and the verification to run both become obvious: open the source for an extraction, demand the label and method for an estimate, trace the claim for a generated figure. A metered fact, a labeled guess, and a generated claim should never again look identical to you, because you will have learned to see past the uniform surface to the three very different machines underneath. That seeing, applied to every value before it reaches the assurer, is what keeps an unsupported number out of a published, assured disclosure, which is the entire job and the whole reason this diagnostic is worth making automatic.

Key Takeaways

  • Three AI jobs put values and claims into a disclosure, and each fails differently: extraction by misreading, estimation by laundering, generation by fabrication.
  • Extraction is the most assurance-friendly job because a source exists by definition; push as much work toward extraction as you can, and verify by opening the source.
  • The most dangerous extraction errors are plausible ones, like a digit transposition, so trace a sample to source and reconcile totals every time.
  • Estimation is legitimate and unavoidable in Scope 3; the failure is not estimating but laundering, letting a guess look identical to measured data.
  • The assurer lives by primary versus secondary data: measured or supplier-reported (primary) must never look the same as estimated or averaged (secondary) in the file.
  • A defensible estimate confesses, carrying a label, a method, and an uncertainty; the same number unlabeled is a misstatement.
  • Generation fabrication hides numbers and false targets inside fluent prose; good writing is not a source, so trace every claim and figure independently of how well the sentence reads.
  • The diagnostic is one habit: for any AI value, name the job, hunt its specific failure, and apply its specific verification before it reaches the assurer.