AI for ESG & Sustainability Reporting
Proficient · M7 · lesson 7 of 24 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Defensible Estimation and Uncertainty Disclosure
📖
now learning

Defensible Estimation and Uncertainty Disclosure

15 min

Half of a company's Scope 3 footprint sits in a single category, purchased goods and services, and more than half of that category has no primary data behind it at all. The suppliers did not respond, the meters do not exist, the activity happened three tiers up a chain the company cannot see. The carbon accountant cannot leave the category blank, because an omitted material category is its own assurance failure. She cannot fabricate, because a made-up number is the fastest way to fail the engagement and make a greenwashing headline. So she has to estimate, openly, and then do the thing most teams skip: tell the reader and the assurer exactly how rough the estimate is. This lesson is about estimating the unmeasurable in a way that is labelled, reproducible, and honest about its own uncertainty, because the assurer will probe that uncertainty whether or not you disclosed it.

Estimation Is Mandatory, Not a Shortcut

Start by clearing away the guilt, because it pushes people toward an impossible standard and, paradoxically, toward worse choices. Estimation in a greenhouse-gas inventory is not a corner cut; it is built into the standard. The GHG Protocol, the global greenhouse-gas accounting standard most frameworks point back to, explicitly provides estimation methods because primary data does not exist for large parts of any value chain. Scope 3, the value-chain emissions split across fifteen categories, averages roughly 75% of a company's total footprint, and it is built substantially on estimates precisely because 79% of reporters cannot reliably obtain supplier data and 62% struggle with internal data quality. An inventory with zero estimates is not purer; for most companies it is impossible, or it means silently dropping the hard categories, which is the worse sin.

So the question is never whether to estimate. It is whether the estimate declares itself. A defensible estimate is one that is labelled as an estimate, names its method, states its inputs and the provenance of its factor, and discloses its uncertainty. A laundered estimate is the same arithmetic with all of that stripped off, dropped into the inventory in a cell that looks measured. The number can be identical. One is good professional practice an assurer respects; the other is a misstatement. The entire craft of this lesson is building the first and never the second, and the part teams most often skip, the uncertainty, is the part the assurer cares about most.

The Estimation Methods, and When Each Is Defensible

Two workhorse methods carry most Scope 3 estimation, and each is defensible because every input is stateable. A spend-based estimate multiplies how much you spent on a category by an emission factor expressed per unit of currency, for example kilograms of CO2e per euro spent on packaging. It is the coarsest method, because money is a poor proxy for physical emissions, but it is fully disclosable: the spend, the factor, the factor's source, and the method's known roughness are all on the record. An average-data estimate multiplies a physical quantity, such as mass or distance, by an industry-average factor per unit; it is finer than spend-based because it uses a physical activity measure, and it too is fully stateable. Both sit below primary data, the supplier-reported or metered figure that is the gold standard, and both are tagged secondary.

The choice of method is itself a disclosure, not a hidden detail. Method choice changes the number and changes its uncertainty, so the assurer wants to see which method you used for which category and why. The defensible pattern is a hierarchy: use primary data where you can get it; fall to average-data where you have a physical quantity but no supplier figure; fall to spend-based only where you have neither, and where the spend is the best handle you have. AI can help you apply any of these methods quickly, and that speed is genuinely useful. The danger is that AI applies them silently and uniformly, hiding the method behind a clean total, which is exactly the move that turns a legitimate estimate into a laundered one.

The line is not estimated versus measured. The line is labelled versus laundered. The same number is good practice when it declares its method and uncertainty, and a misstatement when it pretends to be measured.

Quantifying the Uncertainty You Will Be Asked About

Here is the part most reporters underbuild. An estimate without an uncertainty statement invites the reader to treat it as exact, which is itself a misrepresentation when the figure is coarse. Uncertainty in a Scope 3 estimate comes from two places, and naming both is what makes your disclosure credible. There is activity-data uncertainty: how well the input quantity is known, which is high when you are using spend as a proxy and lower when you have a measured physical quantity. There is emission-factor uncertainty: how representative the factor is for your specific activity, which is high when the factor is a broad industry average applied to a specific supplier and lower when it is well-matched and recent.

You can express the result qualitatively or quantitatively, and both are acceptable if honest. A qualitative band tags the estimate high, medium, or low confidence, with a one-line reason: a spend-based figure for an opaque category might be low confidence because both the proxy and the factor are coarse, while an average-data figure for a well-characterised material might be medium confidence. A quantitative range states the estimate as a central figure with a plus-or-minus band, for example 2,300 tonnes CO2e plus or minus 30%, derived from the known coarseness of the method and factor. The quantitative form is stronger where you can support it, but a disciplined qualitative band beats a false-precise single number every time. The point is the same in both forms: the uncertainty travels with the number so the reader knows exactly how much weight it can bear.

Communicating Uncertainty Without Undermining the Number

Many teams hide uncertainty because they fear it makes the inventory look weak. The opposite is true in an assurance context. A figure stated with its confidence band tells the reader precisely how to use it: solid enough to include in the total, not solid enough to anchor a precise target. A figure stated as a single confident value, when it is actually a coarse estimate, gives false precision, and false precision is the thing an assurer punishes, because it misleads the reader about how much is really known. Disclosing uncertainty is therefore a service, not a confession. Assurers consistently respond better to an estimate that knows its own limits than to one that pretends it has none, because the first shows professional judgment and the second shows either ignorance or concealment.

What the Assurer Probes, and How to Be Ready

An assurer reading your inventory under a limited assurance engagement, the most common form today, performs procedures sufficient to conclude that nothing has come to their attention suggesting the figures are materially misstated. Even at this lighter level, estimates are where they push, because estimates are where misstatement hides. They will ask, of any estimated category: what method did you use, why that method, what is the factor and where is it from, how much of this category is estimated versus measured, and how uncertain is the estimate? Under reasonable assurance, the higher level the market is trending toward, they push harder and test more, and an estimate that cannot answer these questions becomes a finding.

The way to be ready is to build the answers into the estimate before the question is asked. Every estimate carries its four parts: the value, the label, the named method with factor provenance, and the uncertainty. When the assurer asks how much of Category 1 is estimated, you point to a clearly separated estimated portion and answer in a sentence. When they ask how uncertain it is, you point to the confidence band already attached. The estimate that was built to be questioned survives the question; the estimate built to look finished does not. The discipline is to assume every estimate will be probed, because under either assurance level the material ones will be.

A Worked Example: Estimating With Uncertainty Disclosed

A company is closing Scope 3 Category 1, purchased goods and services, for its chemical inputs. It buys from twelve suppliers; five respond with supplier-reported emissions, seven do not. The seven non-responders represent EUR 3.2 million of chemical spend. The accountant uses AI to help estimate the missing seven, and the difference between two builds is the whole lesson.

Before, the laundered estimate. She asks the AI to "complete the chemical emissions for all twelve suppliers." The model returns one clean figure for the category, 5,800 tonnes CO2e, with the seven missing suppliers filled using a generic factor and folded invisibly into the total. There is no split between reported and estimated, no method note, no uncertainty, no statement that seven of twelve did not respond. The category reads as fully measured. When the assurer asks how much is estimated and how uncertain, the file cannot answer, because it erased its own gap. The 5,800 tonnes looks finished and is indefensible.

After, the defensible estimate with uncertainty. She separates the known from the unknown. The five responders' supplier-reported figures are recorded as primary, totalling, say, 2,400 tonnes CO2e. For the seven non-responders, she builds an explicit spend-based estimate: EUR 3.2M times a named spend-based factor of, say, 1.05 kg CO2e per euro, giving roughly 3,360 tonnes CO2e, tagged secondary, labelled as an estimate, method stated as spend-based with the named factor's provenance. She assigns it an uncertainty: low-to-medium confidence, expressed quantitatively as plus or minus 40%, because spend is a coarse proxy and the factor is a broad average. The file now shows the category total of roughly 5,760 tonnes, split into 2,400 measured and 3,360 estimated, with the method, the factor source, the uncertainty band, and the disclosed fact that seven of twelve suppliers did not respond. The total is near the laundered figure. The difference is that this version can answer every question the assurer asks, in a sentence and a click, and the estimated portion is visibly separate from the measured one. The number survives.

That is the discipline in one comparison. The same AI, the same EUR 3.2M gap, nearly the same final tonnage, produces either a laundered fill that erases its uncertainty or a labelled estimate that quantifies it. The accountant chooses which by refusing to let the model complete or fill anything silently, and by insisting every estimate arrive with its method, its factor provenance, and its uncertainty already attached.

Three Ways the Estimate Gets Laundered in Practice

The line between a labelled estimate and a laundered one is easy to state and easy to cross under real pressure, so it helps to see the crossings as concrete situations rather than abstractions. Each one feels reasonable in the moment, which is exactly why it is dangerous.

The deadline crossing. A material category is incomplete the night before filing. The honest move, a labelled estimate with method, factor provenance, and an uncertainty band, takes an extra hour to set up. The unlabelled fill takes a minute and produces a number that looks just as finished on the page. Under deadline pressure the minute wins, and an estimate ships with no label and no band. Nobody decided to deceive anyone; they decided to save an hour, and the saving quietly crossed the line into misstatement. The defence is to make the labelled path the default path, so that under pressure the honest move is also the fast move, because the AI was instructed from the start to produce estimates already wearing their labels and bands.

The tidiness crossing. A reviewer or a manager prefers a clean single number to a number with caveats attached, and asks for the uncertainty band and the method note to be stripped so the report reads better. The estimate was honest when it was built; it becomes a misstatement when its band and method are deleted in the name of presentation, because deleting them converts a known-coarse figure into a false-precise one. The defence is to treat the label, method, and uncertainty as inseparable from the number, not as optional commentary, so that removing them is understood as changing the number's meaning, not tidying its appearance. A number that lost its band did not get cleaner; it got less true.

The inheritance crossing. An analyst inherits a prior-year estimate with no recorded method and no band, carries it forward, and presents it this year as if it were solid. The laundering here is not fresh invention; it is the carrying-forward of an unknown as if it were an assertion of quality the analyst cannot actually support. The defence is to treat any inherited number with no method and no uncertainty as unproven until re-derived, and to estimate it transparently this year rather than inherit a false confidence. In all three cases the crossing was ordinary, almost administrative, which is exactly why the discipline has to be a standing habit rather than a heroic act reserved for obvious temptations.

The Reproducibility Test You Can Run on Yourself

Before any estimate leaves your hands, run the test the assurer will run, because it is the cleanest single check on whether your estimate is defensible. Ask: could a competent stranger, handed only your file and no access to you, rebuild this number and arrive at the same figure? For a disclosed spend-based estimate the answer is yes, because the stranger has the spend, the named dated factor, and the arithmetic. For a disclosed average-data estimate the answer is yes, because the stranger has the physical quantity, the named factor, and the arithmetic. For a laundered fill the answer is no, because some input was invented, the method was never stated, or the factor has no source, so the stranger hits a dead end. If you cannot hand the calculation to a colleague and have them reproduce it, it is not yet a defensible estimate, whatever label you have put on it. Reproducibility is the property that does the heavy lifting, and it is the same property that lets the uncertainty band mean something, because a band attached to an irreproducible number is decoration, not disclosure.

Why AI Launders by Default, and How to Stop It

The reason this matters so acutely with AI is that a generative model launders estimates by default and does so fluently. Asked to "complete the category," it fills gaps with plausible numbers, presents them in the same uniform format as the real data, omits the method, and never flags that it estimated or how uncertain the result is, because completing the table is what it was asked to do and admitting uncertainty was not in the instruction. The model is not deceiving anyone on purpose; it is doing exactly what it was told, and what it was told to do is the laundering. The fix is to change the instruction. Never ask the model to complete, fill, or finish a category. Ask it to estimate explicitly, declare the method, attach the factor's provenance, and state the uncertainty, so that every estimate it produces arrives already labelled and already carrying its confidence band. The same speed, redirected, produces defensible estimates instead of laundered ones.

Key Takeaways

  • Estimation is mandatory, not a shortcut: the GHG Protocol provides estimation methods because primary data does not exist for much of any value chain, and Scope 3, roughly 75% of a footprint, is built substantially on estimates given the 79% supplier-data barrier.
  • A defensible estimate declares itself: it is labelled, names its method, states its inputs and factor provenance, and discloses its uncertainty; a laundered estimate is the same arithmetic with all of that stripped off and presented as measured.
  • Spend-based and average-data are the two workhorse methods; method choice changes the number and its uncertainty and is itself a disclosure the assurer wants to see, following a hierarchy from primary to average-data to spend-based.
  • Uncertainty has two sources, activity-data uncertainty and emission-factor uncertainty, and can be expressed as a qualitative confidence band or a quantitative plus-or-minus range; a disciplined band beats a false-precise single number.
  • Disclosing uncertainty is a service, not a confession: a confidence band tells the reader how much weight the number can bear, while false precision is the thing an assurer punishes because it misleads about what is really known.
  • Even under limited assurance, the most common level, the assurer probes estimates first; build the answers in before the question is asked, with each estimate carrying value, label, method with provenance, and uncertainty.
  • In the worked Category 1 example, a labelled spend-based estimate of roughly 3,360 tonnes at plus or minus 40%, kept separate from 2,400 tonnes of measured data, survives every assurer question, while a single laundered 5,800-tonne figure cannot.
  • AI launders by default when told to complete or fill; redirect the instruction to estimate explicitly with method, provenance, and uncertainty, and the same speed produces defensible estimates instead.