Estimating Without Fabricating
A carbon accountant is closing a Scope 3 category with a hole in it. Two-thirds of the suppliers responded with real numbers; one-third never answered, and the deadline is in three days. She has two choices that produce the exact same final figure. She can build a transparent estimate for the missing third, label it as an estimate, name the method, note the uncertainty, and disclose the gap. Or she can let the AI quietly fill the missing rows with a plausible average, drop them into the inventory unlabelled, and ship a category that looks complete and measured. The number is identical either way. One of these is good professional practice that an assurer will accept and even respect. The other is fabrication, and it ends careers. The entire difference is a label, a method, and an honest line about what is not known.
Estimation Is Not the Enemy
It is worth saying plainly at the outset, because the fear of fabrication can push people toward an impossible standard: estimation is not only allowed, it is unavoidable and expected. The GHG Protocol, the global standard most reporting frameworks point back to, explicitly provides estimation methods precisely because primary data does not exist for large parts of a value chain. Scope 3, which averages roughly 75% of a company's footprint across 15 defined categories, is built substantially on estimates, because 79% of reporters cannot reliably get supplier data and 62% struggle with internal data quality. An inventory with no estimates in it would not be more honest; for most companies it would be impossible, or it would mean silently excluding the categories where data is hard, which is its own assurance failure.
So the goal of this lesson is not to teach you to avoid estimation. It is to teach you the one thing that separates a legitimate estimate from a fabrication, because they can produce the same number and a careless eye cannot tell them apart. That one thing is disclosure: an estimate that declares itself, its method, and its uncertainty is a defensible part of an inventory; an estimate dressed up as measured activity data, with no label and no method, is a misstatement. The skill is building the first kind and never the second, especially when an AI tool makes the second kind so easy and so tempting.
The Anatomy of a Defensible Estimate
A defensible estimate is not a number with a confession attached as an afterthought. It is a small, complete package, and every part of it does work the assurer needs. There are four parts.
The value: the estimated figure itself, in its proper unit. The label: an explicit marker that this is an estimate, not measured activity data, carried in the file so it can never be mistaken for a metered or supplier-reported figure. The method: how the estimate was built, named specifically, for example spend-based using a named per-currency factor, or average-data using a named average factor applied to a physical quantity, or a proxy from a comparable site, with the inputs and the factor's provenance stated. The uncertainty: an honest characterisation of how rough the estimate is, whether a qualitative band (high, medium, low confidence) or a quantitative range, so the reader knows not to treat it as precise. Together these four turn a bare number into a number an assurer can weigh: they can see it is an estimate, see exactly how it was made, and see how much to trust it. Strip any one of the four and the estimate weakens; strip the label and the method and you no longer have an estimate at all, you have a fabrication wearing the clothes of a measurement.
Spend-Based and Average-Data as Honest Estimates
The two workhorse estimation methods are legitimate precisely because they can be fully disclosed. A spend-based estimate multiplies how much you spent on a category by an emission factor expressed per unit of currency; everything about it is stateable, the spend, the factor, the factor's source, and its known coarseness. An average-data estimate multiplies a physical quantity, such as mass, by an industry-average factor per unit; again, every input is nameable. Neither method is a guess in the dark. Each is a documented procedure with declared inputs and a declared factor, which is exactly why each can be labelled, reproduced, and assured. The estimate is honest not because it is accurate, but because it shows its whole working and admits its own roughness.
Reproducibility is the property that does the heavy lifting here, so it is worth stating as a test. Ask of any estimate: could a competent stranger, handed only your file and no access to you, rebuild this number and arrive at the same figure? For a disclosed spend-based estimate the answer is yes, because the stranger has the spend, the named factor, and the arithmetic. For a disclosed average-data estimate the answer is yes, because the stranger has the quantity, the named factor, and the arithmetic. For a fabrication the answer is no, because some input was invented, omitted, or unsourced, so the stranger hits a dead end. This reproducibility test is the same one an assurer applies, and it is the cleanest single check you can run on your own estimates before anyone else does: if you cannot hand the calculation to a colleague and have them reproduce it, it is not yet a defensible estimate, whatever label you put on it.
It is also worth being clear that uncertainty is not an admission of weakness that undermines the number; it is part of what makes the number usable. A figure stated as a single confident value invites the reader to treat it as exact, which is itself a kind of misrepresentation when the figure is a coarse estimate. A figure stated with a confidence band tells the reader exactly how to use it: solid enough to include in the total, not solid enough to base a precise target on. Disclosing uncertainty is how an estimate communicates its own appropriate use, and a reader who knows a number is low-confidence will treat it correctly, whereas a reader who is given false precision will misuse it. Honest uncertainty is therefore a service to the reader, not a confession, and assurers consistently respond better to an estimate that knows its own limits than to one that pretends it has none.
The line is not estimated versus measured. The line is labelled versus laundered. A disclosed estimate is good practice. The same number, stripped of its label and method, is a misstatement.
The Bright Line: Where Estimation Becomes Fabrication
Fabrication is not a separate, exotic act. It is most often an ordinary estimate with its honesty removed. Walk the bright line carefully, because it is crossed quietly. An estimate becomes a fabrication the moment it is presented as something it is not. The most common crossing is the unlabelled estimate: a perfectly reasonable spend-based figure dropped into the inventory in a cell that looks identical to a measured one, with no flag that it was estimated. The number might be fine; the presentation is a lie about its nature. The second crossing is the methodless estimate: a number offered as an estimate but with no statement of how it was built, so it cannot be reproduced or judged, which makes it indistinguishable from a guess. The third is the invented input: where the model does not estimate from a real basis at all but fills a gap with fabricated activity data, inventing the missing supplier rows or the missing quantities outright, then estimating on top of fiction.
The reason AI makes this dangerous is that a generative model crosses all three lines by default and does so fluently. Asked to "complete the category," it will fill gaps with plausible numbers, present them in the same uniform format as the real data, omit the method, and never once flag that it estimated, because completing the table is what it was asked to do and admitting uncertainty is not. The model is not malicious; it is doing exactly what it was told, and what it was told to do is the fabrication. The discipline is to never ask the model to complete or fill, but to ask it to estimate explicitly, declare the method, declare the uncertainty, and flag the gap, so that every estimate it produces arrives already wearing its label.
A Worked Example: One Data Gap, Two Ways
A company is closing Scope 3 Category 1 (purchased goods and services). For its packaging spend, thirty suppliers were surveyed; twenty responded with usable data, ten did not. The accountant must account for the ten non-responders, who represent EUR 1.4 million of packaging spend, and she uses AI to help.
Before (the laundered fill, what almost shipped): The accountant asks the AI to "complete the packaging emissions for all thirty suppliers." The model returns a single clean figure for the whole category, 2,300 tonnes CO2e, with the ten missing suppliers silently filled using a generic average and folded invisibly into the total. There is no line saying ten of thirty suppliers did not respond, no method note, no uncertainty, no label distinguishing the estimated portion from the measured portion. The category reads as if all thirty suppliers reported real data. When the assurer asks "how complete is this, and how much is estimated," the file cannot answer, because it erased its own gap. The 2,300 tonnes looks finished and is indefensible, and the silent fill, if discovered, reads as fabrication.
After (the labelled estimate, the defensible build): The accountant separates what is known from what is not. For the twenty responders, she records their reported figures as measured, tagged primary. For the ten non-responders, she builds an explicit estimate: EUR 1.4M of their packaging spend multiplied by a named spend-based factor with provenance, producing an estimated portion that is labelled as an estimate, with the method stated (spend-based, named factor) and the uncertainty noted as low-to-medium confidence given the coarse method. The file now shows the category total, the split between measured and estimated, the fact that ten of thirty suppliers did not respond, the method used to estimate them, and the uncertainty on that portion. The total might still be near 2,300 tonnes. The difference is that this version tells the truth about itself: here is what we measured, here is what we estimated, here is how, here is how rough it is, and here is the gap we are working to close. When the assurer asks how much is estimated, the accountant answers in one sentence and points to the labelled portion. The number survives. The career survives.
That is the whole discipline in one comparison. The same AI, the same EUR 1.4M gap, the same final tonnage, produces either a laundered fill that erases its own uncertainty or a labelled estimate that declares it. The accountant chooses which, by refusing to let the model complete or fill anything silently, and by insisting every estimate arrive with its label, its method, and its uncertainty already attached.
Three Ways the Line Gets Crossed in Practice
The bright line is easy to state and easy to cross under pressure, so it helps to see the three crossings as concrete situations rather than abstractions, because each feels reasonable in the moment.
The deadline crossing. A category is incomplete the night before filing. The honest move, a labelled estimate with method and uncertainty, takes an extra hour to set up. The unlabelled fill takes a minute and produces a number that looks just as finished. Under deadline pressure the minute wins, and an estimate ships with no label. Nobody decided to fabricate; they decided to save an hour, and the saving quietly crossed the line. The defence is to make the labelled path the default path, so that under pressure the honest move is also the fast move, because the model was instructed from the start to produce estimates already wearing their labels.
The tidiness crossing. A reviewer or a manager prefers a clean, single number to a number with caveats attached, and asks for the qualifications to be stripped so the report reads better. The estimate was honest when it was built; it becomes a fabrication when its method note and uncertainty band are deleted in the name of presentation. The defence is to treat the label, method, and uncertainty as inseparable from the number, not as optional commentary, so that removing them is understood as changing the number's meaning, not tidying its appearance.
The inheritance crossing. An analyst inherits a prior-year figure with no provenance and carries it forward, not knowing whether it was measured or estimated, presenting it this year as if it were solid. The fabrication here is not fresh invention; it is the laundering of an unknown into an assertion of quality the analyst cannot actually support. The defence is to treat any inherited number with no tier and no method as unproven until re-sourced, and to estimate it transparently this year rather than inherit a false confidence. In all three cases the crossing was ordinary, almost administrative, which is exactly why the discipline has to be a standing habit and not a heroic act reserved for obvious temptations.
Working Rules for Estimating Without Fabricating
A handful of rules keep your estimates on the right side of the bright line. Never ask a model to complete or fill a category; ask it to estimate explicitly and to flag what it estimated. Require every estimate to carry its four parts: value, label, method, and uncertainty. Keep the estimated portion visibly separate from the measured portion, so no estimate ever masquerades as measured data. State the method specifically enough that someone else could reproduce the estimate, because an estimate nobody can reproduce is indistinguishable from a guess. Attach the factor's provenance to the estimate, because a spend-based or average-data estimate is only as defensible as the named, dated factor inside it. Disclose the gap honestly: a category where one-third of suppliers did not respond should say so, not hide the hole inside an average. Treat invented activity data, the fabricated missing rows, as the brightest line of all, never to be crossed, because estimating on top of invented inputs is fabrication twice over. And remember that an estimate's honesty does not depend on its accuracy: a rough estimate that declares its roughness is defensible, while a precise-looking number that hides its nature is not. Follow these and AI helps you achieve completeness without fabrication, which is the whole point of the goldmine. Ignore them and AI helps you fabricate faster.
Key Takeaways
- Estimation is not the enemy; it is allowed, unavoidable, and expected, because Scope 3 is roughly 75% of the footprint and built substantially on estimates where primary data does not exist.
- The one thing that separates a legitimate estimate from a fabrication is disclosure: the same number is good practice when labelled with its method and uncertainty and a misstatement when dressed up as measured data.
- A defensible estimate is a four-part package: the value, an explicit estimate label, the named method with its factor provenance, and an honest characterisation of uncertainty.
- Spend-based and average-data estimates are legitimate because every input is stateable and the procedure is reproducible; the estimate is honest not because it is accurate but because it shows its whole working.
- Estimation becomes fabrication at three crossings: the unlabelled estimate (looks measured), the methodless estimate (cannot be reproduced), and the invented input (fabricated activity data estimated on top of fiction).
- AI crosses all three lines by default when asked to complete or fill a category, doing so fluently; the fix is to ask it to estimate explicitly, declare method and uncertainty, and flag the gap.
- Keep the estimated portion visibly separate from the measured portion, disclose data gaps honestly rather than hiding them in an average, and never estimate on top of invented rows.
- When an assurer asks how much of a category is estimated and by what method, you should answer in one sentence and point to a labelled portion; a category that erased its own gap is indefensible no matter how finished it looks.
Skill.re