โ†
AI for ESG & Sustainability Reporting
Aware ยท M4 ยท lesson 4 of 19 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
AI in the GHG Inventory (Scope 1, 2, and 3)
๐Ÿ“–
now learning

AI in the GHG Inventory (Scope 1, 2, and 3)

15 min

A carbon accountant is closing the annual inventory at 11pm. Scope 1 and Scope 2 are done and reconciled. Scope 3 is the problem. It is 75% of the footprint, it spans fifteen categories, and the activity data simply is not there for half of them. An AI assistant offers a tempting shortcut: "I can estimate the missing categories from industry averages and have your total in five minutes." The number it produces would look exactly like every other number in the inventory. That is precisely what makes it dangerous. An assurer cannot tell a measured tonne from an invented one by looking at the cell. They can only tell by pulling the thread back to the evidence, and if there is no evidence, the inventory unravels.

The Three Scopes in Plain Terms

The greenhouse gas inventory is organised by the GHG Protocol, the global standard nearly every framework points back to, into three scopes. The scopes are not about how important an emission is; they are about who controls the source, which decides who counts it. Getting the boundaries right is the first place AI can quietly mislead you, so define them precisely.

Scope 1 is direct emissions from sources the company owns or controls: fuel burned in your own boilers, furnaces, and vehicle fleet, and process emissions from your own operations. If you light it or you run it, it is Scope 1. Scope 2 is indirect emissions from the energy you purchase and consume: the electricity, steam, heat, or cooling you buy. You did not produce the emissions; the power plant did, but they happened because of your demand, so you report them. Scope 3 is every other indirect emission across your value chain, upstream and downstream, that you do not own or control: the emissions embedded in the goods and services you buy, in business travel, in the transport of your products, in the use of your products by customers, and in their disposal. Why you care about the distinction: each scope has different data sources, different reliability, and a different assurance burden. Scope 1 and 2 are largely yours to measure. Scope 3 is mostly other people's data, which is exactly why it is both the largest number and the riskiest.

The 15 Scope 3 Categories

Scope 3 is not one number; it is fifteen defined categories, eight upstream and seven downstream. Upstream covers purchased goods and services, capital goods, fuel and energy activities not in Scope 1 or 2, upstream transportation and distribution, waste generated in operations, business travel, employee commuting, and upstream leased assets. Downstream covers downstream transportation and distribution, processing of sold products, use of sold products, end-of-life treatment of sold products, downstream leased assets, franchises, and investments. You do not have to report every category; you report the ones that are material to your business and you document why the others are excluded. Why you care: an undocumented category exclusion is one of the most common assurance findings, and it is exactly the kind of gap an AI tool will paper over silently if you let it estimate without telling you.

How the Numbers Are Actually Built

Every line in a GHG inventory is built from the same simple equation, and understanding it tells you exactly where AI can help and where it must not. The equation is activity data multiplied by an emission factor equals emissions. Activity data is the measure of the thing that happened: litres of diesel burned, kilowatt-hours of electricity consumed, tonnes of steel purchased, kilometres flown. An emission factor is the conversion rate that turns that activity into greenhouse gas, for example kilograms of CO2 equivalent per litre of diesel, sourced from a named, dated, authoritative database. Multiply the two and you get the emissions for that line. Sum the lines and you get the inventory.

This equation is the whole game. A defensible number requires defensible activity data and a defensible emission factor, each traceable to a source. The two ways an inventory fails are: the activity data was wrong, missing, or invented; or the emission factor was wrong, outdated, or hallucinated. AI touches both inputs, helpfully on some tasks and catastrophically on others, and the difference is always the same: does the AI produce a number that traces to evidence, or a number that merely looks right?

One more concept belongs here before we let AI near the inventory, because it decides which numbers you even count: the organizational and operational boundaries. The organizational boundary defines which entities are inside your inventory, usually by an equity-share or control approach, and the operational boundary defines which activities you include within those entities. These boundaries are choices you make and document. They are not something an AI infers, and a model that quietly includes or drops a site, a joint venture, or a leased asset can shift your total without anyone noticing. When the assurer reconciles your reported total to your corporate structure, an undocumented boundary call is the kind of discrepancy that opens an entire engagement back up. So the first rule of an AI-assisted inventory is the same as the first rule of a manual one: the boundary is human, written, and stable, and every number sits inside it.

An assurer cannot see the difference between a measured tonne and an invented one in the spreadsheet cell. They can only see it by following the thread to the evidence. If the thread ends in the model's imagination, the number is a misstatement.

Where AI Genuinely Accelerates the Inventory

AI earns its place in the GHG inventory on the tasks that are high-volume, repetitive, and verifiable against a source. The key word is verifiable: every one of these gains comes with a way to check the output against evidence, which is what keeps it inside the assurance perimeter.

The first and best use is extraction. Activity data arrives as a mess of utility bills, fuel invoices, travel records, and supplier spreadsheets in inconsistent formats, in different languages, with different units, and with the number you need buried somewhere on page two next to a marketing logo. AI is genuinely good at reading a stack of PDFs and pulling out the numbers, units, dates, and account references into structured fields, while preserving where each value came from. The accountant still checks the extraction against the source document, but the model does the tedious transcription that used to eat days, and it does not get bored on page 300. The second use is classification and mapping: AI can map a line of spend or a supplier to the right Scope 3 category, or flag which categories are likely material for your sector, giving you a faster first-pass scoping. The third is reconciliation support: AI can compare this year's extracted figures against last year's and flag the implausible jumps, the doubled meter reading, the missing site, the unit that switched from kWh to MWh, so a human can investigate before the error reaches the total. The fourth is emission-factor lookup with provenance, but only when the workflow forces the model back to a named database rather than letting it recall a number from memory. In each case AI accelerates the work and a human verifies the result against a source. That is the pattern that keeps the inventory assurable.

Where AI Must Not Estimate Silently

Now the danger zone, and it is concentrated almost entirely in Scope 3. Because Scope 3 averages roughly 75% of a company's total footprint and the data is the hardest to get, with 62% of reporters citing internal data quality and 79% citing supplier-data availability as top barriers, it is exactly where the temptation to let AI "fill the gap" is strongest and most fatal.

The cardinal failure is the silent estimate: the AI encounters a category with no activity data, quietly substitutes an industry average or a plausible-looking figure, and drops it into the inventory in a cell that looks identical to a measured value. Now your inventory contains a number that nobody measured, with no flag, no method note, and no source. When the assurer pulls the thread, it ends in nothing. This is the move that fails an engagement and can trigger a restatement and a greenwashing headline. The related failures are the hallucinated emission factor, a plausible but invented conversion rate with no database behind it, and fabricated activity data, where the model "completes" a partial dataset by inventing the missing rows. All three share the same DNA: a number that looks right and traces to nothing.

Estimated and Disclosed Versus Made Up

This is not a ban on estimation. Real inventories are full of estimates, because for many Scope 3 categories primary data genuinely does not exist yet, and the GHG Protocol expressly allows estimation methods such as spend-based and average-data approaches. A spend-based estimate multiplies how much money you spent on a category by an emission factor expressed per unit of currency; an average-data estimate multiplies a physical quantity, like mass, by an industry-average factor per unit. Both are legitimate, both are weaker than primary supplier data, and both must be visible as estimates in the file so the assurer can weigh them. The line is not estimation versus measurement. The line is labelled versus laundered. A defensible estimate is marked as an estimate, names its method (for example spend-based using a named factor set), and carries its uncertainty, so the assurer sees exactly what it is and can judge it. A laundered estimate is dressed up to look like measured activity data, with no label and no method, so the reader cannot tell it apart from a metered figure. The first is good practice. The second is a misstatement. AI is happy to produce either; your job is to make sure it only ever produces the first, by forcing every estimate to declare itself.

Before and After: A Scope 3 Category, Two Ways

Watch a single category, purchased goods and services (Scope 3 Category 1), go two different ways with the same AI tool. The company buys steel, packaging, and electronic components, and primary supplier data is thin.

Before (the silent shortcut, what almost shipped): The accountant asks the AI to "complete the purchased goods emissions for the year." The model returns a single tidy figure: 48,200 tonnes CO2e. It looks like every other number in the file. There is no note on how it was built, no split by material, no flag that supplier data was missing for two-thirds of spend, and no emission factor cited. In fact the model had primary data for steel only and silently filled packaging and components with a generic sector average it recalled from training, which may or may not be current or correct. Dropped into the inventory, this number is indistinguishable from a measured one and completely indefensible. The assurer asks "how did you get 48,200 tonnes," and the answer is "the AI completed it," which is no answer at all.

After (the labelled, traceable build): The accountant forces the work into the equation and demands provenance on every piece. For steel, there is primary data: 6,400 tonnes purchased, multiplied by a supplier-specific emission factor traced to the supplier's own verified disclosure, giving a primary-data line clearly tagged as primary. For packaging, there is no supplier data, so the accountant builds a spend-based estimate: EUR 2.1M of packaging spend multiplied by an environmentally-extended input-output factor from a named, dated database, tagged as a secondary, spend-based estimate with its uncertainty noted. For electronic components, also no supplier data, an average-data estimate using mass and a named average factor, again tagged secondary with method and uncertainty. The AI helped with all of it: it extracted the spend, looked up candidate factors with their sources, and drafted the calculation, but every factor was verified back to its database and every estimate declared itself. The category total now comes with a basis: which lines are primary, which are estimated, by what method, with what factors, and where the gaps are. The number might even be close to 48,200 tonnes. The difference is not the answer. The difference is that this one survives the engagement and the other one ends a career.

That is the whole discipline in one comparison. The same model, pointed at the same data, produces either a laundered black-box total or a labelled, traceable build. The accountant chooses which, by refusing to accept any number that does not trace to evidence and forcing every estimate to wear a label.

Working Rules for AI in the Inventory

A handful of rules keep AI on the right side of the line. Use AI freely for extraction, classification, reconciliation flagging, and provenance-forced factor lookup, because each is verifiable against a source. Never accept a Scope 3 total the model "completed" without seeing how each category was built. Force every emission factor back to a named, dated database, and reject any factor the model cannot source, because a hallucinated factor silently corrupts everything downstream of it. Insist that every estimate declares its method and uncertainty and is tagged as secondary, so it can never masquerade as measured data. Keep the activity-data extraction linked to its source document, so any line can be traced back to the bill or invoice it came from. Treat an implausible jump as a finding to investigate, not a number to accept, and let the model help you spot it. And remember the boundary discipline: a missing category should be a disclosed, reasoned exclusion, never a silent AI estimate that hides the gap. Follow these and AI gives you a faster inventory with a cleaner trail. Ignore them and AI gives you a faster path to a restatement.

It helps to hold the whole picture in one frame. Scope 1 and Scope 2 are where AI is mostly a productivity tool: the data is yours, the sources are clean, and extraction and reconciliation simply save you time. Scope 3 is where AI is both your biggest opportunity and your biggest exposure, because it is where the data is missing and the temptation to fabricate is highest. The skill is not using AI everywhere with equal trust. It is calibrating your trust to the scope: lean on AI hard for the verifiable transcription and lookup work, and put a hard human gate in front of any Scope 3 number that the model would otherwise complete on its own. The accountant who internalises that calibration closes the inventory in weeks instead of months and hands the assurer a file where every tonne, primary or estimated, declares exactly what it is. That accountant is faster and more defensible at the same time, which is the only position worth being in when the engagement starts.

Key Takeaways

  • The GHG inventory has three scopes defined by control: Scope 1 is direct emissions you own or operate, Scope 2 is purchased energy, and Scope 3 is everything else across your value chain.
  • Scope 3 splits into 15 defined categories (8 upstream, 7 downstream); you report the material ones and must document why others are excluded, because a silent exclusion is an assurance finding.
  • Every line is activity data multiplied by an emission factor; a defensible number requires both inputs to trace to a source, and AI failures come from corrupting one or the other.
  • AI genuinely accelerates verifiable tasks: extracting activity data from messy documents, classifying spend into categories, flagging implausible year-on-year jumps, and looking up factors with provenance.
  • Scope 3 is the danger zone because it is about 75% of the footprint with the worst data (62% cite internal data quality, 79% cite supplier-data availability as barriers), which is where the temptation to let AI fill gaps is strongest.
  • The cardinal failure is the silent estimate: AI substitutes an industry average or invented figure into a cell that looks identical to measured data, with no flag, method, or source.
  • The line is not estimation versus measurement; it is labelled versus laundered. A disclosed estimate naming its method and uncertainty is good practice; an estimate dressed up as measured data is a misstatement.
  • An assurer cannot distinguish a measured tonne from an invented one in the cell; force every factor to a named database, tag every estimate as secondary, and keep extraction linked to its source so the thread always leads to evidence.