AI for ESG & Sustainability Reporting
Capable · M20 · lesson 20 of 24 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Spend-Based vs. Activity-Based Methods
📖
now learning

Spend-Based vs. Activity-Based Methods

15 min

A carbon accountant has two ways to calculate the emissions from the steel her company bought this year, and the choice between them is worth fifteen points of accuracy and an entire conversation with the assurer. She can take the EUR 4.2 million the company spent on steel and multiply it by an emission factor expressed per euro: fast, requires only the invoice total, and gives a coarse number. Or she can take the 6,400 tonnes of steel physically purchased and multiply it by an emission factor expressed per tonne: more accurate, but it requires her to actually know the tonnage and source a physical factor. An AI assistant, asked to "calculate the purchased-goods emissions," will almost always reach for the first method, because it is the one the spend data alone can support. Whether that shortcut is defensible depends entirely on the category, and knowing the difference is the skill that separates a fast inventory from a fast restatement.

The Three Methods on One Shelf

Every emissions calculation follows the GHG Protocol's core equation, the global standard most reporting frameworks point back to: activity data multiplied by an emission factor equals emissions. The three Scope 3 calculation methods differ only in what you use as the activity data and which factor matches it. Understanding them as three points on a single quality ladder is the foundation for choosing well.

The spend-based method uses money as the activity data. You take how much you spent on a category, in a currency, and multiply it by an emission factor expressed as a mass of CO2e per unit of currency, typically drawn from an environmentally-extended input-output database. Its great virtue is that you almost always have spend data, because finance has it, so it is fast and complete. Its weakness is that it is coarse: it assumes every euro spent in a category emits the same, so two suppliers, one efficient and one dirty, look identical if you paid them the same amount, and price inflation can masquerade as emissions growth. The activity-based method uses physical units as the activity data: tonnes, litres, kilowatt-hours, kilometres, multiplied by a factor expressed per physical unit. It is more accurate because it reflects the actual physical quantity, not the money, but it is data-hungry, because someone has to know and source the physical quantity, which finance does not automatically hold. The supplier-specific method goes further still: it uses the supplier's own reported, measured emissions for the goods you bought, which is the most accurate because it reflects that specific supplier's actual performance, but it is the most data-hungry, because it depends on the supplier measuring and disclosing, which most do not yet do.

The Ladder of Accuracy and Effort

Read top to bottom, the three methods are a ladder where accuracy and data effort rise together: spend-based at the bottom (least accurate, least effort), activity-based in the middle, supplier-specific at the top (most accurate, most effort). There is no free lunch. Every step up the ladder buys accuracy with data you have to go and get. This is exactly why the methods exist as a set rather than a single rule: a mature inventory uses different rungs for different categories, climbing to supplier-specific data where it matters and the data exists, and resting on spend-based estimates where it does not yet, while being honest about which rung each number sits on.

The reason the lowest rung is so coarse repays a closer look, because it is the rung AI reaches for. A spend-based factor is built by taking the total emissions of an entire economic sector and dividing by the total economic output of that sector, giving an emissions intensity per unit of currency. That single number is then applied to your spend as if your particular purchases emit at the sector average. The hidden assumptions are heavy. It assumes your suppliers are average for the sector, which the efficient ones are not. It assumes the prices you paid reflect physical quantities consistently, which they do not when one supplier charges a premium and another runs a discount. And it assumes the sector definition matches what you actually bought, which gets shakier the more specialised your purchase. None of this makes the method illegitimate. It makes it a blunt instrument, excellent for a first sweep and poor for a precise answer, and understanding why it is blunt is what lets you judge when its bluntness is acceptable.

The middle and top rungs earn their accuracy by replacing those assumptions with facts. Activity-based calculation throws out the price assumption entirely: it does not matter what you paid for the steel, only how many tonnes you bought, so a premium supplier and a discount supplier with the same tonnage now get the same physical emissions, which is correct. Supplier-specific calculation goes further and throws out the sector-average assumption too: it uses the actual emissions that actual supplier actually caused, so an efficient supplier finally shows up as efficient in your numbers. This is why the top of the ladder is where decarbonisation becomes visible. Only when you are using physical quantities and supplier actuals can your inventory reflect a supplier genuinely cleaning up its operations, which is the entire point of measuring in the first place.

The method you choose is not a private calculation detail. It is part of the disclosure. An assurer reads the method as carefully as the number, because the method tells them how much the number can be trusted.

Method Choice Is Itself a Disclosure

The single most important idea in this lesson is that the method is not a hidden step on the way to a number; it is a thing you disclose and the assurer tests. When you report a Scope 3 category, you are expected to state how you calculated it, and that statement is part of what is being assured. A spend-based number and a supplier-specific number for the same category are both legitimate, but they make very different claims about reliability, and presenting one as if it were the other misstates the quality of your disclosure exactly the way mislabeling primary as secondary data does.

This matters because method choice carries the same data-quality weight as the primary-versus-secondary tier. A spend-based estimate is, by its nature, secondary data: it is derived from money and an average factor, not from measuring the specific activity. An activity-based or supplier-specific figure can rise toward primary data when it uses the actual physical quantity and the supplier's actual emissions. So when you choose a method, you are simultaneously setting the data-quality tier of that number, and both must be visible in the file. The assurer's question is never just "what is the number"; it is "by what method, and is that method defensible for this category." A number whose method is undisclosed or indefensible for its category is a finding, regardless of the value.

When AI's Spend-Based Shortcut Is Acceptable

Because AI defaults to spend-based, you need a clear sense of when that default is genuinely fine and when it is not. The spend-based shortcut is defensible, and often the right professional choice, in well-defined situations.

It is acceptable for immaterial categories, where the emissions are small relative to the footprint and the effort of a more accurate method would not change any decision or any number that matters. Spending weeks getting tonnage data for a category that is 0.3% of your footprint is poor use of effort; a labelled spend-based estimate is the proportionate, professional choice. It is acceptable as a first-pass screening across all categories, to find out which ones are material before you invest in better data, which is one of the smartest uses of the method: cast a fast spend-based net, see where the big numbers are, then climb the ladder only where it counts. It is acceptable as a transparent placeholder for a category where better data genuinely does not exist yet, as long as it is labelled as spend-based with its method and uncertainty, and ideally flagged as a target to improve. In all three cases, the spend-based number is honest about what it is, proportionate to its materiality, and a deliberate choice rather than the path of least resistance. That is the difference between using the shortcut and being used by it.

When It Is Not Acceptable

The shortcut becomes indefensible when accuracy actually matters and better data is reasonably available. It is not acceptable for a large, material category that drives the footprint, where a coarse spend-based number could be materially wrong in either direction and an assurer will expect you to have climbed to activity-based or supplier-specific data, or to explain convincingly why you could not. A spend-based estimate for the single biggest line in your Scope 3 inventory, when the physical data was obtainable, is the kind of choice that draws a finding.

It is not acceptable when you are tracking reduction over time, because spend-based numbers move with price and procurement volume, not with real-world decarbonisation. If a supplier genuinely cut its emissions but you paid the same, a spend-based method shows no improvement; conversely, inflation can make your footprint appear to grow while nothing physical changed. A target tracked on a spend-based baseline can be quietly meaningless. And it is never acceptable to use the shortcut silently, presenting a spend-based number with no method label so it reads like a precise, physically grounded figure. That is not a method choice; it is a concealment, and concealment of method is exactly what an assurance engagement is designed to surface. The bright line is the same throughout the inventory: the method may be coarse, but it must be disclosed and defensible for the category it is used on.

There is a subtler failure worth naming, because it catches careful people: the method mismatch across years. Suppose you reported a category spend-based last year and switch it to activity-based this year because you finally obtained the tonnage. That is a genuine improvement, but it also means this year's number and last year's number were built by different methods, so any change between them mixes a real-world change with a method change. If you present the drop as if it were pure decarbonisation, you have misled the reader, even though both individual numbers are defensible. The professional handling is to disclose the method change explicitly and, where possible, restate the prior year on the new method so the comparison is like-for-like. An assurer reviewing a trend will look precisely for an unexplained method switch hiding inside an apparent improvement, because it is one of the easier ways for a footprint to flatter itself without anyone technically lying.

Notice that every one of these unacceptable cases shares a root: the spend-based number was allowed to make a claim it could not support, whether a claim of precision on a material line, a claim of improvement on a trend, or a claim of being measured when it was estimated. The acceptable cases all share the opposite property: the spend-based number only ever claims to be what it is, a fast, coarse, labelled estimate used where coarseness does not matter. That symmetry is the whole rule. You are not banned from the bottom rung. You are banned from letting the bottom rung pretend to be a higher one.

A Worked Example: Two Categories, the Right Method for Each

A manufacturer is building Scope 3 Category 1 (purchased goods and services). Two lines dominate the decision: steel, which is large and material, and office supplies, which is tiny. The accountant uses AI on both.

The naive uniform approach (what almost shipped): The accountant asks the AI to "calculate purchased-goods emissions from our spend data." The model applies the spend-based method to everything, because spend is what it has: steel EUR 4.2M times a per-euro factor, office supplies EUR 90k times a per-euro factor, and so on down the list, producing one tidy total with no method labels. The office-supplies line is fine this way. The steel line is the problem: steel is the largest, most material category in the inventory, the physical tonnage was readily available from procurement, and a spend-based number for it is both coarse and, worse, useless for tracking the decarbonisation the company has actually been pushing its steel supplier to achieve. When the assurer asks "why did you use spend-based for your single biggest category when tonnage was available," there is no good answer.

The method-matched approach (the defensible build): The accountant runs the spend-based method first as a screen across all categories, sees that steel dominates, and climbs the ladder where it counts. For steel, she pulls the physical quantity, 6,400 tonnes, and applies an activity-based factor per tonne with provenance, and because the supplier now reports verified actuals, she uses the supplier-specific figure where available, tagging it primary. For office supplies, immaterial at EUR 90k, she keeps the spend-based estimate, labelled secondary, spend-based, with its method noted, because climbing the ladder there would be effort wasted. The file now shows steel calculated supplier-specific or activity-based and tagged primary, office supplies calculated spend-based and tagged secondary, each method matched to the category's materiality and disclosed on its face. When the assurer reviews method choice, every line answers for itself. The same AI did the arithmetic in both versions. The difference is that the second one chose the method deliberately, climbed the ladder where accuracy mattered, and disclosed every choice.

That is the discipline in one comparison. AI will reach for the spend-based shortcut by default because it is the method the easy data supports. Your job is to decide, per category, whether that default is proportionate and defensible, to climb to activity-based or supplier-specific data where materiality demands it, and to make the method visible on every line so the assurer can test the choice rather than discover it.

One closing habit ties the whole skill together: run the spend-based and the better-data figure side by side on a material category before you commit. If they land close, you have evidence that the coarse method happened to work here and a documented basis for accepting it where data was thin. If they diverge sharply, you have just caught, at your own desk, the exact discrepancy an assurer would otherwise have caught at the engagement, and you can climb to the better number before it matters. Either way you end up with a defensible position and a documented reason for the method you used. The point is not that spend-based is bad and physical data is good. The point is that on the lines that move the footprint, you should know how far apart the two methods sit, and you should choose with that knowledge rather than letting an AI default choose for you in the dark.

Key Takeaways

  • Three Scope 3 methods sit on one accuracy-and-effort ladder: spend-based (currency times a per-currency factor, fast and coarse), activity-based (physical units times a per-unit factor, accurate and data-hungry), and supplier-specific (the supplier's own measured emissions, most accurate and most data-hungry).
  • Every step up the ladder buys accuracy with data you have to go and get; a mature inventory uses different rungs for different categories and is honest about which rung each number sits on.
  • Method choice is itself a disclosure: you state how each category was calculated and the assurer tests it, so an undisclosed or indefensible method is a finding regardless of the value.
  • Method choice sets the data-quality tier: a spend-based number is secondary data by nature, while activity-based and supplier-specific figures can rise toward primary, and both the method and the tier must be visible.
  • AI defaults to spend-based because spend is the data it can always get; that default is acceptable for immaterial categories, for first-pass screening, and as a labelled placeholder where better data does not yet exist.
  • The shortcut is not acceptable for large material categories where better data is reasonably available, for tracking reduction over time (spend moves with price, not decarbonisation), or ever used silently with no method label.
  • The smartest use of the spend-based method is a fast screen across all categories to find the material ones, then climbing the ladder only where it counts.
  • When an assurer asks why you used a given method for a category, the answer must show that the method is proportionate to materiality, defensible for that category, and disclosed on the line; "the AI calculated it from spend" is not such an answer.