Labeling Primary vs. Secondary Data
The assurer opens the Scope 3 inventory and points to a single cell: 1,180 tonnes of CO2e for a key supplier's emissions. "Is this measured or estimated?" she asks. The carbon accountant pauses. He thinks the number came from the supplier's own questionnaire response, but he is not certain, because three rows down sits another 1,180 that he knows he built from a spend-based average, and the two cells are visually identical. Same font, same decimals, same column. There is nothing in the file that says one is a supplier-reported figure and the other is an estimate the company assembled. That ambiguity is not a cosmetic problem. To an assurer it is the problem, because the entire reliability of the inventory rests on knowing which numbers came from the real world and which the company manufactured, and a file that cannot tell them apart cannot be assured.
The Distinction the Assurer Lives By
Carbon accounting runs on a hierarchy of data quality that the assurance profession treats as foundational, and the first and most important cut in that hierarchy is between primary and secondary data. Primary data is a value that comes from direct measurement or direct report of the actual activity: a supplier telling you its actual emissions, a meter reading your actual electricity, a fuel invoice recording the actual litres you burned, a weighbridge recording the actual tonnes you shipped. It is grounded in something that physically happened and was observed. Secondary data is a value derived without measuring the specific activity: an industry-average factor applied to your spend, a proxy from a comparable site, a database average for a material class, an estimate built from a model rather than a measurement. It stands in for the real number because the real number was not available.
Neither tier is forbidden. Real inventories, especially Scope 3, are unavoidably a blend of both, because for many of the 15 Scope 3 categories defined by the GHG Protocol, the global standard most reporting frameworks point back to, primary data simply does not exist yet. The point is not to purge secondary data. The point is that a primary figure and a secondary figure carry very different weight, and the assurer must be able to see which is which at a glance. A supplier-reported, measured tonne is strong evidence. A spend-based estimate is a reasoned placeholder. They can sit in the same inventory, but they cannot sit in the same file looking identical, because then the reader cannot weigh the inventory's reliability at all.
Why the Tier Changes the Weight of a Number
The reason the assurer cares so much is that the tier of a datapoint tells her how much to trust the total built from it. An inventory that is 80% primary data is a fundamentally more reliable artifact than one that is 80% secondary, even if the two produce the same headline number, because the primary one is anchored to measured reality and the secondary one is anchored to assumptions. The mix of primary and secondary data is itself a disclosure about quality. Frameworks ask reporters to describe their data quality precisely because the tier mix is the honest signal of how solid the footprint is. If you cannot state your primary-versus-secondary split, you cannot describe your data quality, and if you cannot describe your data quality, the assurer cannot scope or sign the engagement. The label is not bureaucracy. It is the load-bearing information about how good your number actually is.
It helps to see the tier as a property of a datapoint rather than a property of a number. Two cells can hold the value 1,180 and be completely different objects. One is a fact about the world that a supplier measured and reported; pull its thread and you reach a meter, an invoice, a verified disclosure. The other is the output of a calculation the company performed in the absence of measurement; pull its thread and you reach an assumption, an average, a model. The numeric coincidence is irrelevant. What the assurer is testing is the thread, and the tier label is the signpost that tells her which kind of thread she is about to pull. Strip the signpost and she has to interrogate every cell from scratch, which is slow, expensive, and exactly the friction that turns a smooth engagement into a painful one.
This is also why the tier mix tends to improve over time in a well-run program, and why being able to show that improvement is valuable. A company that moves a category from a spend-based estimate this year to supplier-reported actuals next year has genuinely strengthened its inventory, and the rising primary share is the evidence of that progress. But you can only tell that story if you tagged the tiers in both years. A company that flattened its data has no way to show that it improved, because it threw away the very metadata that would prove it. Honest tiering is therefore not only a defensive measure against findings; it is the record that lets you demonstrate a maturing data program to a board, an investor, or an assurer.
A measured tonne and an estimated tonne are not the same evidence. If they look the same in your file, your file is lying about its own reliability, and the assurer's whole job is to catch that.
Where AI Quietly Erases the Distinction
AI is genuinely useful in building a GHG inventory: it extracts activity data from messy documents, parses supplier responses, looks up factors, and assembles estimates. But every one of those tasks touches the primary-versus-secondary boundary, and the default behaviour of a generative model is to flatten it. The model is built to produce a clean, uniform, finished-looking output. Asked to "complete the supplier emissions table," it will happily return a tidy column of numbers in which a supplier's measured disclosure, a spend-based estimate it computed, and an industry average it recalled all appear as the same kind of number, with no tier label on any of them. The output looks more finished precisely because the model erased the distinction that mattered.
This is the silent failure mode for data labelling. It is not that the model lies about a value. It is that the model strips the metadata, the where-did-this-come-from, that tells you the value's tier, and presents everything at the same apparent confidence. A measured figure and an estimate come out wearing the same uniform. The danger is doubled because the result looks better than a properly labelled file: cleaner, more complete, more authoritative. An untrained reviewer sees a finished table and relaxes. A trained one sees a table that has thrown away the single most important piece of quality information and asks the question the assurer will ask: which of these did anyone actually measure?
The flattening is not limited to building a table from scratch. It creeps in at every step where AI touches the data. When AI parses a stack of supplier responses, it can blend a supplier's hard figure with a value it inferred from context, returning both as plain numbers. When AI reconciles this year to last, it can carry forward a prior estimate as if it were a fresh measurement. When AI summarises a category for the disclosure narrative, it can describe estimated figures in the same confident prose it uses for metered ones. In each case the model is optimising for a clean, decisive output, and a tier label reads to it as noise to be smoothed away. The professional has to push back against that optimisation at every touchpoint, treating the tier as sacred metadata that survives every transformation, because the moment it is dropped once, it is gone, and reconstructing it later means re-tracing every datapoint to its origin.
How to Keep the Tier Attached to Every Datapoint
The discipline is to treat the tier as part of the datapoint itself, not as a separate note you might add later. Every value that enters the inventory carries, alongside its number and unit, a tier label and a basis. The number is never naked.
Capture the tier at the moment of capture, not at the end. When a value comes from a supplier's reported actuals, it is tagged primary at the instant you record it, with a pointer to the response it came from. When you build a spend-based or average-data estimate because primary data was missing, it is tagged secondary at the instant you compute it, with its method named. Retrofitting tiers onto a finished, flattened table is slow, error-prone, and exactly the situation in the opening scene, where the accountant can no longer remember which cell was which. The tier is cheapest and most reliable when it is captured at the source.
Force AI output to carry the tier. Never ask the model to "complete the table." Ask it to produce, for every line, the value, the unit, the tier (primary or secondary), and the basis (the supplier response for primary, or the named method and source for secondary), and to flag any line where it does not have a basis rather than filling it. A model told to label its tiers cannot launder an estimate into the primary column without you seeing it do so. A model told only to complete the table will flatten everything, every time.
Make the file show the tier visibly. In the inventory itself, a primary figure and a secondary figure must be distinguishable on sight: a tier column, a flag, a colour, anything that means an assurer scanning the file can see the mix without interrogating each cell. The cardinal sin is the opening scene, two identical-looking cells where one is measured and one is estimated. The remedy is that the tier is never invisible.
Be able to state the mix. Because the tier is captured per datapoint, you can roll it up: for any category, scope, or the whole inventory, you can say what share is primary and what share is secondary. That single statement is what lets you describe your data quality honestly and is among the first things an assurer will ask for. If you can produce it on demand, you are ready. If you cannot, the inventory is not yet assurable, no matter how complete the numbers look.
A Worked Example: One Supplier Table, Two Ways
A company is building Scope 3 Category 1 (purchased goods and services) for three key suppliers: a steel supplier, a packaging supplier, and an electronics supplier. The accountant uses AI to assemble the table.
Before (the flattened table, what almost shipped): The accountant asks the AI to "complete the supplier emissions for these three." The model returns a clean table: steel 6,200 tonnes, packaging 1,180 tonnes, electronics 940 tonnes, total 8,320 tonnes. Every cell looks the same. What the table hides: the steel figure is the supplier's own measured, verified disclosure (strong primary data); the packaging figure is a spend-based estimate the model computed from EUR 2.1M of spend times an input-output factor (secondary); and the electronics figure is an industry average the model recalled with no specific source at all (secondary, and weakly sourced). Three completely different qualities of evidence, presented as one uniform column. When the assurer asks "what share of this is primary," the accountant cannot say, because the file threw the information away. The table that looked most finished is the one that fails first.
After (the tiered, labelled build): The accountant instructs the AI to produce, per supplier, the value, unit, tier, and basis, and to flag any line lacking a basis. The result is a table with a tier column. Steel: 6,200 tonnes, primary, basis = supplier's verified disclosure dated this year, with a pointer to the response. Packaging: 1,180 tonnes, secondary, basis = spend-based estimate, EUR 2.1M times a named input-output factor with provenance, uncertainty noted. Electronics: 940 tonnes, secondary, basis = average-data method using a named average factor, flagged because no supplier-specific data exists yet. Now the same total, 8,320 tonnes, comes with its quality on its face: roughly 75% of it is primary (the steel), 25% is secondary and labelled, and the weakest line is flagged as a collection target for next year. When the assurer asks "what share is primary and what is estimated," the accountant answers in one sentence and points to the column. The number is identical. The defensibility is total.
That is the whole lesson in one comparison. The same AI, the same three suppliers, the same 8,320 tonnes, produces either a flattened column that hides its own reliability or a tiered table that declares it. The accountant chooses which, by refusing to let any value enter the inventory without a tier and a basis, and by never asking the model to merely complete the table.
Why Mislabeling Is an Assurance Finding
It is worth being precise about why this is not just good housekeeping but a live audit risk. When an assurer reviews a GHG inventory, a central part of the engagement is testing whether the reported data quality matches the actual data quality. If your file presents a spend-based estimate as if it were a measured, supplier-reported figure, you have misstated the quality of that number, and that is a finding in its own right, separate from whether the number happens to be accurate. A correct number with a wrong tier is still a finding, because the tier is part of what you are asserting when you publish.
The two failure directions both bite. Labelling secondary data as primary overstates your data quality and your reliability, which is the more dangerous direction because it flatters the footprint and invites a restatement when discovered. Labelling primary data as secondary understates it, wastes the strength of evidence you actually have, and can make your inventory look weaker than it is. Either way, the tier is an assertion, and an inaccurate assertion is a finding. The professional habit is to make the tier exactly true for every datapoint, so that when the assurer tests data quality, the file's claims about itself are already correct. That is the difference between an engagement that goes smoothly and one that reopens.
There is a practical reason the assurer treats a single mistagged material line so seriously, and it is worth internalising because it explains why the standard is exacting rather than approximate. An assurance engagement works by sampling: the assurer cannot re-test every number, so she tests a selection and reasons from it to the whole. When a sampled line turns out to be mistagged, the inference is not merely that one line is wrong; it is that the control which was supposed to keep tiers accurate is unreliable, which casts doubt on every line she did not sample. One mistag therefore does not cost you one line; it can cost you the assurer's confidence in the file's self-assessment, prompting expanded testing and a harder engagement. This is why accurate tiering is a system property, not a per-line nicety: the value of the labels comes from their being trustworthy in aggregate, and a single visible failure undermines that trust disproportionately.
Key Takeaways
- Primary data comes from direct measurement or direct report of the actual activity (a supplier's actuals, a meter, an invoice); secondary data is derived without measuring the specific activity (spend-based, average-data, proxy, model estimate).
- Neither tier is forbidden, and real Scope 3 inventories are unavoidably a blend; the rule is that a primary figure and a secondary figure must never look identical in the file.
- The tier changes the weight of a number: an 80%-primary inventory is more reliable than an 80%-secondary one even at the same headline total, and the primary-versus-secondary mix is itself a disclosure about data quality.
- AI's default failure is flattening: asked to complete a table, a model returns a uniform column in which measured figures, computed estimates, and recalled averages all look the same, stripping the tier metadata.
- The flattened output is dangerous precisely because it looks more finished and more authoritative; a trained reviewer asks the question it hides, which of these did anyone actually measure.
- Capture the tier at the moment of capture, force AI to output value, unit, tier, and basis for every line, make the tier visible in the file, and be able to state the primary-versus-secondary mix on demand.
- Mislabeling is an assurance finding in its own right: a correct number with a wrong tier is still a misstatement of data quality, and labelling secondary data as primary is the more dangerous direction because it flatters the footprint.
- When an assurer asks what share of a category is primary versus estimated, you should be able to answer in one sentence and point to a tier column; if you cannot, the inventory is not yet assurable however complete it looks.
Skill.re