Building a Provenance-Tagged Scope 3 Workflow
A carbon accountant opens a blank workbook in January and knows the shape of the year ahead. By autumn this file has to become a Scope 3 inventory that is roughly three-quarters of the company's entire footprint, built across fifteen defined categories, assembled from suppliers who mostly will not answer, and read line by line by an external assurer who can ask of any number: show me where this came from. She has powerful AI tools that can draft questionnaires, parse responses, suggest emission factors, and compute totals in seconds. She also knows that every one of those tools will, if she lets it, produce a clean number with no source attached. The whole job of the year is to build the inventory at AI speed while making sure that every single datapoint in it carries its origin on its back. This is the provenance-tagged Scope 3 workflow, and this lesson builds it end to end.
What the Workflow Has to Produce
Before any step, fix the target in your mind, because it governs every design choice. The deliverable is not a number. It is a number plus its complete lineage, repeated for every datapoint, assembled so that a stranger could reconstruct the whole inventory from raw evidence without you in the room. The GHG Protocol, the global greenhouse-gas accounting standard that nearly every reporting framework points back to, divides a company's emissions into three scopes: Scope 1 is direct emissions from owned sources, Scope 2 is purchased energy, and Scope 3 is everything else in the value chain, upstream and downstream, split into fifteen categories. Scope 3 is the hard one because the company does not own the activity that creates the emissions; a supplier does, or a customer does, and the data lives outside the company's walls.
Every figure in a Scope 3 inventory is, at bottom, an activity data times emission factor calculation. Activity data is the measure of the thing that happened, such as tonnes of steel purchased, litres of fuel burned in outbound logistics, or kilometres flown by employees. The emission factor is the coefficient that converts that activity into emissions, such as kilograms of CO2 equivalent per tonne of steel. Multiply the two and you get the emissions for that line. The inventory is thousands of these multiplications. Provenance means that for every line, you can show where the activity data came from, where the factor came from, and which method tied them together. Build the workflow so that provenance is captured at the moment each number enters, not bolted on at the end, because provenance reconstructed after the fact is exactly the thing an assurer distrusts.
The Fifteen Categories, So You Map the Whole Chain
You cannot tag what you have not mapped, so the workflow starts by laying the company's value chain against all fifteen Scope 3 categories. Upstream are: Category 1 purchased goods and services; Category 2 capital goods; Category 3 fuel- and energy-related activities not in Scope 1 or 2; Category 4 upstream transportation and distribution; Category 5 waste generated in operations; Category 6 business travel; Category 7 employee commuting; and Category 8 upstream leased assets. Downstream are: Category 9 downstream transportation and distribution; Category 10 processing of sold products; Category 11 use of sold products; Category 12 end-of-life treatment of sold products; Category 13 downstream leased assets; Category 14 franchises; and Category 15 investments. The point of naming all fifteen is completeness: you decide which apply to your business and which do not, and the deciding is a documented act, not a silent omission. An undocumented exclusion, a category you quietly dropped because it was hard, is one of the most common assurance findings there is.
Step One: Map the Value Chain With a Documented Boundary
The first real work is drawing the boundary, and the boundary is two decisions, not one. The organizational boundary defines which entities count as the company: do you consolidate emissions by equity share, by financial control, or by operational control? The operational boundary defines which activities and scopes you include within those entities. For Scope 3, the operational boundary decision becomes the category-by-category question: which of the fifteen are relevant and material, and which are excluded, and on what stated basis? This is where AI earns its first keep. You can give a model your industry, your business description, and your procurement categories, and ask it to propose which Scope 3 categories are likely material and to draft the rationale for each. That draft is a starting point, and a genuinely useful one, because it surfaces categories a busy team might overlook, such as Category 11 use of sold products for a company that sells fuel-burning equipment.
But the boundary decision stays human and it stays documented. The model proposes; the accountant decides and records the decision with its reasoning, and crucially records the exclusions with their reasoning too. If you exclude Category 14 franchises because the company has no franchises, you write that down. If you exclude Category 8 upstream leased assets because it is captured elsewhere, you write that down and say where. The output of step one is not a number; it is a boundary memo: organizational consolidation approach, the fifteen categories each marked relevant or excluded, and a one-line documented basis for every call. This memo is the spine the assurer follows. Every later step hangs off it.
A category you excluded silently is not a smaller inventory. It is an assurance finding waiting to be written. The boundary is defensible only when every exclusion is on the record with its reason.
Step Two: Collect Supplier Data, Tagged Primary or Secondary
With the boundary drawn, the workflow turns to the material categories and goes after data. The single most important tag in the whole inventory gets applied here: every datapoint is either primary data, meaning it comes from the actual supplier or activity, a supplier-reported figure or a metered reading, or secondary data, meaning it is an estimate built from averages, industry factors, or spend. Primary and secondary data cannot look identical in the file. An assurer weighs a supplier-reported tonne of CO2e differently from a tonne estimated off spend, and the file has to let them see which is which at a glance. This is the distinction the whole engagement turns on.
AI compresses the collection bottleneck without dissolving this tag. A model can draft a supplier questionnaire tailored to a category, triage the incoming responses, and parse a messy supplier spreadsheet or PDF into structured fields. The discipline is that the parser preserves provenance on every field it extracts: which supplier, which document, which line of that document, what was reported, in what unit, on what date. A parsed datapoint arrives tagged primary, with its source document and location attached, or it does not enter the inventory. When a supplier does not respond, the workflow does not let the model quietly invent a number; it flags the gap, and that gap becomes a candidate for explicit, labelled estimation later, tagged secondary. The rule is simple and absolute: the model may help you read and structure what a supplier said, but it may never manufacture what a supplier did not say and present it as primary.
What a Provenance Record Actually Contains
Make the provenance record concrete, because vagueness here is where audits fail. For each activity datapoint, capture: the value and its unit; the tier, primary or secondary; the source, named specifically, such as a supplier name and the document or system the figure came from; the date of the data and the date it was collected; the method by which it was derived if it is secondary; and the human who accepted it into the inventory. For each emission factor, capture: the factor value and unit; the database it came from, named and dated, such as a specific version of a recognised factor library; and the reason this factor was chosen for this line. A line in the inventory is complete only when both its activity datapoint and its factor carry full records. A bare number with no record is not a small problem to fix later; it is a hole in the audit trail, and the workflow is designed so that such a number cannot enter in the first place.
Step Three: Select Emission Factors With Provenance
Now each line needs its factor, and this is the step where AI is most seductive and most dangerous at once. Ask a general model for the emission factor for, say, purchased aluminium, and it will give you a confident, specific, plausible number. It may be right. It may also be a hallucinated factor: a fabricated coefficient with the shape of a real one and no source behind it, which then silently corrupts every line it touches. The workflow refuses to accept any factor on the model's word. Instead it forces every suggested factor back to a named, dated, authoritative database, and records that source on the line. The model is allowed to help you find and retrieve a factor from your approved factor library; it is never allowed to be the source of the factor itself.
The practical pattern is to ground the model on your own factor database rather than the open web, so that when it proposes a factor it proposes one that exists in your evidence base, with its identifier and version, and you can click straight through to the source. If the right factor genuinely is not in your library, the gap is surfaced, not papered over with an invented number, and a human goes and finds an authoritative factor and adds it with provenance. The test for any factor on any line is the one the assurer will apply: can I follow this number to a named, dated source and see that it says what you say it says? If yes, the line stands. If no, the line is not finished, however confident the number looks.
Step Four: Compute, and Log Every Step
With activity data tagged and factors sourced, the computation is the easy part: activity data times factor, line by line, summed by category and into a Scope 3 total. AI can do this instantly and can show its working, which is exactly what you want it to do. The requirement is that the calculation is transparent and reconstructable: each line shows its activity value, its factor, the product, and the records behind both inputs, so that the arithmetic can be re-run by anyone and lands on the same total. A computed total with no visible working is just another unsourced number; a computed total where every line traces to its inputs and every input to its source is an assurable inventory.
Logging is not a final tidy-up; it is the running record of the whole build. Every step leaves a trace: the boundary memo from step one, the supplier responses and their tags from step two, the factor sources from step three, the computations from step four, and crucially every human decision and override along the way, with who made it and why. When the analyst overrides an AI suggestion, accepts an estimate, or excludes a line, that decision is logged. The cumulative log is the basis-of-preparation: the document that lets the assurer reconstruct the inventory from raw evidence. The discipline that makes this possible is capturing the trace as you go, because a log assembled from memory at year-end is both incomplete and, to an assurer's eye, suspect.
A Worked Example: One Category, End to End, With Numbers
Walk Category 4, upstream transportation and distribution, for a mid-size manufacturer, to see the workflow produce a traceable line rather than a bare figure.
Step one, boundary. The boundary memo marks Category 4 as material and relevant: the company ships a lot of inbound freight, so the category is clearly in. Documented basis: significant inbound logistics spend across road and sea freight. In scope.
Step two, data, tagged. The team surveys its top freight carriers. Two of three respond with actual tonne-kilometre data. Carrier A reports 4,200,000 tonne-km of road freight; the AI parses this from Carrier A's emissions report, tags it primary, and records the source as Carrier A 2025 freight report, page 3, collected March 2026. Carrier B reports 1,800,000 tonne-km of sea freight, tagged primary, source recorded the same way. Carrier C, representing roughly EUR 600,000 of freight spend, never responds; the workflow flags this as a gap, tagged for secondary estimation, not silently filled.
Step three, factors with provenance. For road freight the analyst selects a factor of 0.10 kg CO2e per tonne-km from a named, dated factor library, recording the database, version, and identifier on the line. For sea freight she selects 0.015 kg CO2e per tonne-km from the same library, recorded the same way. The AI proposed both; the analyst confirmed each against the source and accepted them with provenance. Neither factor entered on the model's word alone.
Step four, compute and log. Road: 4,200,000 tonne-km times 0.10 kg = 420,000 kg, or 420 tonnes CO2e, tagged primary. Sea: 1,800,000 tonne-km times 0.015 kg = 27,000 kg, or 27 tonnes CO2e, tagged primary. Carrier C's gap is closed with an explicit spend-based estimate: EUR 600,000 times a named spend-based factor, producing an estimated portion that is tagged secondary, labelled as an estimate, with its method and uncertainty noted. Category 4 now totals roughly 447 tonnes of primary plus a clearly separated estimated portion, every line carrying its tag, its source, and its factor provenance.
Now compare what almost shipped. The fast, dangerous version asks the AI to "calculate our upstream freight emissions" and accepts a single clean figure of, say, 480 tonnes CO2e, with the carriers, the factors, and the gap all dissolved into one number with no tags and no sources. That figure looks finished and is indefensible: it cannot say what is primary, what is estimated, where its factors came from, or that one carrier never responded. The provenance-tagged version may land near the same total, but it can answer every one of those questions in a sentence and a click. The total is similar. The defensibility is not even close.
Why the Discipline Is the Speed, Not the Tax on It
It is tempting to read all this tagging as overhead, a slow careful process bolted onto fast AI. That reading is backwards. The provenance discipline is what lets you use AI aggressively, because it makes AI's output safe to trust. Without it, every AI-produced number is a liability you have to re-verify by hand under deadline, which is slower and more frightening than doing it carefully the first time. With provenance captured at the moment of entry, the AI drafts, parses, suggests, and computes at full speed, and the human's job narrows to the decisions that genuinely need judgment: the boundary, the factor confirmations, the estimate labels, the overrides. The same workflow that satisfies the assurer is the one that lets you move fast, because a number you can defend is a number you never have to redo. That is the goldmine in one sentence: faster and more defensible are, here, the same move.
Key Takeaways
- The deliverable of a Scope 3 workflow is not a number but a number plus its complete lineage, repeated for every datapoint, assembled so a stranger could reconstruct the inventory from raw evidence without you in the room.
- Scope 3 is roughly 75% of a footprint across fifteen GHG Protocol categories, and every line is an activity data times emission factor calculation; provenance means showing where the activity data, the factor, and the method each came from.
- Step one draws the boundary: an organizational consolidation approach plus an operational boundary that marks each of the fifteen categories relevant or excluded, with a documented basis for every call, because an undocumented exclusion is an assurance finding.
- Step two collects supplier data and tags every datapoint primary (supplier-reported, metered) or secondary (estimated), preserving the source document and location on each parsed field; the model may structure what a supplier said but never manufacture what a supplier did not say.
- Step three selects factors and forces every one back to a named, dated, authoritative database; a model may help retrieve a factor but is never the source of it, because a hallucinated factor silently corrupts every line it touches.
- Step four computes transparently, line by line, and logs every step and every human decision as it happens, building the basis-of-preparation that lets the assurer reconstruct the inventory from raw evidence.
- In the worked Category 4 example, a provenance-tagged build lands near the same total as a single AI-generated figure but can answer what is primary, what is estimated, where factors came from, and which carrier never responded, while the bare figure cannot.
- The provenance discipline is not a tax on AI speed; it is what makes AI output safe to trust, so faster and more defensible become the same move.
Skill.re