Chain-of-Thought for Emission Calculations
A carbon accountant pastes one number into the inventory: 412.8 tonnes of CO2e for the diesel a logistics fleet burned last year. The assurer reads it, then asks the only question that matters. "Show me how you got there." If the answer is a single figure with no working behind it, the number is not finished. It is exposed.
The Number Is Not the Deliverable
Most people think the output of an emission calculation is the result: a tidy figure that slots into a cell in the GHG inventory. Inside an assurance engagement, that belief is the source of almost every painful finding. The deliverable is not the number. The deliverable is the number plus the working that produces it, where every step is visible, checkable, and tied back to a source an external assurer can independently confirm. The figure is the last line of a proof. A proof with only its last line is not a proof at all.
This is why a raw AI output is so dangerous in disclosure work. Ask a general-purpose model "what are the emissions from 153,000 litres of diesel?" and it will hand you a confident single number. It looks done. It is not done, because you cannot see the activity data it assumed, the emission factor it pulled, the density and energy conversions it ran, or whether any of those came from a real, authoritative source or from the model's fluent imagination. A single number from an AI is the worst of both worlds: it carries the authority of precision and none of the evidence.
Chain-of-thought is the technique that fixes this. In data-science circles it means prompting a model to reason step by step before answering, which improves accuracy on multi-step problems. For a disclosure professional, the value is different and larger. When you force the model to lay out activity data, then factor, then each conversion, then the arithmetic, then the result, with a citation on every step, you are not just improving accuracy. You are generating the audit trail as a by-product of the calculation. The reasoning trail is the audit trail. That reframing is the whole lesson.
The Anatomy of an Emission Calculation
Before you can make a model expose its working, you have to know what the working contains. Almost every emission calculation, however simple or complex, is the same shape: activity data times an emission factor, with conversions in between to make the units line up. Strip away the jargon and it is multiplication with bookkeeping.
The activity data is the physical quantity of the thing that caused the emissions: litres of diesel, kilowatt-hours of electricity, tonne-kilometres of freight, kilograms of a purchased material. It must trace to a record, a fuel invoice, a utility bill, a supplier statement, a meter reading. If you cannot point to where the activity data came from, the calculation has no foundation, no matter how elegant the rest of it is.
The emission factor is the published coefficient that converts a unit of activity into a mass of greenhouse gas, for example kilograms of CO2e per litre of diesel. A factor is only assurable if it traces to a named, dated, authoritative source: a government dataset, the GHG Protocol, a recognised database release, with its version and the exact row identified. A factor without provenance is the single most common AI failure mode in this work, the plausible-but-invented number, and chain-of-thought is built to surface it.
The conversions are the quiet middle of the calculation where errors hide. Diesel is often invoiced in litres but factors may be expressed per kilogram or per unit of energy, so you convert by density and net calorific value. Electricity may be metered in kilowatt-hours but a factor may be per megawatt-hour. A misplaced factor of a thousand, or a density applied upside down, can move a result by orders of magnitude and slip past anyone reading only the final figure. Exposing the conversions is exactly where chain-of-thought earns its keep.
The result is the last line of a proof. If the only thing you publish is the last line, you have not published a calculation, you have published a claim.
Forcing the Chain Out of the Model
A model will not expose its working unless you instruct it to, and you must instruct it precisely, because a vague "show your work" invites the model to invent a plausible-looking trail rather than a checkable one. The discipline is to define the structure of the chain and demand a source on every step, and to forbid the model from filling gaps silently.
Here is the kind of instruction that turns a black-box answer into an audit trail. Notice that it does not ask the model to be clever. It asks the model to be transparent and to refuse rather than guess.
Calculate the CO2e for the diesel below. Show every step as a numbered line: (1) activity data with its source, (2) the emission factor with its database name, release year, and row reference, (3) every unit conversion with the conversion constant and its source, (4) the arithmetic, (5) the result with units. If any input is missing or you are not certain of a factor or constant, stop and say so. Do not estimate a factor. Do not fill a gap with an industry average without labelling it as an unverified assumption.
Three things in that prompt do the heavy lifting. First, it fixes the shape of the chain, so you get the same five-part structure every time and can scan it fast. Second, it demands provenance per step, so a hallucinated factor has nowhere to hide: the model must name the database, year, and row, and a fabricated citation is far easier to catch than a fabricated number. Third, it gives the model an exit, an explicit instruction to stop and flag rather than invent, which converts the dangerous "confident guess" into a useful "I cannot complete step two."
Why the Trail Is the Control, Not the Convenience
It is tempting to treat the exposed chain as a nicety, something you read once to feel reassured and then discard. Resist that. The chain is a control in the audit sense: it is the artifact that lets a second person, the reviewer, the assurer, the future version of you reopening the file a year later, reconstruct the number without you in the room. The test of a good chain is reconstructability. Hand the activity data, the factor citation, and the conversions to a competent colleague who has never seen your work, and they should arrive at the same result. If they cannot, the chain is incomplete, and an incomplete chain is an assurance finding waiting to be written.
A Worked Example: Diesel, Step by Step
Watch a single calculation move from an opaque AI answer to a traceable, assurable chain. A logistics fleet burned diesel last year. The procurement file shows fuel-card records totalling 153,000 litres for the reporting boundary. We need the Scope 1 CO2e.
The opaque version. Ask a model casually and it returns something like: "The emissions from 153,000 litres of diesel are approximately 412.8 tonnes CO2e." It is fluent, it is roughly the right size, and it is unassurable. You cannot see the factor, you cannot see whether it used a per-litre or a per-kilogram factor, you cannot see the source, and you cannot tell whether 412.8 is the honest product of real inputs or a number the model regressed toward because it looked typical. If the assurer asks for the basis, you have nothing.
The exposed chain. Now the same calculation with the working forced out, each line checkable and cited:
- Activity data. 153,000 litres of diesel. Source: fuel-card consumption report, FY reporting period, reconciled to the procurement ledger. Boundary: owned and leased fleet within the operational control boundary.
- Emission factor. 2.66 kg CO2e per litre of diesel (combustion, well-to-tank excluded for this Scope 1 line). Source: national government GHG conversion factors, 2025 release, "Fuels, diesel (average biofuel blend), litres" row. Factor type and release noted so a reviewer can pull the exact row.
- Conversion. None required: the factor is already expressed per litre and the activity data is in litres, so the units line up directly. This line matters precisely because it is empty. Stating "no conversion needed, units already aligned" is itself a checkable claim, and it stops a reviewer wondering whether a density step was forgotten.
- Arithmetic. 153,000 litres times 2.66 kg CO2e per litre = 406,980 kg CO2e.
- Result. 406,980 kg CO2e = 406.98 tonnes CO2e, rounded to 407.0 tonnes for the inventory line, with the unrounded figure retained in the working.
Look at what just happened. The exposed chain produced 407.0 tonnes, not the 412.8 the opaque answer offered. The difference is not rounding. It is the difference between a number you can defend and a number that drifted. With the chain visible, a reviewer can immediately spot that the opaque answer must have used a slightly different factor, perhaps one that included an upstream component that does not belong on this Scope 1 line, and ask the right question. With only the opaque figure, that error sails into the inventory unchallenged and surfaces eighteen months later as a restatement.
Reading the Chain as a Reviewer
When the chain is in front of you, your job is to attack each line in turn. Is the activity data tied to a real record, and is it the right record for the boundary? Is the factor real, current, and the correct row for the fuel and the scope, and did the model cite a release that actually exists? Is each conversion necessary, in the right direction, and using the right constant? Does the arithmetic actually multiply out? Is the result in the units the inventory expects, and is the rounding sensible? Five lines, five attacks. A calculation that survives all five is one you can sign. The point of chain-of-thought is that it gives you five things to attack instead of one thing to trust.
A Multi-Step Chain Where Conversions Bite
The diesel example had no conversion, which is exactly why it was a gentle warm-up. The technique earns its reputation on calculations where the units do not line up and a hidden conversion error can multiply the answer. Take purchased natural gas, invoiced by the utility in cubic metres but with a factor published per unit of energy. Here is the exposed chain.
- Activity data. 48,500 cubic metres of natural gas. Source: utility invoices for the reporting period, summed and reconciled to the gas-meter readings on the site facilities file.
- Conversion. Convert volume to energy using the gross calorific value stated on the utility invoice: 48,500 cubic metres times 10.9 kWh per cubic metre = 528,650 kWh. The calorific value is taken from the supplier's own invoice, not assumed, and that source is named so a reviewer can confirm it.
- Emission factor. 0.183 kg CO2e per kWh (gross CV basis, natural gas). Source: national GHG conversion factors, 2025 release, "Gaseous fuels, natural gas, kWh (gross CV)" row. The basis matters: a gross-CV factor must be paired with gross-CV energy, and the chain states both so the mismatch cannot hide.
- Arithmetic. 528,650 kWh times 0.183 kg CO2e per kWh = 96,743 kg CO2e.
- Result. 96,743 kg CO2e = 96.74 tonnes CO2e for the Scope 1 stationary-combustion line.
The danger line is the second one. If the model had grabbed a net-calorific-value factor and applied it to gross-CV energy, or inverted the kWh-per-cubic-metre constant, the result could swing by ten percent or by a factor of ten, and a reader of the final figure alone would have no way to know. Because the conversion is exposed with its constant and its source, a reviewer checks one multiplication and one matching of bases, and the error has nowhere to live. This is the entire argument for the technique in a single line of a single calculation.
Where the Technique Pays, and Where It Can Mislead
Chain-of-thought pays the most exactly where the stakes are highest: multi-step calculations with conversions, the kind that dominate a real GHG inventory. A single-line electricity calculation barely needs it. A freight emission built from tonnes, distances, mode-specific factors, and load assumptions needs it badly, because that is where a hidden unit error or a silently swapped factor can move the answer by a factor of ten and no one notices until the assurer does.
But understand the technique's limit, because misunderstanding it is its own failure mode. A model writing out a plausible chain is not the same as a correct chain. The model can produce a beautifully structured five-line trail in which line two cites a database release that does not exist and line three applies a conversion backwards. Chain-of-thought does not verify the steps. It exposes them so that you can verify them. The model generates the trail; the human checks it against the actual source. A cited factor is a lead to check, never a fact to trust. The whole value of the technique collapses if you read the neat structure and assume the neatness implies correctness. Exposed working that no one checks is theatre, not assurance.
This is also why the per-step citation requirement is non-negotiable. A chain that shows arithmetic but not sources still leaves you unable to confirm the factor was real. The citation is the hook that lets you leave the model entirely, open the actual government dataset or the licensed database, and confirm the row. The arithmetic you can recompute in your head; the provenance you can only confirm by leaving the model and going to the source. Both have to be in the chain.
Filing the Chain Into the Assurance File
An exposed chain that lives only in a chat window has solved nothing. The assurer does not read your chat history. The chain has value only when it is captured into the working papers next to the inventory line it supports, so that the file itself answers "show me how you got there" before anyone has to ask. In practice that means copying the five-part chain, with its citations intact, into the basis-of-preparation or the calculation workbook, tagging which step was AI-drafted and which step a human verified, and recording the reviewer who confirmed each citation against the actual source. The model produced the trail fast; the human's confirmation and signature are what make it evidence. A chain in the file with a named verifier is an audit trail. The same chain in a deleted chat is a rumour.
There is a quiet efficiency win here that is easy to miss. Teams often treat the calculation and the documentation as two jobs: do the math, then later, painfully, write up how you did it for the assurer. Chain-of-thought collapses them into one. If the working is exposed and cited as the calculation is performed, the documentation is already written. You are not reconstructing the trail after the fact from memory, which is where errors and gaps creep in. You captured it at the moment of calculation, when the inputs were in front of you. That is the rare case where the assurance-friendly way is also the faster way.
Key Takeaways
- The deliverable of an emission calculation is not the result, it is the result plus the checkable working that produces it. A single number with no trail is a claim, not a calculation.
- Every emission calculation is the same shape: activity data times an emission factor, with unit conversions in between. Knowing the shape tells you what the chain must contain.
- Chain-of-thought forces the model to lay out activity data, factor, conversions, arithmetic, and result as separate, checkable, cited lines. The reasoning trail becomes the audit trail.
- The prompt must fix the five-part shape, demand provenance on every step (database name, release year, row), and explicitly instruct the model to stop and flag rather than invent a factor or fill a gap silently.
- Conversions are where errors hide. Exposing them, including stating when no conversion is needed, catches the misplaced thousand and the upside-down density before they reach the inventory.
- The test of a good chain is reconstructability: a competent colleague with the same inputs and citations should reach the same number without you in the room.
- Chain-of-thought exposes the steps, it does not verify them. A neatly structured trail can still cite a factor that does not exist. The human checks each line against the actual source; a cited factor is a lead, not a fact.
- The technique pays most on multi-step calculations with conversions, exactly where a hidden unit error can move a result by orders of magnitude and slip past anyone reading only the final figure.
Skill.re