Reconciling Extracted Data to the Prior Period
The Scope 1 figures are nearly ready. The analyst has extracted the gas consumption for every facility, paired each with a verified emission factor, and the inventory is reconciling itself toward a clean close. Then one row stops her cold: the Rotterdam site reads 3.1 times last year's gas. No new building, no harsh winter, no process change anyone mentioned. The number is sourced, the factor is real, the arithmetic is correct, and it is almost certainly wrong. Last year's figure just did its job: it flagged the error that every other check waved through. This lesson is about using the prior period as the cheapest error detector you own.
Last Year's Number Is a Free Control
Every other verification in this chapter checks a number against its source: quote-the-source confirms the citation, recompute confirms the arithmetic, range-check confirms the magnitude is possible. Prior-period reconciliation does something none of them do. It checks a number against what your own organisation actually did last year, and that comparison catches a whole class of errors the source-level checks structurally cannot see. A figure can be perfectly extracted, perfectly sourced, and perfectly multiplied, and still be wrong because the wrong document was extracted, a meter was double-counted, a unit was misconverted upstream, or a facility's data got swapped with another's. None of those leave a fingerprint on the individual number. All of them leave a fingerprint when you set this year next to last year.
Think of it as the difference between proofreading a sentence and fact-checking it. The source-level checks proofread: they confirm the number is spelled correctly, so to speak, that it is real, well-formed, and possible. Prior-period reconciliation fact-checks: it confirms the number is true to the world it describes. A sentence can be perfectly proofread and completely false, and a figure can be perfectly sourced and completely wrong, and in both cases only the check that looks outward catches it.
The reason this control is so valuable is that it is almost free and you already own it. You do not have to build a benchmark or buy a dataset. Your prior-year inventory, the one that was assembled, reviewed, and in most cases externally assured, is a known-good reference point sitting in a drawer. Comparing against it costs a subtraction and a moment of judgment, and it routinely catches the expensive errors, the order-of-magnitude swings and the silent double-counts, that would otherwise sail into the inventory and surface only when an assurer asks why your emissions tripled.
This is also why assurers themselves lean so heavily on year-over-year analysis. When an external assurance team plans an engagement, one of the first things they do is compare the current numbers to the prior period and ask about every material movement. If you have not run that comparison before they arrive, you are handing them surprises in your own data. If you have, you walk in already able to explain every swing. The prior-period check is not just an internal hygiene step; it is a rehearsal of the exact analysis the assurer will perform.
It is worth being precise about why source-level checks are structurally blind to this whole class of error, because the blindness is not a weakness in those checks, it is a property of what they examine. Quote-the-source, recompute, and range-check all evaluate a number against its own internal evidence: its citation, its arithmetic, its magnitude. They ask, in effect, "is this number self-consistent and possible?" They never ask, "is this number consistent with the rest of reality?" A figure can be flawless on every internal measure and still be the wrong figure, because the error happened in the selection of what to read, not in the reading itself. The wrong meter, the wrong period, the wrong facility, the duplicated line: each produces an internally perfect number that is externally wrong. The only check that looks outside the number, at the world it is supposed to describe, is the comparison to what that world looked like last year. That is why it is not optional and not redundant. It is the one check that asks a fundamentally different question.
Three Ways to Compare to Last Year
There are three complementary comparisons, and each catches something the others miss. Run all three and you have a net that is hard for a bad number to slip through.
Year-Over-Year Deltas
The first and most direct is the year-over-year delta: this year's value minus last year's, expressed as a percentage change. You compute it for every facility, every fuel, every Scope 3 category, every line that has a prior-year counterpart. Then you set a threshold, a swing beyond which a number must be explained before it is accepted, often something like plus or minus a fixed percentage, tightened for material lines. A delta inside the threshold is presumed fine. A delta outside it is not rejected, it is flagged for explanation. The point is not that a big change is wrong; the point is that a big change must be understood and documented, never silently accepted.
Ratio and Intensity Checks
The second comparison is more powerful because it normalises for real growth. A raw delta cannot tell the difference between "emissions rose 20% because we genuinely produced 20% more" and "emissions rose 20% because of an extraction error." An intensity ratio, emissions per unit of output, per square metre, per employee, per tonne produced, per unit of revenue, strips out the legitimate scale change and leaves the part that should be stable. If your absolute energy rose with production but your energy per unit produced suddenly halved or doubled, that intensity swing is the signal, and it is one a raw total would have hidden. Intensity checks are how you catch errors that are camouflaged by real business change.
The "Explain This Jump" Test
The third is less a calculation than a discipline: for any flagged movement, demand a real-world explanation before the number advances. A jump is acceptable only when it ties to something that actually happened, an acquisition, a new facility, a fuel switch, a production increase, a methodology change, a grid factor revision. "The data says so" is not an explanation. "We brought the Hamburg plant online in March, which adds roughly this much" is. The test forces every material swing to connect to a fact in the business, and an unexplained swing is held back, not published.
The deeper value of the explain-this-jump discipline is that it converts a vague unease into a specific, falsifiable claim. "This number looks high" is a feeling that an analyst under deadline pressure will talk themselves out of. "This site's gas is up 210%, and the only candidate explanation, a fourteen-month billing cycle, fully accounts for it" is a claim you can check and, once checked, either confirm or refute. The discipline does not merely flag; it forces the resolution of the flag into a documented finding. And the explanation, once written down, is not busywork. It is precisely the artifact the assurer asks for. When they point at the same swing you already flagged and ask what drove it, the answer is in the file, dated, reasoned, and tied to a fact, rather than reconstructed on the spot under questioning. The explanation you write to clear your own flag is the explanation you hand the assurer, which is why the discipline pays for itself twice.
An unexplained 3x swing is not a data point. It is a stop sign. The number does not advance until the jump ties to something that actually happened.
Why This Catches AI Extraction Errors Specifically
Prior-period reconciliation earns its place in a chapter on AI-assisted extraction because the failure modes of automated extraction are exactly the kind that produce clean-looking but wrong totals. When a model extracts across hundreds of documents, the errors it makes are systematic and quiet: it reads a sub-meter as a total, it double-counts a line that was also summed elsewhere, it grabs a twelve-month figure where you wanted a quarter, it swaps a unit, it attaches one facility's data to another's row. Every one of these can pass quote-the-source, because the value it found is real, and pass recompute, because the arithmetic on that value is correct. What none of them survives is a comparison to last year, because the resulting total moves in a way the real world did not.
This is the deep reason year-over-year is the cheapest error detector you own: it does not care how the error was made. It does not need to understand the model, the OCR, or the prompt. It only needs to know that emissions do not triple without a reason, and that when they appear to, a human goes and finds out why. An automated pipeline that extracts faster than ever still has to pass through this one stubborn human question, and that question is what keeps speed from turning into a silent restatement.
There is a tempting counter-argument worth dismantling: surely, as extraction tooling improves, these errors will simply stop happening, and the check becomes unnecessary. The opposite is true. Better tooling makes extraction faster and more confident, which means errors are introduced at greater scale and with more fluent presentation, not fewer of them and not more obviously. A pipeline that processes 340 documents in an hour can introduce a systematic period error across dozens of facilities in that same hour, and every one of those wrong numbers will arrive looking exactly as clean as the right ones. Speed multiplies the consequence of a systematic error rather than eliminating it. The prior-period check is the control that does not get less necessary as the tooling improves; if anything it gets more necessary, because it is the last point where a human asks whether the fast, confident, scaled output actually matches what the organisation did.
A Worked Example: The Tripled Gas Bill
Return to Rotterdam. Watch the prior-period check do what the source checks could not.
Before, without the prior-period check. The pipeline extracts Rotterdam's gas consumption from the year's invoices, finds a real figure on a real bill, pairs it with the correct natural gas factor, computes the emissions, and the number reconciles into the inventory. Quote-the-source passes, the bill exists and says what the model read. Recompute passes, the multiplication is right. Range-check passes, the magnitude is within the plausible band for a facility of that type. By every source-level test, the number is clean, and it goes toward a disclosure that will be externally assured.
After, with the prior-period check. The year-over-year delta flags Rotterdam at plus 210%, far beyond the threshold. Nobody rejects the number; somebody asks why. The intensity check sharpens it: gas per square metre tripled, and Rotterdam did not triple its floor area or its process load, so genuine growth is ruled out. The "explain this jump" test sends the analyst back to the documents, where she finds it: the model extracted a bill covering a fourteen-month period, the utility had shifted its billing cycle and issued a long catch-up invoice, and the extraction pulled the whole span into a twelve-month inventory line. The real annual figure is close to last year's. The error was invisible at the source, the long bill was genuine and correctly read, and only the comparison to last year exposed it. The fix is documented, the corrected figure carries a note explaining the billing-cycle anomaly, and the assurer, who will run the same year-over-year analysis, finds the explanation already waiting.
Notice what made this work. The error was not in the number's source, its arithmetic, or its magnitude, so the first three checks were always going to pass. It was an error of scope, fourteen months where twelve belonged, and scope errors are precisely what a prior-period comparison is built to surface. That is why this check is not redundant with the others; it covers a blind spot they all share.
Doing Prior-Period Reconciliation Well
A few practices make the difference between a reconciliation that catches errors and one that becomes a rubber stamp. Set thresholds deliberately and tighter for material lines. The lines that drive your footprint deserve a smaller tolerance than trivial ones, because an error there matters more and a real change there is more likely to be genuinely explainable. Compare like with like. If your boundary changed, an acquisition, a divestment, a methodology update, the prior-year number may need restating onto the same basis before the comparison is meaningful; comparing a new boundary to an old one generates false alarms that train people to ignore the flags. Document every explanation, not just every flag. The explanation is what the assurer wants, so capturing "Rotterdam plus 210% explained by a fourteen-month catch-up invoice, corrected to twelve months" turns the check into assurance evidence, not just an internal note.
One more practice deserves emphasis because it is the one most often skipped: investigate the swing you can explain too easily. A flag that is instantly explained by a story that happens to be convenient, "oh, that must be the new line we added," is exactly the flag that lets a real error through, because the convenient story was accepted without being checked against the numbers. The Hamburg plant might explain a rise, but does it explain a rise of this exact size? If the plant accounts for a 40% increase and the data shows 210%, the convenient explanation is hiding a second, unexplained error underneath it. The discipline is to confirm that the explanation accounts for the whole of the movement, not just its direction. A swing is only cleared when the explanation and the magnitude actually meet.
And keep the human in the loop where judgment lives. AI can compute every delta and every intensity ratio instantly and flag the outliers for you, which is genuinely useful and a good use of the tool. What AI cannot do is decide whether a given jump is a real-world event or an error, because that requires knowing what actually happened in the business this year. The division of labour is the same one that runs through this whole program: the machine flags, the human explains and decides, and the file records who decided what. A reconciliation where the AI both flags and clears its own outliers is no reconciliation at all.
Key Takeaways
- Last year's number is the cheapest error detector you own. Your prior-year inventory is a known-good reference you already have, and comparing against it catches expensive errors for the cost of a subtraction.
- It catches errors the source-level checks cannot. A figure can pass quote-the-source, recompute, and range-check and still be wrong from a double-count, a swapped facility, or a scope error, and only the comparison to last year exposes it.
- Run three comparisons. Year-over-year deltas flag big swings, intensity ratios strip out real growth to expose camouflaged errors, and the "explain this jump" test forces every movement to tie to a real event.
- An unexplained large swing is a stop sign, not a data point. A flagged number is not rejected, it is held until the jump connects to something that actually happened in the business.
- Intensity ratios catch what raw totals hide. Emissions per unit of output distinguishes a genuine 20% production rise from a 20% extraction error that a raw delta would wave through.
- This is a rehearsal of the assurer's own analysis. Assurers plan engagements around year-over-year movement, so running it first means you walk in able to explain every swing instead of handing them surprises.
- Compare like with like. If the boundary or methodology changed, restate the prior year onto the same basis first, or the false alarms will train people to ignore the flags.
- The machine flags, the human explains. AI can compute every delta and ratio instantly, but only a human knows whether a jump is a real event or an error, and the file must record who decided.
Skill.re