Verification Techniques for Non-Quants
A disclosure manager who has never written a line of code opens an AI-built emissions calculation. It is forty rows of activity data, factors, and conversions, and the total at the bottom reads 1,240,000 tonnes of CO2e. She does not need to rebuild the spreadsheet to know something is wrong. The company is a mid-sized manufacturer. A million-plus tonnes would put it in the league of a small national economy. In ten seconds, without touching the math, she has caught the error. That is the skill this lesson teaches: how to verify what AI built without being the person who could have built it.
You Do Not Need to Be a Quant to Catch the Error
There is a quiet fear among disclosure professionals handed AI-built calculations: that checking them requires the same technical depth that produced them, that you must be able to redo the work to be allowed to judge it. That fear is wrong, and acting on it is dangerous, because it leads people to wave through numbers they could have caught simply because they felt unqualified to look. The truth is the opposite. The most powerful checks in disclosure are not deep recomputations. They are fast, structural sanity tests that a non-quant can run in minutes, and they catch the errors that matter most: the ones large enough to misstate the disclosure.
This matters because the assurer is not a quant either, in the sense that they do not blindly trust complexity. An assurer's most feared questions are often the simplest: does this number make sense, does it tie to last year, why did this jump. The disclosure professional who can ask those questions of an AI output before the assurer does is exactly the professional this program is building. Your judgment about whether a number is plausible, whether it ties to what you know about the business, is not a weaker check than the math. For catching material error, it is frequently the stronger one.
There is a second reason the non-quant is well placed to verify, and it is about distance. The person who built the calculation, or the model that built it, is the worst-placed to spot its errors, because they are inside the logic that produced them. A fresh reader who knows the business but not the formulas brings exactly the outside perspective that catches the howler the builder's eye slides past. This is why assurance is performed by an independent party and not by the preparer marking their own work. When you check an AI-built calculation, you are playing that independent role, and your outsider's intuition about whether a number belongs is a feature of your position, not a deficiency in your training. The model can produce forty internally consistent rows that sum to an absurd total; it takes a human who knows the company is not the size of a small nation to see that the total is absurd.
What follows is a small toolkit of five checks. None requires statistics. All require what you already have: knowledge of your business, last year's report, and a willingness to be suspicious of a number that wants to be trusted. Run them on every AI-built calculation before it moves.
The Five Checks Every Disclosure Professional Needs
1. The order-of-magnitude check
Before anything else, ask: is this number even the right size? Not the right value, the right size, the right number of zeros. The manager in the opening did this. A mid-sized manufacturer's total footprint might plausibly be tens of thousands of tonnes; a number in the millions is off by orders of magnitude and signals a unit error, a double-counting, or a misplaced factor. The order-of-magnitude check asks whether the answer lands in the right ballpark given everything you know about the company's size, sector, and prior numbers. It will not catch a number that is ten percent wrong, but it catches the thousand-fold errors instantly, and those are precisely the unit slips and misplaced conversions that AI calculations produce. Train yourself to glance at the magnitude before you read the digits.
2. The unit check
Most large calculation errors are not arithmetic errors. They are unit errors: litres treated as kilograms, kilowatt-hours read as megawatt-hours, a factor expressed per tonne applied to a figure in kilograms. The unit check is simply reading the units across a calculation and confirming they cancel and combine into the units the answer should have. If you multiply litres by kilograms of CO2e per litre, the litres cancel and you are left with kilograms of CO2e, which is right. If the factor was per kilogram and the activity was in litres, the units do not cancel, and that mismatch is a flashing light. You do not need to know the chemistry. You only need to read the units like a sentence and check that they make grammatical sense. A calculation whose units do not resolve to the answer's units is wrong, full stop, regardless of how clean the arithmetic looks.
3. The ratio or intensity check
Absolute numbers are hard to judge; ratios are easy. Take the emissions figure and divide it by something you understand: revenue, headcount, units produced, floor area. Now you have an intensity, tonnes of CO2e per million of revenue, or per employee, and intensities have known, sense-checkable ranges. If your AI-built inventory implies an emissions intensity ten times your sector's typical range, or ten times your own prior year, the absolute number is suspect even if you cannot say exactly why. The ratio check works because it converts an unfamiliar absolute into a familiar relative, and your professional intuition about relatives is sharp even when your intuition about absolutes is not. It is the single most useful check for spotting a number that is wrong but not wildly wrong.
4. Recompute one line
You cannot rebuild forty rows, and you do not have to. Pick one line, the most material one, or one at random, and recompute just that single line by hand or on a calculator. Take its activity data, its factor, do the one multiplication, and see if you get what the AI output says for that line. This is sampling, the same logic an assurer uses: you cannot test everything, so you test one thing well and let it tell you whether to trust the rest. If the one line you recompute matches, your confidence in the whole rises; if it does not, you have found a thread to pull. Recomputing one line costs two minutes and is the most concrete check on the list, because it touches the actual arithmetic without requiring you to reproduce all of it.
5. Compare to prior
Last year's number is the cheapest fraud-and-error detector you own. Put this year's AI-built figure next to last year's audited figure and ask whether the change is explainable. A footprint that moved two percent in a stable business is unremarkable. A footprint that doubled, or halved, demands a reason, and "the AI calculated it" is not a reason. Real changes have real causes: an acquisition, a divestment, a fuel switch, a methodology change, a correction. If you cannot name the cause of the change, the change is unverified, and an unverified jump is exactly what an assurer will stop on. Comparing to prior turns the previous, already-assured report into a free check on the current one.
You do not verify an AI calculation by being able to rebuild it. You verify it by asking whether it is the right size, whether the units make sense, whether the ratio is sane, whether one line recomputes, and whether it ties to last year. None of those needs a quant. All of them need you.
A Worked Example: Running the Checks on an AI Output
An AI has produced the Scope 1 and 2 inventory for a mid-sized manufacturer. The headline total is 1,240,000 tonnes CO2e. The disclosure manager runs the five checks in order, and watch how little math she needs.
Order of magnitude. A mid-sized manufacturer with a few hundred employees and one main site should land in the tens of thousands of tonnes, not the millions. 1,240,000 is roughly a hundred times too big. The check fails immediately, and she now suspects a unit error somewhere in the chain. She has not opened a single formula.
Unit check. She scans the largest line, electricity. The consumption is in kilowatt-hours, but the factor is the grid-average factor expressed per megawatt-hour, and the calculation multiplied kilowatt-hours by the per-megawatt-hour factor without converting. The units do not cancel: she is implicitly treating kilowatt-hours as megawatt-hours, inflating that line by a factor of a thousand. There is the error, found by reading the units, not the numbers.
Ratio check, after the fix. She has the calculation corrected so the electricity line uses consistent units, and the total drops to about 12,400 tonnes. She divides by revenue and gets an emissions intensity that sits squarely in her sector's normal range and close to last year's. The relative number now passes the sniff test that the absolute one could not give her on its own.
Recompute one line. She takes the corrected electricity line, the consumption in megawatt-hours times the per-megawatt-hour factor, and does the one multiplication on a calculator. It matches the AI's corrected line. One line recomputed, confidence in the rest restored.
Compare to prior. Last year's audited Scope 1 and 2 total was 11,900 tonnes. This year's corrected 12,400 is a four percent increase, and she can name the cause: a slightly higher production volume, documented in the operations file. The change is explainable, so it is verified. Had the figure stayed at 1,240,000, the comparison to prior would have caught it even if the order-of-magnitude check had not, because a hundred-fold jump screams for a cause that does not exist.
Five checks, perhaps fifteen minutes, no statistics, and a thousand-fold error caught and corrected before it reached the inventory. That is the entire argument for treating verification as a disclosure skill rather than a quant skill.
The Order the Checks Run In Matters
Run the five checks in the sequence above, because the order is itself a piece of efficiency. Order of magnitude comes first because it is the fastest and catches the biggest errors, so there is no point recomputing a line of a number that is a hundred times too big. The unit check comes second because once magnitude flags a problem, units are where the problem usually lives. Only after those structural checks do you spend time on the ratio, the recomputed line, and the prior-year tie, which are sharper but slower. A non-quant who runs the cheap checks first stops most bad numbers in the first thirty seconds and reserves the more effortful checks for numbers that have already earned a second look. This is the same triage logic an assurer uses, attack the cheapest, highest-yield test first, applied to your own review before the engagement.
Notice too that the checks reinforce one another. In the worked example, the order-of-magnitude failure pointed at a unit error, the unit check found it, the ratio check confirmed the corrected number was sane, the recomputed line confirmed the arithmetic, and the prior-year tie confirmed the change was explainable. No single check proved the number; together they triangulated it. A figure that passes all five from different directions, size, units, ratio, arithmetic, and history, is far more trustworthy than one that survives a single deep review, because the five checks fail in different ways and a real error tends to trip at least one of them.
Why These Checks Beat Deep Review for Catching Material Error
It is counterintuitive that fast, shallow checks catch more material error than slow, deep review, but for the errors that threaten a disclosure, they do. The errors that matter in an inventory are the large ones: the unit slip that inflates a line a thousand-fold, the double-count that adds a whole site twice, the misplaced factor that shifts a category by a factor of ten. These large errors are exactly what the order-of-magnitude, unit, ratio, and prior-year checks are tuned to catch, and they catch them in minutes. A deep line-by-line review is slower, exhausting, and ironically more likely to miss the large structural error precisely because the reviewer is buried in detail. The non-quant checks work at the level where material error lives: the level of size, shape, and sense.
The Limits, and the Discipline That Makes Them Work
Be honest about what these checks do not do. They do not confirm a number is exactly right; they confirm it is not obviously wrong. A figure can pass all five and still contain a subtle error: a factor that is real but slightly outdated, an estimate mislabeled as measured, a boundary quietly drawn too narrow. The five checks are a first line of defense that catches the gross errors cheaply, not a substitute for the provenance work, the factor verification, the primary-versus-secondary labeling, the chain-of-thought tracing that the rest of this program teaches. Pass the five checks and you have earned the right to do the deeper work on a number that is at least the right size and shape.
It is worth naming the specific blind spot of each check, so you do not over-trust any one of them. Order of magnitude misses errors smaller than a factor of about ten. The unit check misses an error where the units happen to be consistent but a wrong-but-same-unit factor was used. The ratio check misses an error that moves the absolute figure and the denominator together. Recomputing one line tells you nothing about the thirty-nine lines you did not touch. The prior-year comparison misses an error that was also present, undetected, last year. Each gap is exactly why you run all five rather than picking a favourite: their blind spots do not overlap, so the set covers far more than any single check while no member of the set is sufficient alone.
The discipline that makes the checks reliable is running them every time, in order, before the number moves, and documenting that you did. A check you run only when a number "feels off" is a check you will skip exactly when you most need it, because the dangerous errors are the ones that feel fine. Make the five checks a standing gate on every AI-built calculation, record the order-of-magnitude judgment, the unit confirmation, the intensity, the recomputed line, and the prior-year comparison in the working papers, and you have done two things at once: caught the gross errors, and produced a piece of the verification trail the assurer will want to see. The cheapest checks in the toolkit, run consistently and recorded, become part of the assurance file.
Key Takeaways
- You do not need to be able to rebuild an AI calculation to verify it. The most powerful checks are fast structural sanity tests a non-quant can run in minutes, and they catch the errors large enough to misstate the disclosure.
- Order-of-magnitude check: is the number even the right size, given the company's scale, sector, and prior figures? It catches thousand-fold unit slips instantly.
- Unit check: read the units across the calculation like a sentence and confirm they cancel into the answer's units. A calculation whose units do not resolve is wrong regardless of clean arithmetic.
- Ratio or intensity check: divide the figure by revenue, headcount, or output to convert an unfamiliar absolute into a familiar relative, which your professional intuition can judge. Best for catching a number that is wrong but not wildly wrong.
- Recompute one line: sample a single material line and redo its one multiplication. You cannot rebuild forty rows, but one recomputed line tells you whether to trust the rest.
- Compare to prior: last year's audited figure is the cheapest error detector you own. If you cannot name the cause of the change, the change is unverified, and an unverified jump is what an assurer stops on.
- Fast shallow checks beat slow deep review for material error, because the errors that threaten a disclosure are large and structural, exactly what these checks are tuned to catch while a line-by-line reviewer gets buried.
- The checks confirm a number is not obviously wrong, not that it is exactly right. Run them every time, in order, before the number moves, and record them, and the cheapest checks become part of the assurance trail.
Skill.re