Catching the Hallucinated Factor or Figure
An analyst asks an AI assistant for the emission factor for natural gas combustion. Back comes a clean, confident answer: "0.18416 kgCO2e per kWh, from the UK DEFRA 2024 conversion factors." It looks perfect. It has a database name, a year, five decimal places. The analyst pastes it into the inventory and moves on. There is just one problem: DEFRA's 2024 natural gas factor is not that number, and when the assurer later asks to see the DEFRA workbook, the row does not exist. The figure was fluent, specific, and entirely invented. This lesson is about the move that would have caught it in thirty seconds.
The Fabrication and the Misread Both Look Fluent
There are two ways an AI-surfaced number can be wrong, and the dangerous thing they share is that both look exactly as confident as a correct one. The first is the hallucinated factor: a plausible but invented coefficient with no real source, sometimes wrapped in a fabricated citation that names a real database and a real-sounding year. The second is the misread figure: a real number that the model pulled from the wrong place, the standing charge instead of the consumption, the prior-year comparison instead of the current period, a value off by a factor of ten because a decimal moved. One is a number that never existed. The other is a real number in the wrong job. Both arrive in the same calm, articulate prose, and that fluency is precisely the danger.
Why this matters so much in disclosure: an emission factor is a published coefficient that converts activity data into greenhouse gas emissions, and it is supposed to come from a named, dated, authoritative database, the UK Government conversion factors, the US EPA factors, the IEA electricity factors, an Ecoinvent dataset, a supplier-specific factor. The whole credibility of an emissions number rests on the factor tracing to such a source. A hallucinated factor silently corrupts every calculation it touches and leaves no trace until an assurer pulls the thread. By then it may be in a published, assured disclosure, which is where a quiet error becomes a restatement and a headline.
The reflex you are building in this lesson is simple to state and hard to skip under deadline pressure: do not trust a surfaced number because it sounds right. Trust it only after you have checked it against the source. Fluency is not evidence. A citation the model wrote is not evidence. The only evidence is the source document or the factor database itself, opened with your own hands.
The Four Cross-Checks That Catch It
There is no single test that catches every wrong number, because the failures live in different places: in the citation, in the arithmetic, in the magnitude, and in the vintage. So you run four complementary cross-checks, each aimed at one of those places, and a number earns its way into the disclosure only after it has survived all four. The order matters less than the completeness, but a sensible sequence is to range-check first because it is free, then open the source, then reproduce the math, then confirm the database and year. Below, each technique is laid out with the specific failure it is built to catch.
Quote the Source
The first and cheapest technique is to refuse to accept a number that is not accompanied by an exact, quotable source you can reopen. When the AI surfaces a figure, your immediate question is not "is this plausible?" but "where exactly did this come from, and can I see it?" For an extracted figure, that means the file, page, and line. For an emission factor, it means the database name, the table, the row, the version, and the year. And then you actually go and look.
This is where most fabrications die instantly. A hallucinated factor cannot survive you opening the named database and searching for the row, because the row is not there. A misread figure cannot survive you opening the named bill and reading the line, because you will see the model grabbed the currency total, not the kilowatt-hours. The act of reopening the source is the single most powerful verification you have, and it is almost free. The trap is skipping it because the answer already looked complete.
A citation the model wrote is not a source. The source is the document you can open and the row you can point to. Until you have looked, you have a claim, not a number.
Build the habit into your prompts, too. Ask the model to give you the factor and the exact source in a form you can verify, and to say "I cannot find a sourced factor" rather than produce one if it does not have a real reference. A model instructed to cite or refuse will fabricate less often than one instructed simply to answer. But never let the instruction replace the check. The instruction reduces the rate of fabrication; only your own eyes on the source eliminate it for the number in front of you.
Recompute
The second technique catches the errors that quote-the-source can miss: the ones inside the arithmetic rather than the citation. When an AI hands you not just a factor but a finished emissions figure, redo the multiplication yourself. Activity data times emission factor equals emissions; it is one line of arithmetic, and doing it independently catches a startling range of failures.
Recomputing catches the misplaced decimal that turns 1,841 into 184 or 18,416. It catches the unit mismatch, a factor expressed per cubic metre applied to a quantity in kilowatt-hours, which produces a number that is wrong by whatever the conversion between those units is. It catches the model that quietly used a different activity figure than the one you thought you gave it. And it catches the case where the factor is real and correctly sourced but was multiplied wrong. You are not checking whether the model can do arithmetic in general; you are checking this specific number, because this specific number is the one going into the disclosure.
The discipline here is to never let a calculated emissions figure into the inventory without having reproduced it from its stated inputs at least once. If you cannot reproduce it, that is not a rounding quibble to wave away, it is a signal that one of the inputs is not what you think it is, and you stop until you understand why the numbers do not meet.
Range-Check
The third technique is the fastest of all and requires no documents at all: ask whether the number is even possible. A range-check, sometimes called a sanity check or a reasonableness test, compares the surfaced figure against what you already know to be roughly true, and it is astonishingly good at catching gross errors before you spend time on the precise ones.
Emission factors live in known bands. Grid electricity factors sit in a familiar range that varies by country and year but does not swing by orders of magnitude overnight. A natural gas combustion factor of "1.8 kgCO2e per kWh" should make you stop instantly, because it is roughly ten times too high, the decimal has moved. A factor returned as "0.0018" should make you stop for the opposite reason. You do not need to remember the exact value to know the magnitude is wrong, and magnitude errors are exactly what a misplaced decimal or a unit mix produces.
Range-checks also apply to the activity data and the resulting emissions. If a single office's monthly electricity suddenly reads as enough to power a steel mill, the figure is wrong regardless of how cleanly it was sourced. The range-check is your first filter precisely because it is so cheap: a number that fails it never needs the more expensive checks, because you already know it is wrong. A number that passes it has earned the right to be checked properly.
You do not need to be a specialist to build the rough bands you check against, and you do not need to memorise factor tables. The bands come from three cheap sources you already have: last year's numbers for the same activity, the order of magnitude you would expect from physical intuition about the activity, and a quick look at the published range of the relevant factor family. A natural gas combustion factor expressed per kilowatt-hour sits in the low fractions of a kilogram of CO2e; a grid electricity factor varies widely by country but lives in a band you can learn in an afternoon. The point of the range-check is not precision, it is the catch of the gross error, the decimal that moved, the unit that flipped, the figure that is wrong by a factor of a thousand. Those are the errors that, left in, distort a footprint most violently, and they are the easiest of all to catch, because they fail the simplest possible test: could this number even be true?
Demand the Database Name and Year
The fourth technique is specific to emission factors and it closes a gap the others leave open. Even a factor that is real, correctly multiplied, and in the right range can still be wrong for your disclosure if it comes from the wrong database or the wrong year. Factors are revised. The grid decarbonises, methodologies change, and a 2019 electricity factor applied to 2026 consumption overstates or understates emissions in a way an assurer will flag immediately. So you demand, every time, the database name and the year, and you confirm both are the right ones for the reporting period and the geography.
This is also where "a named tool produced it" must never be mistaken for provenance. A factor that appears inside a reporting platform's dropdown is not automatically trustworthy because the platform is reputable. You still need to know which database it draws from and which vintage. The obligation to disclose a sound basis does not transfer to the software. The question is always the same: which named, dated, authoritative source does this factor trace to, and is it the correct one for this activity, this place, and this year?
When you cannot get a clean answer, the factor does not go in. "The AI suggested it and it seemed standard" is not a basis. A factor with no confirmable database and year is, for assurance purposes, indistinguishable from a fabricated one, because neither can be reopened and confirmed.
The year matters more than people new to inventory work expect, because emission factors are not constants of nature, they are periodically republished estimates that move as the underlying reality moves. Electricity grid factors fall as renewables displace coal and gas, so the same kilowatt-hour of consumption maps to fewer kilograms of CO2e each year in a decarbonising grid. Apply a 2019 factor to 2026 electricity and you overstate the footprint, which sounds conservative until you remember that an overstatement is still a misstatement, still wrong, and still something an assurer flags. Apply an old fossil-fuel factor after a methodology revision and you can be wrong in either direction. The discipline of confirming the year is the discipline of matching the coefficient to the period it is meant to convert, every single time, with no exceptions for factors that "everyone uses."
A Worked Example: Two Wrong Numbers, Two Catches
Watch the four techniques catch two different failures in a single afternoon of inventory work.
The fabricated factor. An analyst asks for the emission factor for diesel combustion and the model returns "2.68787 kgCO2e per litre, from the UK Government 2025 conversion factors, fuels table." It is plausible, it is in the right range, the arithmetic downstream will check out. Quote-the-source is what catches it: the analyst opens the actual 2025 UK Government conversion-factor workbook, goes to the fuels table, and finds that while there is a diesel factor near that magnitude, the precise figure and the table reference the model gave do not match the published row. The model blended a real-looking citation with a not-quite-real number. Because the analyst opened the source rather than trusting the fluent citation, the wrong factor is replaced with the correct, confirmed one, and the source reference now points to a row that actually exists.
The misread figure. The same analyst feeds a utility bill to the model and asks for the month's electricity consumption. Back comes "4,820 kWh." It goes toward the calculation. Range-check is what catches it: the analyst knows this facility runs around 48,000 kWh a month, and 4,820 is an order of magnitude low. Reopening the bill, quote-the-source confirms the diagnosis, the model read a sub-line that was one of several meters, not the total, off by a factor that happened to look like a clean number. The full consumption is 48,210 kWh across all meters. A figure that would have understated this facility's footprint tenfold is caught not by suspicion of the model but by a number that simply could not be true.
Neither catch required special tooling. One needed the analyst to open a database; the other needed the analyst to know roughly what the answer should be. Both took minutes. Both prevented a wrong number from entering a figure an assurer would later read. That is the entire value of treating every surfaced number as a claim to be checked rather than an answer to be pasted.
Building the Reflex Into the Workflow
These techniques only protect you if they fire every time, not just when something feels off, because the whole problem is that fabrications and misreads do not feel off. Relying on a sense that something is wrong is exactly the trap, since a hallucinated factor is engineered, by the way these models work, to feel right. The way to make the techniques reliable is to bake them into the workflow rather than leave them to vigilance. Make "quote the source" a required field, so a number without a confirmable origin cannot be entered at all. Make "recompute" a standing step before any calculated figure is accepted. Make "range-check" the first glance every figure gets. Make "database name and year" mandatory metadata on every emission factor.
This is also the right place to remember the cardinal rule of the whole program: accountability stays with the human discloser. The model is a fast, fallible assistant that surfaces candidate numbers; the analyst is the one who confirms each against its source and signs. "The AI gave me that factor" is not a defence to an assurer or a regulator, because the assurer does not assure the model, they assure your disclosure. The four techniques are how you earn the right to put your name on a number the AI surfaced: you checked the source, you reproduced the math, you confirmed the magnitude, and you named the database and the year. A number that has survived all four is a number you can defend.
Key Takeaways
- The fabricated factor and the misread figure both look fluent. One is a number that never existed, the other a real number in the wrong job, and both arrive in the same confident prose, so fluency can never be your test.
- Quote the source. Refuse any number not accompanied by an exact, reopenable source, then actually open it. A hallucinated factor cannot survive you searching the named database for a row that is not there.
- Recompute. Redo activity data times emission factor yourself for any calculated figure. This catches misplaced decimals, unit mismatches, and silently swapped inputs that a citation check would miss.
- Range-check first. Ask whether the number is even possible before checking precision. Factors live in known bands, and a magnitude error from a moved decimal or a unit mix is caught instantly and for free.
- Demand the database name and year. A real, well-multiplied, in-range factor can still be wrong if it is from the wrong source or vintage. Confirm both match the reporting period and geography every time.
- A citation the model wrote is not a source. The source is the document you can open and the row you can point to. Until you have looked, you hold a claim, not a number.
- A named tool is not provenance. A factor inside a reputable platform still needs its underlying database and year confirmed; the obligation to disclose a sound basis does not transfer to the software.
- Bake the checks into the workflow. Make source, recomputation, range-check, and database-and-year mandatory steps, because fabrications and misreads do not feel wrong, so vigilance alone will miss them.
Skill.re