Data Governance, Factors, and Provenance
The assurer's request is one line long: "Show me the emission factor you used for purchased electricity in Germany, the version, the date, the source, and prove it is the same one behind the number in the report." A junior analyst opens the reporting platform and finds the figure, but not the factor. The factor lived in a spreadsheet that has since been overwritten. Nobody changed the answer dishonestly; they just cannot prove what the answer was built on. In that moment, a correct number becomes an unsupported one, and the whole disclosure above it starts to wobble. This is the layer that decides whether reporting AI is assurable at all.
The Foundation Everything Else Sits On
Every lesson in this program has pushed toward a single discipline: every figure you publish must trace to evidence, and "the AI estimated it" is not evidence. That discipline is not something AI provides. It is something the data foundation beneath the AI provides, and if that foundation is weak, nothing built on top of it is defensible. You can have the best system prompts, the cleanest workflows, the most careful analysts, and still fail assurance because the factor library is a mess, the golden source is ambiguous, or lineage was never captured. This is why data governance sits at the enterprise-transformation level: it is the load-bearing wall, and AI is drywall hung on it.
The uncomfortable truth for a sustainability leader is that this is the least glamorous part of the whole enterprise and the part that most reliably ends careers. Boards want to talk about AI drafting the report. Assurers, and increasingly regulators, want to talk about whether the data behind the report is governed. Roughly 73% of large global companies now obtain external assurance on sustainability data (a figure to verify against the current IFAC study, but the direction is clear), and GHG emissions are the most-assured category. Every one of those engagements eventually reaches the data foundation. Get it right and AI accelerates a defensible system. Get it wrong and AI accelerates the production of numbers you cannot defend.
AI is only as assurable as the data foundation it sits on. Ungoverned data plus fast AI equals fast, unsupported numbers.
Factor Libraries With Version Control
The emission factor is where AI's hallucination risk meets disclosure's evidence requirement most directly, and it is where the data foundation earns its keep. A hallucinated factor is a plausible but invented number with no source; it silently corrupts a footprint because it looks exactly like a real factor in the output. The defense is not vigilance, which does not scale, but a governed factor library that the AI is forced to draw from and that an assurer can inspect.
A governed factor library has properties an ad-hoc spreadsheet does not. Every factor entry has a stable identifier, a value, a unit, a named authoritative source, a publication date, and an effective date range. Crucially, it is version controlled: when a source revises a factor, the library adds a new version rather than overwriting the old one, because last year's disclosure was built on last year's factor and must remain reconstructable. The library records which version was current for which reporting period, so a figure can always be tied back to the exact factor that produced it.
Version control is what turns a factor revision from a crisis into a lookup. When a factor changes, you do not lose the past; you gain a new version and a clear record of the transition. AI can safely suggest factors only when it retrieves from this governed library, returning the identifier and version rather than a number from memory. That is the mechanism behind "cite or refuse" at the factor level: the model resolves to a library entry or it declines, and the library is what makes the citation real.
It is worth being precise about why an emission factor, of all the numbers in a disclosure, deserves this much machinery. A factor is a multiplier: it converts activity data into emissions, so an error in a single factor does not stay small. Apply a wrong electricity factor and every facility in that market is misstated at once; apply a wrong material factor and every product using that material carries the error. Factors are also deceptively easy to get slightly wrong, because plausible values cluster close together, a real factor and an invented one can differ by a believable-looking amount, and the wrong one will not announce itself. And they change: grids decarbonize, methodologies are refined, sources issue corrections. A number that is both high-leverage, easy to fake, and prone to revision is exactly the number you must never leave to memory, a spreadsheet, or a model's confident guess. The governed library is the structural answer to a structural risk.
Who Owns the Library
A library without an owner rots. Someone must be accountable for adding new sources, retiring superseded ones, approving version changes, and documenting the rationale for factor selection where alternatives exist. This is a governance role, not an IT task, because choosing between two plausible factors is a disclosure judgment the assurer may test. The owner is also the person who can answer, in a walkthrough, why a particular factor was chosen over another. An orphaned library is an assurance finding in waiting.
The ownership question sharpens when the enterprise runs across many geographies and business units, each with a local preference for a particular factor source. Left ungoverned, this produces a quiet fragmentation: one region uses one grid-emissions dataset, another uses a different one, and the consolidated footprint mixes methodologies in ways nobody decided on purpose. The library owner's job is to make those choices deliberate, either standardizing on a single authoritative source per activity and geography, or documenting and defending a considered exception where local data is genuinely better. What the owner cannot allow is drift, factors entering the library because someone found them convenient, with no record of why. The discipline is not that every factor comes from the same place; it is that every factor's origin was a decision, made by an accountable person, that can be explained. That is the difference between a library and a pile.
The Golden Source and Master Data
The second pillar of the foundation is the golden source: the single, authoritative record for each piece of governed data. When the same fact, a facility's energy consumption, a supplier's reported emissions, an entity's inclusion in the consolidation boundary, exists in three systems with three slightly different values, you have no golden source, and every disclosed figure that touches that fact is ambiguous. The assurer's question "which number is right?" has no clean answer, and ambiguity reads as weakness in the control environment.
Establishing a golden source means designating, for each governed data domain, the one system of record whose value is authoritative, and defining how other systems reconcile to it. Closely related is master data: the stable reference data that everything else hangs on, the entity structure and consolidation boundary, the facility register, the supplier master, the unit and category taxonomies. If two facilities are recorded inconsistently, or an entity's boundary status is unclear, the errors propagate upward into every aggregate. Master data governance is unglamorous and decisive: it is why an assurer can trust that the sum of the parts is the whole.
At enterprise scale, the golden source and master data are also what make one fact base serve many frameworks. ESRS, ISSB, and CBAM can all draw from the same authoritative records only if those records are unambiguous. The moment master data forks, the frameworks diverge, and the same underlying emission becomes two different disclosed numbers. Governance here is what keeps the disclosures consistent with each other, which is itself something an assurer and a regulator check.
The reason master data feels invisible until it fails is that it works by absence: when it is right, nobody notices, because the numbers simply add up. When it is wrong, the symptom shows up far from the cause. A duplicated facility in the register surfaces as a Scope 1 total that is a few percent too high, and the analyst chasing that discrepancy may spend days in the calculation layer before discovering the problem was a naming inconsistency two layers down. This distance between cause and symptom is why master data governance cannot be an afterthought bolted on when a discrepancy appears; it has to be established up front, with a single golden record for each facility, entity, and supplier, and a discipline that new records are created deliberately rather than accreting through copy-paste and import. An assurer who sees clean, deduplicated master data reads it as a sign that the control environment is mature. One who sees the same facility under three names reads it, correctly, as a warning about everything built on top of it.
Lineage: The Thread From Raw Data to Published Number
Lineage is the recorded path from a raw input to a published figure, every transformation, factor application, estimate, and human decision along the way. It is the single most important artifact for assurance, because the core assurance test is reconstructability: can someone rebuild this number from the evidence, without you in the room? If the answer is no, the number is unsupported regardless of whether it is correct.
The failure in the opening scene was a lineage failure. The number existed; the thread back to the factor did not. Lineage governance requires that every disclosed figure carry a traceable path: which raw records fed it, which golden-source values were used, which factor version was applied, which estimation method was used where primary data was missing, and which human accepted or overrode the result. When AI performs a step, extraction, factor suggestion, drafting, lineage records that too, honestly, so the file shows where the machine was in the loop.
The design principle that makes lineage work is the same one that makes the Scope 3 architecture work: lineage is captured as data flows, not reconstructed afterward. Reconstruction after the fact is guesswork under deadline, and it is the most common way a correct number fails assurance. Captured lineage turns the assurer's request list into a set of queries. It also turns a restatement from a forensic investigation into a targeted correction: you can find exactly which figures a revised factor or a corrected input touched, and fix only those.
Lineage is also the honest record of where AI sat in the process, and that honesty is worth more to an assurer than a claim that AI was not used at all. An engagement is not hostile to AI; it is hostile to unexplained numbers. A lineage record that says "this value was extracted by an AI model from a named invoice, reconciled by a named analyst against the prior period, and signed off" is more defensible than a value with no record, whether or not a machine was involved. The instinct to hide AI's role, to present a figure as if a human derived it by hand, is exactly wrong: it removes information the assurer needs and, if discovered, poisons trust in the whole file. Lineage that faithfully marks AI involvement lets the assurer weight the control appropriately and tests as designed. Transparency about the machine is a feature of a mature foundation, not an admission of weakness.
How AI Raises the Stakes on the Foundation
AI does not change what the data foundation must do; it raises the cost of getting it wrong. A human analyst working by hand produces figures slowly, and slowness is an accidental control: there is time to notice an odd factor or an implausible input. AI removes that friction. It can apply a wrong factor to ten thousand datapoints as fast as to one, and it can generate a fluent, confident narrative around an unsupported number that looks entirely persuasive. Speed without a governed foundation is not efficiency; it is the industrialization of unsupported numbers.
This is why the enterprise that wins does not choose between speed and governance. It grounds the AI on the governed foundation, retrieval over the factor library and the golden source, not the open web or model memory, so that the speed operates inside the guardrails rather than around them. The factor library makes factor citations real. The golden source makes inputs unambiguous. Lineage makes the whole chain reconstructable. AI then accelerates a system that is defensible by construction, which is the only kind of acceleration a regulated discloser can safely buy.
There is a sequencing lesson here that senior leaders miss at their peril. The instinct, under board pressure, is to buy the AI first and sort out the data later, because the AI is the visible, exciting purchase and the data foundation is invisible plumbing. That order is backwards and expensive. AI deployed on an ungoverned foundation does not sit idle waiting for the data to be fixed; it actively produces figures, at scale, that will have to be unwound when the foundation is finally built and the numbers turn out to be unreconstructable. The disciplined sequence is to establish the version-controlled factor library, designate the golden source, govern the master data, and turn on lineage capture first, then ground the AI on all of it, and only then let it accelerate. The foundation is not the boring prerequisite you get to after the AI; it is the thing that makes the AI purchase worth anything at all. Buy them in the wrong order and you have paid for speed toward a restatement.
Worked Example: The Overwritten Factor, Prevented
Return to the opening scene and rebuild it on a governed foundation. Same company, same assurer, same one-line request: show me the electricity factor for Germany, its version, date, and source, and prove it is the one behind the number.
In the ungoverned version, the factor lived in an overwritten spreadsheet and the thread was lost. In the governed version, the analyst opens the reporting figure and follows its lineage. The lineage points to a specific factor-library entry with a stable identifier and a version number. That entry names the authoritative source, its publication date, and the effective period; the library shows this version was the one current for the reporting year. The golden source confirms the underlying electricity consumption came from the designated system of record, reconciled and unambiguous. The human sign-off is recorded against the figure. The analyst answers the assurer's request in minutes, by query, and the assurer moves on. Nothing about the answer changed; everything about its defensibility did.
Now extend it. Mid-year, the source revised the German electricity factor. On the ungoverned foundation, this is a slow-motion disaster: which figures used the old factor? On the governed foundation, the library added a new version, lineage records which figures used which version, and the team can identify and restate exactly the affected figures for comparability, walk the assurer through the change, and close cleanly. The foundation did not just answer one question; it made the whole system resilient to change. That is the payoff of getting the least glamorous layer right: it is the layer that decides whether everything above it is defensible, and once it is solid, AI can move as fast as the board wants without carrying the enterprise toward a restatement.
Key Takeaways
- The data, factor, and provenance foundation is the load-bearing layer: AI is only as assurable as the data beneath it, and ungoverned data plus fast AI produces fast, unsupported numbers.
- Run a governed factor library with version control: stable identifiers, named dated sources, effective periods, and new versions on revision rather than overwrites, so every figure ties back to the exact factor that produced it.
- Give the factor library a named owner accountable for adding sources, approving version changes, and documenting factor-selection rationale, because choosing between factors is a disclosure judgment an assurer can test.
- Establish a golden source, one authoritative system of record per governed data domain, so the assurer's "which number is right?" has a clean answer and disclosures stay internally consistent.
- Govern master data (entity and boundary structure, facility register, supplier master, taxonomies) because errors there propagate into every aggregate and fork the frameworks apart.
- Capture lineage as data flows, never reconstruct it afterward: the recorded path from raw input to published figure is the artifact that passes the reconstructability test at the heart of assurance.
- AI raises the stakes rather than changing the requirements: it removes the accidental control of slowness, so a wrong factor or unsupported number scales instantly unless the foundation constrains it.
- Ground AI on the governed foundation via retrieval over the library and golden source, so speed operates inside the guardrails and accelerates a system that is defensible by construction.
Skill.re