Assurance, GHG Protocol, and California SB 253 and SB 261
Picture the room where it all gets decided. An external assurer sits across the table from your carbon accountant, a request list open on a laptop. She points at one line in your published report, a single Scope 1 figure, and says: "Show me how you got this. Walk me from the meter reading to the number on the page." Everything this program teaches, every provenance tag and every labeled estimate, exists for this moment. The assurance engagement is the perimeter every disclosed number sits inside, and it is the last thing AI can talk its way past.
Assurance Is the Spine of Everything
Across this whole regulatory chapter, one fact sits underneath all the others. 73% of large global companies now obtain external assurance on at least some of their sustainability disclosures, up from 51% in 2019. That jump, measured by IFAC together with AICPA and CIMA, is the single most important trend for anyone using AI in reporting. It means the era of the unaudited sustainability narrative is over. Every disclosed figure is now, increasingly, an audited figure, read by an independent professional whose job is to test whether it traces to evidence.
This is why the program is assurance-first rather than speed-first. A glossy AI-generated report that cannot survive the assurer's request list is worse than no report, because it creates a public commitment the file cannot back. The discipline that makes an AI output traceable is the same discipline that makes it assurable. Speed and defensibility are not opposites here; they are produced by the same habits.
An external assurer does not read your report for style. She reads it for evidence, and "the AI generated it" is the one answer that fails every time.
This single sentence reorganizes how you should think about every AI tool in the reporting stack. A model's fluency, the very quality that makes it useful for drafting, is worthless and even dangerous at the moment of assurance, because the assurer is not grading the prose. She is asking whether the number behind the prose can be rebuilt from evidence. A team that optimizes for how good the report reads is optimizing for the wrong audience. The audience that decides whether the report stands is the one holding the request list, and that audience is moved only by traceability.
It is worth noting which category gets tested hardest. Greenhouse gas emissions are the most-assured category of sustainability data. The number an AI is most tempted to estimate, a Scope 3 figure built on missing supplier data, is precisely the number most likely to land under an assurer's eye. That is not a coincidence to plan around. It is the center of gravity to plan from.
Sit with the irony for a moment, because it is the whole reason this lesson closes the regulatory chapter rather than opening it. The frameworks you have studied, CSRD with its ESRS, ISSB with S1 and S2, CBAM with its embedded emissions, all converge on the same hard quantity: greenhouse gas emissions, and especially the value-chain emissions of Scope 3. That is the number hardest to gather, most dependent on data you do not control, and therefore most tempting to let a model "fill in." It is also the number an assurer is most likely to examine closely. The temptation and the scrutiny point at the same spot. Any approach that treats AI as a way to paper over a Scope 3 gap is aiming the riskiest tool at the most heavily tested number, which is precisely backwards. The assurance-first approach aims the discipline there instead: the harder and more tested the number, the more rigorously its provenance is built.
Limited vs Reasonable Assurance
Assurance comes in levels, and the difference between them is one of the most practically important distinctions in this chapter. The two you must know are limited assurance and reasonable assurance.
What Limited Assurance Means
Limited assurance is the lower level, and it is where most sustainability engagements sit today. In a limited-assurance engagement, the assurer performs procedures (primarily inquiry and analytical review) sufficient to conclude that nothing has come to their attention suggesting the information is materially misstated. The conclusion is expressed in the negative: "nothing came to our attention." It is a real check, but a lighter one. The assurer is not exhaustively re-performing your work.
What Reasonable Assurance Means
Reasonable assurance is the higher level, the kind long applied to financial-statement audits. Here the assurer performs much more extensive, deeper procedures and expresses a positive conclusion: in their opinion, the information is fairly stated. It demands far more evidence, far more testing, and a far more robust trail behind every number.
The trajectory matters as much as the definitions. Most sustainability engagements today are limited assurance, but the direction of travel is toward reasonable assurance over time. The practical implication for an AI workflow is blunt: build your evidence trail to the higher bar now. A provenance habit that satisfies a limited-assurance reviewer today may need to satisfy a reasonable-assurance reviewer tomorrow, and it is far cheaper to build the trail once than to reconstruct it after the bar rises. Treat every disclosed figure as if a reasonable-assurance reviewer will one day re-perform it.
The difference between the two levels is easiest to feel through what the assurer actually does. Under limited assurance, the assurer leans on inquiry, asking you how something was done, and analytical review, checking whether the numbers move the way you would expect. If a figure looks reasonable and your answers are coherent, a limited engagement may not dig deeper. Under reasonable assurance, the assurer does not take the coherent answer at face value; they re-perform, sampling source documents, recalculating, and testing the controls behind the number. A trail that survives "tell me how you did this" is not the same as a trail that survives "I am going to redo this myself from your evidence." Building to the reasonable bar means the second sentence holds, and the first one comes free. Building only to the limited bar means you may pass today and scramble when the engagement deepens.
The GHG Protocol Is the Backbone Standard
When the assurer asks how you built an emissions number, she is asking whether you followed a recognized standard. For greenhouse gas accounting, that standard is the GHG Protocol, the backbone methodology that defines how organizations measure and report emissions. It establishes the concepts the rest of the program leans on: the split into Scope 1 (direct emissions), Scope 2 (purchased energy), and Scope 3 (the value chain, across its 15 categories), the rules for organizational and operational boundaries, and the distinction between activity data and emission factors.
The reason the GHG Protocol matters for AI work is that it is the shared language between you and the assurer. When you tag a number as a Scope 3 Category 1 figure built from activity data and a named emission factor, you are speaking in terms the assurer recognizes and can test. An AI output that produces an emissions number without reference to the Protocol's structure, no clear scope, no boundary, no separation of activity data from factor, is an output that has skipped the very framework the assurance rests on. The Protocol is not bureaucratic overhead. It is the grammar that makes your numbers checkable.
It helps to be concrete about the two halves the Protocol keeps separate, because their separation is what makes a number reconstructable. Activity data is the measure of the thing that happened: litres of fuel burned, kilowatt-hours of electricity purchased, tonnes of a material bought. An emission factor is the conversion that turns that activity into emissions: the kilograms of carbon dioxide equivalent per unit of activity. Multiply the two and you get the emissions figure. When these are recorded together but kept distinct, an assurer can test each independently: is the activity data sourced and accurate, and is the factor a recognized, dated value rather than something invented. When an AI tool collapses them into a single number with no trace of either, the assurer has nothing to test, and a number with nothing to test is a number that fails. The boundary, which facilities and operations are included, is the third piece, because the same activity can belong inside or outside the inventory depending on how the organizational and operational boundaries were drawn.
The US Thread: California SB 253 and SB 261
Assurance and the GHG Protocol are global, but the regulatory map is not only European. The United States has its own thread, and it runs through California. Two state laws matter.
SB 253: The Climate Corporate Data Accountability Act
California SB 253 is the emissions-disclosure law. It requires large companies doing business in California to disclose their greenhouse gas emissions across Scope 1, Scope 2, and Scope 3. This is significant because it pulls Scope 3, the hardest and most value-chain-dependent scope, into a mandatory US disclosure regime, reaching companies far beyond those captured by EU rules. If you operate in the US market, SB 253 can put your full emissions inventory, including the difficult Scope 3 categories, under a disclosure obligation.
SB 261: Climate-Related Financial Risk
California SB 261 is the climate-risk-disclosure law. Rather than emissions totals, it focuses on a company's climate-related financial risk and how the company is managing it. It is the risk-narrative companion to SB 253's emissions numbers. Together, SB 253 and SB 261 mean a US-operating company can face both an emissions-data obligation and a climate-risk-narrative obligation, echoing the structure you already saw in IFRS S2 and the ESRS.
The strategic point for an AI-aware professional is convergence again, from a different direction. SB 253's Scope 1, 2, and 3 emissions are built on the same GHG Protocol backbone, drawn from the same fact base, and increasingly subject to the same assurance expectations. The US thread is not a separate universe requiring a separate inventory. It is one more framework your clean fact base maps into, and one more reader of the same evidence-backed numbers.
This is the moment to connect the whole chapter. You have now seen four regulatory surfaces, CSRD, ISSB, CBAM, and the California laws, and underneath all of them sit two constants: the GHG Protocol as the way emissions are measured, and external assurance as the way numbers are tested. The frameworks differ in scope, in materiality lens, and in legal mechanics, but they all read numbers that were built the same way and they all increasingly expect those numbers to be assured. That is why a professional who masters the assurance perimeter and the Protocol grammar is not learning four separate skills. They are learning the one discipline that every framework rewards, and then mapping it into whichever obligation a jurisdiction imposes. The frameworks are the many; the fact base and the assurance perimeter are the one.
A Worked Example: Surviving the Request List
Return to the assurer and the single Scope 1 line. Watch the two ways this goes.
In the first company, the number was produced fast. Someone prompted an AI tool to "calculate our Scope 1 emissions from these fuel records," and the model returned a clean total. It was pasted into the report. Now the assurer asks to walk from the meter reading to the page. The carbon accountant opens the file and finds, behind the number, only the AI's output. There is no record of which emission factor was used or where it came from, no separation of the metered activity data from the factor, no boundary note explaining which facilities were included. The model may even have applied a factor it generated rather than one from a recognized database. The accountant cannot reconstruct the number. Under a limited-assurance engagement this becomes a finding; under reasonable assurance it would be fatal.
In the second company, the same AI tool was used, but inside the discipline this program teaches. The Scope 1 figure traces to metered fuel consumption, each input sourced. The emission factor is named, dated, and drawn from a recognized database, recorded alongside the activity data and kept separate from it, in line with the GHG Protocol. The boundary is documented: these facilities in, those out, with the rationale. When the assurer asks to walk from the meter to the page, the accountant does exactly that, in minutes. The AI accelerated the work. It did not replace the evidence. The number survives the request list because it was built to.
The difference between the two companies is not the AI tool. They used the same one. The difference is whether the team understood that the assurer, the GHG Protocol, and the disclosure law are the perimeter every number lives inside, and built accordingly. That understanding is what L1 of this program exists to give you, and what every later level turns into hands-on practice.
There is a simple mental test you can apply to any AI-assisted number before it ever reaches a report, and it is the same test the assurer will apply later: can someone who was not in the room reconstruct this figure from the evidence alone. If the answer is yes, the activity data is sourced, the factor is named and dated, the boundary is documented, the number will survive the request list. If the answer is no, you have a liability dressed as a disclosure, regardless of how confident the AI sounded or how clean the total looks. Running that test yourself, before the assurer ever asks, is the cheapest insurance in the entire reporting cycle, because the cost of fixing a traceability gap before publication is a fraction of the cost of a finding or a restatement after.
That test also reframes what "using AI well" means in disclosure. It does not mean producing the most polished narrative the fastest. It means producing numbers and claims that a stranger can rebuild from your evidence, with AI accelerating the parts of that work it can do safely, the structuring, the drafting, the organizing, while every figure stays anchored to a source the assurer can pull. The teams that thrive under a 73-percent-assured, increasingly-reasonable-assurance world are not the ones who generate fastest. They are the ones who build every number so that the answer to "walk me from the meter to the page" is always, simply, yes.
Key Takeaways
- Assurance is the spine: 73% of large global companies now obtain external assurance on at least some sustainability disclosures, up from 51% in 2019, so every disclosed figure is increasingly an audited figure.
- Greenhouse gas emissions are the most-assured category, so the Scope 3 number AI is most tempted to estimate is the one most likely to land under an assurer's eye.
- Limited assurance (the lower level, most common today) gives a negative conclusion that nothing came to attention; reasonable assurance (the higher level) gives a positive opinion and demands far more evidence. The trend runs toward reasonable, so build the trail to the higher bar now.
- The GHG Protocol is the backbone standard: Scope 1, 2, and 3, the 15 Scope 3 categories, organizational and operational boundaries, and the split of activity data from emission factors. It is the shared grammar that makes your numbers checkable.
- An AI emissions output with no scope, no boundary, and no separation of activity data from factor has skipped the very framework the assurance rests on.
- California SB 253 requires large companies doing business in California to disclose Scope 1, 2, and 3 emissions, pulling the hardest scope into a mandatory US regime.
- California SB 261 requires disclosure of climate-related financial risk, the narrative companion to SB 253's numbers, echoing the IFRS S2 and ESRS structure.
- The US thread is one more framework your clean, GHG-Protocol-based fact base maps into. The difference between a number that survives the request list and one that does not is whether it was built inside the assurance perimeter from the start.
Skill.re