โ†
AI for ESG & Sustainability Reporting
Visionary ยท M5 ยท lesson 5 of 16 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Identifying Novel, Defensible Use Cases
๐Ÿ“–
now learning

Identifying Novel, Defensible Use Cases

15 min

A vendor demo ends, the room is quiet, and your head of innovation turns to you with the sentence that decides the next two years: "We could point this at almost anything. Where do we point it first?" On the screen is a model that will happily draft a target, cluster a materiality matrix, estimate a Scope 3 category, or reconcile an emission factor. All of it looks equally impressive. None of it is equally defensible. Your job, as the transformer running an enterprise sustainability-data program, is to answer that question with a framework, not a hunch, because the wrong first bet does not just waste budget. It teaches your assurer that your AI produces numbers nobody can trace.

The Two Questions That Screen Every Use Case

Most AI opportunity lists rank use cases on a single axis: value. How much time will it save, how much cost will it strip out of the reporting cycle, how many analyst-weeks it returns. That axis is real, and at the enterprise level it is large. But ranking on value alone is exactly how a sustainability function ends up with a beautiful pilot that its external assurer refuses to rely on. In a regulated, externally-assured disclosure, a use case has to clear a second axis before the first one matters at all: can its output survive assurance?

So the screen is two questions, asked in order. First: does this create real value, meaning it attacks a genuine bottleneck in the reporting cycle, not a cosmetic one. Second, and this is the gate: can the output be made assurable, meaning every figure or claim it produces can be traced to evidence, labeled by method, and reconstructed by someone who was not in the room. A use case that scores high on value and low on assurability is not a good use case with a caveat. It is a liability with a demo attached.

The reason the order matters is the asymmetry that governs all disclosure. The report is easy to generate and hard to defend. AI collapses the generation cost toward zero, which makes the generation side of the ledger look spectacular and leaves the defense side exactly where it was. If you screen on value first and assurance second, you will systematically over-invest in the half of the problem that was never the constraint. Screen on assurability first, and value second, and you will invest where the two axes point the same direction, which is the only place an enterprise sustainability program should be spending.

There is a second reason the order is not merely tidy but load-bearing, and it is about what your assurer learns from your choices. Every workflow you deploy teaches the assurance partner something about how your AI behaves. A first bet that is fast but untraceable teaches them that your models emit numbers no one can source, and that lesson colors how they scope every future engagement: they test more, they rely less, they widen their sampling. A first bet that is fast and fully traceable teaches the opposite, that your AI produces evidence alongside its output, and that lesson buys you reliance you can spend for years. The order of the screen is therefore not just a spending discipline. It is a relationship strategy with the one external party who can stall your entire disclosure. Choosing the value-first path is choosing to spend your assurer's trust on the exact use cases most likely to lose it.

A use case that produces a number no one can trace is not innovation. It is a restatement you have not filed yet.

The Four Quadrants of the Screen

Plot value on one axis and assurability on the other and you get four quadrants. Naming them turns a vague "should we do this" conversation into a portfolio decision your governance body, your CFO, and your assurance partner can all read.

Fund now: high value, high assurability

These are use cases where AI attacks a real bottleneck and the output naturally carries its own evidence. Emission-factor lookup that returns a named, dated, authoritative source alongside every factor. Activity-data extraction that preserves the page and cell location of each field it pulled from an invoice or a supplier file. Supplier-questionnaire triage that flags which responses are primary and which are estimated. In each of these the assurable version is not more expensive than the fast version. Traceability is a property of how you build the workflow, not a tax you pay afterward. This quadrant is where your first enterprise bets belong.

Redesign then fund: high value, low assurability as pitched

This is the most dangerous quadrant because it is the most seductive. The use case attacks a genuine, expensive bottleneck, so the value is real, but the way it is being pitched produces output an assurer cannot rely on. AI "filling the gap" in a Scope 3 category with an industry average and presenting the result as measured activity data is the archetype. The bottleneck is real: Scope 3 averages roughly 75% of a footprint, and 79% of reporters cite supplier-data availability as a top barrier. The value is enormous. But the pitched version launders an estimate into a measured number, which is the single fastest way to fail an assurance engagement. These use cases are not rejected. They are sent back to be redesigned so the estimate is labeled, method-tagged, and disclosed with its uncertainty. Redesigned, they move into "fund now." Funded as pitched, they become the restatement.

Park: low value, high assurability

Perfectly traceable, perfectly assurable, and not worth doing. AI that reformats a table you could reformat by hand, or that answers a question no one in the cycle was blocked on. There is nothing wrong with these except that they consume the scarce thing an enterprise transformation runs on, which is credibility and attention. Park them. Do not let a governance body spend a meeting approving something safe and pointless while the value-and-assurability wins wait.

Refuse: low value, low assurability

The easy ones. Fully generative narrative that invents claims to fill a section, "AI sets your targets," anything where the output is both hard to defend and not solving a real constraint. These are not portfolio decisions. They are declines.

What Makes an Output Assurable

The screen only works if you can judge the assurability axis honestly, and that requires a concrete definition rather than a feeling. An output is assurable when it satisfies four properties, and you can test any proposed use case against them before a single euro is spent.

Provenance. Every figure traces to a named, dated, authoritative source. An emission factor points to a specific database and version. An activity figure points to a specific document, page, and field. "The model produced it" is not provenance. If the use case cannot attach a source to each number it emits, it fails here.

Method transparency. The output declares how each number was derived: measured, supplier-reported primary data, or estimated by a stated method (spend-based, average-data) with its assumptions visible. A primary figure and an estimate must never look identical in the file. If the use case blurs measured and estimated into one indistinguishable stream, it fails here even if every input was real.

Reconstructability. Someone who was not in the room, using only the file, can rebuild the number from raw data to published figure. This is the assurer's actual test under a limited or reasonable assurance engagement. If reproducing the output requires the original analyst to explain what the model "was thinking," it fails here.

Human accountability. A named person reviewed and signed the output, and the file records that they did. AI assists, the human decides, the file proves it. A use case that removes the human sign-off in the name of efficiency has not become more efficient. It has become indefensible, because "the model recommended it" is never a defense to an assurer or a regulator.

Run a candidate use case against these four and the assurability axis stops being subjective. A use case that can satisfy all four, cheaply and by design, is high on the axis. One that can satisfy them only with heroic manual effort bolted on afterward is low, and one that cannot satisfy them at all belongs in the refuse quadrant no matter how large the value looks.

A practical way to score the axis without turning it into a debate is to rate each of the four properties as one of three states: satisfied by design, satisfiable with rework, or not satisfiable at all. Then apply two simple rules. Any property that lands on "not satisfiable" caps the whole use case at refuse, because a workflow that structurally cannot produce provenance, or structurally cannot record a human sign-off, is indefensible no matter how strong the other three are. Any property that lands on "satisfiable with rework" caps the use case at redesign then fund, because the value may be real but the workflow as pitched is not yet assurable. Only a case where all four are satisfied by design earns a place in fund now. This turns a vague conversation into a scored decision that a governance body and an assurer can both audit, and it removes the temptation to average a fatal flaw away against three strong properties. In disclosure, assurability is not an average. It is a set of gates, and a single failed gate is a fail.

Defensible and Indefensible, Side by Side

Abstractions are easy to nod along to and hard to apply under deadline pressure, so hold two concrete pairs in your head. In each pair the value is comparable. The difference is entirely on the assurability axis, which is the axis that decides.

Pair one: the emission factor

Defensible: an AI-assisted factor lookup that, when you give it an activity and a region, returns three candidate factors, each with its source database, publication year, and the exact category match, and refuses to answer when it cannot find a real one. The analyst chooses, the file records the choice and the source. Indefensible: the same convenience, except the model produces a single "best" factor as a bare number because that is faster and cleaner in the report. It is faster. It is also a hallucinated-factor incident waiting to be discovered, because a plausible but invented factor silently corrupts every downstream calculation and there is nothing in the file to catch it.

Pair two: the ESRS narrative datapoint

Defensible: AI drafts a narrative datapoint from a supplied evidence pack, and every quantitative claim in the draft carries a footnote pointing at the source figure, so the reviewer can check each claim against the file before it ships. Indefensible: AI drafts the same narrative from a thin prompt, produces fluent prose about progress against a target, and in doing so softens a negative impact the company is obligated to disclose or references a target the company never actually set. The prose is better. It is also a greenwashing finding and, potentially, a material misstatement, because the model optimized for a smooth narrative and the smooth narrative was not true.

Notice the pattern across both pairs. The indefensible version is not sloppy or low-quality. It is often more polished. That is precisely why it is dangerous at enterprise scale: the failure mode is fluent, confident, and invisible until an assurer pulls the thread. The screen exists to catch it at the use-case selection stage, before it is built, deployed, and relied upon across the report.

A Worked Example: Screening a Portfolio

Picture the innovation intake for a large undertaking still in CSRD scope after the Omnibus, more than 1,000 employees and more than EUR 450M turnover, obtaining limited assurance on its sustainability statement and trending toward reasonable. The head of sustainability data has six candidate use cases on the table and one quarter of budget. Watch the screen do its work.

Candidate A: AI-assisted emission-factor lookup with provenance. Value: high, it touches every calculation in the inventory. Assurability: high, provenance and method transparency are native to how it is built. Verdict: fund now. This is a textbook quadrant-one bet and a strong first move because it strengthens the assurance posture of everything downstream.

Candidate B: AI that estimates the missing Scope 3 categories to "complete" the inventory. Value: very high, it attacks the largest and hardest part of the footprint. Assurability as pitched: low, because it presents estimates as if they were measured data. Verdict: redesign then fund. Send it back with a hard requirement that every estimate is labeled, method-tagged, and disclosed with its uncertainty, and that estimated categories are visibly distinct from measured ones. Redesigned, it becomes one of the highest-value bets you own. Funded as pitched, it is the restatement.

Candidate C: fully generative drafting of the strategy narrative from a one-line brief. Value: moderate, it saves some writing time. Assurability: low, it invents claims and cannot trace them. Verdict: refuse, or narrow it to drafting from a supplied evidence pack with per-claim footnotes, which turns it into a different, defensible use case.

Candidate D: AI that reformats the prior-year datapoints into the current template. Value: low, it is a small clerical task. Assurability: high, it is fully traceable. Verdict: park. Safe, tidy, and not worth a governance slot this quarter.

Candidate E: supplier-response parsing that tags each datapoint primary or secondary and preserves the source location. Value: high, it attacks the 79% supplier-data bottleneck. Assurability: high, provenance and method transparency are built in. Verdict: fund now.

Candidate F: an AI that drafts and "auto-approves" narrative to hit the filing deadline without a reviewer. Value: high on speed. Assurability: it removes human accountability entirely. Verdict: refuse outright. There is no redesign that saves a use case whose whole premise is deleting the sign-off.

The portfolio that emerges is not the one that ranked highest on value. It is A and E funded now, B sent back for a redesign that will make it fundable and formidable, C narrowed, D parked, F declined. That is a defensible innovation agenda: it spends where value and assurability agree, it rescues the high-value case by fixing its assurability rather than abandoning the value, and it refuses the cases that would teach your assurer to distrust your AI. Presented to a governance body and an assurance partner, that portfolio is a credential. Presented to the same room, a value-only ranking that funded B-as-pitched and F is the beginning of a very bad engagement.

Running the Screen as an Enterprise Discipline

At the transformer level the screen is not a one-off exercise. It is a standing intake control, because novel use cases will keep arriving from vendors, from analysts, from a board that read the same AI headlines you did. Three practices keep the screen honest over time.

First, put the assurance liaison in the room when use cases are screened, not after they are built. The cheapest moment to discover that an output is indefensible is before it exists. An assurer who sees the intake understands your controls and is far more comfortable relying on the workflows that pass. Second, require every intake to state its assurability case in the four-property language, provenance, method transparency, reconstructability, human accountability, so the axis is argued in the assurer's own terms rather than in marketing terms. A use case that cannot describe how it satisfies the four properties has failed the screen by failing to answer it. Third, keep a written record of what you refused and why. When the board asks why you did not chase the flashy generative use case a competitor announced, the record is your answer, and it is a strong one: you declined it because it could not be made assurable, and in a regulated disclosure an unsupported number is not efficiency, it is a misstatement.

Done this way, the search for novel use cases stops being a race to adopt whatever the newest model can do and becomes what an enterprise sustainability program actually needs: a disciplined way to find the places where AI makes the disclosure both faster and more defensible, and to refuse, in writing, the places where it would only make it faster.

Key Takeaways

  • Screen every novel use case on two axes in order: does it create real value, and can its output survive assurance. Assurability is the gate; value only matters once the gate is passed.
  • The four quadrants are fund now (high value, high assurability), redesign then fund (high value, low assurability as pitched), park (low value, high assurability), and refuse (low value, low assurability).
  • The most dangerous quadrant is the seductive one: high-value cases pitched in an indefensible way, such as AI filling a Scope 3 gap with an average presented as measured data. Redesign them to label the estimate; never fund them as pitched.
  • An output is assurable only if it has provenance, method transparency, reconstructability, and human accountability. Test any candidate against these four before spending.
  • The indefensible version of a use case is usually more polished, not sloppier. Fluent, confident, and untraceable is the failure mode that ends careers, so the screen must catch it at selection.
  • Traceability is a property of how a workflow is built, not a tax added afterward. The best first bets are ones where the assurable version costs no more than the fast version.
  • Run the screen as a standing intake control: assurance liaison in the room, every case argued in the four-property language, and a written record of what you refused and why.
  • "The model recommended it" is never a defense to an assurer or a regulator. A use case that removes the human sign-off has not become efficient. It has become indefensible.