AI for ESG & Sustainability Reporting
Capable · M13 · lesson 13 of 24 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Keeping the Analyst Accountable
📖
now learning

Keeping the Analyst Accountable

15 min

The assurer is in the room, and she has one question that the whole supplier pipeline has been building toward: "Walk me through this Scope 3 figure. Who collected the underlying data, which numbers are measured and which are estimated, and what did you change from the raw supplier responses, and why?" The analyst opens the file. If the file answers in minutes, with names, tiers, and a clear record of every override and its reason, the engagement moves on. If the file is a folder of spreadsheets and AI outputs with no record of who decided what, the analyst is now reconstructing months of work from memory under questioning, and every gap in his memory is a gap in the assurance. This lesson is about the documentation that makes the first outcome happen instead of the second. Across the supplier pipeline, drafting, parsing, gap-flagging, the AI assisted at every step, but the accountability never moved. The analyst signs, the AI assists, and the file proves it. That sentence is the entire chapter, and this lesson is how you make the file actually prove it.

Accountability Does Not Transfer to the AI

The single most important fact about using AI in a disclosure pipeline is that accountability for the disclosed figure stays with the human, completely and without exception. AI drafted the questionnaire, but the analyst chose what to ask. AI parsed the responses, but the analyst confirmed each datapoint against its source. AI flagged the gaps, but the analyst decided how each was handled. At no point did responsibility for the resulting number move to the model, and "the AI estimated it" or "the model recommended it" is not a defense to an assurer or a regulator. It never has been and it never will be. The model is a tool the analyst operated, exactly as a calculator is, and the analyst owns the output the way he would own a figure he typed himself.

This matters in practice because it changes what the file has to contain. If accountability stayed human, the file must show the human's decisions, not just the AI's outputs. An AI output sitting in a folder is not documentation of a decision; it is a draft the analyst either accepted, rejected, or modified, and the act of accepting, rejecting, or modifying is the thing that needs recording. The assurer is not assessing whether the AI is good. The assurer is assessing whether the analyst exercised competent judgment over the AI, and that judgment is only visible if it was written down. A pipeline that produces beautiful AI outputs and records none of the human decisions over them has automated the work and erased the accountability trail, which is the worst of both worlds: the speed of AI with none of the defensibility.

What the File Must Prove

The assurance file for the supplier pipeline has a specific job: it must let someone reconstruct how each figure came to be, including the human judgment at every step, without the analyst in the room. That is the reconstructability standard, and it is the test an assurer applies. To meet it across the supplier pipeline, the file has to answer a handful of questions for any datapoint the assurer might sample.

Who Collected What

For each datapoint, the file should show its origin: which supplier it came from, who on the team collected and confirmed it, and when. This is the provenance from the parsing lesson, now extended to include the human who handled it. The assurer asks "where did this come from" and the file answers with both the source and the person, so a question about any number routes to a named owner rather than a shrug.

Which Figure Is Primary, Which Estimated

For each datapoint, the file should carry its tier, primary or secondary, with the basis, exactly as the parsing and gap lessons required. The accountability addition is that the file shows the analyst confirmed the tier, not just that the AI proposed it. When the assurer asks "how much of this category is primary," the file answers with a tier on every datapoint that a human verified, so the primary-data share is a confirmed figure rather than an AI guess nobody checked.

What Was Overridden, and Why

This is the field most often missing and most important. Wherever the analyst changed an AI output, rejected a flagged value, accepted a partial after follow-up, chose one estimation method over another, or overrode any AI suggestion, the file should record what was changed and the reason. An override with no reason is a decision the assurer cannot evaluate; an override with a clear reason is competent judgment on display. The override log is where the analyst's accountability becomes legible, because it is the concentrated record of every place a human exercised judgment over the machine. A pipeline with no override log is claiming either that the AI was never wrong or that nobody checked, and neither claim survives an assurer.

The analyst signs, the AI assists, the file proves it. An AI output is not documentation; the human decision over it is. A pipeline that records the outputs and not the decisions has automated the work and erased the accountability.

The Anatomy of an Override Entry

Because the override log carries so much weight, it is worth being precise about what a single good entry contains, since a log of vague entries is barely better than no log at all. A defensible override entry has five parts, and each answers a question the assurer will ask. The what: the specific change made, stated concretely, for example "supplier figure of 12,000 tCO2e accepted despite being flagged as an order-of-magnitude increase over prior year." The why: the reason for the decision, the actual judgment, for example "supplier confirmed a genuine increase driven by a new production line commissioned mid-year." The evidence: a pointer to what supports the reason, for example "supplier email of this date and the production-line capacity figures attached," so the reason is not merely asserted but backed. The who: the named analyst who made and owns the decision, so the override has a responsible human attached to it. And the when: the date, which lets the assurer confirm the decision was contemporaneous with the work rather than reconstructed later. Strip any of the five and the entry weakens: a why with no evidence is an assertion, a what with no why is an unexplained change, and an entry with no named analyst is an orphaned decision nobody will stand behind under questioning.

The same five-part discipline applies whether the override is an accepted implausible value, a rejected AI extraction, a partial accepted after follow-up, a tier reclassified from the AI's proposal, or an estimation method chosen over an alternative. Each is a place where the human stepped in and changed what the machine produced, and each deserves an entry that a stranger could read and understand without further explanation. The test for a good entry is the same reconstructability test that governs the whole file: could a competent colleague, reading only this entry, understand what was changed, why, on what evidence, and by whom, without asking you? If yes, the override is documented; if they would have to come and ask you, it is not yet documented, however much you remember about it.

The Signature Is Not a Formality

At the end of the pipeline the analyst signs off on the figures, and it is tempting to treat that signature as a closing formality, a box ticked so the report can move forward. It is not. The signature is the moment the analyst formally takes ownership of every figure and every judgment the file records, and it is meaningful precisely because the file behind it is complete. An analyst who signs a file with documented provenance, confirmed tiers, and a full override log is asserting something they can stand behind: that they exercised competent judgment over the AI at every step and that the record proves it. An analyst who signs a folder of unexamined AI outputs is asserting the same thing with nothing behind it, which is the position no professional wants to be in when the assurer starts sampling.

This is why the documentation discipline is not bureaucratic overhead but self-protection. The file the analyst builds as they work is the file that defends them when the figures are challenged, and the challenge can come long after the work, from an assurer in the engagement, a regulator reopening a filing, or a colleague inheriting the inventory next year. In every case the analyst's defense is the same: not their memory, which fades, and not the AI, which cannot be held accountable, but the file, which holds. The signature means "I take responsibility for this," and the file is what makes that responsibility a defensible position rather than an exposed one. The two go together: a signature with no file is a liability, and a file with no signature is incomplete, because somewhere a named human has to take ownership of the whole.

Document as You Go, Not in a Panic at the End

The most common reason accountability documentation fails is timing. The analyst means to write it up, but does the work first and the documentation last, and by the time the assurer is scheduled the decisions are months old and half-remembered. Reconstructing an override log from memory is slow, incomplete, and exactly the panic the file is supposed to prevent. The discipline is to capture the decision at the moment it is made, when the reason is fresh and obvious, so the file assembles itself as a byproduct of the work rather than as a separate project at the end.

This is also where AI can help with accountability rather than threaten it, if pointed correctly. The same tools that draft and parse can maintain a running log: when the analyst confirms a datapoint, the system records who and when; when the analyst overrides an AI suggestion, the system prompts for the reason and stores it alongside the change; when an estimate is applied to a gap, the system logs the method, the uncertainty, and the analyst who approved it. AI used this way becomes the scribe that keeps the accountability trail current, turning the documentation from a dreaded end-of-cycle reconstruction into an automatic, contemporaneous record. The caution is that the log records human decisions, not AI actions presented as decisions: the entry "analyst confirmed against source" is accountability; the entry "AI extracted value" is not, because the latter records what the tool did, not who took responsibility for it.

A Worked Example: Two Files, One Assurer

Two analysts ran the same AI-assisted supplier pipeline for the same Scope 3 category. The assurer samples one datapoint from each: a supplier figure that was originally flagged as implausible, then overridden and accepted. Watch the two files answer.

Before (the undocumented pipeline): The first analyst used AI to draft, parse, and flag, and accepted the AI's outputs efficiently. The flagged implausible value was investigated and accepted, but nothing was written down about why. The file is a folder: the supplier responses, the AI's parsed records, the AI's gap flags, and the final inventory. When the assurer asks "this value was flagged as implausible, why was it accepted," the analyst tries to remember. He thinks the supplier confirmed it was a real increase due to a new product line, but he is not certain, the email is somewhere, and he cannot point to where the decision was recorded because it was not. The assurer cannot evaluate a decision that left no trace, so it becomes a finding: an unsupported override on a material datapoint. The AI made the work fast; the missing documentation made it indefensible.

After (the accountable pipeline): The second analyst ran the same tools but logged decisions as she went. The same value was flagged as implausible; she queried the supplier, who confirmed a genuine increase from a new production line and provided supporting figures. She recorded the override in the log: value flagged implausible against prior year, supplier queried on this date, confirmed as a real increase from new product line, supporting evidence attached, override approved by named analyst. When the assurer asks the same question, she opens the override log, shows the entry, shows the supplier confirmation and the supporting evidence, and the datapoint stands as a documented, competently judged exception. Same AI, same flagged value, same final number. One file produced a finding; the other produced a one-minute answer and a clean sample.

The difference was not diligence in the work, both analysts investigated the value, but diligence in the record. The first analyst's judgment was invisible because it was never written; the second analyst's judgment was legible because the override and its reason were captured at the moment of the decision. Accountability that is not documented is, to an assurer, accountability that did not happen.

Working Rules for Keeping the Analyst Accountable

A handful of rules make the file prove what it needs to prove. Hold that accountability for every disclosed figure stays human and never transfers to the AI, because "the model recommended it" is not a defense to an assurer or a regulator. Document the human decision, not the AI output, because an AI output is a draft and the analyst's acceptance, rejection, or modification of it is the thing that needs recording. Make the file answer, for any sampled datapoint, who collected it, which tier it is and that a human confirmed the tier, and what was overridden and why. Keep an override log as the concentrated record of every place the analyst exercised judgment over the machine, because an override with no reason is a decision the assurer cannot evaluate. Capture each decision at the moment it is made, when the reason is fresh, so the file assembles itself as a byproduct of the work rather than a panicked reconstruction at the end. Use AI as the scribe that maintains the running log, recording who confirmed what and prompting for the reason on every override, so the accountability trail stays current. Keep the distinction sharp between logging a human decision and logging an AI action, because only the former is accountability. Follow these and the file answers the assurer's hardest question in minutes and the analyst signs with confidence. Ignore them and you have a fast pipeline that produces a beautiful number nobody can defend, which in disclosure is not a saving but a liability.

Key Takeaways

  • Accountability for every disclosed figure stays with the human and never transfers to the AI; "the AI estimated it" or "the model recommended it" is not a defense to an assurer or a regulator.
  • Because accountability stayed human, the file must show the human's decisions, not just the AI's outputs; an AI output is a draft, and the analyst's acceptance, rejection, or modification of it is the thing that needs recording.
  • The assurer is not assessing whether the AI is good but whether the analyst exercised competent judgment over it, and that judgment is only visible if it was written down.
  • The file must meet the reconstructability standard: for any sampled datapoint it answers who collected it, which tier it is with a human-confirmed basis, and what was overridden and why, all without the analyst in the room.
  • The override log is the most important and most often missing record, the concentrated trace of every place a human exercised judgment over the machine; an override with no reason is a decision the assurer cannot evaluate.
  • Capture each decision at the moment it is made, when the reason is fresh, so the file assembles itself as a byproduct of the work rather than a slow, incomplete reconstruction from memory at the end.
  • AI can be the scribe that keeps the accountability trail current, recording who confirmed what and prompting for the reason on every override, but the log must record human decisions, not AI actions dressed up as decisions.
  • Accountability that is not documented is, to an assurer, accountability that did not happen; a fast pipeline that records outputs and not decisions has the speed of AI with none of the defensibility, which in disclosure is a liability, not a saving.