Designing the Analyst-AI Handoff
An assurer flips to a Scope 3 line in the working file and asks the carbon accountant a question that sounds simple and is not: "The AI suggested this category mapping and this factor. Where, exactly, did you take over?" The accountant knows she took over. She reviewed it, she changed two things, she signed off. But the file does not say so. There is a clean final number and a vague memory of having checked it, and between the model's output and the published figure there is a fog. That fog is the problem this lesson removes. In a regulated disclosure, it is not enough that a human took over from the AI. The file has to show exactly where the human took over, what the human decided, and that the human, not the model, owns the result. The artifact that makes that visible is the handoff log, and it is the difference between "we reviewed it" and "here is the moment we did, who did it, and what changed."
Why the Handoff, Not the Output, Is the Whole Game
The iron rule of the program is that every figure you publish must trace to evidence, and "the AI estimated it" is not evidence. That rule has a structural consequence most teams discover too late: the moment a figure passes from a model to a human is the single most important event in the whole workflow, because it is the moment accountability transfers. Before the handoff, a number is a candidate, a suggestion, a draft with no standing in the disclosure. After the handoff, it is the company's figure, owned by a named person, defensible to an assurer. The handoff is where a suggestion becomes a representation.
And yet the handoff is almost always invisible. The model produces output, a human looks at it, the human pastes it into the workpaper, and the only trace left behind is the final number, which looks identical whether it was checked with care or waved through in a hurry. The assurer cannot see the review; they can only see the result. So when they ask "where did you take over," a vague answer is not just unsatisfying, it is an assurance weakness: it suggests the team cannot distinguish a verified number from an unverified one, which is precisely the distinction the engagement exists to test.
Designing the handoff means refusing to let that moment stay invisible. It means deciding, in advance and on purpose, exactly where in the workflow the AI stops being trusted to proceed and a human must decide, defining what that human must decide, and recording the decision so the file shows the transfer of accountability as a concrete, dated, attributed event. The handoff stops being a fog and becomes a line you can point to.
Drawing the Judgment Boundary
Every AI-assisted reporting workflow has a judgment boundary: the line at which the work stops being something a model can safely carry forward on its own and becomes something a human must decide. The boundary is not a single point in the whole report; it recurs at every stage where AI output feeds an assured number or a disclosed judgment. Drawing it well is the core design skill.
The boundary falls wherever crossing it without a human would let an unsupported or unowned result enter the disclosure. Some boundaries are obvious. The model can extract a figure from an invoice, but a human must confirm it matches the source before it enters the inventory: the boundary is at "confirmed against source." The model can shortlist an emission factor, but a human must confirm it in the named database: the boundary is at "confirmed in source." The model can cluster stakeholder inputs, but a human must own the materiality conclusion: the boundary is at "the determination." The model can draft a narrative datapoint, but a human must verify every claim and confirm no impact was softened: the boundary is at "verified against evidence."
The pattern across all of them is the same. On the model's side of the boundary is acceleration: finding, drafting, shortlisting, extracting, clustering, work that produces candidates. On the human's side is judgment: confirming, deciding, owning, signing, work that produces representations. The design question for each stage is one sentence: what is the last thing the model is allowed to do, and what is the first thing a human must do before the result can move on. The answer is the boundary, and the handoff is what happens when you cross it.
Before the handoff a number is a suggestion the company owes no one. After the handoff it is a representation the company must defend. Designing the handoff is deciding, on purpose, where that change happens, and proving in the file that it did.
What the Human Must Decide at the Boundary
A handoff is not a glance. If the human's role at the boundary is undefined, the handoff is theater: the number passes through a person who adds nothing the assurer can rely on. So the second half of designing the handoff is specifying what the human must actually decide, concretely enough that doing it produces evidence.
At a verification boundary, the human must decide whether the AI output is correct against its source, and the decision is binary and recorded: confirmed, or not confirmed and corrected to X, with the source checked named. At a factor boundary, the human must decide that the factor is the right one, confirmed in the named database with its full provenance, and that decision carries the database, version, year, unit, and geography. At a materiality boundary, the human must decide which topics are material and why, and the decision is the documented basis the assurer will test. At a narrative boundary, the human must decide that every claim traces to evidence and that nothing was softened, omitted, or invented, and the decision is recorded as a sign-off with any changes noted.
The discipline that turns each of these from a vague "I checked it" into real evidence is to make the decision explicit, attributed, and consequential. Explicit: it is a stated conclusion, not an implied one. Attributed: a named person made it, not "the team." Consequential: the decision could have gone the other way, and when it did, the file shows the correction. A handoff where the human could only ever say "looks fine" and never "no, corrected" is not a real boundary; it is a rubber stamp, and an assurer can smell a rubber stamp from across the table.
There is a subtler trap here worth naming, because it catches careful people. The danger is not only the human who waves output through; it is the human who genuinely reviews but reviews the wrong thing. A reviewer who checks that an AI-suggested factor is in the right range, moves sensibly from last year, and carries plausible units has done real work and learned almost nothing the assurer cares about, because a hallucinated factor is engineered to pass exactly that check. The decision the boundary demands is not "is this plausible" but "did I confirm this in the source." Designing the handoff means specifying the second question, not the first, so that the human's effort lands where it produces evidence rather than where it produces reassurance. The whole point of defining what the human must decide is to stop a conscientious review from being aimed at the wrong target.
Defining the Handoff Log
The artifact that records the boundary crossing is the handoff log. It is the structured record, kept alongside the data, that captures each moment AI output became human-owned. It is what lets you answer the assurer's "where did you take over" with a row instead of a memory. A workable handoff log captures, for each handoff, a defined set of fields.
| Field | What it records | Why the assurer cares |
|---|---|---|
| Item | The specific datapoint, factor, figure, or claim handed off | Ties the handoff to a precise line in the disclosure, not the report in general. |
| AI output | What the model produced: the suggested value, factor, mapping, or draft text | Shows what was proposed before the human touched it, so the change is visible. |
| Source checked | The evidence the human verified against: document and location, or database, version, year, unit, geography | Proves the verification was against real evidence, not the model's own claim. |
| Human decision | Confirmed as-is, or corrected to a stated value, or rejected, with reasoning | This is the moment accountability transferred; it must be a real, consequential conclusion. |
| Decided by | The named person who made the decision | Accountability is human and individual; the model decided is never a defense. |
| Date | When the decision was made | Places the handoff in the timeline and supports reconstructability. |
| Status after | The label the item now carries: primary or secondary, estimated, verified, signed off | Ensures the downstream file knows the item's standing, so an estimate is never mistaken for measured data. |
Two design principles make the log assurable rather than decorative. First, it must preserve the AI output, not overwrite it. The value of the log is that it shows the before and the after, so the assurer can see what the human changed. A log that records only the final number has erased the very event it exists to capture. Second, it must be kept at the moment of the handoff, not reconstructed afterward. A log written after the fact is a memory wearing the costume of a record, and reconstructing it under engagement pressure is exactly the scramble the log exists to prevent. The cost of a row is a few seconds at the moment of decision; the cost of not having it is an afternoon of archaeology in front of an assurer.
A common objection is that this is too much overhead for a high-volume workflow: thousands of extracted figures cannot each carry a hand-written paragraph. The objection misreads what the log is. The log is not an essay; it is a structured record, and for the common case it can be terse. Item, AI output, source confirmed (a document reference or a database row), decision (confirmed), decider, date, status. Most rows are confirmations and take seconds. The rows that matter most, and where the few extra words belong, are the ones where the human disagreed with the model: a correction, a rejection, a flagged gap. Those are rare, they are the most valuable evidence in the file, and they deserve a sentence of reasoning. A handoff log is not heavy because it captures every decision; it is light because most decisions are routine and only the consequential ones need narrative. Designing it well means making the routine row almost free and the exceptional row genuinely informative.
It also matters where the log lives. A handoff log kept in a separate document that someone updates "later" is the reconstructed-memory failure waiting to happen. The log belongs with the data, attached to the workpaper or the record itself, so the act of moving a number forward and the act of logging its handoff are the same act, not two acts one of which gets skipped. The closer the log sits to the work, the more it behaves like a control and the less it behaves like paperwork. The goal is a workflow in which the handoff is recorded as a byproduct of doing the work correctly, rather than an extra chore bolted onto the end of it.
A Worked Example: The Handoff Made Visible
Watch a single Scope 3 line move through a workflow with no handoff log, then through one with it.
Without the log
The analyst asks the AI to map a purchased-goods spend line to a Scope 3 category and suggest a factor. The model returns "Category 1, purchased goods and services; spend-based factor 0.42 kg CO2e per euro." The analyst thinks it looks reasonable, notices the spend line is actually capital equipment so it belongs in Category 2, quietly changes the category, keeps the factor, and pastes the result into the inventory. The number is now in the file. Six months later the assurer asks: was this category mapping the model's or yours, and where did the factor come from. The analyst remembers changing the category but cannot prove it was her judgment rather than the model's; she half-remembers checking the factor but cannot say against what. The honest answer is "I reviewed it," and the assurer hears a number that might be the model's unverified output dressed as a human-owned figure. The thread is now a finding, not because the number is wrong, but because the handoff is unprovable.
With the log
Same start: the model returns "Category 1; spend-based factor 0.42 kg CO2e per euro." Now the analyst records a handoff row. Item: purchased-goods spend line, supplier X, capital equipment. AI output: Category 1, factor 0.42 kg CO2e per euro. Source checked: procurement record confirms the line is capital equipment, not consumables; factor confirmed in the named factor database, version 2026.1, reference year 2024, unit kg CO2e per euro, geography EU. Human decision: rejected the model's Category 1 mapping and corrected to Category 2, capital goods, on the basis of the procurement record; confirmed the factor as appropriate for the corrected category. Decided by: the analyst, named. Date: recorded. Status after: secondary data, spend-based estimate, verified and labeled.
When the assurer asks "where did you take over," the answer is a row. Here is what the model proposed, here is the source I checked, here is the call I made and why, here is my name and the date, here is the label the number now carries. The correction from Category 1 to Category 2 is not a vulnerability now; it is the strongest evidence in the file, because it proves a human exercised judgment the model got wrong. The same diligence that was invisible in the first version is the centerpiece of the second. Nothing about the analyst's actual work changed. What changed is that the handoff became a fact instead of a fog.
The lesson generalizes past Scope 3. Every place an AI output feeds an assured number, the handoff log turns the invisible act of taking over into a visible, dated, attributed event. It is the mechanism that lets a team capture AI's speed and still answer, instantly and from the record, the one question that decides every engagement: where did the human take over, and what did the human decide.
Designing Handoffs Into the Workflow, Not Onto It
The final discipline is to build handoffs into the workflow from the start rather than bolting documentation on at the end. A handoff log assembled the night before the assurer arrives is a tell; a handoff log that filled itself row by row as the work happened is a control. Designing handoffs in means, for each stage on your reporting-cycle map that is tagged AI-assisted-with-verification, writing down three things before any AI runs: where the boundary is, what the human must decide there, and which log row that decision produces. Then the workflow is built so that crossing the boundary without filling the row is not the easy path; ideally it is not possible.
This also clarifies who owns what. The handoff log makes the division of labor concrete and contractual: the AI assists, the human decides, and the file proves which is which. It protects the analyst as much as the assurer, because when a number is later questioned, the analyst can point to exactly what they decided and on what basis, rather than carrying a vague liability for everything the model touched. And it scales: a workflow whose handoffs are logged consistently produces an audit trail an assurer can reconstruct without anyone in the room, which is the standard the whole L3 capstone is built toward. The handoff log is not extra work layered on a good workflow. It is the spine of one.
Key Takeaways
- The handoff is where accountability transfers: before it, a number is a suggestion the company owes no one; after it, a representation the company must defend. It is the single most important event in an AI-assisted workflow.
- In a regulated disclosure it is not enough that a human took over; the file must show exactly where the human took over, what they decided, and that they, not the model, own the result.
- The judgment boundary is the line where AI work stops being safely carried forward and a human must decide. It recurs at every stage where AI output feeds an assured number or a disclosed judgment.
- What the human decides at the boundary must be explicit, attributed, and consequential: a stated conclusion, by a named person, that could have gone the other way and shows the correction when it did.
- The handoff log records each boundary crossing with the item, the AI output, the source checked, the human decision, the decider, the date, and the status after, so "where did you take over" is answered with a row, not a memory.
- The log must preserve the AI output (showing before and after) and be kept at the moment of the handoff, never reconstructed afterward, because a reconstructed log is a memory in the costume of a record.
- A correction logged at the boundary is not a vulnerability; it is the strongest evidence in the file, because it proves a human exercised judgment the model got wrong.
- Design handoffs into the workflow, not onto it: for each amber stage, define the boundary, what the human must decide, and the log row it produces, so the trail fills itself and the audit trail reconstructs without anyone in the room.
Skill.re