Documenting the Human Decision
The screening tool flagged the family at the eightieth percentile of its risk score. The intake supervisor, a fifteen-year veteran, read the referral, pulled the prior history, and made a series of phone calls. By the end of the afternoon she had decided not to assign the case for investigation. The grandmother who had called was anxious but not reporting abuse; the prior reports were old and unsubstantiated; the school had no concerns; the score, she judged, was driven by poverty markers the model treated as risk. She closed the referral with a one-line note: "Reviewed, screened out." Six weeks later the family came to attention again under different and more serious circumstances, and a review was opened into the earlier decision. The reviewers found a high risk score and a screen-out, and almost nothing in between. The supervisor's actual reasoning, the calls she made, the factors she weighed, her assessment that the score reflected poverty rather than danger, none of it was in the record. To the reviewers, and to anyone who came after, it looked like a human had glanced at a high-risk flag and waved the family through. The decision may have been entirely sound. It was undocumented, and so it was indefensible.
Why the Human Decision Itself Must Be Documented
This program has built, carefully, toward a single hard-won principle: an AI risk signal is one audited input under mandatory human review, never a verdict. You have learned to read a signal as one input, to put a human between the signal and any action, and to audit the screening tool for the bias these tools are known to encode. This lesson closes the loop on the second deep use case by addressing what happens after the human reviews and decides: the decision has to be documented, and documented in a specific way, or every protection built upstream becomes invisible and therefore unprovable.
The trap is intuitive once you see it. We are used to documenting outcomes. The referral was screened in or screened out. The case was substantiated or not. The benefit was approved or denied. But when an AI signal is in the workflow, documenting only the outcome is precisely what makes the human review disappear. A record that shows a high risk score and a screen-out, with nothing in between, is indistinguishable from a record where the human rubber-stamped the model. The reasoning is the evidence that a human decided. Without it, the decision reads as the score's decision, which is the exact failure the entire screening-support discipline exists to prevent.
Put the stakes in terms of the people involved. A decision to screen a child-welfare referral in or out determines whether an investigator shows up at a family's door. A decision to substantiate determines whether a parent carries a finding that can affect employment and custody for years. Each of these is bound by due process, which means the affected family has the right to understand and challenge the basis of the decision, and an oversight body has the authority to review it. When the basis of the decision is a human judgment that was never written down, the family cannot challenge what they cannot see, and the reviewer cannot distinguish careful judgment from negligence. Documenting the human decision is what makes the decision contestable, and contestability is what due process requires.
If you document only the outcome and not the reasoning, the record cannot tell the difference between a human who decided and a human who rubber-stamped the model.
What a Decision Record Must Contain
A defensible decision record for a screening-informed call is not a longer version of "reviewed, screened out." It is a structured account of how a human turned a set of inputs, one of which was an AI signal, into a judgment. It captures several distinct elements, each answering a question a reviewer or an advocate will eventually ask.
What the Signal Said, Stated Plainly
The record states what the AI tool produced: the score, the tier, or the flag, in the tool's own terms. This is not an endorsement of the signal; it is an honest statement of one input the human had. A record that hides the score is as indefensible as one that obeys it. If the supervisor in the opening scene had written "the tool placed this family at the eightieth percentile," she would have put the signal on the table where it could be weighed and, here, set aside for stated reasons.
What the Human Reviewed Beyond the Signal
The record names the other inputs the human gathered and weighed: the prior history, the calls to the school and the grandmother, the absence of current concerns, the specifics of the present referral. This is the heart of mandatory human review made visible. It is the proof that the human did not act on the signal alone but brought independent information and professional judgment to bear. In the opening case this is exactly what existed in the supervisor's afternoon and vanished from the record: the calls, the history, the school contact, all the work that made her decision a judgment rather than a reflex.
How the Human Weighed the Signal Against Everything Else
The record explains the reasoning that connected the inputs to the decision, and critically, how the human treated the AI signal within that reasoning. Did the signal corroborate independent concerns, or did it conflict with them? When the human's judgment diverges from the signal, as it did here, the record must say so and say why. "The tool scored this family high; I judged the score to reflect poverty markers rather than safety concerns, because the prior reports were old and unsubstantiated, the school reports no concerns, and the present referral describes anxiety rather than abuse" is a defensible divergence. It shows a professional engaging the signal critically, which is the posture the program demands, rather than ignoring it or obeying it.
The Decision and the Named Person Who Owns It
The record states the decision and the name of the human who made it. This is the cardinal rule made concrete: AI informs, a person decides, and the person has a name attached to the call. The accountability does not transfer to the tool. "The model said so" is never the reason; the reason is the named human's documented judgment, and the named human owns the outcome whichever way it goes.
Run the opening scene again with a real decision record. The supervisor writes: tool placed the family at the eightieth percentile; I reviewed three prior reports (all old, all unsubstantiated), spoke with the school (no concerns) and the reporting grandmother (anxious, no allegation of abuse); the present referral describes household stress and poverty, not maltreatment; I judge the elevated score to reflect poverty-correlated factors rather than child-safety risk; decision: screen out; decided by [name], [date]. Now when the review opens six weeks later, the reviewers see a documented professional judgment they can evaluate on its merits. They may agree or disagree, but they can no longer mistake a careful decision for a rubber stamp, and the supervisor's defensible reasoning is on the record where due process requires it to be.
The Divergence Case Is the One That Matters Most
There is a specific moment in screening-support work that deserves its own attention because it is where documentation most often fails and where it matters most: the case where the human's decision diverges from the AI signal. There are two directions of divergence, and both must be documented with care.
The first direction is the opening scene: the tool says high risk, the human decides not to act. The danger people feel here is that overriding a high score will look reckless if the case later goes wrong. That fear creates a perverse incentive to defer to the score, to screen in a family the worker would otherwise screen out, simply to be covered if the model turns out right. This is automation bias operating through fear, and it is corrosive: it lets the score quietly become the decision because no one wants to be the human who disagreed with it. The antidote is documentation. A worker who can write a clear, sourced rationale for diverging from the score is protected by that rationale and is free to exercise the judgment the cardinal rule requires. The documented reasoning is what makes it safe to be right when the model is wrong.
The second direction is subtler: the tool says low risk, and the human, on independent grounds, decides to act anyway, or the tool says high and the human agrees. Even agreement must be documented as a judgment, not a deferral. A record that says only "high score, screened in" in the case that goes well is building the same bad habit as "high score, screened out" in the case that goes badly. In both, the human's independent reasoning is absent, and the practice of treating the signal as one input among many erodes one undocumented decision at a time. The discipline is the same regardless of direction: the record shows the human reviewed independent information and reached a reasoned judgment, whether or not that judgment matched the model.
Why This Protects the Worker and the Family Together
It is worth naming that documenting the human decision protects two parties at once, which is unusual and is why it is worth the time. It protects the family, because the basis of a consequential decision becomes visible and challengeable, satisfying due process. And it protects the worker, because a decision made under caseload pressure with real professional judgment becomes defensible against a later review that has the unfair advantage of hindsight. A reviewer looking back at a case that went badly is tempted to read every prior decision as a mistake. A documented rationale written at the time, before the outcome was known, is the worker's evidence that the decision was reasonable on the information available. The supervisor in the opening scene may have made the right call; without the record she has no way to show it, and the review can only see the bad outcome and the high score she appeared to ignore.
Making the Decision Record Routine, Not Heroic
A decision record that depends on a worker writing a thoughtful paragraph at the end of every screening, on top of a crushing caseload, will be the first thing to go when the unit is short-staffed and the referrals are stacking up. The principle that runs through this whole program holds here too: a control that requires heroism will not survive contact with the real workload. The decision record has to be made routine by design.
Routine does not mean a free-text box that is usually left blank or filled with "reviewed." It means a structured prompt built into the screening workflow that asks the specific questions: what did the tool indicate, what else did you review, how did you weigh them, what did you decide, and if you diverged from the signal, why. A structured form lowers the effort because it converts a blank page into a set of answerable questions, and it raises the floor because it makes the absence of a rationale visible rather than letting "screened out" pass as a complete record. The structure is also what makes the records readable later in aggregate, so a supervisor or an equity auditor can review divergence patterns across a unit rather than parsing a hundred idiosyncratic notes.
This connects directly to the governance and audit-trail practice the program builds alongside it. The decision record is not a separate document floating free; it attaches to the case in the case-management system with a timestamp and the decider's name, captured as part of the same provenance trail that records how every AI-touched record was produced. Documenting how the signal was used, a practice introduced earlier in the program, reaches its full form here: the signal, the human review, the reasoning, the decision, and the accountable person are all logged together, so that the screening-support workflow can be reconstructed end to end for any case a court or an advocate asks about.
What Good and Bad Records Look Like Side by Side
The difference is concrete enough to see in a single comparison. A bad record: "High risk per tool. Screened out." A good record: "Tool flagged eightieth-percentile risk. Reviewed three prior referrals, all over four years old and unsubstantiated; contacted school (no current concerns) and reporting grandmother (anxiety, no abuse allegation); present referral describes housing instability and stress, not maltreatment. Judged the elevated score to reflect poverty-correlated factors the tool treats as risk rather than child-safety danger. Decision: screen out. [Name], [date]." The first record is four words and defends nothing. The second is four sentences and defends everything: it shows the signal honestly, shows the independent review, shows the reasoning for the divergence, names the decider, and would let a reviewer, an advocate, or the family understand and contest the call. Four sentences, written once at the moment of decision, are the difference between a defensible program and the indefensible gap that opened in the lesson's first paragraph.
Documenting the Decision Across the Other Screening Contexts
The child-welfare screen-out in the opening scene is the sharpest illustration, but the same discipline governs every place an AI signal informs a consequential call, and it is worth tracing it into the other contexts so the practice does not read as a single special case. The principle is identical; the inputs and the stakes shift.
Take an eligibility setting. An AI tool flags an application as a likely-fraud or high-error case and routes it for additional review, a pattern with a documented history of harm when it runs unchecked, as Michigan's MiDAS system showed when it issued tens of thousands of automated fraud determinations that were later found to be wrong at a staggering rate. An eligibility worker who reviews such a flag and decides whether to approve, deny, or request more documentation is making a screening-informed decision, and it must be documented the same way: the tool flagged this application; I reviewed the income verification, the household composition, and the prior case history; the discrepancy the tool detected is explained by a reported and verified change in employment; I find no basis for a fraud referral and approve the benefit; decided by [name]. Without that record, a denial that leaves a family without food traces back to a flag with no visible human reasoning, which is exactly the indefensible automated determination the field has learned to fear.
Take a resource-prioritization setting, where a tool ranks clients for a scarce service such as a housing slot. A case manager who moves a client up or down from the tool's ranking on the basis of independent knowledge, an acute safety concern the tool could not see, must document that judgment, because a person who was deprioritized has the same right to understand why. In each context the structured record answers the same questions: what the signal said, what the human independently reviewed, how the human weighed them, and who decided. The consistency is the point. A worker who internalizes the four-element record in one context carries it into all of them, and an agency that builds the structured prompt into one screening workflow can build it into the rest, so that the documented human decision becomes the uniform shape of accountable AI-assisted practice across the agency rather than a habit confined to child-welfare intake.
Key Takeaways
- When an AI signal is in the workflow, documenting only the outcome (screened in or out, substantiated or not) makes the human review disappear: a record showing a high score and a decision with nothing in between is indistinguishable from a rubber stamp. The reasoning is the evidence that a human actually decided.
- A defensible decision record contains four elements: what the signal said in the tool's own terms, what independent information the human reviewed beyond the signal, how the human weighed the signal against everything else, and the decision plus the name of the accountable human who made it.
- The divergence case, where the human's judgment differs from the AI signal, is where documentation matters most. A clear, sourced rationale for diverging is what makes it safe for a worker to override a wrong score, protecting them against fear-driven automation bias that would otherwise let the score become the decision.
- Even agreement with the signal must be documented as an independent judgment, not a deferral. "High score, screened in" with no reasoning erodes the one-input discipline just as much as an undocumented override does.
- Documenting the human decision protects the family and the worker at once: it makes the basis of a consequential decision visible and challengeable as due process requires, and it gives the worker contemporaneous evidence that the decision was reasonable on the information available, defending against hindsight review.
- The decision record must be made routine by design through a structured prompt built into the screening workflow, not a free-text box that gets filled with "reviewed." Structure lowers the worker's effort, raises the documentation floor, and makes divergence patterns readable across a unit.
- The decision record attaches to the case in the case-management system with a timestamp and the decider's name, joining the same audit trail that records how every AI-touched record was produced, so the full screening-support workflow can be reconstructed for any court or advocate.
- The contrast is concrete: "High risk per tool. Screened out." defends nothing; four sentences naming the signal, the independent review, the reasoning, and the decider defend everything, and they are written once at the moment of decision rather than reconstructed from memory after a case goes wrong.
Skill.re