Mandatory Human Review
The screen showed a number. The intake worker had finished entering the referral, a call from a school counselor worried about a seven-year-old who had come to class hungry three days running, and the screening tool returned its output: a risk score in the high band, with a recommendation to assign for investigation. It was 4:50 PM on a Friday. The worker had nine other referrals waiting, a supervisor who measured the unit by how fast the queue cleared, and a tool that had just handed her a clean, confident answer. The fastest thing she could do was accept the score, route the case, and go home. The discipline that this lesson is about is the set of structures that make the fastest thing not the thing that happens. Because the score is not a decision. It is an input. And between that input and any action that touches a family, there must be a person who looks, who weighs, who can disagree, and who signs their name to the call. That person is the point of mandatory human review, and the entire reason it exists is that under exactly the pressure of that Friday afternoon, it is the first thing that gets skipped.
What Mandatory Human Review Actually Is
Mandatory human review is the non-negotiable process that places a qualified person between any AI risk signal and any action taken on a case. It is not a suggestion to be careful. It is not a culture of professionalism. It is a defined, required step in the workflow, one that a case cannot pass through without a human looking at the signal, weighing it against the full picture, reaching an independent judgment, and recording that they did so. The word that matters most in the phrase is "mandatory." A review that happens when the worker has time, or when the case feels important, or when the supervisor remembers to ask, is not mandatory review. It is optional review, and optional review under caseload pressure becomes no review.
This connects directly to the cardinal rule of the field: AI informs, humans decide. A risk-screening tool in child welfare, or a fraud-detection flag in benefits, or any predictive model that surfaces a signal about a person, produces an output that looks like a conclusion. It is not a conclusion. It is one piece of information, generated by a system that learned patterns from historical data, that a human must integrate with everything else known about the case before any consequential step is taken. CPS (child protective services) decisions to investigate, to substantiate, or to remove are among the most consequential a government makes, bound by due process. The same is true of a decision to deny SNAP (the Supplemental Nutrition Assistance Program, federal food assistance), to terminate Medicaid, or to sanction a TANF (Temporary Assistance for Needy Families) recipient. None of these may rest on a model's output. The model can point. Only a person can decide.
Consider what the alternative looks like in practice. A unit receives 40 referrals a week. A screening tool scores each one. If the high-band scores are routed to investigation automatically, with the human reviewing only the ones that "look wrong," then for the cases that look right the model has effectively made the decision. The worker who clicks accept on 35 of 40 scored referrals in a shift is not reviewing. They are ratifying. Mandatory human review is the design that prevents ratification from masquerading as judgment. It requires that the human do the work of judgment on every signal, not only on the ones that draw attention.
A review you perform only when you have time is not mandatory review. It is optional review, and optional review under caseload pressure becomes no review at all.
Why the Pressure Makes It Fail
The reason mandatory human review needs to be engineered, rather than simply expected, is that every force in a caseworker's day pushes against it. Understanding those forces is the first step in building a process that survives them.
The first force is volume. A worker carrying 25 to 30 families, or an intake worker processing dozens of referrals a week, does not have unlimited attention. Each genuine review of a risk signal, done properly, takes time: reading the full referral, pulling the case history, weighing the score against what the record actually shows, and forming an independent judgment. If that takes 20 minutes per case and a worker has 15 scored cases in a day, that is five hours of review on top of every other task. When the day does not contain five extra hours, something gives. What gives is the depth of the review.
The second force is the authority of the number. A risk score arrives looking objective, quantitative, and confident. It says 8.4, or "high risk," or "92nd percentile." Against a number that precise, a worker's hesitation can feel like second-guessing the science. This is automation bias: the documented human tendency to over-trust the output of an automated system and to under-weight one's own judgment when it conflicts with the machine. Automation bias is strongest exactly when a worker is tired, rushed, or unsure, which describes a large share of a caseworker's week. The number does not announce that it is one input among many. It presents as an answer, and a tired person reaches for answers.
The third force is incentive misalignment. If a unit is measured on how quickly it clears the queue, and accepting the model's recommendation is faster than overriding it, then the metric rewards ratification and punishes the careful override that the work requires. A worker who slows down to genuinely review every signal will look slower than a colleague who accepts the score and moves on. Over a quarter, the careful worker's numbers look worse. Unless the agency measures the quality of review rather than the speed of disposition, the incentive structure quietly dismantles the very process the agency claims to require.
The fourth force is diffusion of responsibility. When a model produces the signal and the worker passes it along, it can feel as though no single person made the call. The worker thinks the model decided. The supervisor thinks the worker reviewed. The system thinks a human was in the loop. If the process is not designed so that one named person owns the judgment and records it, accountability evaporates into the seams between people and the machine. A decision that everyone touched and no one owned is the precondition for the kind of harm that no one can later explain.
What a Real Review Requires
If mandatory human review is to be real and not a rubber stamp, it has to require specific work. A review that consists of looking at the score and clicking accept is not review. Here is what a substantive review of a risk signal actually involves, step by step.
Read the Whole Case, Not the Score
The reviewer reads the full referral and the relevant case history before looking at the score, or at minimum integrates the score into a full reading rather than letting it frame everything else. The order matters. A worker who sees "high risk" first reads the entire case through that lens, noticing the facts that confirm the score and discounting the ones that do not. This is confirmation bias compounding automation bias. The discipline is to understand the case on its own terms, then ask what the signal adds, rather than treating the signal as the headline and the case as supporting detail.
Weigh the Signal Against the Record
The reviewer asks a specific question: does the signal correspond to what the actual record shows, or does it conflict with it? A high score on a referral that, on full reading, describes a one-time event with no corroborating history is a signal that conflicts with the record, and the conflict is information. A model score is built on statistical patterns across a population. The case in front of the worker is a specific family. When the population pattern and the individual facts diverge, the worker's job is to notice the divergence and to weight the specific facts, because the family will live with the decision, not with the average.
Reach an Independent Judgment
The reviewer forms a judgment that they would defend on its own terms, in a court or to an advocate, without reference to the score. The test is simple and demanding: if the score disappeared, would the decision still stand on the facts? If the answer is no, if the only reason for the action is the number, then the action rests on the model, and the cardinal rule has been broken. The score may legitimately raise the worker's attention, prompt a closer look, or surface a pattern worth checking. It may not, by itself, justify a consequential step.
Record the Review and the Reasoning
The reviewer documents that the review happened, what the signal said, what the worker found, and why the decision was made. This is not bureaucratic overhead. It is the only proof that a human was actually in the loop, and it is what makes the decision defensible to a court, an advocate, or an oversight body later. A decision recorded as "screened high, assigned for investigation" documents ratification. A decision recorded as "screening tool returned high band; on review, the referral describes a single missed-meal episode with a plausible explanation and no prior history; assigned for a welfare check rather than full investigation based on the specific facts" documents judgment. The next lesson in this chapter develops this documentation discipline in full; here the point is that the recording is part of the review, not separate from it.
Walk these four steps against the Friday-afternoon intake. A real review of that hungry-child referral would read the full school report, check whether the family has any prior history, ask whether a single hard week explains the hunger, and decide on the specific facts whether the right next step is an investigation, a lighter-touch welfare check, or a referral to food assistance. The score said high. The judgment might still land on investigation, or it might land somewhere else. What mandatory human review guarantees is that a person made that call on the facts, and can say why.
Building It So It Cannot Be Skipped
Knowing what a real review requires is not enough, because the pressures described above will erode any process that depends on individual willpower. The review has to be engineered into the workflow so that skipping it is harder than doing it. This is a design problem, not a motivation problem.
Make the review a required gate, not an optional step. The case-management workflow should not allow a scored case to advance to a consequential action without a recorded review. If the system lets a worker route a high-band referral to investigation without capturing the reviewer's independent reasoning, the system has made review optional and the agency should not be surprised when it is skipped. A required field that asks "what did your independent review find, and how does it relate to the signal?" turns the review from a hope into a structural condition of moving forward.
Separate the surfacing of the signal from the taking of the action. A design that shows the worker the full case first, then reveals the score, then requires a judgment, fights automation bias by changing the order of attention. A design that leads with a large "HIGH RISK" banner and a one-click "accept and assign" button engineers ratification. The interface is not neutral; it either supports judgment or undermines it.
Build in the time. If a real review takes 20 minutes and the workload assumes two, the process is set up to fail and the failure is the agency's, not the worker's. An agency that mandates human review without staffing for the hours it requires has mandated a fiction. The honest version either reduces the volume routed for human review, adds the staff to review it properly, or accepts that the review will be shallow and stops calling it mandatory. The genuine time that AI returns elsewhere in the workflow, by drafting notes faster, for example, is one place those review hours can come from, but only if the agency deliberately routes the saved time to review rather than to a higher caseload.
Measure the quality of review, not the speed of disposition. If the only metric is queue-clearance speed, the agency is paying people to ratify. Supervisory review should sample scored cases and ask whether the recorded reasoning shows genuine independent judgment or rubber-stamping. When a unit's records consistently read "screened high, assigned," that pattern is itself a finding: the review has collapsed into ratification, and the agency is operating a system where the model decides while a human's name is on the file.
Protect the override. A worker who looks at a high score and concludes, on the facts, that the case does not warrant the recommended action must be able to act on that conclusion without it counting against them. If overriding the model is professionally risky, if the worker who disagrees with the score and turns out to be right gets no credit but the worker who disagrees and turns out to be wrong gets blamed, then workers learn to defer to the model defensively. Deferring to the model to protect yourself is the opposite of mandatory human review. The agency has to make the considered override the safe choice, because the override is where human judgment does its most important work.
The History That Makes This Non-Negotiable
Mandatory human review is not a precaution against a hypothetical failure. It is a response to real harms that occurred when the human was removed from the loop or reduced to a rubber stamp, and the history is the reason the field treats this as non-negotiable rather than as best practice.
The Dutch childcare-benefits scandal is the clearest cautionary tale. A government fraud-detection system flagged families, disproportionately families with dual nationality or immigrant backgrounds, as likely benefits fraudsters. Tens of thousands of families were wrongly accused, ordered to repay large sums, and driven into financial ruin. The system's outputs were treated as findings rather than as signals requiring genuine human scrutiny, and the bias encoded in the system reproduced itself at scale across real lives. The harm was enormous, the victims were among the most vulnerable, and the failure was precisely the failure that mandatory human review is designed to prevent: a machine signal converted into a consequential action without a person doing the work of independent judgment on each case.
Michigan's MiDAS system tells a similar story in the unemployment-benefits context. An automated system accused tens of thousands of people of fraud, with a false-positive rate later found to be extraordinarily high, and imposed penalties on them with minimal human review. People lost money, faced collections, and suffered serious financial harm based on automated determinations that were frequently wrong. Again, the through-line is the same: the human review that should have stood between the system's output and the action taken against a person was absent or hollow.
In child welfare specifically, the long debate over the Allegheny Family Screening Tool centered on exactly this question. The tool produces a risk score to support screening decisions, and the central design commitment, and the central point of contention, is that the score is meant to inform a human screener, never to make the screening decision. The legitimate version of such a tool depends entirely on the human review being real. Strip the genuine review away, let the score drive the decision, and the tool becomes a mechanism for encoding historical inequity into present decisions about which families get investigated.
It is worth naming what these cases share at a more practical level, because the lesson is not only historical. In each, the volume was high, the signal looked authoritative, and the people on the receiving end were among the least able to push back: immigrant families, unemployed workers, parents already under scrutiny. Those are the exact conditions of an ordinary human-services unit on an ordinary day. The scandals were not exotic failures of unusually bad systems; they were the predictable result of ordinary pressures acting on tools whose outputs were trusted as decisions. That is why a caseworker reading this should not file these examples under cautionary tales about other people's agencies. The 4:50 PM Friday referral in the opening of this lesson is the same situation in miniature, and the only thing standing between that situation and the next headline is whether the human review between the signal and the action is real.
The pattern across all three is consistent and instructive. The technology was not the proximate cause of the harm. The removal or hollowing-out of human judgment was. Each of these systems could have functioned as a decision-aid under genuine human review. What turned them into engines of harm was the treatment of their outputs as decisions. That is why the field's response is not "do not use these tools" but "never let the tool decide, and build the human review so it cannot be skipped." The non-negotiability of mandatory human review is written in the harm that followed every time it was negotiated away.
Key Takeaways
- Mandatory human review is the required, defined step that places a qualified person between any AI risk signal and any consequential action on a case. It is engineered into the workflow, not left to individual diligence, because under caseload pressure the optional version becomes no review at all.
- A risk score is an input, never a conclusion. The cardinal rule holds: AI informs, humans decide. Decisions to investigate, substantiate, remove, or deny benefits are bound by due process and may never rest on a model's output.
- Four forces erode review under pressure: sheer volume, the false authority of a precise-looking number (automation bias), incentives that reward fast ratification over careful judgment, and diffusion of responsibility across the human and the machine.
- A real review requires four things: reading the whole case rather than the score, weighing the signal against what the record actually shows, reaching a judgment that would stand on the facts even if the score vanished, and recording the review and the reasoning.
- The override is where human judgment does its most important work. A worker who concludes on the facts that the recommended action is wrong must be able to act on that conclusion safely, or workers will defer to the model defensively to protect themselves.
- To make review unskippable, design it as a required gate, separate the signal from the action, staff the hours it genuinely takes, measure the quality of review rather than the speed of disposition, and protect the considered override.
- The history makes this non-negotiable: the Dutch childcare-benefits scandal, Michigan's MiDAS unemployment system, and the Allegheny Family Screening Tool debate all show that harm followed when a machine signal was converted into action without genuine human judgment on each case.
- An agency that mandates review without staffing the time, fixing the incentives, and protecting the override has mandated a fiction: a system where the model effectively decides while a human's name sits on the file.
Skill.re