The Audited Screening-Support Workflow
The call came into the hotline at 9:14 on a Monday morning: a teacher reporting that a six-year-old had come to school in the same clothes three days running, had fallen asleep at her desk twice, and had mentioned that there was no food at home over the weekend. The intake screener took the report, entered it into the system, and saw a number appear beside it. The agency's screening tool had scored the case, and the score was high. For a moment the screener felt the pull that every agency that deploys these tools must guard against: the pull to let the number decide. To treat the score as the answer to the question the law actually asks, which is whether this report should be screened in for investigation, screened out, or routed to a voluntary service. But the score is not the answer. The score is a signal, one input, generated by a model trained on history, and history in this field includes the inequities that produced the very disparities the agency is trying not to repeat. This lesson is about the workflow that lets an agency use that signal without being used by it: the audited screening-support workflow, where a signal is surfaced, a human reviews it, and a human decides, with equity auditing built in before any family is harmed.
The Most Fraught Use Case in the Field
Documentation is the goldmine because it is the safest, most humane place for AI to help. Screening support is the second well, and it is handled carefully because it is the opposite: it is the place where AI in human services has the longest and most painful track record of harm. Predictive risk-screening, the use of a model to flag early indicators of possible abuse or neglect, is real and deployed, and it is ethically fraught precisely because the decisions it touches, whether to investigate a family, whether to remove a child, are among the most consequential a government makes.
The history is the reason for the caution, and it must be taught, not skipped. The debate over the Allegheny Family Screening Tool, one of the most studied predictive risk models in child welfare, surfaced the central concern: a model trained on historical case data learns the patterns in that data, and if the historical system referred, investigated, and substantiated some communities at higher rates for reasons of poverty or race rather than actual risk, the model can learn to reproduce those disparities and present them as objective risk. The harm is not hypothetical in adjacent domains either. The Dutch childcare-benefits scandal saw a fraud-detection system wrongly accuse tens of thousands of families, disproportionately families with immigrant backgrounds, of benefits fraud, with devastating consequences. Michigan's MiDAS system falsely flagged tens of thousands of people for unemployment-benefits fraud. The lesson these episodes teach is consistent: an automated risk or fraud signal can encode and amplify inequity, and when it is treated as a verdict rather than an audited input, real families are harmed at scale.
This is why screening support is built so carefully. The workflow exists not to make AI risk-screening more powerful but to wrap it in the human review and equity auditing that the history proves are non-negotiable. Every risk signal is one audited input under mandatory human review, never a verdict. That sentence is the entire ethic of this lesson compressed into a line.
A risk score is not a finding. It is a model's guess, trained on a history that includes the very inequities the agency is trying not to repeat. It is one input, audited, under mandatory human review, and never a verdict.
What the Signal Is, and What It Is Not
To use a screening signal responsibly, a worker has to understand precisely what it is. A predictive risk score in child welfare is a number, often on a scale, that a model produces from the data available about a referral and the people named in it: prior contacts, prior services, demographic and administrative data, and the content of the current report. The model has learned, from thousands of past cases, which combinations of these features were historically associated with later outcomes the agency tracks. The score is a statistical estimate of association, not a measurement of what is happening in a home right now.
Several things follow from that, and each one is load-bearing for the workflow. First, the score is about correlation in past data, not causation and not present reality. A high score does not mean a child is being harmed; it means the case resembles past cases that, in the historical data, were associated with the tracked outcome. Second, the data the model learned from carries the imprint of how the historical system behaved, including any bias in who was reported, investigated, and substantiated. A model trained on biased outcomes can produce biased scores while appearing perfectly objective. Third, the score cannot see context the data does not contain: the protective grandmother who just moved in, the new job, the safety plan working, the explanation for the school absences that has nothing to do with neglect.
So the signal is genuinely useful as one of several inputs: it can prompt a reviewer to look more carefully, to gather more information, to not screen out a report too quickly. And it is genuinely dangerous as a verdict: treated as the answer, it can pull an agency toward investigating the families the historical data already over-investigated, deepening the disparity it was supposed to help with. The workflow is designed to capture the first use and forbid the second.
The Workflow: Signal Surfaced, Human Reviews, Human Decides
The audited screening-support workflow has three movements, and like the documentation workflow, their order and their boundaries are the safeguard. Signal surfaced. Human reviews. Human decides. Around all three sits a continuous equity-auditing practice that this lesson introduces and the next lesson develops in depth.
Signal Surfaced
In the first movement, the model produces its signal and presents it to the reviewer in a way designed to inform rather than to anchor. How the signal is presented matters enormously, because presentation shapes how much weight a tired reviewer gives it. A well-designed system surfaces the score alongside the information that lets a human interpret it: what the signal is, roughly what features drove it where the system can show that, the known limitations of the model, and an explicit statement that the score is an input and not a recommendation to screen in. A poorly designed system shows a big red number and an implied instruction. The workflow requires the former. The signal is surfaced as a prompt to look, not as a conclusion to adopt.
Human Reviews
In the second movement, a trained human reviewer examines the full referral, the score among the inputs but never above them. The reviewer reads the actual report, considers the actual allegations, weighs the context the model cannot see, and applies professional judgment and the agency's screening criteria, which are grounded in statute and policy. The score may legitimately prompt the reviewer to gather more information before deciding, to call the reporter back, to check a detail in the record. What it may not do is substitute for that review. Mandatory human review means exactly this: no case moves from signal to decision without a qualified person examining the whole picture. The review is not a rubber stamp on the number; it is the place where a human asks whether the number is telling the truth about this particular family or merely repeating a pattern from the data.
Human Decides
In the third movement, the human makes the decision that the law assigns to humans: screen in, screen out, or route to a voluntary or alternative response. The cardinal rule governs absolutely here. The decision to investigate a family, and any decision downstream about removal, is among the most consequential a government makes, and an algorithm must never make it. "The model scored it high" is never a sufficient reason for a consequential decision. The reviewer, the supervisor, and ultimately the legal process own the call. The score informed the decision; it did not make it, and the record will show a human reasoning, not a number executing.
Signal surfaced, human reviews, human decides. The score can tell a reviewer to look harder. It can never tell a reviewer what to conclude.
Equity Auditing, Built In Before Harm
Mandatory human review protects the individual case. It is necessary, but it is not sufficient, because a single reviewer cannot see the pattern across thousands of cases that reveals whether the tool itself is biased. A reviewer can give a fair hearing to one family and still be part of a system in which the tool systematically scores one community higher than another for reasons unrelated to actual risk. Catching that requires equity auditing, the continuous, systematic testing of the tool's behavior across groups, and it has to be built into the workflow before harm, not bolted on after a scandal.
Equity auditing in this context means asking, regularly and rigorously, a set of hard questions about the tool's behavior in operation. Does the model score referrals from some racial, ethnic, or socioeconomic groups systematically higher than others, after accounting for actual risk? Are screen-in rates for AI-assisted decisions diverging across groups in ways the agency cannot justify on the basis of safety? Is the tool surfacing poverty as if it were neglect, flagging the markers of being poor, unstable housing, missed appointments, reliance on public benefits, in ways that pull scrutiny toward families whose problem is a lack of resources rather than a risk to a child? These are not one-time questions. They are a standing practice, because a model's behavior in the world drifts, the population changes, and a tool that looked fair at launch can become unfair in operation.
The history makes the stakes of skipping this concrete. The agencies and systems that caused the most harm did not set out to discriminate; they deployed a tool, trusted it, and failed to audit its disparate impact until tens of thousands of families had already been harmed. Equity auditing before harm is the discipline that turns "we did not intend to discriminate" into "we tested for discrimination continuously and acted on what we found." That is the difference the people in the system experience as either protection or catastrophe. Equity is first, not an afterthought, and in this workflow it is a continuous practice wrapped around every decision the tool touches.
Documenting the Human Decision
The screening-support workflow, like the documentation workflow, must be defensible to a court and an advocate, and that defensibility comes from documentation and disclosure. What gets recorded is not just the decision but the reasoning, and specifically how the signal was used. A defensible record shows that a signal was generated, that a qualified human reviewed the full referral, what the human considered beyond the score, and the human reasoning that led to the decision. It shows that the score was an input among inputs, not the driver.
This matters for due process. A family has the right to challenge a determination about them, and that right is hollow if the basis of the decision is an opaque number no one can explain. If an advocate asks why a family was screened in, the agency must be able to answer with human reasoning grounded in the actual allegations and policy, not with "the algorithm said so." A record that can only point to a score is a due-process problem; a record that shows a human weighing the evidence, with the signal as one transparent input, is defensible. Transparency and disclosure about how AI was used in the decision keep the work accountable to the people it affects.
Documenting how the signal was used also feeds the equity audit. The per-decision record of how scores related to human decisions, across groups and over time, is the raw material the audit needs to detect disparate impact. So the documentation stage is not only a due-process safeguard for the individual family; it is also what makes the systemic equity check possible. The two protections, individual and systemic, are connected through the discipline of recording the human reasoning behind every screening decision the tool touched.
When the Honest Answer Is "Not Here, Not Yet"
A workflow this carefully built has to include the possibility that, for a given agency or a given decision, the right answer is not to use the tool at all. The program teaches a written kill-criteria discipline for the most sensitive decisions, and screening is where it bites hardest. There are conditions under which deploying or relying on a risk-screening signal is not responsible, and a mature agency names them in advance rather than discovering them in a crisis.
Consider the conditions that should give an agency pause. If the tool cannot be audited for disparate impact, because the vendor will not provide the access, the data, or the transparency the audit requires, the agency cannot meet its equity obligation and should not rely on the tool. If the agency lacks the capacity for genuine mandatory human review, so that in practice the score becomes the decision because reviewers have no time to review, the workflow's central safeguard does not exist and the deployment is unsafe. If an equity audit reveals disparate impact the agency cannot explain or correct, continuing to use the tool while the harm runs is indefensible. And for the most consequential decisions, the workflow holds an absolute line: a risk score may inform whether to investigate, but the decision to remove a child is reserved for the worker, the supervisor, and the court, and no score, however high, substitutes for that human and legal process.
This is not anti-technology. It is the discipline that makes responsible use possible. The agencies that earn the trust of the families they serve and the courts they answer to are the ones that can say, truthfully, that they used the signal where it helped, audited it continuously for harm, kept every consequential decision human, documented the reasoning, and were willing to stop using the tool when the conditions for safe use were not met. That posture, careful, audited, human-governed, and honest about its own limits, is what separates screening support that protects children from screening support that repeats history.
Key Takeaways
- Screening support is the field's most ethically fraught AI use case, the careful second well next to the documentation goldmine, because the decisions it touches, whether to investigate a family or remove a child, are among the most consequential a government makes.
- The history is the reason for the caution: the Allegheny Family Screening Tool debate, the Dutch childcare-benefits scandal, and Michigan's MiDAS system all show that an automated risk or fraud signal can encode and amplify inequity and harm families at scale when treated as a verdict rather than an audited input.
- A predictive risk score is a statistical estimate of association in past data, not a measurement of present reality. It can reflect bias in who was historically reported and investigated, and it cannot see protective context the data does not contain.
- The workflow has three movements whose order and boundaries are the safeguard: signal surfaced (presented to inform, not to anchor), human reviews (the full referral, the score among inputs but never above them), and human decides (the screen-in, screen-out, or route decision the law assigns to humans).
- Mandatory human review protects the individual case, but it cannot see the cross-case pattern, so equity auditing, the continuous, systematic testing for disparate impact across groups, must be built in before harm rather than bolted on after a scandal.
- Equity auditing asks whether the tool scores some groups higher without a safety justification and whether it surfaces poverty as if it were neglect; it is a standing practice because a tool that looked fair at launch can drift into unfairness in operation.
- Documenting the human decision, specifically how the signal was used and the human reasoning behind the call, is both a due-process safeguard (a family's right to challenge a determination is hollow if the basis is an unexplainable number) and the raw material the equity audit needs to detect disparate impact.
- Responsible use includes a written kill-criteria discipline: if the tool cannot be audited, if genuine human review is not resourced, if an audit reveals uncorrectable disparate impact, or for the decision to remove a child, the right answer may be not to rely on the tool, because no score ever substitutes for the human and legal process.
Skill.re