AI for Social Work & Human Services
Aware · M5 · lesson 5 of 18 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
AI in Risk Screening — and Its Perils
📖
now learning

AI in Risk Screening — and Its Perils

15 min

On a Tuesday morning in Allegheny County, Pennsylvania, a child-protective-services (CPS) worker opens her dashboard before her first home visit. A number sits at the top of a family's intake record: 18 out of 20. The number was generated overnight by a statistical model that ingested more than 100 data points about the family, including prior CPS contacts, public-benefits history, birth records, and neighborhood service-utilization data. The model does not know the family. It has never been in their kitchen, heard the father's voice, watched the mother redirect a toddler, or seen whether the older child's backpack was packed that morning. But the number is there, and it will travel with everything else the caseworker carries into that visit. How she uses it, and whether the tool that produced it was tested for the ways it can quietly encode and amplify decades of racial and economic inequity, is the question this lesson is about.

What Predictive Risk Screening Is

Predictive risk screening (PRS) in human services is the use of a statistical or machine-learning model to assign a numerical risk score to a case, a referral, or a family, based on a combination of administrative data and service records. The purpose is to surface cases that may carry a higher probability of a negative outcome, so that workers and supervisors can direct attention and resources accordingly. In child welfare, the most common application is scoring incoming CPS (child protective services) referrals to estimate the likelihood of future maltreatment, particularly of children who would otherwise not be investigated or visited.

The appeal is real. In a county child-welfare agency handling thousands of referrals a year, with a caseworker workforce that is perpetually understaffed and carrying caseloads above recommended levels, a tool that helps triage attention is not a frivolous idea. If a model can identify, among a hundred new referrals arriving on a Friday afternoon, the twenty families most likely to experience a serious harm before Monday, and if the workers act on that signal, lives may be saved. That is the genuine promise. It is what moved Allegheny County, Pennsylvania, to deploy its Allegheny Family Screening Tool (AFST) in 2018, and it is what has moved dozens of other child-welfare agencies and benefits programs to explore similar tools.

But the promise comes packaged with a set of structural and ethical dangers that are not hypothetical. They are documented. They have harmed real families. They have been litigated. And understanding them is not optional for a caseworker, supervisor, or agency leader who is expected to use these tools or to advocate for the families they serve.

How the Model Is Built

A predictive risk model in child welfare is typically built by taking a historical dataset of cases, identifying which cases resulted in a confirmed adverse outcome (a substantiated maltreatment report, a subsequent removal, an emergency placement), and then training a statistical algorithm to find the data features most associated with that outcome. The resulting model is then applied prospectively: when a new referral arrives, its data is fed to the model, and the model produces a score representing its prediction.

The input data almost never includes protected characteristics like race directly. The agency's legal team would not permit it. But the data does include features that are proxies for protected characteristics: number of prior CPS contacts (which varies by race because of differential investigation rates, not differential abuse rates), history of receiving public benefits such as SNAP (the Supplemental Nutrition Assistance Program, commonly called food stamps) or TANF (Temporary Assistance for Needy Families), neighborhood of residence, prior involvement with the child-welfare or juvenile-justice system, and housing instability indicators. These features are statistically predictive of future CPS involvement. They are also, in the United States, profoundly correlated with race and poverty, for reasons that have nothing to do with individual conduct and everything to do with which families are most surveilled and contacted by public systems in the first place.

The result is a model that can be racially and economically discriminatory without ever naming race or income as inputs. This is what advocates and researchers mean when they say a risk model can "encode inequity." The inequity is already in the training data, because the historical system that produced the training data was not equitable.

The Allegheny Family Screening Tool: A Case Study in Both Promise and Peril

Allegheny County's AFST is the most studied and debated child-welfare risk-scoring tool in the United States, for reasons that make it genuinely valuable as a case study. It was deployed with more transparency than most tools of its kind, with published methodology and ongoing independent evaluation. It attracted rigorous criticism from scholars, advocates, and affected families. And it illustrates, concretely, the tension between a tool that may have improved some targeting of services and a tool that may have imported and amplified the biases of the administrative systems it learned from.

The AFST scored incoming referrals on a 1-to-20 scale. Scores at or above 14 triggered an automatic call-in for investigation; scores in the middle range were presented to a screener as one input among others; lower scores could support a decision to screen out a referral. The data feeding the model came from the county's integrated data system, which aggregated records from child welfare, the courts, the Allegheny County Department of Human Services, and several other public agencies.

The critique that emerged from researchers including Virginia Eubanks, whose book "Automating Inequality" documented the AFST and similar systems, and from scholars like Dorothy Roberts, was focused on a specific structural problem: the model was more likely to score families highly if they had more contacts with public systems. Families who accessed more public benefits, who had more prior CPS contacts (regardless of whether those contacts resulted in substantiation), and who lived in neighborhoods with more aggressive CPS investigation patterns would generate higher scores. Because Black families, low-income families, and single-parent families have disproportionately more contact with the public systems that feed the data, they were disproportionately likely to score high, irrespective of their actual risk to their children.

The Allegheny County office and the tool's developers, Rhema Vaithianathan and Emily Putnam-Hornstein, pushed back on some of these critiques, arguing that the model was tested and that the score did predict future CPS involvement, which was its stated purpose. But that defense also illustrates the deepest problem: if the ground-truth outcome the model is predicting is "future CPS involvement," and if CPS involvement is itself racially disparate, then a model that accurately predicts future CPS involvement will accurately predict a racially disparate pattern, and calling it accurate does not make it fair or free of inequity. A model can be both statistically predictive and systemically discriminatory. These are not mutually exclusive.

The Allegheny debate did not end with a clear verdict. The AFST continues to operate in modified form, with additional human-review requirements and updated protocols. But it produced a body of scholarship, advocacy, and public accountability that every agency considering a predictive risk tool should engage with before deploying one. The lesson it teaches is not "never use AI in risk screening." The lesson is: a predictive risk tool requires independent equity auditing, radical transparency, robust mandatory human review, and ongoing public accountability if it is to be used at all.

What an Equity Audit Means for a Risk Tool

An equity audit of a predictive risk tool asks a set of specific questions about the tool's behavior across demographic groups, before deployment and on an ongoing basis afterward. It is not a one-time check performed before a press release. It is a continuous discipline that the AUTHORING-KIT.md for this program describes plainly: equity auditing is a continuous practice, not a one-time check.

The core questions an equity audit must ask include:

  • Disparate scoring rates: Does the model score families of a particular race, ethnicity, or income level systematically higher than families with comparable objective risk profiles? Not just more CPS contacts, but more contacts adjusted for surveillance intensity by neighborhood?
  • False positive disparity: Is the model more likely to flag a family as high-risk when the family will not, in fact, experience a maltreatment event? A model that generates disproportionate false positives for Black or low-income families is causing harm even when those families are not ultimately substantiated.
  • Feature proxy analysis: Which features are driving scores most strongly, and what are those features proxies for? Prior CPS contact, benefit receipt, and neighborhood variables are obvious candidates. Has the agency analyzed whether removing or reweighting these features reduces disparity without unacceptably degrading predictive performance?
  • Ground-truth validity: Is the outcome the model is predicting actually "child maltreatment" or is it "future CPS involvement"? These are not the same thing. A model that predicts future CPS involvement will encode investigation patterns; a model that predicts verified adverse outcomes requires a much harder and more defensible outcome definition.

Most agencies deploying these tools in 2026 cannot answer all of these questions fully, because they do not have access to the vendor's model documentation, training data, or feature-importance records. That gap is, itself, a governance failure.

When Risk Screening Goes Fully Wrong: Benefits Fraud Detection

The child-welfare context is the field's most emotionally visible risk-screening domain, but it is not the only one where AI-driven scoring has caused documented, large-scale harm to vulnerable people. The benefits-administration context provides two cases that are among the most important in the history of public-sector AI: the Dutch childcare-benefits scandal and Michigan's MiDAS system. Both cases involve algorithmic fraud-detection systems that assigned risk scores to benefits recipients, produced large-scale false-positive findings, and caused catastrophic harm to thousands of families before the systems were stopped. Neither involved child welfare directly, but both carry lessons that apply with full force to risk-screening in any human-services context.

The Dutch Childcare Benefits Scandal

Between approximately 2013 and 2019, the Dutch Tax and Customs Administration (Belastingdienst Toeslagen) used an algorithmic risk-assessment system to flag applications for childcare-benefits subsidies as potentially fraudulent. The system assigned risk scores to applications based on a set of features that included nationality and dual citizenship, which functioned as proxies for ethnicity. Families with scores above a threshold were automatically flagged for investigation and, in many cases, required to repay large sums of benefits the agency alleged had been incorrectly paid.

The result was a disaster of historic scale for the Dutch welfare state. An estimated 26,000 families were wrongly required to repay thousands of euros in benefits. Many were pushed into serious debt, with consequences including family breakdown, loss of housing, and documented cases of severe mental health consequences. The algorithmic system produced these outcomes with minimal individualized human review: the score was treated as sufficient basis for action, in violation of the due-process and proportionality principles that Dutch administrative law requires. The scandal resulted in the resignation of the entire Dutch cabinet in January 2021, the largest political accountability event in the Netherlands since World War II, triggered specifically by an algorithm's failure of due process and equity.

The key lessons for human-services practitioners are several. First, the model used nationality and dual citizenship as direct inputs, which are proxies for ethnicity, and the resulting disparate impact fell heavily on families with immigrant backgrounds. Second, the system operated without meaningful individual review: the score was the outcome, not one input among many. Third, the victims had no practical mechanism to challenge the determination in a timely way, which is a due-process failure that amplified the scale of the harm. Fourth, the harm was not discovered by internal oversight but by journalists, advocates, and eventually a parliamentary inquiry, which means the agency's internal governance had already failed before the external reckoning came.

Michigan's MiDAS: Automated Accusation at Scale

Michigan's MiDAS (Michigan Integrated Data Automated System) was an automated unemployment-insurance fraud-detection system deployed between 2013 and 2015. The system used algorithmic analysis of unemployment claims to identify potential fraud cases, and when it flagged a case, it automatically issued a Notice of Determination, imposed a mandatory repayment obligation, and assessed penalties of up to four times the allegedly overpaid amount, all without a human reviewer examining the individual case before the determination was issued.

The results were devastating. MiDAS falsely accused approximately 40,000 Michiganders of fraud, in cases later determined by courts and auditors to have no valid basis. The false-positive rate in one audit was found to be approximately 93 percent: of the cases the system flagged as fraud, more than nine in ten were wrong. People lost their unemployment benefits, had wages garnished, had tax refunds seized, and faced debt burdens that pushed them toward bankruptcy, job loss, and in documented cases, the extreme mental health consequences of being publicly accused of fraud without justification. The State of Michigan ultimately paid a settlement of more than $20 million to resolve class-action litigation and was required to reform its fraud-detection systems under court oversight.

MiDAS is not simply a story about a bad algorithm. It is a story about what happens when a system removes mandatory human review from a high-stakes determination and substitutes algorithmic scoring for individual judgment. The fundamental due-process right at stake in every benefits determination is that a person has the right to notice, a right to contest the determination before or promptly after adverse action, and a right to a fair hearing before a neutral decision-maker. MiDAS violated all three by treating the algorithmic output as the decision itself. "The system flagged it" is not a legal basis for denying benefits, garnishing wages, or assessing fraud penalties. The human worker, the supervisor, and the fair-hearing examiner own those determinations.

The Iron Rule: A Risk Signal Is One Audited Input, Never a Verdict

The three cases above, the AFST, the Dutch childcare-benefits scandal, and MiDAS, establish a pattern so consistent that it should function as law in any human-services agency. When an algorithmic risk score is treated as the determination, or when the human review required between score and action is perfunctory rather than genuine, the outcome is unjust, often illegal, and frequently catastrophic for the families who have the least capacity to withstand a wrong government decision.

A risk signal is one audited input under mandatory human review, never a verdict. "The model scored it high" is not a sufficient reason for any consequential action in human services. The caseworker, the supervisor, and the court own every call.

This rule is not a counsel of AI-pessimism. It is a statement about the nature of the consequential decisions in this field. Child-welfare decisions, including the decision to investigate a referral, to substantiate a report, to seek emergency placement, or to open a case for ongoing services, are among the most consequential decisions any government makes about people's lives. Benefits determinations, including the decision to deny, reduce, or terminate food, shelter, or cash assistance, affect whether a family eats and stays housed. These decisions are bound by constitutional and statutory due-process requirements. They require individual review. They require the possibility of challenge. They require a human accountable for the outcome.

A machine-learning model can inform that human. It cannot replace that human. When it tries to, the documented outcome is not efficiency. The documented outcome is injustice.

What Mandatory Human Review Means in Practice

Mandatory human review is not a checkbox. It is a structural requirement that a qualified human being, with access to the case record and the decision-making authority to agree or disagree with the model's output, must exercise genuine independent judgment before any consequential action is taken based on a risk score. The word "mandatory" means it cannot be skipped, abbreviated, or rendered pro forma by caseload pressure. The word "human" means a person, not a review process that routes all flagged cases to an action queue with a rubber-stamp workflow. The word "independent" means the reviewer must be capable of and permitted to reach a different conclusion than the model.

What does this look like operationally? It looks like a workflow where:

  • The risk score is presented to the caseworker alongside the case record, not instead of it. The worker reads the record, makes a home visit or a collateral contact, and forms their own assessment.
  • The caseworker documents their own professional judgment, explicitly noting where they agree or disagree with the model's signal and why. The documentation belongs to the worker, not to the score.
  • A supervisor reviews cases at or above a defined score threshold before any determination is made, and that review is a genuine second-opinion review, not a mechanical sign-off.
  • When the model's score is high and the worker's assessment is low, or vice versa, that discrepancy is explicitly discussed and documented. A discrepancy between human and model judgment is information, not an error to be resolved by defaulting to the model.
  • The final decision, whatever it is, is documented in the worker's own words, citing the specific observations, contacts, and case-record facts that supported it. "Score was 18" is not documentation of a human judgment. It is documentation of a model output.

This is the standard that protects families, protects workers, and produces case records that can withstand scrutiny from a court, an advocate, or a fair-hearing examiner. It is also the standard that is most frequently compromised under caseload pressure: when a unit is carrying 150 percent of the recommended caseload, the temptation to let a high score substitute for a full review is real and understandable. The answer is not to accept the compromise. The answer is to name the caseload pressure as the structural risk it is, and to build explicit workflow protections against it.

What a Risk Score Actually Measures, and What It Does Not

Risk scores in child welfare are statistical predictions about population-level patterns. The score for a given family says, in effect: "Among all families in our historical dataset with data features like those of this family, a certain percentage went on to have a subsequent confirmed maltreatment event within a defined time window." That is a probabilistic statement about a group of families. It is not a statement about this family.

This distinction matters profoundly in casework, because the caseworker is not serving the group. The caseworker is serving the individual family in front of them. A score of 18 out of 20 does not mean there is a 90 percent chance that this specific child will be maltreated. It means that the data features associated with this family are associated with higher rates of maltreatment in the historical dataset. The family in front of the worker may have every protective factor the model cannot see: a stable extended-family network, a new job, a parent who recently completed substance-use treatment, a set of genuine strengths that no administrative database records. Or the family may have risks the model cannot see, because those risks are not in any administrative database at all.

The model is working from the parts of the family's life that left a digital trace in public systems. Those traces are incomplete, unrepresentative, and skewed toward contact with agencies that disproportionately monitor certain populations. A family that has never accessed public benefits, never had a prior CPS contact, and lives in a high-resource neighborhood may have serious underlying risk that generates no data signal at all. A family that has accessed many services, had prior CPS contacts that were screened out, and lives in a neighborhood with high investigation rates may generate a high score with no current risk at all. The model sees the data. It does not see the family.

This is not a reason to dismiss risk-scoring tools. It is a reason to understand what they are and are not. At their best, they are tools that surface cases for additional attention that might otherwise be overlooked. They are not diagnostic instruments. They are not verdicts. They are, at most, one signal among many that a skilled, experienced, professionally trained caseworker uses in forming a judgment that belongs to them, not to the model.

Surveillance Inequality and the Training Data Problem

One of the most important structural problems with predictive risk tools in child welfare and benefits is a problem that exists before a single prediction is made: the training data reflects a world in which surveillance and system contact are not evenly distributed across the population. Families in poverty, particularly Black and Indigenous families, are more likely to have prior CPS contacts, more likely to have prior system involvement, and more likely to live in neighborhoods where CPS investigation rates are high, not because they maltreat their children at higher rates but because they are investigated at higher rates. This pattern is well-documented in the research literature and was extensively discussed in the Allegheny County AFST debate.

When a model is trained on this data and uses "prior CPS contact" as a predictive feature, it is, in effect, using past surveillance as a predictor of future surveillance. The resulting predictions will correlate with race and class because past surveillance correlates with race and class. The families most harmed by any individual bias in the model are the families who already bear the greatest burden from the system's historic inequities. This is what it means to "amplify" inequity: the model does not create the disparity, but it encodes it, systematizes it, and applies it at scale to every new intake.

No amount of technical sophistication in the model's architecture resolves this problem. It is a data problem, and therefore a governance and equity problem. The only partial mitigations are: rigorous equity auditing of the model's outputs across demographic groups, with ongoing monitoring; transparency about what the model uses as features and what those features are proxies for; and the absolute requirement that human professional judgment remain the decision-making authority for every individual case.

This Is Awareness Level: What Comes Later

This lesson is designed for Level 1 (L1): the AI-Aware Human-Services Professional. The goal here is to build a clear, honest, equity-grounded understanding of what predictive risk screening is, what it has done in documented history, and what the non-negotiable rules are for working with it. That is different from the hands-on practice of reading, using, and documenting a risk signal responsibly, which is the work of L2, and from the operational discipline of running an audited screening-support workflow with mandatory human review and ongoing equity monitoring, which is the work of L3.

At L2, you will work with risk signals in practice: how to read a score, how to document your independent judgment, how to record where you agree and disagree with the model, and how to write a case note that reflects human decision-making rather than model transmission. At L3, you will build the full audited screening-support workflow, including the equity-audit discipline and the documentation standard that makes a screening-informed determination defensible to a court and an advocate.

At L1, the goal is the foundation: understanding why this is the field's most ethically fraught AI use, what history has demonstrated about the cost of getting it wrong, and why the rules that govern it are not suggestions. They are protections for families. They are, in many cases, legal requirements. And they begin with the caseworker who sits down at a dashboard on a Tuesday morning and sees a number next to a family's name.

That number is information. It is one piece of information among many. The judgment belongs to the caseworker, the supervisor, the court, and the family's right to due process. It does not belong to the number.

Before we close, it is worth noting that these concerns are not an argument for the status quo without AI. The status quo, without risk-screening tools, has its own forms of inequity: unstructured human judgment is also susceptible to racial bias, to fatigue, and to the effects of caseload overload. The research literature on unstructured risk assessment in child welfare does not paint a picture of a fair and consistent baseline. The argument here is not that AI is worse than human judgment in every case. The argument is that AI in risk screening requires more rigorous governance, more explicit equity auditing, and more structural protection for mandatory human review than most agencies currently provide, because the consequences of getting it wrong are too severe and too irreversible to treat as an acceptable error rate.

The families in Allegheny County, the 26,000 Dutch families wrongly required to repay benefits, and the 40,000 Michigan workers falsely accused of unemployment fraud are not abstractions. They are the documented cost of deploying algorithmic scoring without adequate equity governance and without the iron protection of mandatory, genuine, documented human review. Their experiences are the reason this field holds the rules it holds. And those rules start with this: the risk signal is one audited input. The human being is the decision-maker. That is not a slogan. It is the floor.

Key Takeaways

  • Predictive risk-screening tools in child welfare and benefits assign numerical scores to cases using administrative data. The genuine promise is that they can help triage attention in under-resourced agencies. The documented danger is that they can encode and amplify the inequities already present in their training data, particularly through proxy variables that correlate with race and poverty.
  • The Allegheny Family Screening Tool (AFST) debate established that a model can be statistically predictive of future CPS involvement and simultaneously racially disparate in ways that reflect differential investigation patterns rather than differential rates of actual maltreatment. Predictive accuracy and equity fairness are not the same thing.
  • The Dutch childcare-benefits scandal, in which an algorithmic fraud-detection system wrongly accused approximately 26,000 families and contributed to the resignation of the Dutch cabinet in 2021, and Michigan's MiDAS system, which produced a 93 percent false-positive fraud-accusation rate against 40,000 people, demonstrate what happens when algorithmic scoring substitutes for genuine individualized human review.
  • The iron rule for risk signals in human services is: a score is one audited input under mandatory human review, never a verdict. "The model scored it high" is never a sufficient basis for investigation, substantiation, removal, or benefits denial. The caseworker, supervisor, and court own every call.
  • Equity auditing of a risk-screening tool is a continuous practice, not a one-time check at deployment. It must examine disparate scoring rates across demographic groups, false-positive disparity, feature-proxy analysis, and the validity of the outcome measure the model is predicting.
  • Mandatory human review means that a qualified professional must exercise genuine independent judgment, documented in their own words citing specific case-record facts, before any consequential action is taken on a risk score. Review that is perfunctory, under caseload pressure, or that simply routes all flagged cases to an action queue without real independent assessment is not mandatory human review in any meaningful sense.
  • The training-data problem in predictive risk screening is structural: administrative data reflects surveillance patterns, not maltreatment patterns. Families with more public-system contact generate higher scores because they have more contact, not necessarily because their children are at greater risk. No technical improvement to the model resolves a problem that originates in the data's relationship to real-world inequity.
  • This lesson is awareness-level. Deeper hands-on practice with risk signals comes at L2, and the full audited screening-support workflow, including equity auditing as an operational discipline, is built at L3. The foundation established here is why the rules exist, what history proves about the cost of violating them, and why every caseworker who works with a risk score needs to understand those stakes before they ever open a dashboard.