โ†
AI for Social Work & Human Services
Aware ยท M8 ยท lesson 8 of 18 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Equity and Bias in Human-Services AI
๐Ÿ“–
now learning

Equity and Bias in Human-Services AI

15 min

Maria is a mother of three in Allegheny County, Pennsylvania. She has called for help twice in the past five years: once when her electricity was shut off and a neighbor reported the children were cold, and once when she asked a food pantry worker to connect her to a family support program. She has never been found to have abused or neglected her children. But when a new report comes in and a caseworker opens the screening system, a number appears beside her name: a risk score in the high range, generated by an algorithm before the worker has made a single phone call. The number comes from a predictive tool that drew on county administrative data. It knows about her past utility shutoff. It knows about the pantry referral. It knows she receives public benefits. It does not know that she is a loving parent who was navigating a temporary financial crisis. This lesson is about what that number means, where it comes from, why it can be wrong in patterned and inequitable ways, and what the field has learned, often the hard way, about building systems that do not quietly punish people for being poor.

How AI Tools Encode the Inequities in Their Training Data

To understand why human-services AI (artificial intelligence) can embed and amplify bias, you need to understand, at a plain-language level, what these systems are actually doing. A predictive risk-screening tool in child welfare does not reason about a family the way a thoughtful caseworker does. It does not observe the home, listen to the children, read the expressions on parents' faces, or weigh the context of a neighborhood's resources. Instead, it computes a score by finding statistical patterns in historical data and applying those patterns to a new case.

The logic runs like this. The system is trained on thousands of past cases. It learns which combinations of factors, recorded in the administrative databases the agency uses, appeared in cases that were later substantiated for maltreatment. Then, when a new referral arrives, it looks at the current family's data and asks: how similar is this family's profile to the profiles of families in the historical data that were substantiated? The more similar, the higher the score.

This approach has an immediate and serious problem: the historical data it learns from does not represent objective reality. It represents the history of which families the system investigated, which it substantiated, and which it left alone, and that history is deeply shaped by the structural inequities of American poverty, race, and public-service delivery. Poor families, and disproportionately Black and Indigenous families, have always come into more contact with child protective services (CPS) not because they maltreat their children at higher rates but because poverty is surveilled in ways that wealth is not. A wealthy parent who is struggling with addiction can seek private treatment quietly. A poor parent in the same situation may interact with a public health program, a court, and a benefits office, all of which generate records.

When the algorithm trains on this history, it learns a pattern that is in reality a pattern of surveillance and structural disadvantage, not a pattern of genuine child safety risk. The families who appear most often in substantiated cases are those who appeared most often in the system to begin with, which means the families who are poorest, who use the most public services, and who live under the most governmental scrutiny. The model then scores new families using those same patterns, and the families who receive the highest scores are, with striking regularity, the families who are poorest and most surveilled.

A model trained on biased historical data does not find the truth. It finds the history, and if that history encoded inequity, the model will too.

Virginia Eubanks, in her 2018 book Automating Inequality, called this the "digital poorhouse." She documented in detail how the Allegheny Family Screening Tool (AFST) in Pennsylvania was built, what data it drew on, and how the inputs that drove the highest scores were almost entirely markers of poverty and public-service contact rather than direct indicators of abuse or neglect. She found that families were scored higher not because anything about their parenting was flagged but because their lives were more visible to public systems. A family using Medicaid (the public health insurance program for low-income people), receiving SNAP (Supplemental Nutrition Assistance Program, formerly food stamps), or interacting with drug and alcohol treatment services had more data points in the county's system, and more data points meant more opportunities for the algorithm to find correlations with historical substantiation rates, which themselves reflected historical surveillance patterns.

Dorothy Roberts, a law professor and child-welfare scholar at the University of Pennsylvania, has pushed the analysis further. In her 2022 book Torn Apart, she argues that the structural surveillance of Black families by child protective services is not incidental but is the central feature of a system that was built to police Black poverty rather than to support Black families. Her argument is that any predictive tool trained on the data that system generated will, by design, score Black families higher, because Black families have historically been reported more, investigated more, and substantiated more, even holding constant the actual conditions in the home. The algorithm amplifies an existing disparity rather than creating a new one, but that amplification, delivered with the authority of a number and the efficiency of a computer, is powerful and difficult to contest.

What the Algorithm Knows and What It Cannot Know

The specific data inputs of the AFST are instructive because they illustrate, with precision, the gap between what the algorithm measures and what it is asked to predict. The tool used county administrative data across multiple systems, including mental health, drug and alcohol treatment, and child welfare records going back years. A single contact with any of these systems, even a voluntary one, such as a parent asking for help with depression or calling a crisis line, added a data point that could raise the score.

Notice what this means. A parent who never sought help, who remained invisible to public systems, would score low. A parent who sought help, who was engaged with services, who was trying to address challenges, would score higher. The algorithm, in a real sense, punished help-seeking and rewarded invisibility. It scored engagement with public services as a risk indicator because engagement with public services was correlated with eventual substantiation in the historical data, and it was correlated because the families who used the most public services were the most surveilled and therefore the most likely to have their struggles noticed and recorded.

The algorithm cannot distinguish between a mental health record that reflects a parent in crisis who received effective treatment and fully recovered and one that reflects a chronic and untreated condition that endangers children. It cannot see the difference between a SNAP enrollment that represents temporary hardship and one that represents generational destitution. It reduces all of these contexts to a count of system contacts and a probability estimate derived from historical patterns. The caseworker standing in front of that family can make those distinctions. The algorithm cannot.

The Dutch Childcare Benefits Scandal: When a Government Algorithm Wrongly Accused Thousands

The Netherlands offers one of the starkest and most thoroughly documented cases of algorithmic harm in a social-benefit context. Between roughly 2013 and 2019, the Dutch Tax Authority (Belastingdienst) used an automated fraud-detection system to screen families claiming childcare benefits. The system was supposed to identify families who had claimed benefits to which they were not entitled. Instead, it generated tens of thousands of false accusations, demanded repayment of years of benefits, and pushed many families into financial catastrophe. The Dutch government resigned in January 2021 over the scandal.

The system's most documented pattern of failure was a strong bias against families where one or both parents held dual nationality, meaning they held Dutch citizenship plus the citizenship of another country. Families with dual nationality were flagged at far higher rates than Dutch-only families, even when their benefit claims were entirely legitimate. Investigators later found that having a "second nationality" had effectively become a proxy variable, a characteristic that correlated with flag rates not because dual-nationality families actually committed more fraud but because the system had learned, from historical data that reflected earlier discriminatory enforcement patterns, to associate that characteristic with risk.

The consequences were severe and specific. Families received demands to repay tens of thousands of euros they had already spent on childcare. The repayment demands arrived without adequate notice, without clear explanation of how the determination had been made, and without a usable process for challenging them. Parents lost jobs, fell into serious debt, and in some cases lost their housing. Mental health crises and family breakdowns followed. Investigators who later examined the case found that the algorithm had been treating dual-nationality as a substantive risk signal while suppressing the human review that might have caught the pattern. The system was designed for efficiency, and efficiency meant minimizing human judgment. That turned out to mean minimizing the only mechanism that could have caught the bias before it harmed thousands of families.

A parliamentary inquiry found that the system had violated fundamental principles of administrative law, specifically the requirement that each case be decided on its individual facts, not on a statistical profile. The inquiry used the word "institutional" to describe the discrimination, because it was not the result of any single person's prejudice but of a system that had been designed, optimized, and deployed in ways that systematically disadvantaged a protected group with no meaningful avenue for recourse.

The Due Process Failure at the Center of the Dutch Case

The Dutch case is important not only as a story about algorithmic bias but as a story about what happens when a system is designed to minimize human review and to treat an algorithmic output as a determination rather than as an input. Due process, in the administrative law sense, requires that a person affected by a government decision have notice of the basis for the decision, a meaningful opportunity to contest it, and a decision made by a process that is capable of considering their specific facts. The Dutch childcare-benefits system failed on all three counts. Parents did not know what factors the algorithm had used against them. The appeal process was designed in a way that assumed the algorithm was correct and placed the burden of disproof on families who often lacked the records and expertise to mount a challenge. And the system was, by design, incapable of considering the specific circumstances of a family's situation: it had already decided, and the decision came with the apparent authority of a government determination.

This is a cautionary structure for any human-services AI in the United States. The right to notice and a fair hearing is not an abstract principle in American social-benefit and child-welfare law: it is a specific procedural guarantee that governs every determination about eligibility, substantiation, and placement. When an AI system's output is treated as a determination rather than as an input to human judgment, the due process guarantee is compromised not necessarily through deliberate intent but through institutional design that substitutes efficiency for review.

Michigan MiDAS: Automated Fraud Accusations and the Cost of a False Positive

Michigan's Integrated Data Automated System (MiDAS) for unemployment insurance operated from 2013 to 2015 and generated approximately 40,000 fraud determinations. Federal auditors later found that the vast majority of those determinations, roughly 93 percent by some estimates, were false. The system had been designed to flag potential fraud by cross-referencing employer wage records against unemployment claim data. When it found discrepancies, it generated a fraud determination automatically. It then assessed penalties: a claimant found to have committed fraud faced a demand for repayment of benefits received, plus a 400 percent penalty surcharge, plus potential criminal referral.

The discrepancies the system was detecting were, in most cases, the ordinary noise of an imperfect wage reporting system: employers filing wage information on different timelines, rounding differences in reported hours, the routine lag between when a person starts a new job and when that job is reported to the state. These were not fraud. But the system treated them as fraud, and the consequences for the people it accused were serious. Claimants received notices that were difficult to understand, presented demands that were impossible for many to pay, and faced a bureaucratic process for appealing that was, as later court findings confirmed, inadequate to protect their rights.

A federal court in 2016 found that MiDAS had violated claimants' due process rights. The court's core finding was that the state had delegated a consequential determination, whether a person had committed fraud, to an automated system and had done so without adequate procedural safeguards. The state had, in effect, substituted algorithmic certainty for human judgment in a context where the consequences of a false positive were severe and where the algorithm's false positive rate was extraordinarily high. The court ordered a remediation process that required the state to review tens of thousands of cases individually, a process that took years and cost the state considerably more than the system had saved.

The MiDAS case matters for human-services AI professionals for several reasons. First, it demonstrates the harm cost of a high false positive rate in a context where the consequence of a false accusation is not a minor inconvenience but a financial demand that can devastate a family already in crisis. An unemployment claimant who is wrongly accused of fraud and assessed a 400 percent penalty surcharge faces potential financial ruin from a mistake the algorithm made at a rate of nearly nineteen out of twenty. Second, it illustrates how the design choice to automate a determination rather than to use automation as an aid to human review creates both a harm risk and a legal vulnerability: the court did not find that the algorithm was necessarily wrong in its design, it found that using an algorithm to make a determination without adequate human review and procedural safeguards violated the Constitution. Third, it shows that the efficiency gains from automation can be entirely consumed, and then some, by the cost of remediation when the system fails at scale.

The Equity-First Stance: Why Auditing Is Continuous, Not Optional

The cases above, and the broader research literature on algorithmic bias in public services, have produced a set of principles that the field has arrived at through experience rather than theory. These principles are not guidelines for a distant future of more sophisticated AI: they apply right now, to every predictive risk tool, every automated screening system, and every AI-assisted determination in human services. The most important is this: equity auditing is a continuous practice, not a one-time certification.

What does that mean in practice? It means that when an agency deploys a risk-screening tool, the work of evaluating whether that tool produces equitable outcomes does not end at deployment. It begins. A model that passes an equity review in year one may fail in year three as the population changes, as economic conditions shift, as the agency's own practices change in ways that affect the data the model is trained on. The only way to know whether a tool is operating equitably is to measure its outcomes, disaggregated by race, by ethnicity, by income, by neighborhood, by benefit type, and by any other characteristic that is relevant to the equity question, on a regular and ongoing basis.

The technical term for this kind of measurement is disparate impact analysis (DIA). DIA looks at the outcomes a tool produces across demographic groups and asks: are these groups experiencing the tool's outputs at meaningfully different rates? If a risk-screening tool is generating high-score flags for Black families at twice the rate of white families with otherwise similar profiles, that disparity needs to be investigated. It may have an innocent explanation: perhaps the population served by the agency is different in ways that legitimately affect risk. But it may not. And the only way to find out is to look.

The equity-first stance requires that this looking be built into the operation of the tool from the beginning. It requires that an agency deploying a risk-screening tool have a designated process for reviewing disparate impact regularly, a clear threshold at which a disparity triggers further review, a mechanism for raising findings to leadership, and a documented path from a finding to a response. Without these structural elements, an equity audit is a gesture rather than a safeguard.

Every Risk Signal Is One Audited Input Under Mandatory Human Review

The second principle follows directly from the first: no risk signal from an AI tool is, by itself, a sufficient basis for a consequential determination. Every score, every flag, every algorithmic output is an input to human judgment, not a substitute for it. This is the cardinal rule of human-services AI, and it applies with special force to risk-screening tools precisely because those tools are the most likely to embed and amplify historical bias.

In practice, this means that a caseworker receiving a high-risk score for a family from a predictive tool has received one piece of information, not a determination. The caseworker's job is to evaluate that score in the context of everything else they know: the specific circumstances of this family, the quality and recency of the data the tool used, the known limitations of the tool's design, and the equity track record of the tool in the agency's own population. The score may be useful. It may point the caseworker toward questions worth asking. But it cannot answer those questions, and it must never be the reason a family is investigated, a child is removed, or a benefit is denied.

This does not mean that risk-screening tools have no value. The goal is not to eliminate them but to use them correctly. A risk score that consistently flags families with genuine indicators of danger, that is audited regularly for disparate impact, that is treated as one input among many, and that sits under mandatory human review before any consequential action is taken, is a different tool from one that operates as a decision engine. The difference is not in the algorithm. It is in the governance, the training, and the culture of the people who operate it.

Mandatory human review, in this context, means something specific. It means that before a caseworker acts on a risk signal, there is a documented process requiring human consideration of the specific facts of the case. It means the caseworker must be able to explain the basis for their decision in terms that go beyond "the score was high." It means supervisors review cases where a risk signal played a role in a decision to investigate or to not investigate, to look for patterns that might indicate the tool is being misused. And it means that a family has a meaningful opportunity to know that a tool played a role in their case and to challenge the accuracy of the information the tool used.

What Advocates and Researchers Have Found: The Evidence Base

The cases described in this lesson are not isolated incidents. They represent a pattern that academic researchers, civil-rights advocates, and investigative journalists have documented across multiple tools, multiple jurisdictions, and multiple benefit programs. The documentation is substantial enough that the field now has a reasonably clear picture of how algorithmic harm propagates in human services, and why the structural features that produce it are common rather than exceptional.

Eubanks's analysis of the AFST documented that the tool's scores were most heavily driven by indicators of poverty and public-service contact. She found that the county had built a tool that effectively predicted not child maltreatment but child poverty, and that the two were being treated as equivalent in the scoring logic. Her findings were contested by some researchers, and the county made a number of design modifications to the tool in response to criticism, but the underlying concern, that poverty proxies were functioning as risk indicators, was never fully resolved. The tool remained in operation through successive revisions while the debate about its equity implications continued, which is itself instructive: equity debates about deployed tools rarely produce the clean resolution that a decommissioning and replacement would represent. Agencies operate under resource constraints, tools get modified rather than replaced, and the people most harmed by a biased tool are the least likely to have the access and resources to mount an effective challenge.

Roberts's work extends the analysis to the structural level. She argues that the child-welfare system in the United States is, as a matter of documented history and current operation, organized around the surveillance and discipline of Black families, and that predictive tools deployed in that system necessarily inherit and amplify its racial logic. Her argument is not that the people building or operating these tools intend to discriminate: it is that the system within which the tools operate produces discriminatory outcomes regardless of the intentions of any individual actor. This is a structural argument about bias, and it is important for human-services professionals to understand because it points to the limits of individual-level interventions. A caseworker who is committed to equity cannot, by themselves, correct for a systemically biased tool. The correction has to happen at the level of the tool's design, governance, and operation.

Research published by the AI Now Institute, the Brookings Institution, and academic groups studying algorithmic accountability has consistently found several common features in systems that produce harm: a lack of transparency about how the system works and what data it uses; inadequate or absent mechanisms for affected people to challenge determinations; pressure on frontline workers to defer to algorithmic outputs because deviating from a tool's recommendation requires documentation and justification; and limited or no ongoing equity monitoring after deployment. These are not features of a few bad actors. They are features of how automated systems tend to get deployed in public-sector contexts where budgets are limited, timelines are pressured, and the political incentive is to show efficiency gains rather than to build in the oversight infrastructure that makes efficiency sustainable and equitable.

The Surveillance of Poverty and Its Limits as a Safety System

One of the most important analytical contributions of researchers like Eubanks and Roberts is the concept of the surveillance of poverty: the observation that poor families in the United States are subject to a degree of monitoring by public agencies that wealthy families never encounter. A poor family interacts with the SNAP office, the Medicaid system, the housing authority, the public school's social worker, and potentially the courts around child support, domestic violence, or criminal justice. Each of these interactions leaves a data record. A wealthy family with the same underlying stressors, the same substance use, the same relationship conflicts, the same mental health struggles, handles those challenges in private with private resources, and leaves no comparable data trail.

This asymmetry means that when an AI system trains on administrative data, it is training on a data set that systematically overrepresents poor families and underrepresents wealthy ones. The system's picture of "what a family at risk looks like" is, inevitably, built on a picture of what a poor, publicly-engaged family looks like. The system does not have enough data on wealthy families in crisis to learn what their risk signals look like, because wealthy families in crisis mostly do not interact with public systems in ways that leave administrative records.

The practical consequence for any human-services professional evaluating an AI risk-screening tool is this: ask where the training data came from and whose families it represents. If the training data is drawn entirely from the agency's own historical records, it represents a population that is already selected for public-services engagement, which means it is already selected for poverty. A tool trained on that data will be most predictively valid for the population it was trained on, which is the poorest and most publicly-engaged population. Its validity for families who are not currently in the system, including families in crisis who have not yet come to the agency's attention, is unknown. And its scores for currently engaged families will reflect not only genuine safety risk but the accumulated history of surveillance that brought those families into the system in the first place.

Building Equitable Practice: What the Standards Require

For a working caseworker, supervisor, or agency leader, the principles above translate into a set of concrete practices. These are not aspirational guidelines for a more equitable future: they are the operational standards that responsible use of AI in human services requires right now.

Demand transparency about the tool before it is deployed. Before any AI risk-screening tool is used in casework decisions, the agency should require the vendor or developer to provide a clear explanation of what data the tool uses, what population it was trained on, and what the tool's documented accuracy and equity profile look like across demographic groups. A vendor who cannot or will not provide this information is a vendor whose tool cannot be used responsibly. The caseworker who operates a tool without understanding its inputs is not protected if the tool produces a biased outcome: the professional and legal accountability for how the tool is used belongs to the agency and the individual worker, not only to the vendor.

Treat every risk signal as an input, not a determination. This is the cardinal rule restated as a practice: every score, flag, or algorithmic output is one piece of information in a human judgment process, not a conclusion. Document the basis for decisions in terms that go beyond the score. If the decision to investigate, or not to investigate, or to recommend a service, or to flag a case for supervisory review, is based primarily on an algorithmic output, that is a practice problem that supervision and training need to address.

Maintain and review disaggregated outcome data. An agency operating a risk-screening tool should regularly produce and review outcome data disaggregated by race, ethnicity, neighborhood, and benefit type. If the tool is generating high-risk flags for Black families at a materially different rate than for white families with similar profiles, that disparity needs investigation. The investigation may conclude that the disparity is explained by legitimate factors, or it may conclude that the tool is operating in a biased way. Either way, the investigation needs to happen and its findings need to reach leadership.

Build the appeal and challenge mechanism before deployment. Every family whose case is influenced by an AI tool's output should have a clear, accessible mechanism for knowing that a tool played a role and for challenging the accuracy of the information the tool used. This is a due process requirement, not a courtesy. A family that receives an adverse determination based in part on an AI tool's assessment of their history in public systems should know what data the tool relied on and should have a meaningful opportunity to correct errors in that data before the determination is finalized.

Create protected space for professional judgment. Caseworkers who believe an algorithmic output is wrong or misleading need institutional support to act on that judgment without penalty. The documented pattern in systems where algorithmic bias has caused harm is that frontline workers often noticed the problem but felt unable to override the tool because deviating from the algorithmic recommendation required extra documentation, additional supervisory approval, or created a record of non-compliance with agency practice. Agencies need to actively cultivate a practice culture in which a worker's informed professional judgment can override any algorithmic output and in which that override is treated as a sign of good professional practice, not a red flag.

Audit continuously, not just at deployment. A tool's equity profile at deployment is not its equity profile in year three. As the served population changes, as economic conditions shift, and as the tool is retrained on updated data, its disparity patterns may change. Ongoing equity monitoring is not a luxury: it is the mechanism by which an agency discovers that a tool that was operating acceptably has begun to produce problematic outcomes before those outcomes become a pattern of harm.

The goal of all of these practices is not to eliminate AI from human services. The documentation burden that drives caseworker burnout is real, and AI tools that help with summarization, drafting, and information retrieval can genuinely help. The goal is to ensure that the tools that inform the most consequential decisions in the field, the decisions about children's safety, family integrity, and people's access to the benefits they need to survive, are used in ways that protect rather than harm the people they serve. History shows that without deliberate, structural equity safeguards, these tools tend to amplify the very inequities the system is supposed to address. The equity-first stance is not idealism. It is the practical requirement for operating these tools responsibly in a field where the stakes are human lives.

Key Takeaways

  • AI risk-screening tools in human services learn patterns from historical administrative data, and that data reflects the history of which families were surveilled, investigated, and substantiated, a history that is systematically shaped by poverty and racial inequity rather than by the actual distribution of risk.
  • The Allegheny Family Screening Tool (AFST) debate, documented by Virginia Eubanks and Dorothy Roberts, illustrates how poverty proxies such as public-benefit enrollment, mental health service contacts, and drug and alcohol treatment records can function as risk indicators in a tool that is nominally measuring child safety, scoring help-seeking behavior as a danger signal.
  • The Dutch childcare-benefits scandal shows that algorithmic bias in public-benefit systems can scale rapidly, harm thousands of families before anyone intervenes, and produce institutional discrimination that violates administrative due process even when no individual actor intended to discriminate. The Dutch government resigned in January 2021 over a system that generated a high proportion of false accusations, disproportionately targeting dual-nationality families.
  • Michigan's MiDAS system generated approximately 40,000 unemployment fraud determinations, roughly 93 percent of which were later found to be false. A federal court held that automated fraud determination without adequate human review and procedural safeguards violated claimants' constitutional due process rights.
  • The surveillance asymmetry between poor and wealthy families means that any AI tool trained on public-agency administrative data is trained on a dataset that overrepresents poor families and does not have comparable data on the crises of wealthy families, making the tool's risk assessments structurally unequal across income levels.
  • Every AI risk signal is one audited input under mandatory human review, never a determination. A caseworker who acts on a high score without independently evaluating the specific facts of a family's situation has substituted algorithmic output for professional judgment, which is both a practice failure and, in contexts where due process applies, a legal vulnerability.
  • Equity auditing is a continuous practice. A tool that passes an equity review at deployment may produce disparate outcomes by year three as the population, the data, and the economic context change. Disaggregated outcome monitoring, regular review, and a documented process for responding to findings are the structural requirements for responsible deployment.
  • Building equitable AI practice requires structural commitments: transparency about tool design and data, disaggregated outcome monitoring, protected space for caseworkers to exercise professional judgment that overrides algorithmic outputs, accessible challenge mechanisms for affected families, and ongoing equity audits with a documented path from findings to agency response.