โ†
AI for Social Work & Human Services
Aware ยท M6 ยท lesson 6 of 18 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
AI Terminology Every Caseworker Should Know
๐Ÿ“–
now learning

AI Terminology Every Caseworker Should Know

15 min

It is Tuesday afternoon and Marisela, a child-protective-services supervisor in a mid-size county agency, is sitting across a conference table from the agency's director, a program compliance officer, and an attorney from the state's fair-hearing office. The agency has just received an administrative complaint from a family who was denied emergency rental assistance after an AI-assisted eligibility tool flagged their application as ineligible. The family's advocate is asserting that the denial violated due-process requirements and that the AI system's risk logic constitutes an impermissible automated determination. The compliance officer turns to Marisela and asks: "What does the agency's AI tool actually do, and did the worker make the eligibility determination or did the system?" Marisela knows the tool. She uses it herself. But when she opens her mouth to answer, she realizes she is reaching for words she does not quite have. She says the system "helped" with the decision. The attorney from the fair-hearing office sets down her pen. "Does it inform the decision, or does it make it?" Marisela pauses. The next forty-five minutes belong to the advocate.

That situation turns entirely on vocabulary. Not on whether the agency acted in good faith. Not on whether the caseworker meant to do right by the family. On whether the people in the room can say precisely what the AI did, what the human did, and what the law requires of each. The terms are not jargon. They are load-bearing. Each one names a real thing that happens inside an AI-assisted case, and each one has a consequence attached: a family separated, a benefit denied, an appeal upheld, a record that withstands a court or an advocate.

This lesson defines twelve terms grouped into four clusters. The clusters follow the life of AI in a human-services case: how the models work and fail; how they touch case decisions and what that means for who owns the call; how equity and bias enter the picture; and how documentation, audit trails, and due process keep the work defensible. Every term is defined in the language you already use: the case note, the safety assessment, the eligibility determination, the court report, the fair hearing. This is not a data-science lecture. It is the vocabulary that protects the people you serve and the decisions you make.

How AI Models Work and Fail: The First Cluster

Before any other term makes sense, you need a plain-language picture of what a generative AI model actually does when it produces text. Then you can understand why it sometimes produces text that sounds correct but is factually wrong, and why that failure mode has a specific name with specific consequences in a case context.

Hallucination

Hallucination is the term for the phenomenon in which a generative AI model produces text that is fluent, confident, and factually incorrect or entirely invented. The model did not lie in the sense of knowing the truth and choosing to misrepresent it. The model generated statistically plausible text. "Statistically plausible" and "factually accurate" are two different standards, and in a case context they can be extremely far apart.

Here is a concrete example. A caseworker uses an AI-assisted documentation tool to draft a summary of a home visit. The caseworker's notes say the family's apartment was tidy and the children appeared well-fed. The AI draft adds: "Caseworker observed safety hazards in the kitchen, including unsecured cleaning products within children's reach." That observation was never made. It was never in the notes. The model generated it because, based on its training, observations about household safety are statistically common in home-visit summaries, and it filled the space with what typically appears. That fabricated observation is a hallucination. If the caseworker does not catch it before signing the note, it enters the case record as a documented fact. If the case goes to court, a family's attorney will cite it. If a removal decision is later evaluated, the fabricated observation is in the record as evidence that was never gathered.

Hallucination is not a bug that will be patched in the next software version. It is a structural consequence of how generative language models work. They are trained to predict what text should come next, and prediction produces plausible output whether or not the plausible output is true. The only reliable control for hallucination in human-services documents is human verification before the document is signed. The job has shifted from "produce the draft" to "verify the draft against what actually happened."

In the specific context of human services, hallucination takes three forms that matter most: invented observations (a detail that was never present at a home visit or interview appears in the AI-drafted note), misapplied policy (the model cites a rule or eligibility threshold that does not apply to this family's program or jurisdiction), and fabricated history (the model attributes a prior incident, prior finding, or prior contact to the family's record when no such record exists). Each form can harm a real family. A fabricated observation in a court report can be the deciding factor in a removal. A misapplied policy rule can generate a wrong eligibility denial that leaves someone without food or shelter. Fabricated history can follow a family through the system long after the document that introduced it is forgotten.

Hallucination is structural, not fixable. Every AI-drafted case document must be verified against what was observed and documented before it is signed.

Grounded Generation and Retrieval-Augmented Generation (RAG)

Grounded generation is the practice of providing a generative model with the actual source material it should draw on before generating output. Instead of drafting from general training patterns, the model receives the actual case notes, the actual policy text, and the actual record before it writes anything. Retrieval-augmented generation (RAG) is the automated implementation of that practice: a RAG system retrieves relevant documents from a database and includes them in the model's working context before generation begins.

The difference between a grounded model and an ungrounded model is the difference between a caseworker who drafts a note from their own visit notes and a caseworker who drafts a note from general memory of what home visits usually involve. The first worker is grounded on the specific facts. The second worker is generating from statistical patterns. In a legal record, only the first approach is acceptable.

A grounded AI documentation tool operating over the case record will substantially reduce the hallucination rate for the retrieved content. It can summarize what is in the notes without inventing what is not there. But "substantially reduce" does not mean "eliminate." Even a grounded model can misread the notes it was given, weight them incorrectly, or generate a transition sentence that introduces a claim not supported by the source material. Every grounded AI draft still requires human verification. Grounding reduces the verification burden; it does not eliminate it.

When a vendor tells you their AI documentation tool is "grounded on the case record," your follow-up questions should be: grounded on what, exactly? Only the caseworker's notes from this visit, or also the prior service history and prior case notes? Is the policy text current and jurisdiction-specific? What happens when a fact is absent from the record: does the model leave it blank, flag it, or invent a plausible substitute? The answers tell you how much verification work remains after the AI produces its draft.

Grounded generation reduces hallucination by anchoring the model to the actual record; it does not eliminate it. Grounding is the floor, not the ceiling, of verification.

AI in Case Decisions: Who Informs and Who Decides

The second cluster covers the terms that define the boundary between what AI can do and what a human must do in any case decision. These terms are not philosophical preferences. They map to legal obligations, to due-process rights, and to the specific forms of accountability that govern child-welfare and benefits work.

Decision-Aid Versus Decision-Maker

A decision-aid is a tool that provides information, analysis, summaries, or signals that a human decision-maker uses to make a better-informed decision. A decision-maker is the person or authority that actually makes the determination. In human services, the caseworker, supervisor, and court are always the decision-makers. An AI tool is always a decision-aid. These are not interchangeable roles, and the distinction is not a matter of degree.

The decision-maker owns the call. The decision-maker signs the document. The decision-maker can be called to account in a fair hearing or a court proceeding. The decision-maker has professional and legal responsibility for the outcome. A decision-aid has none of those properties. When someone says "the AI made the call" or "the system flagged this family for removal" or "the tool determined they were ineligible," they are using the vocabulary of a decision-maker to describe a tool that can only be a decision-aid. That framing is not just imprecise. It is a due-process problem. Families have the right to challenge decisions made by accountable humans. They do not have a meaningful avenue to challenge a determination attributed to an algorithm.

In 2026, the field's practice and research have firmly established this rule: large language models (LLMs, the AI technology behind the text-generating tools now entering human-services agencies) are decision-aids, not decision-makers, in sensitive public-services cases. This is not a limitation that will be overcome by a more capable future model. It reflects the fundamental structure of accountability in the work. No matter how accurate a risk model becomes, the decision to remove a child remains a human judgment with legal due-process requirements attached. The model can surface information. The caseworker decides what to do with it.

In practice, the decision-aid boundary is most at risk when caseloads are high, when a risk score is framed with high confidence, or when the agency's workflow does not explicitly require a documented human decision at the key steps. When a supervisor reviews a caseworker's action and the case record shows only an AI score with no documentation of how the caseworker evaluated and then exercised independent judgment, the decision-aid has effectively become the decision-maker in the record even if that was never the intent. The audit trail must show the human decision, not just the AI input.

AI is always a decision-aid in human-services cases. The caseworker, supervisor, and court are always the decision-makers. "The system flagged it" is never a sufficient reason for a consequential action.

Eligibility Determination

An eligibility determination is the official decision that a person or family does or does not qualify for a specific benefit, service, or program under applicable law and policy. It is a legal act. It triggers rights. A person who receives an adverse eligibility determination has the right to notice that the determination was made, the reasons for it, and an opportunity to challenge it through a fair hearing. These rights apply whether the determination was made with a paper form and a human policy lookup, or with an AI tool that applied policy logic to the intake data.

In an AI-assisted eligibility workflow, the AI tool may apply the program's policy rules to the applicant's information and return a recommendation, a flag, or a preliminary result. That output is not a determination. The worker who reviews the result, exercises judgment about whether the AI applied the correct policy to the specific facts of this case, and signs or records the decision is making the determination. The AI tool's output is one input in that process, not the final product.

The consequence of conflating AI output with eligibility determination is not abstract. If the AI tool applies the wrong policy rule, or applies the right rule to incorrect input data, or misinterprets an exception that applies to this specific applicant, and the worker treats the AI output as the determination without independent review, the family receives a wrong denial or approval. A wrong denial that is not caught before it reaches the family leaves someone without food, housing, or medical care. A wrong approval that is not caught may create an overpayment that the agency later tries to recoup from a family already in crisis.

When a supervisor asks a worker "how did you make this determination," the answer must describe a human judgment process, not a restatement of what the AI output said. "The system showed ineligible" is not an eligibility determination. "I reviewed the application, checked the income documentation against the program's threshold, confirmed the household composition matched what the applicant provided, and determined the household was ineligible under the income-limit rule" is an eligibility determination made by a person who can be held accountable for it.

An eligibility determination is a legal act made by an accountable human. AI output is an input to that process, never the determination itself.

Substantiation

Substantiation is the formal finding in a child-protective-services investigation that abuse or neglect occurred, made to a legally established standard of evidence based on the documented facts of the case. A substantiated finding enters the official record, may be disclosed in background checks for specific purposes under state law, and can affect a person's employment, licensing, and custody. It is among the most consequential findings a government worker makes about a private individual.

In an AI-assisted CPS (child protective services) workflow, an AI tool may analyze case notes, prior history, intake information, and other structured data to flag elevated concern, surface patterns across contacts, or recommend that a case be elevated for supervisory review. None of that is substantiation. Substantiation requires a documented finding by a trained investigator that the evidence meets the legal standard. The AI tool's output is one input the investigator may consider, alongside the interviews, the observations, the history, and the professional judgment that the investigation requires.

The risk of vocabulary confusion here is high. When an AI tool returns a "high-risk" flag on a case, workers under caseload pressure may experience that flag as a de facto substantiation, or at minimum as strong evidence pointing in one direction. Awareness of the decision-aid rule is the discipline that holds the line: the flag is information about what the model calculated from the data it was given. The substantiation decision requires human investigation of the full picture, including the context that the model did not have and could not weigh.

Substantiation is a legal finding made by a trained investigator to a documented evidentiary standard. No AI flag or risk score is a substantiation.

Mandated Reporter Context

A mandated reporter is a person designated by state law as legally required to report reasonable suspicion of child abuse or neglect to the appropriate authorities, typically CPS or law enforcement. The designations vary by state but commonly include teachers, physicians, social workers, and other professionals who work with children. Failure to report when legally required is itself a legal violation.

In an AI context, mandated reporter status introduces specific complications. First, some AI tools deployed in human-services agencies are being used to draft reports, review documentation, or triage incoming contacts in ways that touch the mandated-reporting workflow. The human mandated reporter cannot delegate their legal obligation to an AI tool. An AI system that reviews an intake call and returns a flag does not make a report. The person receiving or acting on that information remains legally responsible for their own reporting obligation. Second, an AI tool operating in a mandated-reporter context must be understood by every person using it: the tool's output does not substitute for the professional's own assessment, and the legal obligation remains personal and professional regardless of what any tool outputs. Third, when AI is used to draft or summarize documentation that becomes part of a mandated report, the grounding and verification requirements that apply to all case documentation apply with even greater force, because the accuracy of a report may determine whether a child receives protection or remains in a dangerous situation.

The legal obligation of a mandated reporter is personal and professional. An AI tool cannot fulfill it, delegate it, or discharge it.

Equity, Bias, and the Screening Vocabulary

The third cluster covers the terms that govern how predictive and screening tools are evaluated for bias and whether their outputs are used responsibly. These are the terms that connect the AI the agency deploys to the populations it serves, and they are the terms most likely to appear in an advocate's challenge or an oversight review of a screening program.

Predictive Risk and Risk Screening

Predictive risk refers to the use of statistical models to estimate the probability that a particular outcome will occur: that a family will experience a subsequent abuse or neglect report, that a benefit recipient will become homeless, that a youth in a diversion program will re-enter the system. Risk screening is the application of predictive risk methods at the intake or case-opening stage to triage resources, prioritize investigations, or flag cases for elevated attention.

Predictive risk tools in child welfare have been used in several jurisdictions for more than a decade, and they remain among the most ethically contested tools in the field. The Allegheny Family Screening Tool, deployed in Allegheny County, Pennsylvania, became a widely studied example of both the promise and the peril: it surfaced risk signals from administrative data at scale, but it also raised serious concerns about whether the data inputs it used encoded the inequities already present in the system, particularly the historical over-surveillance of Black families and low-income families who interact with public systems more frequently simply because of poverty and systemic disadvantage.

When you encounter a predictive risk or risk-screening tool in your agency, the foundational vocabulary question is: what does this score actually measure? A score that uses variables like "prior CPS contacts" as a risk indicator may be measuring risk of future harm, or it may be measuring historical likelihood that a family was investigated, which reflects the agency's prior investigation patterns as much as the family's actual circumstances. Those are different things, and using investigation history as a risk proxy can amplify the agency's past patterns rather than identify the families who most need help.

Every predictive risk score must be understood as one input among many, not as a verdict. The score does not know what the worker knows from the home visit. It does not know the context of this family's life. It does not know whether the data inputs were accurate. The worker's professional judgment, the supervisor's review, and the mandatory human review requirement are the structures that keep the score in its proper place.

A predictive risk score is one input among many under mandatory human review. It is never a verdict about a family and never a substitute for professional assessment.

Equity Audit

An equity audit is a systematic review of an AI tool's outputs to determine whether the tool produces meaningfully different outcomes across demographic groups, particularly groups defined by race, ethnicity, national origin, religion, disability status, or family composition. An equity audit asks: does this tool flag Black families at a higher rate than white families with similar risk indicators? Does it recommend denial of benefits for immigrant households more often than comparable citizen households? Does it score families differently based on inputs that serve as proxies for protected characteristics?

An equity audit is not a one-time event at the point of tool deployment. It is a continuous practice because the tool's behavior can change over time as data drifts, as the population it serves shifts, and as the patterns in the data it draws on evolve. A tool that passed an equity review in 2023 may be producing disparate outcomes in 2026 due to changes in the underlying data. Continuous monitoring is the requirement, not a single sign-off.

In practice, an equity audit for a human-services screening tool typically involves comparing tool outputs across demographic groups on the dimensions most relevant to the tool's use: investigation rates, substantiation rates, removal rates, benefit approval rates. If the tool produces systematically different outcomes for comparable cases across protected groups, the agency faces a choice: identify and address the source of the disparity, modify or replace the tool, or document the disparity and the mitigation measures in place. Continuing to use a tool with a known equity problem and no mitigation plan is not a sustainable posture for an agency that will face advocates, courts, and oversight reviews.

The equity audit is also the practice that closes the loop between the promise of AI in human services ("this tool will help us identify families who need help") and the documented history of risk-screening tools amplifying the patterns of the systems that generated the data they train on. History has taught this lesson: tools built on administrative data from systems that have historically over-surveilled and under-served certain communities will reproduce those patterns unless the audit catches the disparity and the agency acts on it.

An equity audit is a continuous practice, not a one-time check. Disparities found must be addressed, not noted and filed.

Disparate Outcome and Proxy Variable

A disparate outcome is the finding that an AI tool or algorithmic process produces significantly different results for members of different demographic groups, even when the inputs are otherwise comparable. The term does not require intent to discriminate. An AI tool can produce disparate outcomes through completely neutral-looking design choices if those design choices interact with patterns in the training data in ways that disadvantage certain groups.

A proxy variable is a model input that is not itself a protected characteristic but is statistically correlated with one in the population the tool was trained on, such that using it in the model effectively uses the protected characteristic indirectly. Zip code is the most discussed proxy variable in the human-services context: in many U.S. communities, zip code and race are correlated because of decades of residential segregation, meaning a tool that uses zip code as a risk input is effectively using race as a risk input through the correlation. Other commonly identified proxy variables in human-services tools include neighborhood-level poverty indicators, prior involvement with the child welfare system (which reflects historical patterns of who the system has surveilled), school absence rates (correlated with poverty), and housing stability indicators (correlated with income, race, and disability).

The key insight for caseworkers and supervisors is that a tool can produce disparate outcomes without any protected characteristic appearing explicitly as a model input. The pathway from protected characteristic to disparate outcome runs through proxy variables in the training data and the correlation patterns the model learned from them. When an advocate challenges a risk tool on equity grounds, they are often making a proxy variable argument: the tool uses inputs that serve as proxies for race or national origin, which means its outputs are influenced by those characteristics indirectly, and the outcomes it produces fall more heavily on certain groups as a result.

Understanding disparate outcome and proxy variable gives you the vocabulary to ask the right questions about any risk tool your agency deploys: What inputs does the tool use? Have those inputs been analyzed for correlation with protected characteristics in your agency's specific population? Have the tool's outputs been compared across demographic groups to identify disparate outcomes? Is there documentation of how those findings were addressed?

A disparate outcome does not require intent to harm. Proxy variables carry protected characteristics into a model indirectly, through correlation. The equity audit is the tool that finds them.

The fourth cluster covers the terms that define the legal and ethical perimeter around every consequential decision in human services. These are the terms that an advocate, a fair-hearing officer, or a court will use to evaluate whether the agency's use of AI was defensible. They are also the terms that, when understood and practiced, protect the families you serve from the specific ways AI can do harm.

Due Process

Due process is the constitutional and statutory guarantee that a person affected by a government action receives adequate notice of the action, an explanation of its basis, and a meaningful opportunity to challenge it before or after it takes effect. In human-services practice, due process governs the most consequential actions the agency takes: removing a child from a home, substantiating an abuse or neglect report, denying or terminating benefits.

In an AI-assisted context, due process has specific implications. First, a family must be able to understand the basis for a decision well enough to challenge it. If a caseworker tells a family "the system flagged your case" without being able to explain what the system flagged, on what basis, and how that flag was used in the decision, the family cannot meaningfully challenge the decision. Second, notice and explanation of AI use are becoming a required element of due process in many jurisdictions. Families have a right to know when AI was used in a determination that affected them, and a right to have that explained in terms they can act on. Third, the decision must ultimately be traceable to a human judgment that can be questioned: "the algorithm said so" cannot be the final answer in a due-process challenge.

In 2026, the legal landscape around AI and due process in public benefits and child welfare is active. Some states have passed or are considering requirements for disclosure of AI use in government decisions. Fair-hearing officers are increasingly sophisticated about the question of whether an AI tool was used as an information source or as a decision-making substitute. Agencies that have clear documentation of how AI was used, what the human judgment layer looked like, and how the decision can be explained to the affected family are in a defensible position. Agencies that cannot answer those questions are not.

Due process requires that a family can understand and challenge the basis for a decision. "The algorithm said so" is not a sufficient explanation in a due-process context.

Mandatory Human Review

Mandatory human review is the operational requirement that a trained, accountable human being reviews any AI output before it is used in a consequential case decision, and that the review is documented in the record. It is not a soft best practice. It is the structural mechanism that keeps the decision-aid rule from eroding under caseload pressure, time pressure, and the general human tendency to defer to confident-sounding outputs.

Mandatory human review has two components that both must be present. The review itself: a person with relevant professional competence and decision-making authority looks at the AI output, compares it to their own assessment of the case, and exercises independent judgment about what the case requires. The documentation of the review: the record shows that the review happened, who conducted it, what the AI output was, and what the human judgment was. The documentation is not bureaucratic overhead. It is the evidence that the review happened and that the decision is attributable to a person who can be held accountable.

In the context of risk-screening tools, mandatory human review means the risk score is one input the worker reviews alongside the home visit, the interviews, the history, and their professional judgment. It does not mean the worker reads the score and proceeds accordingly without independent assessment. It does not mean the worker notes the score in the record and signs off. It means the worker's professional assessment is documented as a distinct step that can be evaluated separately from the AI output.

The hardest situations for mandatory human review are the ones where caseloads are high and the AI output seems clearly correct. A worker with forty open cases who receives a risk score that matches their preliminary impression of the situation is under real pressure to treat the score as confirmation and move on. The discipline of mandatory human review is most important precisely in those situations, because the cases where the AI output is most plausible are also the cases where the worker is most likely to miss the exception the model could not see.

Mandatory human review is a documented process, not a formality. Both the review and the documentation must be present for the decision to be defensible.

Audit Trail

An audit trail is a time-stamped, sequential log of who did what in a system, and when. In a human-services AI context, the audit trail for an AI-touched case document includes: when the AI tool generated the draft, what inputs it was given, what version of the tool was used, who reviewed the output, what changes the reviewer made, and when the final document was signed or submitted. For a risk-screening tool, the audit trail includes: when the tool generated the risk score, what score was produced, who conducted the mandatory human review, when the review was documented, and what the outcome of the human review was.

The audit trail is the chain of accountability for the work. In a court proceeding or fair hearing, the audit trail is the evidence that the agency followed its own procedures, that a human being made the consequential decision, and that the AI tool played the role of decision-aid rather than decision-maker. The absence of an adequate audit trail is not simply a documentation gap. It is a structural vulnerability that an advocate or court will exploit: if the record does not show that a human reviewed and made the decision, the agency cannot demonstrate that a human did.

In practice, many case-management systems in use across the country as of 2026 were not designed with AI audit trails in mind. CCWIS (Comprehensive Child Welfare Information Systems), Casebook, FAMCare, and other platforms are adding AI-logging functionality as agencies deploy AI tools, but the integration is uneven and the logging requirements vary by state. Caseworkers and supervisors need to understand what their agency's systems are logging, what they are not logging, and what manual documentation steps are required to fill the gaps. The audit trail requirement does not disappear because the system was not built to support it automatically.

The audit trail is the chain of accountability. Without it, the agency cannot demonstrate that a human made the consequential decision rather than the algorithm.

Fair Hearing

A fair hearing is a formal administrative proceeding in which a person can challenge a government agency's determination affecting their rights or benefits. In human-services contexts, fair hearings are available for adverse eligibility determinations (the agency decided a person is not eligible for SNAP, TANF, Medicaid, or another benefit program), for certain child-welfare actions that affect family members' rights, and for other agency decisions governed by administrative law. The fair-hearing process gives the person a right to present evidence, to question the agency's basis for its decision, and to have the matter decided by a neutral hearing officer.

The presence of AI in the decision-making process does not eliminate or reduce a person's fair-hearing rights. It adds complexity to how those rights are exercised. A hearing officer evaluating an AI-influenced eligibility denial needs to understand how the AI tool was used, whether the worker exercised independent judgment, and whether the AI output was based on accurate input data. An advocate representing a family at a fair hearing will ask for documentation of how the AI was used and challenge any framing that suggests the AI made the determination. An agency that cannot explain the role of AI in the decision with specificity is poorly positioned in that proceeding.

The term connects directly to due process (the fair hearing is the primary due-process mechanism in benefits cases) and to the audit trail (the documentation the agency produces in the hearing). Understanding fair hearing as a vocabulary term means understanding that every AI-assisted determination the agency makes is a potential fair-hearing case, and the documentation practices around AI use are the agency's defense in those proceedings.

Every AI-influenced case determination is a potential fair-hearing case. Documentation of the human decision is the agency's defense when families exercise their right to challenge.

Putting the Vocabulary to Work in Three Conversations

These twelve terms are not definitions to memorize. They are tools for participating in conversations that are happening in every agency right now, at supervisor reviews, at leadership meetings, at vendor demonstrations, at fair hearings, and at oversight visits. Consider three conversations where the vocabulary changes what happens next.

The vendor demonstration. A software company comes to the agency to demonstrate a risk-screening tool for intake. The presenter says the tool has been shown to identify at-risk families earlier and is designed to be unbiased. A supervisor who knows the vocabulary asks: Has the tool been tested for disparate outcomes across demographic groups in populations similar to this agency's? What proxy variable analysis was done on the inputs? What does the mandatory human review workflow look like in practice: is there a documented step where the worker records independent judgment, separate from the AI output? Is there an audit trail that captures both the AI score and the worker's subsequent assessment? The vendor's answers to these questions are the difference between a tool that can be defended at a fair hearing and a tool that will create an equity challenge in eighteen months.

The supervisory case review. A supervisor is reviewing a CPS worker's documentation on a family where the risk score was elevated. The case file shows the AI score and a home-visit note drafted with AI assistance. The supervisor who knows the vocabulary asks: Is there documentation of mandatory human review that is separate from the AI outputs? Did the worker's assessment note draw any distinction between what was observed and what the AI added to the draft? Is the substantiation finding traceable to the worker's professional judgment, or to the AI score? Are there any observations in the case note that were not in the worker's raw visit notes? The last question is the hallucination check, and it is the one most often skipped under time pressure. Skipping it is how fabricated observations enter case records.

The fair hearing preparation meeting. The family has appealed a benefit denial and the hearing is in ten days. The eligibility worker, the supervisor, and the agency's legal representative are reviewing the file. The legal representative asks: what was the basis for the eligibility determination? The worker says the AI tool flagged the application as ineligible based on income. The supervisor who knows the vocabulary says: that is not the eligibility determination. The determination is the worker's documented judgment, based on their review of the application and the applicable policy. The AI flag is an input. The record needs to show the worker's reasoning, not just the AI output. If it does not, the advocate is going to argue the agency substituted automated processing for the required human review, and a fair-hearing officer may agree with them. The ten days are spent documenting what should have been documented at the time of the decision.

Each of those conversations has a different outcome with the vocabulary than without it. The supervisor who can name the specific failure mode is the supervisor who can prevent it before it reaches a hearing, a court, or a public accountability review. The caseworker who understands what an audit trail must contain is the caseworker who creates one at the time of the decision rather than reconstructing it under pressure. The worker who can distinguish hallucination from observed fact is the worker whose case notes hold up in court.

Key Takeaways

  • Hallucination is structural, not a software bug: generative AI models produce fluent, confident text that may be factually invented. In human services, hallucinated observations in case notes, misapplied policy in eligibility determinations, and fabricated history in risk assessments can harm real families. Every AI-drafted document must be verified against what was actually observed, documented, and recorded before it is signed.
  • Grounded generation and RAG (retrieval-augmented generation) reduce hallucination by anchoring the model to the actual case record and policy text before generating output. Grounding is the floor, not the ceiling: even a grounded model requires human verification. Ask vendors what exactly the tool is grounded on and what it does when a fact is absent from the record.
  • AI is always a decision-aid in human-services cases, never the decision-maker. The caseworker, supervisor, and court are always the accountable decision-makers. "The system flagged it" is not a sufficient reason for a removal, a substantiation, or an eligibility denial. The audit trail must show a documented human decision, not just an AI output.
  • Eligibility determinations and substantiation findings are legal acts made by accountable humans to defined evidentiary and policy standards. AI tool outputs are inputs to those processes, not the processes themselves. A worker who cannot describe their own reasoning beyond citing the AI output has not made a determination the agency can defend.
  • Predictive risk scores and risk-screening tools must be used as one input under mandatory human review. They are never verdicts about a family. Their outputs reflect patterns in training data, which in many jurisdictions have historically encoded disparities in how the child welfare and benefits systems have surveilled and served different communities.
  • Equity audits are continuous practice, not one-time sign-offs. Disparate outcomes (tools producing systematically different results for comparable cases across demographic groups) can arise through proxy variables (inputs correlated with race, national origin, income, or disability) even when no protected characteristic is used as a direct model input. Tools that pass an equity review at deployment must be monitored continuously and audited on a defined schedule.
  • Due process, the fair hearing, and the audit trail are the legal perimeter around every AI-influenced decision. Families have the right to notice, explanation, and a meaningful opportunity to challenge determinations that affect them. Agencies that cannot explain the role of AI in a decision with specificity, and show the documented human judgment that preceded the determination, are vulnerable in fair-hearing and court proceedings.
  • Mandatory human review must be both practiced and documented. Both components are required: the review itself (a trained person with decision authority exercises independent professional judgment) and the documentation of the review in the record. One without the other does not satisfy the requirement. Mandatory human review is most critical under exactly the conditions that make it hardest: high caseloads, time pressure, and AI outputs that seem obviously correct.