Summarization, Screening, Generation: Three Different AIs
A county child welfare agency sent its staff a one-paragraph announcement in early 2025: "Effective Monday, AI will be available to support your casework. Log in through the case-management portal." That was the entire briefing. Over the following two months, three things happened in parallel inside that same building. The documentation team adopted what turned out to be a transcription and summarization tool: it listened to home visits and produced draft case notes, returning roughly forty minutes per worker per day to direct contact with families. The intake unit, meanwhile, was quietly piloting a risk-screening tool that assigned numeric scores to incoming referrals based on factors pulled from the child welfare information system; supervisors were reviewing those scores in morning meetings and using them to prioritize investigations. And a third tool, connected to the same portal login, was generating draft court reports and placement letters from templates whenever a caseworker clicked "draft." Nobody in the agency understood that these were three completely different kinds of AI, built on three completely different architectures, carrying three completely different risk profiles, and requiring three completely different oversight practices. An internal review triggered by an attorney's challenge to a court report found all three operating simultaneously with a single use policy that said, in its entirety: "AI tools assist workers; all output must be reviewed before use." What had actually happened was that a summarization tool, a predictive risk model, and a generative prose engine had been conflated into one word: AI. That word was doing far too much work, and families were at stake.
Why the Single Word "AI" Is Dangerous in Human Services
In most industries, conflating different AI types produces inefficiency. In human services, it produces inequitable decisions, undermined due process, and harm to the families, children, and individuals the field exists to serve. Understanding why requires starting with what is actually at stake when a caseworker relies on AI output.
The decisions that human-services professionals make are among the most consequential any government makes about a person's life. To remove a child from a home. To substantiate a report of abuse or neglect. To approve or deny benefits that determine whether someone has food, housing, or medical care. To initiate a court proceeding that may define a family's future for years. These decisions are governed by due process, equity statutes, and the professional judgment of people who answer to a court. An algorithm, under any interpretation of current practice and law, cannot make them. AI informs; humans decide. That is the cardinal rule, and it is not a slogan. It is the structural requirement that governs everything else in this lesson.
But the field is complex enough that even caseworkers who hold this rule firmly in mind can be placed at risk by a confusion that precedes it: the confusion between what kind of AI they are using in the first place. A worker who understands that AI is a decision-aid, not a decision-maker, still needs to know whether the tool in front of her is compressing and distilling information she already has, predicting risk from patterns in historical data, or writing new prose from a template. These are not variations on a theme. They are fundamentally different processes with fundamentally different failure modes and fundamentally different stakes. Using them as if they were one thing is not a conceptual imprecision. It is a practical error that can change a family's outcome.
The paperwork burden in human services provides important context here. Caseworkers spend an enormous share of every working day documenting: case notes from home visits, court reports, intake summaries, eligibility determinations, progress records. That burden is the field's most widely cited driver of burnout and turnover. When an agency announces "we have AI," the first hope in most workers' minds is relief from that burden, and that hope is entirely legitimate. AI transcription and summarization tools (often described under names like Magic Notes, and now among the most widely adopted AI tools in the field) do return real hours to direct practice and human connection. That is the goldmine: the use case that is most clearly beneficial, most widely adopted, and most aligned with the field's fundamental mission of being present with the people being served.
But "AI" in an agency announcement may also refer to a predictive risk-screening tool that assigns a numerical score to a family based on patterns in historical case data, or to a generative prose tool that drafts a court report from a case summary. These three things are not the goldmine. They are different wells, with different depths, different water quality, and different risks if you drill in the wrong place. This lesson gives those three things their names, maps their failure modes, and explains why operating them as one category is dangerous when families' lives depend on the precision of your work.
The equity dimension before anything else
Before the technical distinctions, one principle must be fixed in place: equity is not an afterthought to the analysis of AI types. It is the first frame. Human-services agencies have historically served populations that have been subject to systemic inequity, and the AI tools deployed in this field have inherited the data produced by that history. A risk-screening tool trained on decades of child-welfare case data will encode whatever patterns existed in that data, including the ways in which poverty, housing instability, and differential rates of surveillance concentrated in communities of color were recorded as risk indicators. That encoding is not neutral. It is a technical mechanism by which historical inequity becomes predictive output. The lesson will return to this in detail when discussing screening tools. But it must be stated at the outset, before the taxonomy, because every use of the word "AI" in a human-services context should immediately raise the question: which families will this affect, and how do we know it will affect them fairly?
Summarization: The Goldmine, Lowest Risk, But Not Risk-Free
Summarization AI, in the context of human services, takes a longer input and produces a shorter output. The input might be a transcript of a home visit, a set of case notes across a case history, an audio recording of an intake session, or a collection of documents in a case file. The output is a condensed version: a draft case note, a summary of prior contacts, a compressed history for a new worker picking up a case. The model is not producing new information. It is reorganizing and condensing information that already exists.
This is why summarization tools are the goldmine and the lowest-risk category of AI in this field. The decision-aid boundary is clear: the AI condenses what was there, and the caseworker reads, verifies, edits, and signs. The worker's professional judgment enters the loop before the document becomes a legal record. The primary risk, which is real and must be managed, is the risk of dropping or distorting: content that was present in the original but absent from the summary, or content that was accurately captured but whose meaning was subtly changed by paraphrasing. A home visit in which a parent described struggling with utilities payment and a caseworker noted the kitchen was clean might be summarized as "family experiencing financial stress; home visit unremarkable," which compresses two separate observations into a combined characterization that loses meaning. Or the summary might silently drop the utilities comment entirely if the model weighted it as low-signal. Neither of these is invention. Both of them are loss.
What summarization models actually do
A summarization model is a type of language model trained to produce condensed, coherent text from longer inputs. When it encounters a home visit transcript, it is doing something like identifying which sentences carry the most semantic weight and organizing them into a shorter structure. It does not understand the case in any meaningful sense. It identifies patterns in text that, in its training data, were associated with summaries: short, declarative statements organized around the most prominent information in the source. It does not know that a caseworker's notation about a missing child's presence in the home is not low-signal. It does not know that the parent's comment about a relative visiting was not incidental. It organizes by statistical salience, not by clinical judgment.
This means the caseworker using a summarization tool is doing a specific and important job: verifying that the summary has preserved the clinically relevant observations, not merely the statistically prominent ones. That is a professional task, not a clerical one. The note is a legal record. The parent's description of the child's behavior during the visit may matter in a court proceeding. The caseworker who signs the note is attesting to its accuracy. AI-assisted summarization accelerates the drafting of that note; it does not reduce the professional obligation to verify it against what was observed.
The dropping and distorting failure modes
Two specific failure modes occur with enough regularity in summarization tools to deserve names.
Dropping. The model omits a detail that the caseworker observed and documented in their raw notes but that the model judged, by its statistical weighting, to be insufficiently prominent for inclusion. A child's statement about sleeping arrangements. A parent's disclosure of a recent hospitalization. A notation about a pet's condition in a home that is otherwise being assessed for hygiene. These omissions are invisible in the summary: you cannot see what is not there unless you have the original. The safeguard is the verification practice the field requires for exactly this reason: the caseworker reads the summary against what actually happened, not against what sounds right.
Distorting. The model preserves a detail but changes its meaning through compression or paraphrasing. "The child said she did not want me to leave" becomes "child expressed attachment to caseworker." "The father appeared agitated during the visit and became louder as I explained the safety plan" becomes "father engaged with safety plan discussion." These are not fabrications in the technical sense (the model did not invent an event). They are transformations that change what the record says. In a court context, a judge reading "engaged with safety plan discussion" is receiving different information than a judge reading "became louder as the caseworker explained the safety plan." Distortion can favor either the family or the agency's position, depending on how the compression lands. It is not safe because it is not invention. It is a different kind of risk.
The management of these failure modes does not require abandoning summarization AI. It requires a verification discipline proportionate to the stakes: for case notes, at minimum, a field-by-field comparison of the summary against the worker's raw observations before the note is signed. For court reports, a claim-by-claim verification against the documented record. The tool accelerates the draft; the professional standard governs the final record.
Why summarization is still the goldmine
Despite these failure modes, summarization remains the field's most beneficial and least ethically fraught AI use case for a simple reason: it is bounded. The model is working from the caseworker's own input, not from external historical data or statistical patterns about populations. The output is checked by the person who has the most information about its accuracy, before it becomes a record. The benefit is real and significant: forty minutes per worker per day returned to direct practice is not a rounding error. At a unit of twenty workers, that is more than thirteen hours of reconnection per day. Scaled across an agency, it is thousands of hours per year redirected from documentation to the human work the field was built for. That is the humane case for AI in this context, and it is solid.
Risk Screening: The Highest Ethical Stakes
A predictive risk-screening tool is a different kind of AI entirely. It is not summarizing information the caseworker already has. It is making a statistical prediction about a family's future based on patterns in historical case data. The model takes structured inputs, typically fields from the case management system (Comprehensive Child Welfare Information System, or CCWIS, is the federal framework for these systems) and other public or agency records, and it produces an output that is usually a score or a risk tier. The score represents the model's estimate of some outcome: the probability that this referral involves substantiated abuse or neglect, the probability that a family will require removal services within a defined period, the probability that a child's safety is at elevated risk. The caseworker or supervisor then uses this score, alongside their professional assessment and the case record, in deciding how to prioritize or respond.
That description, stated plainly, reveals immediately why this is the highest-stakes AI category in human services. The model is producing a prediction about a family's risk level based on patterns it found in historical data about other families. The families in that historical data were not randomly sampled from the population. They were families who were already in the child welfare system: families who had been reported, investigated, and had records generated about them. Those families were not a neutral cross-section of society. They were disproportionately families experiencing poverty, housing instability, and other stressors that correlate in the United States with race, ethnicity, and immigration status, not because families from those communities are inherently at greater risk, but because the surveillance that creates child welfare records has historically been unevenly distributed. A predictive model trained on this data will find real statistical patterns. The question the field must ask is not whether those patterns exist in the data; it is whether those patterns are equitable predictors or encoded replications of systemic inequity.
The Allegheny County experience and what it means
The Allegheny Family Screening Tool, deployed by Allegheny County, Pennsylvania starting in 2016, is the most extensively studied example of child-welfare AI risk screening in the United States. Its history is not a simple story of failure. It was designed with documented equity considerations, has been subject to ongoing academic scrutiny, and its administrators have engaged publicly with critics. But the debate it generated is instructive precisely because it is a sophisticated tool that has nonetheless produced contested outcomes. Critics identified patterns suggesting that poverty-related factors in the model's inputs produced scores that, in practice, were correlated with race in ways that affected which families received intensive intervention. The tool's defenders noted it was designed as a decision-aid and that caseworkers were trained to treat it as one input. The debate has not been resolved, and the tool continues to be studied.
What the Allegheny experience establishes for this lesson is not that predictive risk tools are categorically wrong. It is that they are irreducibly complex in ways that demand sustained professional engagement. Even a well-designed tool, with documented equity analysis and ongoing monitoring, operating in a context where leadership treats it as a decision-aid rather than a decision-maker, can produce outcomes that raise equity questions years into deployment. The lesson from that history is not "do not use risk tools." It is: use them with open eyes, continuous equity auditing, and the absolute understanding that the score is one audited input, never a verdict.
How screening models actually work and where they break
A predictive risk model takes structured data, usually numerical or categorical fields from a case management system, and produces a numerical output. It learned the mapping between inputs and outputs by training on historical cases where both the inputs and the ultimate outcomes (substantiation, removal, re-referral) were known. For any new case, it applies the learned mapping to the available inputs and produces a score. The score is a probability estimate, or in some implementations a risk tier label, and it is inherently a population-level statement applied to an individual: this family's inputs resemble the inputs of families in the training data who had outcome X with frequency Y.
The fundamental problem with applying population-level statistics to individual families is not unique to AI. It is a general problem in human services risk assessment that long predates machine learning. Actuarial risk tools of various kinds have been used in child welfare, criminal justice, and mental health for decades. The AI version intensifies the problem in one specific way: the model may be discovering statistical patterns that are not interpretable to the caseworker using it. A human evaluator using a structured assessment tool knows which factors are being weighted and can articulate why. A machine learning model using dozens of variables with complex interactions may produce a high-risk score that no one, including the model's developers, can fully explain in terms meaningful to a caseworker, a supervisor, or a family being assessed. That interpretability deficit is particularly important in due process terms: a family has a right to challenge a decision that affects them, and a challenge requires knowing why the decision was made.
The specific failure modes of risk-screening tools include:
- Proxy variable encoding. A model that does not include race as an input can still produce racially disparate outputs if its inputs are correlated with race in the population. The number of prior child welfare contacts, involvement with public assistance programs (Supplemental Nutrition Assistance Program, or SNAP; Temporary Assistance for Needy Families, or TANF; Medicaid), housing instability, and many other factors that appear in CCWIS data are correlated with race in ways that reflect structural inequity. A model that weights these factors heavily will produce scores correlated with race without ever "seeing" race.
- Training data feedback loops. A model trained on historical decisions learns not just about family risk but about agency behavior. If the historical data shows that families in a particular neighborhood were investigated more frequently, the model may learn that neighborhood as a risk indicator. More scrutiny produces more records, and more records produce higher scores, which produce more scrutiny. This is a feedback loop that amplifies surveillance inequity over time.
- Calibration drift. A model trained on data from one time period may become less accurate as the population and the agency's practices change. Risk factors that were predictive in 2015 may not be predictive in 2025. Regular re-calibration and ongoing accuracy monitoring are required, but they are not always performed in busy agencies with constrained resources.
- Alert fatigue and over-reliance. When scores are presented repeatedly in morning meetings and have repeatedly matched subsequent outcomes, supervisors and workers may begin to weight them more heavily than intended. The score becomes the primary frame, and the professional assessment becomes the exercise of documenting why the score was correct. This is a cultural failure that is as dangerous as any technical one. It is the mechanism by which a decision-aid becomes a decision-maker without anyone explicitly making that choice.
What mandatory human review actually means
The field's established norm is that a risk score is one audited input under mandatory human review. This phrase is worth unpacking carefully, because it can be implemented in ways that satisfy it formally while violating it in practice.
A risk score is one input means it is one consideration among the full professional assessment: the caseworker's direct observation of the family, the family's history as interpreted by someone who has read the record rather than a model that has processed its structured fields, the family's own description of their situation, the collateral contacts, the home environment. The score does not replace any of these. It does not summarize them. It is an additional data point with specific, documented limitations that the professional is required to hold in mind when using it.
Mandatory human review means a documented, required step in which a trained professional reads the score alongside the full case information and produces a written assessment of how the score was considered in forming the professional judgment. It is not a formality. It is a due-process requirement: the person affected by the decision has a right to know how it was made, and the record must reflect human reasoning, not algorithmic output. A case note that says "risk score was elevated; worker concurred" does not satisfy this requirement. A case note that says "the tool produced a score of 72 out of 100 based on factors including number of prior contacts and housing instability; worker assessment considered these factors alongside the following direct observations and concluded..." is closer to what is required.
Audited means the tool's performance is monitored on an ongoing basis for accuracy, calibration, and equity. Does the tool produce scores that are accurate, in the sense that cases scored as high-risk are actually more likely to result in substantiated harm? Does it produce scores that are equitable, in the sense that false positives and false negatives are not concentrated in particular demographic groups? These are not one-time questions answered at deployment. They are ongoing operational requirements that must be built into the agency's practice.
Generation: The Prose Writer That Can Invent
A generative AI tool (a large language model, or LLM) produces new text from a prompt. Unlike a summarization tool that condenses what you give it, a generative model produces prose that is statistically consistent with patterns in its training data, regardless of what was actually observed in the case. In a human-services context, generative AI is used for drafting court reports, writing placement letters, composing communications to families, and producing narrative sections of assessments. The output looks like professional prose. It often reads better than what a time-pressured caseworker can produce at the end of a long day. It has the structure, tone, and vocabulary of field practice. These qualities are what make it useful and what make it dangerous in almost equal measure.
The danger is the one failure mode that does not exist in summarization or screening tools: fabrication. A generative model can invent an observation that was never made, cite a contact that never occurred, describe a parent's demeanor based on statistical patterns in its training data rather than anything the caseworker actually witnessed. This is not a bug in a specific product or a failure mode that better engineering will eliminate. It is a structural property of how large language models work. They predict text. They predict the next word, then the next, based on what words are most likely to follow in sequences similar to their training data. When asked to write a court report for a child welfare case, the model produces what court reports for child welfare cases look like: observations about home conditions, descriptions of parent-child interactions, assessments of the family's engagement with services. If the prompt does not contain specific information about what was observed on a specific visit, the model fills in the gaps with plausible, professionally-sounding prose that was never grounded in reality.
The fabrication failure modes in field practice
Three specific fabrication patterns occur in generative AI used for case documentation, and they deserve to be named by field workers who will encounter them.
The invented observation. A caseworker prompts a generative tool with a brief summary of a home visit and asks it to draft the case note. The summary says "visited family, discussed safety plan, observed both children in the home." The model produces a note that includes "children appeared healthy and engaged in age-appropriate activities; home environment was clean and organized; mother demonstrated understanding of safety plan requirements." None of the items after the first clause were in the original summary. They are plausible observations for a case note of this type. They are not observations the caseworker made. If signed without verification, they become part of a legal record. In a subsequent hearing, the parent's attorney may question the caseworker about the specific observations in the note, and the caseworker may not be able to confirm details she did not actually record.
The invented contact or service. A generative tool drafting a court report from a case summary may include references to service contacts, provider visits, or family engagements that are not in the summary but are statistically common in the type of report it is generating. "Family participated in ten sessions of in-home family support services" is a sentence that appears in court reports of a certain type. If the model is drafting a report for a case where in-home services were offered but the family attended only three sessions, and the prompt does not specify the count, the model may generate the statistically common number. The error is not random. It is a plausible error that, in a court, could affect the judge's assessment of the family's engagement with services.
The invented clinical judgment. Generative models asked to assess risk or describe a family's situation will produce assessments that reflect patterns in their training data, not the specific family in the case. "The family demonstrates resilience and motivation to change" is the kind of summary statement that appears in case records for families where services have been provided for a period of time and the case is moving toward closure. If the model generates this statement for a case where the situation is more complex, the professional reading the draft may accept the framing without verifying whether it reflects the actual record. Clinical language produced by a model is particularly insidious because it sounds authoritative and professional. Workers are trained to read critically for factual accuracy; they may read with less skepticism for clinical framing because it sounds like what a professional would write.
Grounded generation versus free generation
The distinction that makes generative AI safer in a case context is the distinction between grounded generation and free generation. A grounded generative tool is one that produces its output from a specific, supplied context: the case notes are in the prompt, the home visit transcript is in the prompt, the policy framework is in the prompt. The model is directed to produce prose only from what is in the supplied documents, and it is instructed not to add information that is not there. A free generative tool is one that produces prose from the prompt and its general training, filling in from statistical patterns wherever the prompt is sparse. The practical difference is significant: a grounded tool's output contains only what was in the documents provided to it. A free tool's output may contain much more, including things that were not in any document and did not happen in any visit.
Retrieval-augmented generation, or RAG, is the technical architecture that supports grounded generation in many agency tools. In a RAG system, the model's output is produced with specific retrieved documents (the case file, the relevant policy sections, the most recent contact notes) serving as the context, rather than the model generating from general patterns. RAG substantially reduces (though does not eliminate) the fabrication risk. A RAG-based case documentation tool that is given the actual transcript of a home visit and the actual case history will produce a draft that stays much closer to what actually happened than a free generative tool given only a brief prompt.
But RAG is not a substitute for verification. Even a well-grounded generative tool can produce distortions: it may weight certain parts of the source documents more heavily than others, it may paraphrase in ways that change meaning, and it may occasionally produce details that are in the vicinity of what is in the documents without being exactly what the documents say. The court-record standard still applies: every factual claim in the output must be verified against the source, because the document that gets signed is a legal record, and the worker who signs it owns it professionally and legally.
The Three Tools in the Same Portal: A Worked Example
Return to the agency from the opening scene, and trace what happens when a caseworker, Maria, opens her case-management portal and sees a single "AI Assist" button. She does not know which tool is behind it at any given moment. Depending on which screen she is on, the button activates a different function. Here is what is actually happening, and here is what is at risk.
The Monday morning home visit
Maria conducts a home visit with a family that has had an open case for four months. She records the visit using the portal's audio capture feature. After the visit, she clicks "AI Assist" on the case note screen. The summarization tool processes the transcript and produces a draft note. The draft includes accurate observations about the home environment and the children. But it omits the parent's comment, midway through the visit, about a recent argument with her partner that resulted in the partner leaving for two weeks. The comment was brief, the parent moved on quickly, and the model's weighting did not flag it as high-signal. Maria, reading the draft at the end of a day with three more visits, does not compare it clause by clause against the transcript. She edits two sentences, adds her initials, and saves. The omission becomes part of the record.
Two months later, at a review hearing, the parent's attorney asks about changes in the household composition during the period. The case record shows the partner present throughout. There is no documentation of the partner's absence. Maria remembers the conversation but cannot verify it from the record. The attorney challenges the accuracy of the documentation. This is the dropping failure mode: real information, absent from the record because the summarization tool weighed it as low-signal and the verification was insufficient.
The intake unit on a busy Tuesday
The intake supervisor, Victor, receives twelve new referrals. He reviews the morning risk scores in the team meeting. Four cases are flagged as high-risk. One case with a score of 78 out of 100 is assigned to an experienced investigator for same-day response. The investigator, pressed for time, knows the score is high and focuses her initial questions on the factors she knows are weighted in the tool: prior contacts, housing stability, substance use history. She does not have time to read the full case history before the visit. The investigation proceeds with the score serving not as one input but as the organizing frame. The family is Black, living in public housing, with two prior contacts neither of which resulted in substantiation. The score was elevated primarily by the prior contacts and the housing instability, factors that are, in this family's history, reflections of systemic circumstances rather than indicators of ongoing risk to the children. The investigation is more intrusive than it would have been without the score framing the worker's attention. The children are found to be safe. The family is left with the experience of having been intensively investigated, again.
Victor's agency has not asked: does this tool produce higher scores, at equivalent actual risk levels, for families in public housing? No equity audit has been performed since deployment. No one has mapped the relationship between the tool's inputs and demographic outcomes. The case was handled under the tool's influence in a way that respected the letter of "one input under human review" while violating its spirit. This is the alert fatigue and framing failure mode, operating in combination with the proxy variable problem.
The court report on Wednesday afternoon
Maria has a court report due Friday for the same family. She opens the draft feature, which is the generative tool. She enters a brief summary: "family has been open four months, safety plan in place, children in school, parent engaged." The model produces a four-page court report. It includes sections on home conditions, family engagement, service participation, and a summary assessment. The service participation section includes a reference to "eight documented contacts with the family support services provider." Maria knows the family attended five sessions. She does not catch the discrepancy because she is reading for tone and structure rather than fact-checking each sentence against the case record. The report is submitted.
At the hearing, the case manager from the family support services provider mentions five sessions. The judge asks about the discrepancy. The report says eight. Maria's casework documentation says five. The court report, a legal document under her signature, contains a number she cannot explain. This is the invented contact failure mode: the model generated the statistically common number for a case at this stage, and it entered the legal record.
What this scenario teaches
Three tools, one portal, one uninformed caseworker, three different failure modes that affected the same family across a single week. The summarization tool dropped a clinically relevant detail. The screening tool introduced a framing bias that affected the conduct of an investigation. The generative tool invented a fact that ended up in a court report. None of these failures required the AI to be defective. They required only that the tools were used without understanding their nature: what each does, how each fails, and what professional practice each requires in response.
The remedy is not to remove the tools. The summarization tool gave Maria real time back and produced an accurate note on most of the visit. The screening tool correctly identified several high-priority cases that week. The generative tool produced a court report that was substantially accurate in its structure and framing. The tools are useful. The failure was in not knowing what kind of useful each one was, and therefore not applying the specific discipline each one requires.
Asking the Right Questions When a Vendor Says "AI"
When an agency director, a supervisor, or a procurement team is evaluating an AI tool for a human-services context, the word "AI" in a vendor's description is not information. It is a prompt to ask specific questions. Those questions are different for each of the three tool types, and a vendor who cannot answer them clearly is selling a product the agency cannot safely operate.
For summarization and documentation tools
The questions that matter for a summarization tool are: What is the tool's input: is it working from audio transcripts, written notes, structured data fields, or all three? How does the tool handle gaps in the input, meaning, if a section of audio is unclear or a detail is missing from the notes, what does the model do? Does it flag the gap, leave it blank, or fill it with a plausible completion? What is the "grounding" architecture: is the output restricted to information present in the input, or can the model draw from its general training? What does the output look like when the input contains conflicting information? Is there a mechanism to compare the output against the source? What does the agency's pilot data show about dropping rates, meaning what percentage of the time was a documented detail absent from the summary?
A vendor who can answer these questions with data from pilots conducted in child-welfare or social services contexts (not just healthcare or legal, though those are adjacent) is a vendor who understands the stakes. A vendor who answers "our model is highly accurate" without quantifying accuracy against a specific ground truth in a human-services context is asking the agency to learn accuracy the hard way.
For risk-screening tools
The questions that matter for a risk-screening tool are substantially different and substantially more demanding. What is the model predicting, precisely: what outcome, over what time horizon, using what inputs? What does validation look like: was the model validated on a population similar to this agency's client population, and what were the accuracy metrics (sensitivity, specificity, positive predictive value, negative predictive value) for protected subgroups, not just the overall population? What equity audit was conducted before deployment, and what were the results? What ongoing monitoring is built into the product or required of the agency? What does the vendor recommend as the documented human review process? How should workers be trained to use the score as one input rather than allowing it to become the primary frame?
The particular importance of the equity audit question cannot be overstated. An agency that deploys a risk-screening tool without equity audit documentation is, in effect, running a live experiment on families with no baseline data for comparison. History, including the Allegheny debates and the documented harm of benefits fraud-detection tools like the Dutch childcare-benefits scandal and Michigan's MiDAS system, shows that this is not a theoretical risk. These tools have harmed real families when deployed without adequate equity analysis and oversight. The burden of proof must be on demonstrating equity before deployment, not investigating harm after.
For generative tools
The questions that matter for a generative drafting tool are: Is the tool's output grounded on supplied documents, or does it generate from general patterns? What specific controls prevent the model from producing text not present in the source documents? What does the tool do when the prompt is sparse or ambiguous: does it flag that it lacks sufficient information, or does it produce plausible-sounding content? Is there a human review workflow built into the tool that requires verification before a document is finalized? What are the tool's documented hallucination rates (meaning the rate at which it produces confident, specific text not supported by the source documents) in human-services documentation contexts?
These questions are asking about the specific failure modes described in this lesson. A vendor who has not thought about them in those terms, and who cannot produce pilot data showing how the tool performs against a court-record standard of accuracy, is not ready for a human-services deployment. The standard is not "does the tool produce good drafts." The standard is "does the tool produce drafts that a caseworker with a full caseload and limited time can verify to a court-record standard." Those are different questions, and the second one is the right one.
The question that applies to all three
Regardless of which tool type is being evaluated, one question applies in every case: who is responsible for what the tool produces when it affects a decision about a family? The answer, in every human-services context, is the same: the caseworker, supervisor, and agency are responsible. The tool is a decision-aid. The vendor's product performance documentation is relevant to procurement and contracting. It does not transfer professional or legal accountability from the agency to the vendor. Agencies that do not understand this before deployment discover it during litigation or administrative review.
Key Takeaways
- The word "AI" in a human-services context describes at minimum three distinct technologies: summarization tools that condense information already present in a case record, predictive risk-screening tools that produce risk scores from historical data patterns, and generative drafting tools that write new prose from prompts. Each has a different architecture, different failure modes, and different required oversight practices. Treating them as one category is a professional and operational error with real consequences for families.
- Summarization tools are the goldmine: the most widely adopted, most clearly beneficial, and lowest-risk AI use case in the field. They return documented time to direct practice. Their failure modes (dropping details and distorting meaning through compression) require a verification discipline proportionate to the stakes of the record. For case notes and court reports, that means reading the summary against what was actually observed and documented, before signing.
- Risk-screening tools carry the highest ethical stakes of the three types because they make statistical predictions about individual families based on patterns in historical data. Those patterns can encode systemic inequity: proxy variables for race and poverty that produce disparate outcomes without ever processing a protected characteristic. Every risk score is one audited input under mandatory human review, documented in the case record as part of the human professional's reasoning, and never the operative basis for a consequential decision.
- Generative AI tools can invent observations, contacts, and clinical judgments that were never part of the case. This is not a product defect: it is a structural property of how large language models work. A generative tool used without grounding on the case record and without verification of every factual claim against the source introduces fabricated information into legal documents. A grounded tool (retrieval-augmented generation, or RAG) substantially reduces this risk but does not eliminate it. Verification remains mandatory.
- Equity is the first frame for every AI tool in human services, not an afterthought. The people served by these agencies have historically been subject to differential surveillance and systemic inequity that is encoded in the data these tools are trained on. An equity audit is not optional for risk-screening tools: it is a prerequisite for deployment, and ongoing equity monitoring is a requirement for continued use. The burden of proof runs from the tool to the families it affects, not the other way around.
- Due process (CPS: child protective services investigations, substantiation proceedings, benefits eligibility determinations) governs every consequential decision in this field. A family has a right to know why a decision affecting them was made, to challenge that decision, and to receive a record that reflects what actually happened. AI-assisted documentation that drops details, distorts meaning, or invents content undermines each of these rights. Court-record-standard verification is a due-process requirement, not an administrative preference.
- When a vendor presents an AI tool for human-services use and describes it simply as "AI," the agency's obligation is to ask the specific questions that correspond to each tool type: for summarization, questions about grounding and dropping rates; for screening, questions about equity audits, validation methodology, and ongoing monitoring; for generative tools, questions about hallucination controls, grounding architecture, and the required human review workflow. A vendor who cannot answer category-specific questions with category-specific evidence is not ready for deployment in a context where families' lives depend on accuracy.
- The professional and legal accountability for every AI-touched document in human services belongs to the caseworker, supervisor, and agency, not to the vendor. The AI is a decision-aid. Signing a document means attesting to its accuracy. An agency that does not build the verification and oversight practices required for each AI type into its workflows before deployment is accepting responsibility for consequences it does not yet understand.
Skill.re