โ†
AI for Social Work & Human Services
Strategic ยท M7 ยท lesson 7 of 19 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Evaluating Human-Services AI Vendors
๐Ÿ“–
now learning

Evaluating Human-Services AI Vendors

15 min

The demo was flawless. The vendor's sales engineer shared her screen, pasted in a sample home-visit transcript, and clicked one button. Eleven seconds later a clean, complete case note appeared: organized into the agency's preferred headings, written in confident professional prose, even formatted to drop straight into the state's child-welfare information system. The room of supervisors and the deputy director leaned in. Someone said it looked better than what their workers produced by hand. The vendor smiled and said the tool was already live in nine agencies. The deputy director was ready to sign. The one person who did not lean in was the agency's quality manager, who had spent the last hour writing a single question on her legal pad: when this tool invents an observation that was never in the transcript, how will we know, who is accountable, and what does the audit trail show a judge? That question, not the demo, is the entire job of evaluating a human-services AI vendor. The demo shows you the tool on its best day with a clean input. The evaluation tests whether the tool can survive an advocate, a fair hearing, and a court on its worst day with a real case.

Why the Demo Is the Trap

A vendor demo is a sales instrument, not an evaluation. It is built to show the tool succeeding. The input is curated, the use case is the one the tool handles best, and the failure modes that matter most in human services are precisely the ones a demo will never surface. This is not because vendors are dishonest. It is because a demo is structurally incapable of showing you what you actually need to know, which is how the tool behaves on the inputs you cannot control, in the situations where a mistake separates a family or denies a person food.

Consider the math of the curated demo. The vendor showed one transcript and one output. Your agency, if it runs a unit of fifteen caseworkers each carrying twenty-five families, will generate hundreds of case notes a month and dozens of court reports. A tool that produces a perfect note on a clean transcript ninety-five percent of the time will still produce a flawed note roughly one time in twenty. At the volume of a single unit, that is dozens of flawed notes a month entering legal records. The demo showed you the ninety-five. Your evaluation has to find the five, and it has to find them before the tool is live, not after a flawed note has already reached a courtroom.

The evaluation also has to account for a difference that vendors rarely volunteer: the gap between a general-purpose AI capability and a human-services-grade one. A model that drafts a fluent note is using the same underlying mechanism that lets it invent a plausible observation when the transcript is ambiguous. Fluency and fabrication come from the same place. The vendor's demo proves fluency. It does not prove the tool resists fabrication, grounds its output in the actual record, surfaces its uncertainty, or produces a trail that an advocate and a court would accept. Those are the properties you are buying, and the demo does not test any of them.

A demo shows you the tool on its best day with a clean input. An evaluation tests whether it survives an advocate and a court on its worst day with a real case.

The Five Evaluation Domains

A defensible vendor evaluation in human services is organized into five domains, and a tool must pass all five, not average across them. Averaging is how a tool with a brilliant interface and no audit trail gets bought. Each domain maps to one of the program's non-negotiables: grounding maps to documentation that stays grounded, the decision-aid boundary maps to the cardinal rule that humans decide, verification support maps to the court-record standard, equity maps to equity-first, and privacy and audit map to the due-process and privacy perimeter. Treat each domain as a gate. A tool that fails any single gate fails the evaluation, regardless of how it scores elsewhere.

Domain One: Grounding and the Fabrication Test

The first domain asks whether the tool grounds its output in the actual source material or generates free-form prose that merely sounds right. Grounding here means the technique sometimes called retrieval-augmented generation (RAG, an approach that connects the model to a specific document set such as the case record before it generates output, so the output is drawn from that set rather than from the model's general training). A grounded tool can tell you which source passage supports a given sentence. An ungrounded tool produces fluent text and leaves you no way to trace any sentence back to a source.

You test this domain by feeding the tool real but deliberately imperfect inputs, the kind your workers actually produce. Give it a fragmentary field note with gaps. Give it a transcript where the audio was unclear in places. Give it a case with an ambiguous detail. Then read the output looking for one specific thing: did the tool invent any observation, fact, or detail that was not in the input? In a structured evaluation, you run twenty to thirty such inputs and count the fabrications. A tool that fabricated an observation on three of thirty imperfect inputs has a ten percent fabrication rate on the inputs that matter most, and at unit volume that is a stream of false statements entering legal records. The demo never showed you this number because the demo never used an imperfect input.

Domain Two: The Decision-Aid Boundary

The second domain asks whether the tool respects the cardinal rule that AI informs and humans decide, or whether it quietly crosses into making the call. This is the most consequential domain in human services because the decisions here, to remove a child, substantiate a report, or deny benefits, are bound by due process and cannot be delegated to a model. A tool that produces a recommendation, a score, or a yes-or-no determination that a tired worker on a crushing caseload will simply accept has crossed the boundary in practice even if the contract says a human decides.

Evaluate this domain by examining what the tool outputs and how it presents it. Does the tool draft documentation and organize information, leaving the determination to the worker? Or does it output an eligibility decision, a risk verdict, or a recommended action dressed up as a suggestion? Ask the vendor directly: where in your tool does a human make the consequential decision, and what stops a worker from rubber-stamping the tool's output? A vendor whose answer is "the worker can always override" has not understood the question. The default matters more than the override option. If the tool's default output is a decision and overriding requires extra effort, the design pushes workers toward accepting the model's call, and under caseload pressure they will. The boundary has to be built into the default, not offered as an exception.

Domain Three: Verification Support

The third domain asks whether the tool helps the worker verify its output to a court-record standard, or whether it makes verification harder. Verification is the job in AI-assisted human services: the work shifted from producing the draft to checking the draft. A tool that produces a perfect-looking draft with no way to check it has not reduced the agency's risk; it has hidden the risk inside fluent prose.

A tool that supports verification shows the worker, for each factual claim in the draft, the source passage it came from, so the worker can confirm the claim in seconds rather than reconstructing the entire visit from memory. It flags claims it is uncertain about rather than stating everything with equal confidence. It distinguishes what it extracted from the record from what it inferred. Test this by timing verification: take a real AI-drafted note and measure how long it takes a worker to verify every claim against the source. If verification of a single note takes longer than writing the note by hand would have, the tool has not returned any time; the hours it saves in drafting it spends in verification, and the agency's only gain is a more dangerous failure mode. The genuine time dividend appears only when the tool makes verification fast, and that is a property you must measure, not assume.

Domain Four: Equity and Bias

The fourth domain asks whether the tool has been tested for disparate outcomes across the populations the agency serves, and whether the vendor can show you that testing. This domain is non-negotiable because the history of predictive and screening tools in this field, from the Allegheny Family Screening Tool debate to the benefits fraud-detection failures such as the Dutch childcare-benefits scandal and Michigan's MiDAS system, proves that these tools can encode and amplify the inequities in their training data. A tool that surfaces any risk signal, score, or prioritization is a tool that can produce disparate outcomes, and the burden is on the vendor to show it does not.

Ask the vendor for their equity testing: have they measured whether the tool's outputs differ systematically by race, by neighborhood, by language, by disability status, or by any protected or proxy characteristic? Can they show you the results disaggregated by group, not just an aggregate accuracy number? An aggregate number hides disparity: a tool that is ninety percent accurate overall can be ninety-five percent accurate for one group and seventy percent for another, and the aggregate would never reveal it. A vendor who cannot produce disaggregated results has not tested for equity, and a tool that has not been tested for equity cannot be trusted to surface a signal about a vulnerable family. For any tool that touches screening or prioritization, no disaggregated equity evidence is an automatic fail of this gate.

Domain Five: Privacy, Security, and the Audit Trail

The fifth domain asks whether the tool protects the most sensitive data any government holds and produces a trail that withstands a court and an advocate. Human-services records contain the most vulnerable people's most sensitive information: child-abuse allegations, medical and behavioral-health history, immigration status, family violence. The personally identifiable information (PII, data that can identify a specific person) in these records is among the most protected in any field, and the agency's obligation to protect it does not transfer to the vendor no matter what the contract says.

This domain has two parts. On privacy and security, ask where the data goes, whether it is used to train the vendor's models, who can access it, how it is encrypted, and what happens to it when the contract ends. On the audit trail, ask what the tool logs: every AI-drafted output, every human edit, every verification, every decision, with timestamps and identities, so that months later an advocate challenging a record or a court reviewing a determination can reconstruct exactly what the AI produced, what the human changed, and who made the final call. A tool with no audit trail produces records that cannot be defended. When an advocate asks "was this observation written by the worker or the machine, and how do you know," the only acceptable answer is one the audit trail can prove. We return to the privacy and security obligations in depth in the next chapter; here the point is that no audit trail is a fail of this gate.

Building the Evaluation Rubric

The five domains become usable when you turn them into a written rubric that scores each candidate tool the same way, so the decision rests on evidence rather than on who gave the best demo. A rubric does three things a demo cannot: it makes the criteria explicit before you see any tool, so the criteria are not bent to fit a favorite; it scores every candidate on the same axes, so the comparison is fair; and it produces a written record of why the agency chose what it chose, which is itself part of the defensible governance an oversight body will expect.

Structure the rubric as the five domains, each scored against concrete evidence rather than impressions. For grounding, the evidence is the fabrication count from your imperfect-input test. For the decision-aid boundary, the evidence is the tool's default output and the vendor's answer about where the human decides. For verification support, the evidence is the measured time to verify a real draft. For equity, the evidence is the vendor's disaggregated testing results. For privacy and audit, the evidence is the data-handling answers and a demonstration of the audit log. Score each domain pass or conditional or fail, and make the gating rule explicit in writing: any single fail fails the tool. A weighted average that lets a strong interface compensate for a missing audit trail is how an agency buys a tool it cannot defend.

Give the rubric to the people who will live with the consequences. The quality manager who asked the right question at the demo belongs on the evaluation team, alongside a frontline caseworker who knows what real inputs look like, a supervisor who knows the verification burden, an equity or civil-rights voice, and someone who understands the privacy and records obligations. A rubric scored only by leadership who saw only the demo reproduces the demo's blind spots. The frontline worker is the one who will notice that the tool's perfect output came from an input no real visit ever produces.

Score every candidate on the same five gates, in writing, before the contract: any single fail fails the tool, and a strong demo never compensates for a missing audit trail.

The Questions That Separate Vendors

Most of the evaluation comes down to a short list of questions that a serious human-services vendor can answer and a tool built only to demo well cannot. These questions move the conversation from the curated demo to the properties you are actually buying, and the quality of the answers separates a vendor who understands this field from one selling a generic capability with a social-services label.

Ask: What is your tool's fabrication rate on imperfect inputs, and how did you measure it? A serious vendor has a number and a method. A vendor who says the tool does not fabricate does not understand the technology and should not be trusted with a case record.

Ask: Show me, for one output, how a worker traces each claim to its source. A grounded tool can demonstrate this in the room. An ungrounded tool will deflect to the quality of the prose.

Ask: Where does a human make the consequential decision, and what is the default if the worker does nothing? The answer reveals whether the decision-aid boundary is designed in or bolted on.

Ask: Can you show me your equity testing disaggregated by group? The presence or absence of disaggregated results tells you whether equity was engineered or assumed.

Ask: Walk me through your audit log for a single case from draft to final decision. If the vendor cannot reconstruct who wrote what and who decided, the tool produces records the agency cannot defend.

Ask: Does our data train your models, and what happens to it when we leave? The answer governs whether the agency can keep its privacy promises to the people it serves.

Ask: When your tool causes a documented error in a real case, what is your obligation and what is ours? The honest answer is that the legal and professional accountability stays with the agency and the worker, and a vendor who claims otherwise is selling a promise no contract can keep. A vendor who acknowledges this plainly is one who understands the field.

The Pilot Before the Purchase

The evaluation does not end with the rubric and the questions. The final step before a full purchase is a bounded pilot, a proof of concept run on real cases under controlled conditions, because some properties only reveal themselves in use. A vendor reluctant to support a real pilot, or who insists the demo is sufficient, has told you something important about how the tool performs outside the demo. The detailed design of that pilot, including the equity gate that must sit in front of any decision to scale, is the subject of the next lesson; here the point is that the pilot is part of evaluation, not a separate activity that follows it.

A pilot turns the rubric's predictions into measured outcomes. The fabrication rate you estimated on twenty imperfect inputs gets tested against hundreds of real ones. The verification time you measured on a few notes gets measured across a unit under real caseload pressure. The audit trail you saw demonstrated gets exercised by an actual records request or supervisory review. The equity evidence the vendor showed gets checked against your agency's own population and outcomes. A pilot scoped to a single unit, run for a defined period, with the rubric's gates as the success criteria, gives the agency the evidence to sign or walk away with its reasoning documented. Skipping the pilot to move faster is how an agency discovers a tool's worst-day behavior in a courtroom instead of in a controlled test.

Throughout the pilot, the cardinal rule holds without exception: the tool informs, the worker and the supervisor and the court decide, and every consequential call stays human and audited. A pilot is not a period during which the boundary relaxes to see what the tool can do on its own. It is a period during which the boundary is tested under load to confirm it holds. An agency that lets the pilot blur the boundary has learned nothing useful about whether the tool can be deployed responsibly, because responsible deployment is defined by the boundary holding.

Key Takeaways

  • A vendor demo is a sales instrument that shows the tool on its best day with a curated input. Evaluation is the discipline of testing whether the tool survives an advocate and a court on its worst day with a real case, and the two are not the same activity.
  • At the volume of a single unit (for example fifteen caseworkers each carrying twenty-five families), a tool that is ninety-five percent accurate still produces dozens of flawed records a month. The evaluation must find the failures the demo hides, before deployment, not after a flawed record reaches a courtroom.
  • Evaluate across five gating domains: grounding and the fabrication test, the decision-aid boundary, verification support, equity and bias, and privacy plus the audit trail. Treat each as a pass-or-fail gate; a single fail fails the tool, because a strong interface cannot compensate for a missing audit trail.
  • Test grounding with real but imperfect inputs and count the fabrications. Test the decision-aid boundary by examining the tool's default output, since under caseload pressure workers accept the default. Test verification support by timing how long it takes to verify a real draft against its source.
  • For any tool touching screening or prioritization, demand equity testing disaggregated by group. An aggregate accuracy number hides disparity, and the documented history of these tools (the Allegheny debate, the Dutch childcare-benefits scandal, Michigan's MiDAS) makes untested equity an automatic fail.
  • The privacy and audit gate protects the most sensitive PII any government holds and produces a trail that proves who wrote what and who decided. No audit trail means records the agency cannot defend when an advocate asks whether an observation came from the worker or the machine.
  • Turn the five domains into a written rubric scored on concrete evidence, applied identically to every candidate, and put a frontline caseworker, a quality voice, an equity voice, and a privacy voice on the evaluation team alongside leadership. A rubric scored only by people who saw only the demo reproduces the demo's blind spots.
  • Finish with a bounded pilot on real cases as part of evaluation, with the rubric's gates as success criteria and the decision-aid boundary held throughout. The legal and professional accountability for any AI-caused error stays with the agency and the worker; a vendor who claims otherwise is selling a promise no contract can keep.