Success Metrics for Human-Services AI
The director had one slide left in the budget presentation, and it held a single number: documentation time down 34 percent across the three units that had adopted the AI tool. The county administrator nodded. It was a clean number, the kind that justifies a renewal. Then a commissioner who had once been a foster parent asked a question that was not on the slide. "That's how much faster they're typing. How are the families doing?" The room went quiet, because no one had measured that. The agency had stood up a sophisticated AI documentation program, tracked the one number that was easy to count, and could not answer the only question that mattered. Worse, when the quality unit went back and pulled records later, they found that two of the three "faster" units had quietly higher rates of court reports returned for correction, and one unit's disparity in screen-in rates between neighborhoods had widened during the same period. The speed number was real. It was also dangerously incomplete, and an incomplete metric in this field does not just mislead a budget meeting. It hides harm to families behind a number that looks like success.
Why a Single Metric Is a Trap in This Field
Every measurement program encodes a theory of what matters. When a human-services AI program measures only documentation speed, it is making a claim, whether it means to or not, that speed is what the program is for. But the program is not for speed. The program, as this entire curriculum has argued, is for giving hours back to human connection while never letting AI make the consequential call, and doing both in a way that protects equity and due process. A metric set that captures only the speed half of that mission will optimize the agency toward speed and away from everything else that makes the program worth running.
This is not an abstract risk. Metrics drive behavior, and they drive it hardest in environments with crushing workloads, which describes every human-services agency in 2026. If a caseworker carrying 28 families learns that the agency tracks and rewards how fast she produces documentation, she will optimize for that, and the fastest way to produce documentation is to trust the AI draft and file it without the claim-by-claim verification that protects families. The metric that was supposed to prove the program worked becomes the mechanism that hollows it out. A speed-only metric in this field does not merely fail to capture harm; it actively incentivizes the behavior that produces harm.
The strategist's job is to build a balanced metric set that measures the program on every axis it claims to serve, so that no single number can be improved by degrading the others. Four axes matter, and they must be measured together: time returned to the mission, documentation accuracy, workforce wellbeing, and equity. Measure all four and the program stays honest. Measure one and the program drifts toward whatever that one rewards.
A metric set is a theory of what matters. Measure only speed, and you will get speed, even when speed is bought by skipping the verification that protects a family.
Axis One: Time Returned to the Mission, Measured Honestly
Time is the axis everyone reaches for first because it is the easiest to count and the easiest to put in a budget slide. It is also the axis most often measured in a way that proves nothing. Documentation time down 34 percent sounds like success, but the number is meaningless until you answer the question the commissioner asked: returned to what?
The honest version of this axis measures not just hours saved but where the saved hours went. If a documentation tool returns roughly a third of the time a worker spent charting, that time can go to one of three places, and only one of them is the program's actual goal. It can go to direct time with families, which is the mission. It can go to verification of the AI drafts, which is the safeguard. Or it can be silently absorbed by a caseload increase, which means the worker is no better off and the agency has converted a wellbeing investment into a productivity ratchet. The same 34 percent looks identical in all three cases on a speed slide and is completely different in what it means.
So the time axis requires at least two paired measures. The first is documentation time per case, which captures the efficiency gain. The second is direct-contact time, the hours a worker actually spends with families on home visits and in person, which captures whether the saved time reached the mission. A program where documentation time fell 34 percent and direct-contact time rose is succeeding on this axis. A program where documentation time fell 34 percent and direct-contact time stayed flat has a leak: the saved hours went somewhere other than families, and the strategist needs to find out where before celebrating. A worked example: an agency that tracked only the first measure reported a third less charting time for two budget cycles before discovering that direct-contact time had not moved at all, because caseloads had risen in step with the efficiency. The program had returned no hours to a single family. The speed number had hidden that for a year.
Axis Two: Documentation Accuracy, the Safeguard Made Measurable
Accuracy is the axis that directly protects families, and it is the one a speed-focused program most often neglects, because measuring it is harder than counting hours. But the central promise of AI-assisted documentation done right is that records become faster and more accurate, and a program that cannot show the accuracy half of that promise has not earned the speed half.
Accuracy in this context has a specific operational meaning rooted in the program's core discipline: catching the hallucination failure modes (invented observations, misapplied policy, fabricated history) before they reach a record. So the accuracy axis measures the verification system, not just the output. Useful measures include the rate of AI-introduced errors caught in verification (a healthy program catches them; a worrying one catches none, which usually means no one is checking), the rate of court reports or determinations returned for correction after filing (which should fall, not rise, in a working program), and the results of periodic record audits where a quality unit traces factual claims in AI-assisted documents back to their sources.
There is a counterintuitive point the strategist must internalize. A rising number of caught errors early in a rollout is good news, not bad. It means workers are verifying and the safeguard is functioning. The dangerous signal is the opposite: a program where AI is heavily used and almost no errors are ever caught, which does not mean the tool is perfect (no LLM, or large language model, the AI that generates text, is) but almost always means verification is being skipped. In the opening story, two of the three fast units had higher rates of reports returned for correction after filing, which is the accuracy axis flashing red while the speed axis flashed green. A balanced metric set would have surfaced that contradiction in the same meeting instead of a year later.
Axis Three: Workforce Wellbeing and Burnout
The forcing function for this entire program is that documentation burden is the field's defining pain and a top driver of burnout and turnover, and turnover raises caseloads for those who remain, which lets more harm slip through. If the program's premise is that AI can relieve that burden, then wellbeing is not a soft secondary metric. It is a primary measure of whether the program achieved its stated purpose.
Wellbeing resists a single clean number, which is exactly why it gets dropped from speed-focused dashboards, and exactly why the strategist must insist on it. Several measures triangulate it. Turnover and retention rates in units using AI compared to baseline and to non-adopting units give a hard behavioral signal: people vote with their resignation letters. Periodic, confidential burnout and workload surveys give a direct read, ideally using an established instrument rather than an ad hoc question so the number means something over time. Time-off and overtime patterns offer an indirect signal of workload pressure. And qualitative input, the good catches and tool failures surfaced in unit meetings, tells the strategist whether workers experience the tool as relief or as one more thing.
The interaction between this axis and the time axis is where the program's integrity lives. A program can show documentation time down and burnout up at the same time, and when it does, the diagnosis is almost always that the returned hours were absorbed by caseload rather than given back. The strategist who watches only the time axis sees success; the strategist who watches both sees the leak. A worked example: a unit reported strong efficiency gains while its confidential burnout score worsened and two experienced workers left within a quarter. The paired reading revealed that the efficiency had been used to justify holding two vacancies open, raising effective caseloads. The speed metric was, in isolation, an active misrepresentation of the program's effect on the people it was supposed to help.
Axis Four: Equity, Measured Continuously and First
Equity is the axis that, in this field, can never be an afterthought, and a metrics program that adds it last has already failed its most important test. Predictive and screening tools can encode the inequities in their training data, and history (the Allegheny Family Screening Tool debate, the Dutch childcare-benefits scandal, Michigan's MiDAS) proves they have. Wherever AI touches a screening, eligibility, or risk signal, the equity axis must be measured continuously, because a disparate outcome that goes unmeasured is a harm that goes unaddressed.
The equity axis measures whether AI-touched processes produce different outcomes across the groups the agency serves, in ways the agency cannot justify. Concretely, that means tracking outcome rates broken down by race, ethnicity, neighborhood, language, and other protected and proxy characteristics: screen-in rates for investigation, substantiation rates, eligibility approval and denial rates, and the rate at which AI signals were overridden by human review across groups. A widening disparity in any of these during a period of AI adoption is the single most important signal a human-services AI program can produce, and it must be visible at the same altitude as the speed number, not buried in an annual report.
In the opening story, one unit's disparity in screen-in rates between neighborhoods widened during the AI adoption period, and no one saw it until later, because equity was not on the dashboard the budget meeting used. That is the failure mode in its purest form: the program tracked the easy number and missed the one that could be quietly separating families along the same lines the field's history warns about. Equity measurement is also continuous, not a one-time check at procurement, because models drift, populations change, and a tool that audited clean at launch can develop a disparity a year in. The strategist builds equity into the standing dashboard, reviews it on the same cadence as every other axis, and treats a widening disparity as an incident that triggers review, not a footnote.
A disparate outcome you do not measure is a harm you cannot address. Equity belongs at the top of the dashboard, on the same cadence as speed, not in an annual appendix.
Building the Balanced Dashboard and Reporting It Honestly
The four axes only protect the program if they are read together, on the same cadence, by the same people who make the budget and staffing decisions. A dashboard that puts speed on the front page and equity in an appendix has, by its layout alone, told the agency what to optimize. The strategist's deliverable is a balanced dashboard where time, accuracy, wellbeing, and equity sit at equal altitude and are reviewed together, so that no axis can be improved by silently degrading another.
Equal altitude is a design choice with consequences. When the same monthly review shows documentation time down 34 percent, direct-contact time flat, burnout up, errors-caught near zero, and a neighborhood disparity widening, the contradiction is impossible to miss and impossible to spin. That single screen would have given the director in the opening story an honest answer to the commissioner's question, and would have caught the accuracy and equity problems in the same cycle they emerged rather than a year later. The value of a balanced set is precisely that it makes the leaks visible at the moment they form.
Reporting honestly means resisting the strong institutional pull toward the clean number. Budget meetings reward the single confident figure, and a strategist will be tempted to lead with speed because it is the number that renews the funding. The discipline is to report the balanced picture even when it is messier, because a program defended on speed alone is a program that will eventually be exposed by a commissioner's question, an advocate's record request, or a quality audit, and lose the trust it took years to build. The honest report, which says here is the time we returned, here is the accuracy we maintained, here is the wellbeing we improved, and here is the equity review that holds up, is also the more durable one. It is the report that survives the next fair hearing and the next budget cycle.
One last discipline anchors the whole set: every performance figure, the agency's own and any vendor's, is a benchmark to verify, not a guarantee to repeat. A vendor that promises a 40 percent documentation reduction is offering a number to test in a pilot with an equity gate, not a result to put on a slide. The strategist who measures all four axes in their own agency, on real cases, with the disparities broken out, is the one who can stand in front of the county administrator, and the commissioner who used to be a foster parent, and answer the only question that matters: not how much faster the workers are typing, but how the families are doing.
Key Takeaways
- A single metric is a trap. Measuring only documentation speed encodes a false theory that the program is for speed, and in a high-caseload environment it incentivizes the verification-skipping that harms families. Metrics drive behavior, so a balanced set is a safeguard, not just a report.
- Measure four axes together so no one can be improved by degrading another: time returned to the mission, documentation accuracy, workforce wellbeing, and equity.
- Time must be measured as paired figures: documentation time per case (the efficiency) and direct-contact time with families (where the hours went). Falling charting time with flat family time means the saved hours leaked into caseload, not the mission.
- Accuracy makes the safeguard measurable: errors caught in verification, reports or determinations returned for correction after filing, and periodic record audits. Counterintuitively, catching more errors early is good news; catching almost none usually means verification is being skipped.
- Wellbeing is a primary metric because burnout from documentation burden is the program's whole forcing function. Triangulate it with turnover, confidential burnout surveys, overtime patterns, and qualitative input. Efficiency up with burnout up almost always means returned hours were absorbed by caseload.
- Equity is measured continuously and at the top of the dashboard, never as an afterthought, because screening and eligibility tools can encode inequity (Allegheny, the Dutch scandal, MiDAS). Track outcome rates by race, ethnicity, neighborhood, and language; a widening disparity is the most important signal the program can produce and triggers an incident review.
- Build a balanced dashboard at equal altitude reviewed on one cadence, so contradictions (speed up, equity worsening) surface in the same meeting rather than a year later.
- Report honestly: resist leading with the clean speed number alone. The balanced story (time returned, accuracy held, wellbeing improved, equity audit that holds up) is the one that survives a commissioner's question, an advocate's record request, and the next budget cycle. Treat every figure, including vendors', as a benchmark to verify in your own agency.
Skill.re