Measuring Transformation Agency-Wide
The county had been running its AI documentation program for eleven months, and the director was about to walk into a budget hearing to defend it. On her laptop was a single number the deputy chief financial officer had asked for: average minutes saved per case note. It was a good number, 31 minutes, and across roughly 4,000 notes a month it added up to real time. But the director knew that if she walked in with only that number, she would lose. A county commissioner would ask the question she could not answer with a stopwatch: are children safer, are families treated more fairly, and are your caseworkers still here in a year, or did you just teach a burning-out workforce to produce more paper faster. A transformation that can only prove speed has not proven it is working. It has proven it is fast, which in this field is not the same thing, and can be the opposite. That night she rebuilt the report around four numbers instead of one, and the hearing went differently.
Why One Number Is a Trap
Every agency that deploys AI starts with a speed metric because speed is the easiest thing to measure and the easiest thing to sell. Minutes saved per note, notes drafted per day, hours returned per caseworker per week: these are real, they matter, and they are the first thing leadership asks for. The trap is not that speed is wrong. The trap is that speed alone, reported alone, quietly redefines the mission of the agency as throughput, and throughput is precisely the wrong north star for work whose failures are a harmed child or a wrongly denied family.
Consider what a speed-only scorecard rewards. If the only number that goes up the chain is documentation hours saved, then the rational move for a pressured supervisor is to push caseworkers to accept AI drafts faster, verify less, and reinvest the saved time into carrying more cases. The metric goes up. The thing the metric was supposed to protect, careful human judgment over consequential decisions, goes down, and nothing in the scorecard shows it. This is the measurement version of the cardinal rule of this entire program: AI informs, humans decide. A scorecard that measures only what AI accelerated, and never what human judgment it may have eroded, is a scorecard built to miss the only failure that matters.
The history of the field makes this concrete. The Dutch childcare-benefits scandal and Michigan's MiDAS (the Michigan Integrated Data Automated System, the state's automated unemployment-fraud detection system) were, in their own internal reporting, efficiency successes for a time. They processed cases faster and flagged more. The catastrophe, tens of thousands of families wrongly accused, was not visible in a throughput metric because the throughput metric was never designed to see it. The lesson for an agency measuring its own transformation is direct: if your scorecard cannot show harm, your scorecard will not show harm, right up until harm becomes a headline and a lawsuit.
A transformation that can only prove it is fast has not proven it is working. In this field, fast and wrong is worse than slow and right.
The Four-Axis Scorecard
An agency-wide measurement program for AI in human services has to report on four axes at once, because the four are in tension and the tension is the point. Measure them together and they hold each other honest. Measure one alone and it will be gamed. The four axes are wellbeing, equity, outcomes, and accountability. Think of them as a four-legged stool: remove any leg and the whole thing tips, and in this work the thing that tips is a family's life.
Axis One: Wellbeing (Of Workers and Families)
Wellbeing has two halves, and a serious scorecard measures both. The first half is worker wellbeing, because the entire justification for AI documentation in this field is returning hours to a burning-out workforce so they can be present with families. The measures here are time returned to direct contact (not just hours saved, but where those hours went), documentation hours per week, caseworker turnover and vacancy rates, and a direct caseworker-reported burden and burnout measure collected on a regular cadence. The distinction between hours saved and hours redirected to families is load-bearing. An agency that saved 31 minutes per note but quietly raised every caseworker's caseload from 18 to 24 families did not improve wellbeing. It converted an efficiency gain into a workload increase and called it a win. The scorecard must show where the saved time actually went.
The second half is the wellbeing of the families served, measured through their own experience: client-reported experience of contact, complaint and grievance rates, and time-to-service for people in crisis. A program that frees caseworker hours should produce more and better human contact, and families should be able to feel it. If documentation hours fell but families report the same rushed, transactional contact, the saved time leaked somewhere other than the mission.
Axis Two: Equity
Equity is first, not an afterthought, and on a scorecard that means it gets its own standing axis with its own data, not a footnote under outcomes. The core practice is disparate-outcome monitoring: tracking screening rates, substantiation rates, removal rates, benefit-denial rates, and fair-hearing reversal rates broken down by race, ethnicity, geography, disability, and language. The question the equity axis answers every quarter is whether any AI-touched decision point shows a widening gap between groups. A predictive screening tool that raises overall accuracy while widening the racial gap in screen-in rates has made the agency's equity problem worse, and only an equity axis with disaggregated data can catch it.
Two technical points matter here. First, equity measurement requires a baseline taken before the AI tool was deployed, because the question is not whether disparities exist (they almost always do, inherited from the data and the history) but whether the AI widened or narrowed them. Without a pre-deployment baseline you cannot attribute a change to the tool. Second, equity monitoring must cover both error directions: a tool can harm a group by over-flagging it (more unwarranted investigations) or by under-serving it (more missed needs), and a scorecard that watches only one direction will miss the other. An equity audit on this cadence is the continuous practice the program insists on, not the one-time check a vendor offers at procurement.
Axis Three: Outcomes
Outcomes are the substantive results the agency exists to produce, and they are the slowest and hardest axis to measure, which is exactly why agencies skip them and report speed instead. For child welfare the outcome measures are the field's established ones: child safety (re-report and re-injury rates after a case decision), permanency (time to reunification or other permanent placement), and placement stability. For benefits the outcome measures are eligibility-determination accuracy, fair-hearing reversal rate (a high reversal rate means the agency's determinations are wrong and getting overturned), and time-to-benefit for eligible applicants. For all programs, documentation accuracy is an outcome, not a process metric, because in this field an accurate record is a due-process protection, and a court report with a fabricated observation is an outcome failure even if it was produced quickly.
The hard discipline on the outcomes axis is patience. Safety and permanency outcomes move on a timescale of months to years, not the monthly cadence of a speed metric. An agency that demands outcome proof in the first quarter will either get noise or get pressured into substituting a faster proxy. The honest scorecard reports outcomes on their real timescale and resists the pull to replace them with the speed numbers that move faster.
Axis Four: Accountability and Verification Integrity
The fourth axis measures whether the program's own guardrails are actually being followed, and it is the axis most agencies forget. It is also the one that protects the other three from being faked. The measures are verification-rate (the share of AI-assisted documents that received a documented human verification before filing), human-review coverage on any screening signal (the share of AI risk signals that received the mandatory human review the policy requires), AI-use disclosure rate, audit-trail completeness, and incident counts (caught hallucinations, near-misses, and harms). Verification integrity is the metric that tells you whether AI is still informing and humans are still deciding, or whether under caseload pressure the verification step has quietly become a rubber stamp.
Here is why this axis cannot be optional. The single most dangerous failure in an AI documentation program is not a visible error; it is the silent decay of the verification habit. A caseworker carrying 24 families, trusting a tool that is right 95 percent of the time, will under pressure begin to skim rather than verify, and the verification step erodes from a safeguard into a formality without anyone deciding to let it. No outcome metric catches this early, because the 5 percent of cases with a fabricated observation or a misapplied SNAP (Supplemental Nutrition Assistance Program, the federal food-assistance benefit) rule are rare and scattered. Only a direct measure of verification integrity, ideally with periodic blind re-checks of supposedly-verified documents, catches the decay before it becomes a harm.
Reading the Axes Against Each Other
The reason to carry all four axes is that the dangerous patterns only appear when you read them against each other. A single axis can look healthy while the program is failing; the cross-reads are where the truth lives. Walk through the patterns a director should be trained to spot.
Speed up, verification down. Documentation hours saved is climbing, and so are notes-per-day, but verification-rate is falling. This is the most common and most dangerous pattern, and it means exactly what it looks like: the time pressure that AI was supposed to relieve is instead being reinvested in more throughput, and the safeguard is the thing being cut. The correct response is not to celebrate the speed number. It is to lower caseloads or production targets until verification recovers, because a fast pipeline producing unverified court records is a due-process liability, not an achievement.
Outcomes flat, equity widening. Overall safety and accuracy numbers are stable, so a speed-only report would show a healthy program, but the equity axis shows the screen-in or denial gap between groups widening. This means the program is concentrating its errors on a particular group while the aggregate stays calm. Aggregates hide disparities; this is precisely why the equity axis must be disaggregated and standing, not folded into an average.
Worker hours saved, family experience flat. Documentation time fell sharply, but client-reported experience and time-to-service did not improve. The saved time did not reach families. Either it was absorbed by caseload increases (check the caseload number on the wellbeing axis) or it leaked into other administrative work. The promise of the transformation, hours back to human connection, is unfulfilled, and the scorecard should say so plainly rather than let the impressive hours-saved number stand in for a benefit families never received.
Everything up, incidents up too. Speed, satisfaction, and outcomes all improving, but caught-hallucination and near-miss counts rising. Counterintuitively, this can be a healthy sign rather than an alarm, provided the incidents are being caught before harm. A rising count of caught near-misses can mean the verification culture is strong and detection is working, not that the tool got worse. This is why incident counts must always be read alongside whether the incidents were caught before filing or after harm. The number alone is ambiguous; the timing tells the story.
Building and Governing the Scorecard
A scorecard is only as honest as the governance around it, and three design choices separate a scorecard that holds an agency accountable from one that flatters it.
Set baselines before deployment, not after. Every measure on every axis needs a pre-AI baseline, captured before the tool goes live, against which change is measured. The most common measurement failure in the field is deploying first and trying to reconstruct a baseline afterward from incomplete records, which makes every later claim of improvement unfalsifiable. If you cannot say what documentation hours, turnover, the equity gap, and the fair-hearing reversal rate were the quarter before AI arrived, you cannot honestly claim the AI changed them. Baseline first.
Separate the people who report the scorecard from the people whose performance it measures. If the same office that runs the AI program also owns the metrics that judge the AI program, the incentive to emphasize speed and bury verification decay is structural, not a matter of anyone's bad intent. The equity and accountability axes in particular should be owned by an independent function (a quality, equity, or oversight office) that reports to leadership and, where appropriate, to a governance board that includes community and advocate voices. This is the measurement expression of the program's governance principle: the scorecard must be defensible to a court and an advocate, which means it cannot be graded entirely by the team being graded.
Report the uncomfortable numbers, on purpose. A scorecard's credibility comes from the numbers it reports that are not flattering. A transformation report that shows only green is not reassuring to a commissioner or a judge; it is a signal that the agency is measuring the wrong things or hiding the right ones. The director in the opening built her budget hearing around four numbers, and one of them, a verification-rate that had dipped under a caseload spike, was the number that won the room, because it showed she was watching the thing that could go wrong and had a plan to fix it. The disciplined version of transparency about AI use, which this program treats as part of due process, includes transparency about the program's own measured weaknesses.
One practical consequence: the scorecard should produce two different reports from the same underlying data. The internal operational report is detailed, disaggregated, and unflinching, built for the people running the program to find problems early. The external accountability report, for leadership, the governance board, the court, and the community, translates the same data into the wellbeing, equity, outcomes, and accountability story without spin and without burying the equity and verification axes beneath the speed axis. Same numbers, two audiences, no contradiction between them. The moment the external story and the internal story diverge is the moment the scorecard has stopped doing its job.
The Cadence That Keeps It Alive
A scorecard that is built once and reviewed once a year is not a measurement program; it is a slide deck that ages badly. The four axes move on different clocks, and the cadence has to respect that or the program will either drown in noise or miss a developing harm. Verification integrity is the fastest clock and should be reviewed weekly or biweekly, because it is the early-warning system: it moves before outcomes do, and a dip in verification integrity this month is the leading indicator of an outcome failure two quarters from now. The wellbeing measures move on a monthly to quarterly clock, fast enough to catch a caseload-driven burnout spike before it becomes a wave of resignations. Equity should be reviewed quarterly with the disaggregated cuts, frequently enough to catch a widening gap but with enough volume in each subgroup that the numbers are not just noise. Outcomes are the slowest clock, reviewed semiannually to annually, on the months-to-years timescale that safety and permanency actually move.
The discipline in the cadence is matching the response to the clock. A weekly verification-integrity dip triggers a fast operational response (a caseload check, a workflow fix), not a panicked rewrite of the whole program. A quarterly equity finding triggers an investigation and, if confirmed, a change to the tool or the workflow before the next quarter. An annual outcome reading is read for trend, not for any single data point. An agency that demands an annual-clock proof on a weekly-clock metric, or reacts to a single noisy outcome quarter as if it were a verified trend, has misread its own cadence, and a misread cadence produces either false alarms that exhaust the staff or missed signals that become harms.
Key Takeaways
- A speed-only scorecard (minutes saved, notes per day) quietly redefines the agency's mission as throughput, which is the wrong north star for work whose failures are a harmed child or a wrongly denied family. Speed reported alone rewards exactly the verification-cutting behavior the program exists to prevent.
- Measure four axes together so they hold each other honest: wellbeing (of both workers and families), equity, outcomes, and accountability/verification integrity. Any single axis reported alone will be gamed.
- Wellbeing must distinguish hours saved from hours redirected to families. An agency that saved 31 minutes per note but raised caseloads from 18 to 24 families converted an efficiency gain into a workload increase, not a wellbeing improvement.
- Equity gets its own standing, disaggregated axis with a pre-deployment baseline, tracking screening, substantiation, removal, denial, and fair-hearing reversal rates by race, geography, disability, and language, in both error directions (over-flagging and under-serving). Aggregates hide disparities.
- Outcomes (child safety, permanency, placement stability, eligibility accuracy, fair-hearing reversal, documentation accuracy) move on a timescale of months to years. The honest scorecard reports them on their real timescale and refuses to substitute faster speed proxies.
- Verification integrity (the share of AI-assisted documents and screening signals that received documented human review) is the axis that detects the silent decay of the verification habit under caseload pressure, which no outcome metric catches early. Use periodic blind re-checks.
- The dangerous patterns appear only in the cross-reads: speed up with verification down, outcomes flat with equity widening, hours saved with family experience flat, and incidents up that may be healthy if caught before harm.
- Govern the scorecard: set baselines before deployment, give the equity and accountability axes to an independent function, report the uncomfortable numbers on purpose, and keep the internal and external reports built from the same data so they never diverge.
Skill.re