Avoiding Metrics That Erode Trust
The dashboard on the wall of the eligibility unit showed one number in large green type: average determination time, now 14 minutes, down from 41 before the AI tool. Leadership loved it. The number went into the board deck, the press release, and the next budget request. What the dashboard did not show was the unit's quiet adaptation to being measured on speed. Workers had learned that the green number rewarded fast approvals and fast denials equally, and that the slow, careful cases, the household with a child receiving SSI (Supplemental Security Income) that made them categorically eligible for SNAP (the Supplemental Nutrition Assistance Program) under a provision the AI tool consistently missed, were the ones that dragged the average down. So those cases got less time, not more. Six months later a fair-hearing officer overturned a cluster of denials that traced back to the same misapplied rule, an advocate filed a complaint, and the 14-minute number, the agency's proudest metric, turned out to be the exact mechanism that had produced the harm. The metric did not measure the work. It deformed it. This lesson is about how that happens, and how to choose metrics in human services that earn trust rather than erode it.
Why the Metric Becomes the Policy
There is a hard truth about measurement in any organization, and it is sharper in human services than almost anywhere else: the metric you choose to celebrate is the behavior you actually instruct. People, and especially people under caseload pressure, optimize for what is measured and rewarded, not for what is intended. A leader who says "quality is what matters" while displaying speed on the wall has, in practice, instructed the floor to chase speed. The stated value loses every time it competes with the measured one.
This is the classic problem that any number used as a target eventually stops being a good measurement, because the moment people know they are judged by it, they manage the number rather than the underlying reality the number was supposed to reflect. In a factory that distortion costs throughput. In human services it costs due process. A determination-time metric does not just describe how fast eligibility decisions are made; it pressures workers to make them faster, which in a field where a wrong denial leaves a real person without food or shelter is a pressure aimed directly at the most vulnerable point in the system.
The deeper reason this matters for an AI program is that AI tools are sold and justified on efficiency. The pilot deck shows determination time falling from 41 minutes to 14, charting time halved, throughput up. Those numbers are real and they are seductive, and they are exactly the numbers most likely to become targets. An AI strategist who lets the efficiency metrics that justified the purchase become the metrics that run the floor has built a machine that converts a genuine time saving into a due-process hazard. The metric is the policy, so the metric has to be chosen as carefully as a policy, which means asking not just "is this number going up" but "what behavior does rewarding this number produce in a tired worker at 5 PM."
Whatever you put on the wall in green type is the instruction the floor actually receives. Speed on the wall is an order to go faster, whatever the mission statement says.
The Metrics That Erode Trust
Certain metrics are especially corrosive in human services because they reward the wrong thing in a setting where the wrong thing harms people. Naming them precisely is the first defense, because they arrive disguised as obvious wins.
Raw Speed as the Headline Number
Determination time, time-per-case, notes-completed-per-day: these velocity metrics are the most tempting and the most dangerous as headline numbers. They are easy to measure, they move impressively when an AI tool is introduced, and they tell leadership a clean story. They also reward the worker who skips the verification step, because verifying every factual claim in an AI-drafted note against the source record takes time, and a speed metric punishes time. A unit measured on raw speed will, predictably and rationally, file AI drafts with less verification, which is precisely the failure mode, an invented observation or a misapplied policy slipping into the record, that the whole verification discipline exists to prevent. Speed as the headline number does not just fail to capture quality; it actively bids against it.
Adoption and Usage Counted as Success
A subtler corrosive metric is treating tool adoption itself as the outcome: login rates, percentage of notes drafted with AI, number of cases touched by the tool. These are useful operational signals, but when they become success metrics they instruct the floor to use the AI more, regardless of whether using it on a given case is wise. There are cases where a careful worker should not lean on the tool at all, the most sensitive removal decisions, a determination turning on a contested nested policy exception, anything on the agency's written kill-criteria list. A metric that rewards usage pressures workers to use the tool even there, eroding the judgment about when not to use AI that is itself a core skill. Adoption is a means; counting it as the end teaches the floor to feed the tool rather than to decide well.
Caseload Throughput as the Prize
The most damaging metric of all is the one that converts the AI time dividend directly into higher caseloads: cases-closed-per-worker, families-per-caseworker rising because each worker is now faster. This metric takes the genuine benefit of the tool, hours returned that should go to verification and to direct time with families, and reabsorbs them into volume. It teaches the workforce that the tool is a productivity ratchet aimed at them, collapses adoption, deepens the burnout the program was meant to relieve, and raises caseloads in a field where caseload is the upstream driver of the harm that slips through. A throughput prize is how an agency turns a wellbeing investment into a wellbeing loss while reporting success the whole way down.
Metrics That Build Trust Instead
The answer is not to abandon measurement. An AI program with no metrics cannot prove its value to a board under a public budget, cannot justify the next investment, and cannot detect its own failures. The answer is to measure the things that actually matter in this field, which are harder to count and far more honest. Trust-building metrics share a property: they reward the behavior you genuinely want even when no one is watching, and they make harm visible rather than hiding it behind a green number.
First, accuracy and verification quality, not just speed. Measure the rate at which AI-assisted documentation is verified before filing, the rate at which verification catches a hallucinated observation or a misapplied rule (a caught-error rate that should be celebrated, not punished), and downstream accuracy signals such as fair-hearing reversal rates and documentation corrections on review. A rising caught-error rate is a healthy unit, not a failing one; it means the verification discipline is working. An agency that measures and praises caught errors builds a workforce that looks for them.
Second, equity outcomes, measured continuously and disaggregated. Because predictive and screening tools can encode the inequities in their training data, and because history (the Allegheny Family Screening Tool debate, the Dutch childcare-benefits scandal, Michigan's MiDAS) proves they have, the program's metrics must include disparate-outcome monitoring across race, ethnicity, language, disability, and geography. Determination outcomes, screening referrals, and denial rates broken down by group surface the bias an aggregate efficiency number hides entirely. Equity is first, not an afterthought, which means it is on the dashboard, not in an annual report no one reads.
Third, wellbeing and the honest time story. Measure documentation hours reduced and, critically, where the saved time actually went, time returned to home visits and direct family contact, alongside turnover and burnout indicators. The wellbeing metric is what keeps the time dividend honest: it makes visible whether the hours went to families and verification or were quietly clawed into caseload. Paired with the explicit, ideally written, commitment that documentation savings will not be converted automatically into caseload increases, the wellbeing metric is the proof the program kept its promise to the workforce.
Balancing Metrics So No Single Number Rules
The protection against any one metric deforming the work is a balanced set in which speed cannot win alone. Determination time displayed next to fair-hearing reversal rate, disaggregated denial rates, and time-returned-to-families tells a true story that raw speed cannot, because a worker who games speed at the cost of accuracy or equity makes the other numbers worse, and the deformation becomes visible instead of hidden. A balanced scorecard is not bureaucratic clutter; it is the structural defense against the single-number distortion that produced the overturned denials in the opening story. The discipline is to refuse to let any efficiency number stand alone on the wall, because alone it becomes the policy, and the policy it becomes is speed over due process.
Never let speed stand alone on the dashboard. Put accuracy, equity, and wellbeing beside it so gaming one number visibly damages the others.
A Worked Comparison: Two Units, Same Tool
Make the contrast concrete with two child-welfare units in the same agency, each given the same AI documentation tool and the same training. Unit A is measured on a single number posted weekly: notes completed per worker per day, with the unit average celebrated in the team meeting and the slowest workers asked to explain themselves. Unit B is measured on a small balanced set reviewed monthly: documentation hours returned, the caught-error rate (with caught errors discussed openly as wins), the fair-hearing reversal rate, denial and referral outcomes disaggregated by group, and where the returned hours went. The tool is identical. The numbers chosen are not.
By the end of two quarters the two units have diverged in exactly the way the lesson predicts. Unit A's notes-per-day average is impressive and rising, which is precisely the problem: workers under the posted-number pressure have learned to accept the AI draft with a light skim rather than a claim-by-claim verification, because verification costs minutes the metric punishes. Two invented observations and one misapplied substantiation standard have already entered court reports, undiscovered, because no one is measured on catching them and catching one would slow the celebrated average. Unit B's notes-per-day number is lower and less impressive on a slide. But its caught-error rate is healthy and visible, its workers verify because the culture and the metrics both reward it, its disaggregated outcomes are being watched for the disparity an aggregate would hide, and its returned hours are documented as time back to home visits rather than absorbed into caseload. Unit A looks better in a one-number board slide and is quietly accumulating due-process risk. Unit B looks slower and is actually safer, more accurate, and more sustainable. The metric, not the tool, produced the difference.
The lesson of the comparison is not that speed is forbidden. Unit B still reports its time figures, because returned hours are real and worth celebrating. The lesson is that the number a unit is judged on becomes the behavior the unit performs, so a unit judged on a lone velocity number performs velocity at the expense of the verification, equity, and judgment the work requires, while a unit judged on a balanced set performs the balance. An agency choosing its dashboard is choosing which of these two units it will produce.
How Metrics Are Reported, Not Just Chosen
Choosing the right metrics is half the work. The other half is how they are reported, because the same number tells different stories depending on its framing and audience, and dishonest framing erodes trust even when the underlying metric is sound. A board, a workforce, an advocate, and a court each need a true account, and a metric reported to impress rather than to inform will eventually meet the audience that catches it.
To leadership and the board, the temptation is to lead with the efficiency win, 41 minutes to 14, and bury the accuracy and equity numbers in an appendix. Doing so trains leadership to value the wrong thing and sets up the program to be measured on speed forever. The honest report leads with the balanced story: here are the hours we returned, here is documentation accuracy, here is the equity review, and here is the audit trail showing humans made every call. That is the credential the program is supposed to earn, and reporting it that way protects the program from its own seductive efficiency numbers.
To the workforce, metrics reported as surveillance, individual speed rankings, public leaderboards of cases-closed, destroy the trust that adoption depends on, because workers correctly read them as a productivity ratchet and a precursor to caseload increases. Metrics reported as shared wellbeing and quality, the unit's caught-error rate, hours returned, turnover trend, build the decision-aid culture instead. The same data point is trust-building or trust-eroding depending entirely on whether it is aimed at the people or shared with them.
To advocates and the court, the relevant metrics are about due process and transparency: how often AI was used, how often its output was verified, how often a determination traced cleanly to an independent human-verified source rather than to an unreconstructable AI output. An agency that can report these from a court-ready audit trail survives a fair-hearing challenge; an agency that reports only efficiency hands the advocate the argument that the decision was effectively automated. Reporting the due-process metrics is not a concession to scrutiny; it is what makes the work defensible when the scrutiny arrives.
Putting It Together: A Trustworthy Scorecard
The practical output of this lesson is a scorecard the strategist can defend to a board, a workforce, an advocate, and the strategist's own conscience. It starts by demoting the efficiency numbers from headline to context: determination time and documentation hours are reported, but never alone and never as the prize. It elevates the numbers that capture whether the work stayed good, accuracy and caught-error rate, fair-hearing reversals, disaggregated equity outcomes, and the honest wellbeing-and-time-return story. And it pairs every number with the framing appropriate to its audience, refusing to weaponize any metric against the workforce or to hide any metric from oversight.
Above all, the scorecard is built on a single test that a strategist can apply to any proposed metric before it goes on the wall: what behavior does rewarding this number produce in a tired caseworker at 5 PM on a Friday with eight unfinished files. If the honest answer is "skip verification," "use the tool where they should not," or "absorb the saved time into more cases," the metric erodes trust and must be balanced or removed. If the honest answer is "verify carefully," "catch the error and be praised for it," "spend the returned hour with a family," then the metric earns trust. In human services, where a deformed metric does not cost a bad quarter but a wrongful denial or a harmed family, that test is not optional. The number you reward is the work you get.
Key Takeaways
- The metric you celebrate is the behavior you instruct. Any number used as a target eventually stops measuring reality and starts being gamed, and in human services that distortion costs due process, not just throughput. In the worked example, a proud 14-minute determination metric was the exact mechanism that produced a cluster of wrongful SNAP denials.
- AI tools are justified on efficiency, so the efficiency numbers (41 minutes to 14, charting time halved) are the ones most likely to become floor-running targets. Letting them do so converts a real time saving into a due-process hazard.
- Three metrics especially erode trust: raw speed as the headline (it bids against the verification step), adoption or usage counted as success (it pressures workers to use AI even where they should not), and caseload throughput as the prize (it reabsorbs the time dividend into volume and deepens burnout).
- Trust-building metrics reward the right behavior even when no one is watching: accuracy and a celebrated caught-error rate, fair-hearing reversal rates, continuous disaggregated equity outcomes, and an honest wellbeing-and-time-return story that shows where the saved hours actually went.
- No single number should stand alone. A balanced scorecard, speed reported next to accuracy, equity, and wellbeing, makes gaming one metric visibly damage the others, which is the structural defense against single-number distortion.
- How a metric is reported matters as much as which metric. Lead with the balanced story to the board, share quality and wellbeing data with the workforce rather than weaponizing it as individual surveillance, and report due-process and transparency metrics from a court-ready audit trail to advocates and the court.
- Equity outcomes belong on the live dashboard, not in an unread annual report, because predictive and screening tools can encode inequity (Allegheny, the Dutch childcare-benefits scandal, MiDAS) and an aggregate efficiency number hides disparate impact entirely.
- Apply one test to any proposed metric: what does rewarding it make a tired worker do at 5 PM on a Friday. If the answer is skip verification, overuse the tool, or absorb time into caseload, the metric erodes trust and must be balanced or removed. The number you reward is the work you get.
Skill.re