โ†
AI for Social Work & Human Services
Visionary ยท M11 ยท lesson 11 of 16 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Running Pilots to Evidence
๐Ÿ“–
now learning

Running Pilots to Evidence

15 min

The deputy director had asked for one number, and the agency AI lead did not have it. The question, in a budget review meeting three floors up, had been simple: "The Magic Notes pilot has been running for four months across the East District intake unit. Did it work?" The honest answer was that nobody could say. The vendor had a slide showing 47 percent time savings on documentation, but that figure came from a controlled demo, not from the agency's own cases. The unit supervisor said her workers loved it. The county auditor, copied on the meeting, wanted to know whether the case notes the tool helped draft were more accurate or less accurate than before, and whether any family had been affected by an error. There was no answer to that either, because the pilot had been launched to "see how it goes," with no baseline measured before it started, no metrics defined, and no plan for what evidence would justify expanding it or shutting it down. The tool was either a success or a liability, and after four months and a real budget line, the agency could not tell which. That gap, between running a pilot and producing evidence, is the difference between a program an oversight body trusts and a program it shuts down.

Why Most Pilots Produce No Evidence

A pilot is supposed to answer a question. Most pilots in human services answer none, because they were never designed to answer one. They were launched because a vendor offered a trial, a grant required innovation, or a director wanted to "be doing something with AI." The tool gets deployed to a willing unit, the early adopters are enthusiastic, and a few months later the only evidence anyone can point to is anecdote: "the workers like it," "it seems faster," "we haven't had any complaints." None of that survives contact with an oversight body, a county auditor, a court, or an advocate asking whether the AI-assisted documentation that helped decide a family's case was accurate.

The first failure is the missing baseline. If you do not measure how long documentation took, how accurate it was, and how workers felt about it before the pilot, you have nothing to compare the pilot against. A claim that the tool "saved time" is meaningless without a before number. An agency that runs an AI note-drafting pilot in a unit where caseworkers carry 24 families each, and never recorded that those workers spent an average of 3.1 hours per day documenting before the pilot, can never honestly state how much time the tool returned. The vendor's 47 percent figure is not your figure. It came from a demo on clean inputs, not from your messy field notes, your CCWIS (Comprehensive Child Welfare Information System, the federal-standard case-management platform), and your court-report requirements.

The second failure is the undefined success criterion. A pilot without a written, advance definition of what success and failure look like will be judged after the fact by whoever has the loudest opinion. If the answer to "did it work?" depends on who is in the room, the pilot did not produce evidence. It produced a debate. The discipline that satisfies oversight is the discipline of deciding, before the pilot starts, exactly what evidence will justify scaling, what evidence will justify stopping, and who decides.

The third failure is the absence of an oversight relationship. Many pilots are run as internal experiments, with the auditor, the governance board, or the court learning about the AI tool only after it has already touched cases. By then the question is no longer "should we pilot this?" but "why did you deploy AI in case decisions without telling us?", and the agency is on the defensive. A pilot that names its oversight contact at the start, and shares the design before launch, converts oversight from an adversary into a witness. The auditor who reviewed and accepted the pilot design is far harder to surprise later, and far more likely to trust the evidence package at the end. In a field where AI touches removal, substantiation, and eligibility decisions bound by due process, the oversight relationship is not a formality; it is the perimeter that keeps the work defensible.

The fourth failure, quieter than the others, is scope creep. A pilot launched as "AI-assisted note drafting in one unit" quietly becomes "the tool everyone in the building uses for everything" because workers in adjacent units hear it is good and start using it off the books. Now the pilot's careful population definition is meaningless, the cases touched are uncounted, and the evidence is contaminated. A disciplined pilot polices its own boundary: the tool is available only to the pilot population, usage is logged, and any spread outside the defined scope is a finding to report, not a quiet win to celebrate.

A pilot that cannot fail is not a pilot. It is a procurement decision wearing a lab coat.

The Pilot Design That Produces Evidence

A pilot that produces evidence is designed backward from the decision it must inform. Start with the decision: at the end of this pilot, the agency will choose to scale the tool, stop the tool, or extend the pilot with changes. Then ask what evidence would make each of those choices defensible to leadership, to a court, and to an advocate. Then design the pilot to generate exactly that evidence. This is the opposite of "let's try it and see."

Concretely, a defensible pilot has six elements fixed in writing before it launches. First, a specific scope: one tool, one workflow, one or two units, a defined number of cases. A pilot of "AI" is unmeasurable; a pilot of "AI-assisted home-visit note drafting in the East District intake unit, 18 caseworkers, for 90 days" is measurable. Second, a baseline period: two to four weeks of measuring the current state before the tool is introduced, so there is a real before number. Third, defined metrics with targets: documentation hours per worker per week, a verification-defect rate (how often an AI-assisted draft contained an error the worker had to catch), worker-reported burden, and at least one equity metric. Fourth, pre-registered decision thresholds: the specific results that will trigger scale, stop, or revise. Fifth, a defined population and consent posture: which cases and families are touched, what disclosure is made, and what is excluded. Sixth, an owner and an oversight contact: a named person accountable for the pilot and a named oversight body that receives the results.

Consider an agency that runs this correctly. Before launch, the East District intake unit logs two weeks of baseline: caseworkers average 15.5 hours per week on documentation, the supervisor's spot-check of 40 notes finds a 6 percent factual-error rate, and workers rate documentation burden 4.2 on a 5-point burnout scale. The pilot is defined: 18 workers, 90 days, AI-assisted note drafting with mandatory verification, on cases that do not involve active court proceedings (excluded to limit due-process exposure during the trial). The pre-registered thresholds say: scale if documentation time drops at least 20 percent and the verified-defect rate does not rise, and the equity check shows no disparity; stop if the defect rate rises or any family is harmed by an unverified error; revise otherwise. Now the pilot can answer the deputy director's question with a number, and the answer will hold up in front of the auditor.

Measuring What Matters, Not What Is Easy

Vendors measure time savings because time savings sell. An agency that only measures time savings has measured the easy thing and ignored the things that protect people. The evidence an oversight body actually needs is broader, and three categories matter most: the benefit, the safety, and the equity.

The Benefit: Time and Burden, Measured Honestly

Time returned is real and it is the reason to do this work, but it has to be measured against the agency's own baseline, not the vendor's demo. The honest benefit metric is documentation hours per worker per week, measured the same way before and during the pilot. Beware the trap of counting the drafting time saved while ignoring the verification time added. A tool that drafts a note in two minutes but requires fifteen minutes of claim-by-claim verification against the field notes has not saved thirteen minutes; the net figure is what matters. An agency that found drafting time fell from 22 minutes to 4 minutes per note, but verification added 9 minutes, should report a net saving of 9 minutes per note, not the 18-minute gross figure the vendor will quote. Over a unit of 18 workers each writing 40 notes a week, 9 minutes saved per note is 108 hours a week returned, which is real and defensible. The gross figure is neither.

Burden and wellbeing are the second benefit metric, because the point of returning hours is to reduce the documentation load that drives burnout and turnover, and turnover that raises caseloads for everyone who stays. A short, repeated worker survey, run at baseline and at the end of the pilot, asking workers to rate their documentation burden and whether they have more time for direct family contact, captures what a time log cannot. The wellbeing number also guards against a perverse outcome: an agency that returns 108 hours a week to a unit and then fills every recovered hour with more cases has not improved wellbeing, it has raised the caseload under cover of efficiency. The benefit the program exists to deliver is hours given back to families and to workers, not hours harvested for throughput. If the time-back is real but the burden survey does not move, the pilot has revealed that the agency's staffing model, not its documentation tool, is the bottleneck, and that is itself useful evidence.

A worked illustration ties the benefit metrics together. The East District pilot reports a net 9 minutes saved per note, 108 hours a week across 18 workers, a burden score that fell from 4.2 to 3.4 on the 5-point scale, and a worker-reported increase of roughly 2 home visits per worker per week. That is a coherent benefit story: time returned, burden eased, and the recovered hours visibly flowing to direct family contact rather than to a higher caseload. An oversight body can read those four numbers and understand both what the tool did and what the agency chose to do with the time, which is exactly the story the program is built to tell.

The Safety: Verification Defects and Near Misses

The safety metric is the one vendors never volunteer and oversight bodies always want: how often did the AI-assisted draft contain an error that, if filed unverified, would have entered the case record as false? This is the verification-defect rate, and it must be measured by sampling. A supervisor reviews a random sample of AI-assisted drafts against the source field notes and the case record, counts the invented observations, misapplied policy citations, and fabricated history that verification caught, and reports the rate. A pilot that returns hours but raises the rate of errors reaching the record has not succeeded; it has traded time for due-process risk, and that trade is unacceptable in a field where a fabricated observation in a court report can separate a family.

Near misses matter as much as filed errors. A near miss is an error the verification step caught before filing. Counting near misses is not a sign the pilot is failing; a pilot with zero near misses recorded is more likely a pilot where nobody is verifying than a pilot with a perfect tool. The near-miss count is evidence that the human-review safeguard is working, which is itself part of what oversight wants to see.

The Equity: Disparity Measured Before Scale

Equity is not a metric you add at the end. A pilot that touches eligibility screening or risk flagging must measure, during the pilot, whether the tool's behavior differs across the populations the agency serves. If an AI tool helps draft assessments, the equity question is whether the verified-defect rate, or the tone and content of drafts, differs by the race, language, or neighborhood of the family. History makes this non-negotiable: predictive and screening tools have encoded the inequities in their training data, from the debate over the Allegheny Family Screening Tool to the benefits fraud-detection failures of the Dutch childcare-benefits scandal and Michigan's MiDAS system, where automated decisions wrongly accused thousands of fraud. A pilot that scales before it has equity evidence is a pilot that may be scaling a disparity. The equity gate belongs in the pilot, before scale, not in a review after harm.

The Discipline of the Stop Decision

The hardest and most important part of running pilots to evidence is the willingness to stop. An agency that launches ten pilots and scales ten pilots is not running pilots; it is running a procurement pipeline with extra steps. The credibility of the entire AI program with oversight depends on the agency's demonstrated willingness to kill a tool that did not meet its pre-registered thresholds. The first time a director shuts down a pilot that the vendor loved and a few enthusiastic workers wanted to keep, because the evidence said the defect rate rose or the equity check failed, the agency earns something it cannot buy: the trust of the auditor, the court, and the advocate, who now believe the agency's evidence because they have seen the agency act on it.

Pre-registering the stop thresholds is what makes the stop decision possible. If the threshold was written down before anyone was attached to the outcome, stopping is following the plan, not a referendum on the people who championed the tool. Consider the agency whose pilot showed an 11 percent net time saving, below its 20 percent target, and whose verified-defect rate held steady but whose worker survey showed verification fatigue setting in by week 8. Because the threshold required at least 20 percent time saving with no rise in defects, the result was not "scale" but "revise": the agency extended the pilot with a redesigned verification workflow rather than scaling a tool that was not yet returning enough time to justify the new burden. That is a pilot producing evidence and the agency acting on it.

The opposite failure is the sunk-cost scale. An agency that spent six months and a budget line on a pilot will feel pressure to declare it a success regardless of the evidence, because stopping feels like admitting waste. The reframe that protects the agency is that a pilot that produced a clear "stop" answer was not wasted; it was cheap. A 90-day pilot that prevented an agency from scaling a biased screening tool across 40 units saved the agency from the far larger cost of a tool that harmed families and triggered an oversight investigation. The pilot did its job. The job was to produce evidence, and "stop" is evidence.

The agency that has never stopped a pilot has never run one. It has only ever procured.

From Pilot to Evidence Package That Satisfies Oversight

The output of a disciplined pilot is not a vendor slide. It is an evidence package: a short, honest document that an oversight body, a governance board, a court, or an advocate can read and trust. The package answers, in order, what question the pilot asked, what was measured and how, what the baseline was, what the pilot found, whether the pre-registered thresholds were met, what the equity evidence showed, what errors and near misses occurred, and what the recommended decision is with its justification. It includes the limitations honestly: the pilot ran in one unit, on non-court cases, for 90 days, so the evidence supports scaling to similar units, not a blanket agency-wide deployment.

The honesty of the package is its power. An evidence package that admits "the time saving was 14 percent, below our 20 percent target, so we recommend revising the verification workflow and re-piloting" is more trusted than a glossy claim of success, because oversight bodies have seen too many glossy claims. The package that names its own limitations is the package a court will accept when an advocate later asks how the agency knew its AI-assisted documentation was reliable. The agency can point to the baseline, the metrics, the defect rate, the equity check, and the decision rule, and say: we measured this, we set the bar in advance, and we acted on the evidence.

This evidence package also becomes the unit of institutional memory. An agency that accumulates evidence packages from pilot after pilot builds a library of what worked, what did not, and why, which makes the next pilot faster and the next budget request defensible. The deputy director's question, "did it work?", becomes answerable not with an anecdote but with a document, and the auditor's question, "was any family harmed?", becomes answerable with a defect rate, a near-miss count, and an equity check. That is the discipline that satisfies oversight, and it is the difference between an AI program that earns the right to scale and one that gets shut down.

Key Takeaways

  • Most pilots produce no evidence because they lack a baseline and a pre-registered definition of success and failure. "The workers like it" does not survive contact with an auditor, a court, or an advocate. Measure the current state before the tool is introduced, or the time-savings claim is meaningless.
  • Design the pilot backward from the decision it must inform: scale, stop, or revise. Fix six elements in writing before launch: specific scope, a baseline period, defined metrics with targets, pre-registered decision thresholds, a defined population and consent posture, and a named owner plus an oversight contact.
  • Measure the benefit honestly by netting verification time against drafting time saved. A tool that drafts in 2 minutes but needs 15 minutes of verification has not saved 18 minutes; report the net figure, not the vendor's gross demo number.
  • Measure safety with a verification-defect rate (how often an AI-assisted draft contained an error verification had to catch) and a near-miss count. Zero near misses usually means nobody is verifying, not that the tool is perfect.
  • Equity is a gate inside the pilot, not a review after harm. Measure whether defect rates or draft content differ by race, language, or neighborhood before scaling, because history (Allegheny, the Dutch childcare-benefits scandal, Michigan's MiDAS) shows these tools can encode and amplify inequity.
  • The willingness to stop is what earns oversight's trust. Pre-registered thresholds make stopping the act of following the plan rather than a referendum on the tool's champions. A pilot that produces a clear "stop" answer was cheap, not wasted.
  • The output of a disciplined pilot is an honest evidence package: the question, the baseline, the metrics, whether thresholds were met, the equity evidence, the defect and near-miss counts, the recommended decision, and the limitations. Its honesty is its power, and it is what a court will accept when an advocate asks how the agency knew its AI-assisted documentation was reliable.