Time-and-Wellbeing Analytics
The unit supervisor sat across from her director with a single sheet of paper and a problem. Six months earlier, her unit of twelve caseworkers had switched on an AI documentation tool that drafts home-visit notes and court reports from the worker's field notes. The workers said it helped. The director wanted to renew the license, which was not cheap, and the county budget office wanted to know what the agency was getting for the money. "They asked me how many hours it saved," the supervisor said. "And I realized I had no idea. I had a feeling. I had a few workers who told me they were getting home on time again. But I had nothing I could put in front of the budget office, and nothing that would survive a hard question from an advocate who wanted to know whether the time was going back to families or just going back to more files." The director nodded slowly. "So what we have," he said, "is a tool everyone likes and no proof it works." That sentence is the reason this lesson exists. A documentation tool that returns hours is the field's most humane use of AI, but a returned hour you cannot measure is a returned hour you cannot defend, cannot fund, and cannot direct toward the families who need it.
Why Measure the Hours at All
It is tempting to treat measurement as a bureaucratic afterthought, the kind of thing a budget office demands and a practitioner endures. In this field, measurement is the opposite of bureaucratic. It is the mechanism by which a humane intervention proves it is humane, and it is the mechanism by which an agency keeps a good tool and discards a bad one. Consider what is actually at stake when an agency cannot measure the time its AI documentation tool returns.
The paperwork burden is the defining pain of human-services work. Caseworkers spend a large share of every day, often half or more, documenting instead of being with families. That burden is a top driver of burnout and turnover, and turnover raises caseloads for the workers who remain, which raises their documentation burden, which drives more turnover. It is a spiral, and it harms families directly: a worker carrying thirty cases instead of twenty has less time for each child, less time for each home visit, less time to notice the thing that a careful eye would catch. When an AI documentation tool genuinely returns hours, it is intervening at the exact pressure point of that spiral. But "genuinely" is the operative word, and only measurement can establish it.
Here is the worked problem. An agency pays a vendor for an AI documentation license across a unit of twelve caseworkers. The renewal cost is real money in a constrained county budget. At renewal time, three questions arrive, and each one requires a number, not a feeling.
- The budget question: Did this tool return enough time to justify its cost? If a worker's loaded cost is roughly forty dollars an hour and the tool saves each of twelve workers four hours a week, that is forty-eight hours a week across the unit, about two thousand dollars a week in recovered labor, against a license that may cost a fraction of that. But that calculation only exists if someone measured the four hours.
- The wellbeing question: Are workers less burned out, and are they staying? Turnover in this field can run very high, and replacing a trained caseworker is enormously expensive in both money and lost institutional knowledge. If the tool reduces the documentation burden that drives burnout, the retention effect may dwarf the direct labor savings. But that only registers if someone tracked it.
- The mission question: Where did the returned time go? This is the question an advocate or an oversight body will ask, and it is the most important of the three. Hours returned from documentation can go to direct time with families, to verification of AI drafts, or to simply absorbing a larger caseload. Only the first two are wins. If the agency cannot show where the time went, it cannot claim the tool served the mission rather than the spreadsheet.
A returned hour you cannot measure is a returned hour you cannot defend, cannot fund, and cannot direct toward the families who need it.
The Three Metric Families
The three renewal questions map onto three families of metrics. Keeping them distinct matters, because an agency that measures only the first family, the time, will produce a number that looks impressive and tells an incomplete and potentially misleading story. The discipline is to measure all three together so that the time number is always read against where the time went and how the people doing the work are faring.
Time Metrics: Hours Returned
The first family measures the raw documentation time before and after the tool. The cleanest version is a per-document time: how long did it take a worker to produce a court-ready home-visit note before the tool, and how long does it take now, including the verification step. Suppose a careful home-visit note took a worker fifty minutes to write from scratch before the tool. With the AI draft plus the mandatory verification pass, it now takes twenty-two minutes: a few minutes to generate the draft and roughly eighteen minutes to verify every claim against the field notes and edit. That is twenty-eight minutes saved per note. A worker producing twenty such notes a week saves about nine hours and twenty minutes a week. Across a unit of twelve, that is more than a hundred hours a week.
The critical methodological point is that the "after" figure must include verification time, not just generation time. An agency that measures only how fast the AI produces a draft is measuring the wrong thing and inflating its own results. The honest measurement is end to end: from the moment the worker starts the documentation task to the moment a court-ready, verified, signed note enters the record. If the verification step is omitted from the measurement, the number is not just optimistic, it is a misrepresentation that will not survive scrutiny, and worse, it quietly incentivizes workers to skip verification to make the number look better.
Wellbeing Metrics: Burden and Burnout
The second family measures the human effect. Time saved is a means; reduced burnout and improved retention are part of the end. These metrics are harder to quantify than minutes, but they are not unmeasurable. The standard instruments are well established: validated burnout surveys administered before deployment and at intervals after, with the same workers, so the change is tracked over time rather than guessed. Self-reported documentation burden, measured on a simple consistent scale, captures the worker's lived experience of the paperwork load. And the hard outcome metric is turnover: how many workers left the unit in the year before deployment versus the year after, and what they said in exit interviews about workload.
Wellbeing metrics require care because they are sensitive and because they can be gamed by leadership pressure. A worker who fears that a negative survey will be read as a complaint about a tool the director championed will not answer honestly. The measurement only works if the responses are anonymous and if leadership signals genuinely that an honest negative answer is useful, not punishable. A wellbeing metric collected under pressure is worse than no metric, because it produces a false positive that hides a real problem.
Destination Metrics: Where the Time Went
The third family is the one most agencies skip and the one that matters most for defensibility. It answers the mission question: where did the returned hours go? There are three possible destinations, and only the measurement can tell them apart.
- To families. The returned hour becomes a longer home visit, an extra contact, time to arrange a service, time to sit with a child. This is the destination that justifies the entire intervention in mission terms.
- To verification. The returned hour becomes the disciplined claim-by-claim check of the AI draft against the source record. This is not a loss; it is the cost of using the tool safely, and it must be funded out of the time the tool returns rather than skipped.
- To more cases. The returned hour is absorbed by a higher caseload, so the worker does the same documentation faster but for more families and ends up no less burdened. This is the destination that quietly defeats the purpose, and it is the one an honest measurement program is designed to detect.
A practical way to capture destination is a periodic, lightweight time-use sample: for a representative slice of the unit's weeks, where did the recovered time actually go. The point is not surveillance of individual workers, which would itself drive burnout and distrust. The point is a unit-level picture honest enough to answer the advocate's question. If the returned hours all flowed into a rising caseload, the agency has not improved wellbeing or mission delivery; it has only changed the shape of the burden. Measuring destination is what keeps the time number truthful.
Building the Baseline Before You Switch It On
The single most common measurement failure is not measuring the wrong thing; it is failing to measure the "before." An agency that switches on the AI tool first and decides to measure its impact later has already lost the comparison, because there is no longer a clean pre-tool figure to compare against. Memory is not a baseline. "It used to take longer" is not a number a budget office or an advocate can use. The baseline must be captured before the tool is deployed, and it must be captured the same way the post-deployment figure will be captured, or the comparison is invalid.
Consider two agencies that both deploy the same tool. The first agency, before switching anything on, spends two weeks having its twelve workers log the actual time each documentation task took, runs the validated burnout survey, records the unit's caseload and the prior year's turnover, and samples how the workers' time was being spent. The second agency switches the tool on, and three months later, when the renewal question arrives, tries to reconstruct what things were like before. The first agency can produce a defensible before-and-after. The second agency can produce an anecdote. When the budget office or an oversight body presses, the anecdote collapses and the good tool is at risk of being cut for lack of evidence, not for lack of value.
A sound baseline has a few non-negotiable properties. It must be measured under the same definition as the later figure, so the "before" note time and the "after" note time both mean end-to-end court-ready documentation. It must cover a representative period, not a single unusually quiet or unusually chaotic week. It must record the confounders that will be needed to interpret the later number, especially the caseload, because a time saving that coincides with a caseload increase tells a very different story than the same saving with a stable caseload. And it should be documented in a way that a person who was not in the room can understand and trust, because the audience for these numbers includes people who will read them adversarially.
Confounders and Honest Attribution
The hardest part of time-and-wellbeing analytics is not collecting numbers; it is attributing a change honestly to the tool rather than to everything else that changed at the same time. A measurement program that ignores confounders will claim credit the tool did not earn, and that overclaim is exactly what an adversarial reviewer will dismantle, taking the tool's real benefits down with it.
Several confounders show up reliably. Caseload is the largest: if average caseload dropped during the measurement window because the agency hired, documentation time per worker may fall for reasons unrelated to the tool. A new state policy that simplified a form changes documentation time independently of AI. A change in supervisory expectations about note length changes it. Seasonal patterns matter, because referral volume and court schedules are not flat across the year. Even the novelty effect matters: workers may be faster in the first enthusiastic month and slower later, or slower at first while learning and faster once fluent, so a measurement taken too early or too late misleads.
The honest response to confounders is not to pretend they do not exist but to record them and reason about them out loud. If caseload fell ten percent while documentation time per note fell forty percent, the agency can reasonably argue that the bulk of the time saving is attributable to the tool, because the caseload change cannot account for a forty percent per-note drop. If a form was simplified mid-window, the agency should note it and, where possible, separate notes produced under the old form from those under the new one. The goal is a claim a skeptical reader will accept: not "the tool saved exactly nine hours and twenty minutes per worker per week," stated as if from a controlled experiment, but "documentation time per court-ready note fell from about fifty minutes to about twenty-two minutes over a representative period in which caseload was stable, and we have accounted for the form change in week six." The second claim is more modest and far more durable.
The discipline in one line: measure the before, include verification time in the after, account for the confounders, and never let the time number stand alone without showing where the time went.
Telling the Story to Each Audience
A measurement program produces numbers, but numbers do not defend themselves. The same data has to be told three different ways to three audiences who care about three different things, and getting the framing wrong can sink a good tool as surely as having no data at all.
The budget office wants return on investment. For them the story is the labor-cost calculation done honestly: hours returned per worker, multiplied by the loaded hourly cost, across the unit, net of verification time, set against the license cost, with the retention savings noted separately because turnover replacement costs are large but harder to attribute precisely. The frame is stewardship of public money, and the honesty about verification time actually strengthens the case, because a number that already accounts for the cost of safe use is harder to attack.
Agency leadership and the workforce want the wellbeing story. For them the story is the burnout and burden trend and the retention figure, told alongside the destination data so that leadership can see whether the time is reaching families. The frame here is a humane and well-staffed agency, and the destination metric is what makes the story credible rather than self-congratulatory: an agency that can show the returned hours went to families and to verification, not to a quietly rising caseload, is telling a story it can stand behind.
Advocates, courts, and oversight bodies want assurance that the tool serves the people and respects due process. For them the destination metric and the verification time are the heart of the story. Their concern is not that the agency saved money; it is that AI in a sensitive public-services context might be cutting corners on the records that families depend on. The measurement that reassures them is the one showing that verification time was preserved and funded, that the time returned reached families, and that the agency held the cardinal line: AI drafts, humans verify and decide. A time saving achieved by skipping verification is not a win to report; it is a liability to disclose. The measurement program is, in the end, an instrument of accountability as much as of efficiency, and the agency that treats it that way will find its numbers believed.
Key Takeaways
- An AI documentation tool that returns hours is the field's most humane use of AI, but an unmeasured hour cannot be defended to a budget office, funded at renewal, or directed toward families, so measurement is not bureaucracy, it is how a humane intervention proves itself.
- Measure three metric families together: time (hours returned), wellbeing (burnout, documentation burden, and turnover), and destination (where the returned hours actually went). The time number alone tells an incomplete and potentially misleading story.
- The "after" time figure must include verification time and be measured end to end, from starting the documentation task to a court-ready, signed note. Measuring only generation speed inflates results and quietly incentivizes skipping verification.
- Destination is the metric most agencies skip and the one that matters most: returned hours that go to families or to verification are wins, but returned hours absorbed by a rising caseload defeat the purpose and must be detected.
- Capture the baseline before deploying the tool, under the same definitions that will be used afterward, over a representative period, recording confounders like caseload. Memory is not a baseline, and a reconstructed "before" collapses under adversarial scrutiny.
- Account for confounders honestly, especially caseload changes, policy or form changes, seasonality, and novelty effects. A modest claim that survives a skeptical reader is worth more than an overclaim that an adversary can dismantle.
- Wellbeing metrics must be anonymous and free of leadership pressure, or they produce false positives that hide real problems. A gamed wellbeing number is worse than no number.
- Tell the story three ways: return on investment for the budget office, the wellbeing and retention trend for leadership and staff, and preserved verification time plus time reaching families for advocates and courts. The honest accounting of verification time strengthens every version of the case.
Skill.re