โ†
AI for Social Work & Human Services
Visionary ยท M16 ยท lesson 16 of 16 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Your 90-Day Transformation Plan
๐Ÿ“–
now learning

Your 90-Day Transformation Plan

15 min

The new deputy director had read the whole program. She understood grounded generation, the decision-aid rule, equity auditing, the four-axis scorecard, the governance board. She believed all of it. And on her first Monday she sat in front of a blank document with a question that none of the theory had answered: what do I actually do this week. The agency had 240 caseworkers, a CCWIS (Comprehensive Child Welfare Information System, the federal-standard case-management platform the agency runs on) that nobody loved, a union that had heard "efficiency" used as a layoff word before, a juvenile court that would scrutinize any AI-touched record, and a budget that renewed in the fall. She did not need a vision. She had a vision. She needed the first ninety days, sequenced so that the easy wins came first, the irreversible mistakes were impossible to make early, and everyone the transformation touched, leadership, the union, the court, the advocates, and the community, could see it coming and get behind it. This lesson is that ninety days.

The Shape of the First Ninety Days

A ninety-day plan is not a miniature version of the multi-year roadmap. It has a different job. The multi-year program decides where the agency is going; the first ninety days decides whether anyone will follow, and it does that by being safe to fail and visible to succeed. Two principles govern the sequencing, and getting them wrong is how most agency AI starts collapse.

Start where the harm ceiling is low and the time payback is high. The program has one use case that is both the most beneficial and the safest: AI-assisted documentation grounded in the record, verified to a court standard, with the decision being made by humans. That is where the first ninety days lives, entirely. The second use case, AI risk-screening, has a high harm ceiling (it touches the decision to investigate or remove) and demands equity auditing and mandatory human review before it can run responsibly. It does not belong in the first ninety days. An agency that opens its AI program with predictive screening is starting at the most dangerous, most contested, most legally exposed point, and one bad screen in week six ends the whole program. Documentation first is not timidity. It is sequencing the irreversible mistakes out of the early window.

Build the guardrail before the tool, not after. The instinct is to deploy the tool fast and write the policy later. Reverse it. The verification standard, the disclosure rule, the audit trail, and the decision-aid boundary all exist on paper and in training before a single live case note is drafted with AI. This is cheap to do early and nearly impossible to retrofit, because once 240 caseworkers have a habit, changing the habit costs far more than forming it correctly the first time. The ninety days is organized into three phases of roughly thirty days each: Foundation (days 1 to 30), Pilot (days 31 to 60), and Evidence (days 61 to 90). Each phase has a gate that must be passed before the next begins.

The first ninety days is not where you prove AI is powerful. It is where you prove the agency can hold the line while using it, on the safest use case, in front of everyone who will judge the rest.

Days 1 to 30: Foundation

The first thirty days produce no AI-drafted case notes at all. They produce the conditions under which AI-drafted case notes will be safe. A director who skips this phase to "show results faster" is trading a thirty-day delay for the program's credibility, which is a terrible trade in a field where one fabricated observation in a court report can end the experiment.

Stand up the small governance group and name an owner. Not the full multi-program governance board yet, but a working group with the people who can say no: legal, a frontline caseworker representative, a supervisor, a union representative, an equity or quality lead, and ideally an advocate or community voice. Name one accountable program owner. The first decision this group makes is the kill-criteria list: the categories of decision where AI will not be used at all in this pilot (anything touching removal, substantiation, or a benefit denial as a decision rather than a draft). Writing down where AI will not go, first, is what makes the court and the advocates trust where it will.

Write the four guardrail documents. The verification standard (every AI-assisted document gets a documented claim-by-claim human check against field notes, policy, and the record before filing); the disclosure rule (how AI use is noted in the record so it is transparent to a court and an advocate); the audit-trail requirement (every AI-touched record, every verification, and every human decision is logged); and the decision-aid boundary stated in one sentence the whole agency can repeat: AI summarizes, organizes, and drafts; humans decide. These are not new inventions. They are the program's non-negotiables turned into agency policy.

Pick the pilot unit and set the baseline. Choose one unit of roughly 10 to 15 caseworkers, ideally one with a supervisor who is respected and at least cautiously willing, not the most skeptical unit and not the most starry-eyed. Before any tool arrives, capture the baseline on all four scorecard axes for that unit: documentation hours per week, turnover and vacancy, the relevant equity figures, the outcome and accuracy figures, and the current (pre-AI) documentation accuracy. The baseline is the single most perishable asset in the whole program. If it is not captured now, before deployment, every later claim of improvement becomes unfalsifiable, and an unfalsifiable claim is the first thing a budget hearing or a court will dismantle.

The day-30 gate: guardrail documents approved by the working group including legal, kill-criteria list signed, pilot unit selected, and a complete four-axis baseline in hand. No baseline, no pilot. This gate is non-negotiable because it is the only one that cannot be recovered later.

Days 31 to 60: Pilot

The second thirty days put the tool in real hands on real cases, on the documentation use case only, inside the guardrails built in phase one. The goal of the pilot is not maximum adoption or maximum speed. It is to learn whether the guardrails hold under real caseload pressure, with a small enough group that a problem is a lesson and not a scandal.

Train verification as the core skill, not the tool. The training that matters is not how to make the AI draft a note; that takes ten minutes and the tool's vendor will happily provide it. The training that matters is how to verify a draft to a court standard: tracing each observation to the field notes, each policy claim to the current manual, each historical reference to the case-management record, and removing anything that cannot be traced. The pilot unit should leave training believing that their job did not become "produce the draft," it became "verify the draft," and that the time AI returns is meant to go to verification and to families, not to a higher caseload.

Do not raise caseloads during the pilot. This is the discipline that separates a wellbeing transformation from a productivity squeeze. If the pilot unit's caseload rises the moment AI saves them time, the agency has taught its workforce that AI means more work, the union's worst fear is confirmed, and the verification step gets cut first under the new pressure. Hold caseloads flat. Let the saved hours visibly go to home visits and to the verification the program depends on. The wellbeing axis of the scorecard should show the time being redirected, not absorbed.

Watch verification integrity weekly, not outcomes. Outcomes (safety, permanency) move too slowly to tell you anything in thirty days, and demanding them now only invites a fake proxy. The metric to watch weekly is verification integrity: the share of AI-assisted documents that received a documented human verification before filing, ideally spot-checked with a few blind re-reads of supposedly-verified notes to confirm the checks are real and not a rubber stamp. If verification integrity holds under real pressure in the pilot unit, the core question of the whole program, can humans keep deciding when AI makes drafting easy, has a real answer. If it decays, that is the most valuable finding the pilot can produce, and it is far cheaper to learn it here than at scale.

Run a mid-pilot incident drill. Deliberately surface a known hallucination case (a seeded draft with a fabricated observation, clearly labeled as a drill) and confirm the verification process catches it and the audit trail records it. A guardrail that has never been tested is a guardrail you do not know you have. Catching a planted error in week seven builds more confidence with the court and the union than any speed number.

The day-60 gate: verification integrity above the threshold the working group set, the incident drill caught and logged, no live case touching a kill-criteria decision, and caseworker-reported burden trending the right way. If verification integrity is below threshold, the answer is not to scale anyway; it is to extend the pilot and fix the workflow, because scaling a decayed verification habit across 240 people multiplies the risk, it does not dilute it.

Days 61 to 90: Evidence

The third thirty days turns the pilot into the evidence base that earns the next phase of investment and the trust of everyone watching. The deputy director in the opening had a budget that renewed in the fall; this phase is what she carries into that room.

Assemble the four-axis evidence package, against the baseline. For the pilot unit, report the change on all four axes relative to the day-30 baseline: wellbeing (hours redirected to families, burden and burnout trend, caseloads held flat), equity (no widening of any tracked gap, with disaggregated data), outcomes and accuracy (documentation accuracy up, no outcome harm signal, understanding that the slow outcomes are still maturing), and accountability (verification integrity, disclosure rate, audit-trail completeness, the incident drill result). This is the same four-axis scorecard the agency will run forever, used here at pilot scale for the first time. Reporting it honestly now, including any number that dipped, is what makes it credible later.

Bring the five audiences the real story, each in their language. Leadership wants the wellbeing and outcome story and a defensible cost picture. The union wants proof that caseloads were held and the saved time went to workers and families, not to layoffs or speedups. The court wants the verification standard, the audit trail, and the disclosure rule, proof that AI-touched records are more accurate and fully traceable. Advocates and the community want the equity data, disaggregated and unflinching, and the kill-criteria showing where AI is forbidden. Each audience gets the same underlying numbers, framed for their concern, with no contradiction between the versions, because the moment the union's version and the court's version diverge, both stop trusting the program.

Decide the scale question with a gate, not a foregone conclusion. The honest end of ninety days is a go, fix, or no-go decision, not an automatic expansion. Go: expand the documentation use case to the next two or three units, carrying the same guardrails and the same baseline discipline. Fix: the wellbeing and accuracy gains are real but verification integrity needs more work, so extend and harden before widening. No-go on screening, still: the documentation use case can scale, but the second use case (risk-screening) does not enter until a separate, equity-audited pilot with its own gate, because its harm ceiling is categorically higher. A program that treats expansion as automatic has stopped governing and started selling.

The day-90 gate: the four-axis evidence package complete and presented to the working group and the five audiences, an explicit go/fix/no-go decision recorded with its reasons, and the plan for the next ninety days written before this one closes. The transformation does not end at day 90. It earns the right to continue.

The Traps That End Programs Early

Most agency AI programs that fail do not fail at the technology. They fail at the sequencing and the politics, and the failures are predictable enough to name in advance.

Leading with screening. The single most common fatal error is opening with predictive risk-screening because it sounds the most impressive. It is the use case with the highest harm ceiling, the most contested history, and the most legal exposure, and one bad screen ends the program. Lead with documentation. Earn the trust on the safe use case first.

Skipping the baseline. Deploying first and reconstructing the baseline afterward makes every improvement claim unfalsifiable and every budget hearing winnable by the skeptic. The baseline is perishable. Capture it in phase one or lose the ability to prove anything.

Letting saved time become more caseload. The fastest way to confirm the union's fear and erode the verification step is to raise caseloads the instant AI saves time. Hold caseloads flat through the pilot. Make the redirected time visible.

Treating disclosure and the audit trail as paperwork. The court does not trust a tool; it trusts a traceable process. An agency that cannot show, for any given record, what AI touched it, who verified it, and who decided, has no defense when an advocate asks. The audit trail is not overhead; it is the thing that keeps the program legal.

Reporting only green. A ninety-day report with no uncomfortable number is not reassuring; it tells leadership and the court the agency is either measuring the wrong things or hiding the right ones. The dip in verification integrity during a caseload spike, reported honestly with a fix, builds more trust than a flawless story nobody believes. The deputy director who opened this lesson finished her ninety days with a report that named one number that had gone the wrong way and what she did about it, and that was the number that earned her the next year.

What the First Week Actually Looks Like

The deputy director's original problem was not strategy; it was the blank document and the question of what to do this week. So make the first week concrete, because a plan that cannot survive contact with a Monday morning is not a plan. Week one is not a tool evaluation and not a vendor meeting. It is three conversations and one list. The first conversation is with legal and the union together, in the same room, to say plainly that the program will lead with documentation, will hold caseloads flat, and will be governed by a working group that includes a frontline caseworker and a union voice. Saying this first, before any tool is chosen, defuses the layoff fear that kills agency AI programs before they start. The second conversation is with the prospective pilot supervisor, to confirm they are cautiously willing and to understand their unit's real caseload and documentation burden. The third is with the quality or equity office, to scope what baseline data already exists and what must be captured fresh. The one list is the first draft of the kill-criteria: the decisions where AI will not go. Nothing in week one touches a live case, and that is exactly the point.

This sequencing matters because the early credibility of an agency AI program is built on what it refuses to do, not on what it deploys. A director who spends week one defining the boundary, naming the people who can say no, and protecting caseloads has spent that week building the trust that every later phase will draw down. A director who spends week one in a vendor demo has started the clock on the suspicion that this is a productivity squeeze dressed up as innovation. The first week sets the frame, and the frame is harder to change than the tool.

Key Takeaways

  • The first ninety days has a different job than the multi-year roadmap: it proves the agency can hold the line while using AI, on the safest use case, in front of everyone who will judge the rest. Make it safe to fail and visible to succeed.
  • Lead with AI-assisted documentation (high time payback, low harm ceiling) and keep predictive risk-screening out of the first ninety days entirely. Opening with screening starts at the most dangerous, most legally exposed point, where one bad screen in week six ends the program.
  • Build the guardrail before the tool: the verification standard, the disclosure rule, the audit-trail requirement, and the one-sentence decision-aid boundary all exist in policy and training before the first live AI-drafted note. Guardrails are cheap to build early and nearly impossible to retrofit across 240 caseworkers.
  • Days 1 to 30 (Foundation) produce no AI notes: stand up a small governance working group with people who can say no, write the kill-criteria list first, write the four guardrail documents, pick a 10 to 15 person pilot unit, and capture the four-axis baseline. No baseline, no pilot.
  • Days 31 to 60 (Pilot) train verification as the core skill (not the tool), hold caseloads flat so saved time visibly reaches families and verification, watch verification integrity weekly rather than slow outcomes, and run a seeded incident drill to prove the guardrail actually catches a hallucination.
  • Days 61 to 90 (Evidence) assemble the four-axis evidence package against the day-30 baseline, bring the same honest numbers to five audiences (leadership, union, court, advocates and community), and make an explicit go/fix/no-go decision rather than treating expansion as automatic.
  • Each phase has a gate that must be passed before the next begins, and the day-30 baseline gate is the only one that cannot be recovered later. If verification integrity is below threshold at day 60, extend and fix rather than scaling a decayed habit across the agency.
  • The program-ending traps are predictable: leading with screening, skipping the baseline, letting saved time become more caseload, treating disclosure and the audit trail as mere paperwork, and reporting only green. Reporting the one number that went the wrong way, with a fix, earns more trust than a flawless story.