โ†
AI for Social Work & Human Services
Strategic ยท M1 ยท lesson 1 of 19 ยท in progress
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Assessing Agency AI Readiness
๐Ÿ“–
now learning

Assessing Agency AI Readiness

15 min

The vendor demo was excellent. The AI documentation tool transcribed a mock home visit, drafted a clean case note in ninety seconds, and the room of supervisors nodded along. The deputy director was ready to sign. Then the quality manager asked a question that stopped the meeting: "Where would the case data actually live, and could we pull a full audit trail if a judge asked?" Nobody from the agency could answer. The vendor could, but the answer involved a cloud region and a log-retention policy the agency had never reviewed. A second question followed: "Which of our case-management systems would this connect to, and do our caseworkers have the time in their day to verify every draft?" The agency ran three different case-management systems across its programs, a decade of accreted county and state platforms, and its caseworkers were carrying thirty families each. The tool was good. The agency was not ready. The deputy director did not sign that day, and that was the right call. Readiness is not whether the tool works in a demo. It is whether your agency can deploy it without harming the people you serve.

What Readiness Actually Means

This lesson opens Level 4, where the work shifts from using AI well as a practitioner to leading AI well as a strategist. The first strategic act is not picking a tool or writing a roadmap. It is an honest assessment of whether the agency is ready, because deploying AI into an agency that is not ready does not produce the time-back-to-families benefit the field needs. It produces hallucinated court reports nobody had time to verify, eligibility errors at scale, and equity harms that surface only after a family is hurt. Readiness assessment is the diligence that earns the right to deploy.

Readiness has a specific meaning here, and it is broader than technology. An agency is ready to deploy a given AI use case when four things are true at once: its systems can support the tool and capture the record of how it is used; its workforce has the capacity and the skill to verify AI output to a court-record standard; its equity and due-process posture can catch and prevent the harms the tool could cause; and its governance can own the decision and answer for it. A gap in any one of the four is a reason to pause, not a detail to fix later. The agency in the opening story had a systems gap and a workforce-capacity gap, and either alone was enough to make signing the wrong call.

Readiness is also use-case specific, not agency-wide in the abstract. The same agency can be ready to deploy AI-assisted case-note drafting, the documentation goldmine, with its grounded generation and verification discipline, and nowhere near ready to deploy AI risk-screening, which carries equity stakes that demand a mature auditing program before a single signal reaches a worker. "Are we ready for AI?" is the wrong question. "Are we ready for this use case, given its benefit and its specific harm risk?" is the question a readiness assessment answers. This is why the assessment precedes the roadmap: you cannot sequence what you have not honestly scored.

Readiness is not whether the tool works in a demo. It is whether the agency can deploy it without harming the people it serves, assessed one use case at a time.

Dimension One: Systems and Data

The first dimension is technical, and the questions are concrete. What case-management systems does the agency run, and will the AI tool integrate with them? Many agencies, like the one in the opening story, run several systems across programs: a state CCWIS (Comprehensive Child Welfare Information System, the federally defined platform standard for child-welfare case management) for child welfare, a separate eligibility system for benefits, and older platforms in housing or aging services. An AI tool that integrates cleanly with one and not the others is ready for one program, not the agency.

The deeper question under integration is grounding. As taught throughout this program, AI is far safer when it generates from the actual record through retrieval-augmented generation (RAG, a technique that connects the model to a specific document set before it generates output) rather than from a bare prompt and the model's memory. A systems assessment asks whether the agency's data is in a state where the tool can be grounded on it: is the case record digital, structured, and accessible to the tool, or is it scattered across scanned PDFs, free-text notes, and paper? An agency whose records cannot ground the model is an agency whose AI will hallucinate more, because grounding is what holds it to the record.

The systems assessment also has to answer the audit and privacy questions the quality manager raised. Can the agency capture the record of how AI is used, the disclosure, source, verification, decision, and timeline that make an AI-assisted document defensible to a court? Where does the data live, who at the vendor can access it, how long are audit logs retained, and does that retention match the case record's lifespan rather than a vendor default? Sensitive PII (personally identifiable information, the data that identifies a specific person, such as a name, address, or case details) about children and families flows into any AI tool, and a systems assessment that does not nail down where it goes and how it is protected has not assessed readiness; it has assessed features. An agency that cannot answer these questions is not ready, no matter how good the demo was.

Dimension Two: Workforce Capacity and Skill

The second dimension is the workforce, and it is the one agencies most often skip because it is the least flattering to examine. The hardest-won lesson of AI in this field is that the human is the error-detection mechanism: a hallucinated observation, a misapplied SNAP (Supplemental Nutrition Assistance Program, the federal food-assistance benefit) rule, a fabricated prior history, all of these are caught only by a worker who verifies the AI draft to a court-record standard. A workforce that cannot or does not verify turns an AI tool from a time-saver into a harm multiplier.

So the workforce assessment asks two distinct questions. The first is skill: do caseworkers know how to verify, as the specific discipline this program teaches, tracing each factual claim to its source rather than reading the draft for general sense? Skill can be built through training, and the existence of a training path is itself part of readiness. The second question is harder: capacity. A caseworker carrying thirty families and losing half the day to documentation does not have idle time waiting to be filled with verification. If the agency deploys AI to cut documentation time and then assigns the saved hours straight back as more cases, it has removed the very slack that made verification possible. The tool will be used to file faster, unverified, and the harm will follow.

This is the trap the opening story's agency was about to walk into. The tool would have saved drafting time, but the workers had no capacity to spend that time verifying, so under caseload pressure they would have filed AI drafts the way they were filing their own rushed notes, except now with hallucinations they were not trained to catch. A readiness assessment scores capacity honestly: what is the current caseload, how much time does documentation consume, and where will the verification time come from? If the answer is "the time AI saves will fund verification and time with families," the agency is on the right path. If the answer is "we will absorb the savings as productivity," the agency is not ready, because it has planned the harm in from the start.

It helps to make the math concrete. Suppose AI drafting saves a worker forty minutes on a court report that used to take an hour. That forty minutes is exactly the budget for tracing each observation to the field notes, each policy citation to the current manual, and each historical claim to a record in the case-management system, the verification this program teaches. An agency that lets the worker keep that forty minutes for verification and direct family contact has converted a time saving into a quality and wellbeing gain. An agency that responds to the saving by raising the worker's caseload from thirty families to thirty-five has spent the same forty minutes on more cases and left zero minutes for verification. The first agency reduced documentation risk; the second increased it while believing it had improved efficiency. The workforce-capacity score is the number that tells these two agencies apart, and an honest assessment refuses to record a capacity the agency does not actually intend to protect.

Capacity also interacts with turnover, which is the field's chronic condition. A unit losing experienced workers to burnout is a unit where the remaining workers carry higher caseloads and where verification skill is constantly being rebuilt in new hires. Deploying an AI tool into that environment without a training path and protected verification time does not stabilize the unit; it adds a new, error-prone step that tired and inexperienced workers are least equipped to police. The readiness assessment treats workforce stability as part of capacity, because a tool that is safe in a fully staffed unit can be unsafe in the same unit six months into a staffing crisis.

Dimension Three: Equity and Due-Process Posture

The third dimension is the one the program treats as first, not last: equity and due process. An agency's readiness for an AI use case depends heavily on whether it can catch and prevent the equity harms that use case could cause, and this dimension is what most sharply separates readiness for documentation from readiness for screening.

For documentation tools, the equity question is real but contained: does the agency have the verification and audit discipline to ensure AI-drafted records do not encode bias through, for example, the loaded characterizations a fabricated history can introduce into how a judge reads a family? For risk-screening tools, the equity question is existential. Predictive screening can encode the inequities in its training data, and history proves it: the Allegheny Family Screening Tool debate over disparate impact, the Dutch childcare-benefits scandal that wrongly accused thousands of families of fraud, Michigan's MiDAS system that falsely flagged tens of thousands for unemployment fraud. An agency is not ready to deploy risk-screening unless it already has, or is building before deployment, a continuous equity-auditing capability: the ability to test the tool for disparate outcomes across the populations it serves, before harm, and on an ongoing basis rather than once.

The due-process side of this dimension asks whether the agency can preserve the rights of the people it serves with AI in the loop. The non-negotiable that AI informs and humans decide is the spine here. Can the agency guarantee that no consequential decision, removal, substantiation, denial of benefits, is made by a model? Does it have the disclosure practice that lets a family know AI was used and challenge a determination? Can it provide the audit trail that a fair hearing requires? An agency whose answer to these is "we would figure that out during deployment" is not ready, because due process is not a feature you add after a family has already been denied without it. The equity-and-due-process posture is assessed before deployment precisely because its failures cannot be undone after.

An agency may be ready to deploy AI documentation and nowhere near ready to deploy AI screening. The equity-and-due-process posture is what separates the two, and it is assessed first.

Dimension Four: Governance and Accountability

The fourth dimension is governance: whether the agency has someone who owns the AI decision and a structure that can answer for it. Technology, workforce, and equity readiness mean little if nobody is accountable for how AI is used across the agency, because without ownership the disciplines erode the moment caseload pressure rises.

Governance readiness asks whether the agency has, or can stand up before deploying, a few specific things. There must be a clear owner of AI use, not the vendor, and not a single enthusiastic supervisor, but an accountable role or body. There must be a written policy that says which AI uses are permitted, which decisions must remain human, what verification is required before a document is filed, and what disclosure is owed to the people served. There must be the legal, equity, and frontline-practice voices at the table when AI decisions are made, because each catches risks the others miss: legal sees the due-process exposure, equity sees the disparate-impact risk, and practice sees whether a workflow will actually survive a caseload. And there must be a way to respond when an AI tool causes harm, because it will eventually produce an error that reaches a case, and an agency without an incident-response path will discover that gap at the worst possible moment.

The governance dimension is also where the readiness assessment itself gets owned. The agency in the opening story had no governance structure, which is why a vendor demo nearly became a procurement decision with no diligence behind it. A mature agency runs the readiness assessment through its governance body, scores each of the four dimensions for the specific use case, and produces a go, no-go, or not-yet decision with the gaps named. That decision, documented, is itself a piece of due-process and oversight protection: it shows that the agency considered the harms before deploying, which is exactly what a court, an auditor, or the public will later ask.

Turning the Assessment Into a Decision

A readiness assessment is only useful if it produces a decision the agency can act on and defend. The four dimensions, systems and data, workforce capacity and skill, equity and due-process posture, and governance and accountability, are scored honestly for the specific use case, and the pattern of scores points to one of three outcomes.

The first outcome is go: the agency is ready across all four dimensions for this use case, and deployment can proceed with the disciplines in place. This is most achievable for the documentation goldmine in an agency with digital records, a trained workforce that has the capacity to verify, and a governance structure that owns the audit trail. The second outcome is not yet: the use case is sound and the benefit is real, but one or more dimensions has a gap that must close before deployment. The opening story's agency was a not-yet: the tool was good, but the systems gap (where does the data live, can we audit it) and the workforce-capacity gap (do workers have time to verify) had to close first. Not-yet is not a rejection; it is a sequenced plan, and naming the gaps is what turns a stalled procurement into a readiness roadmap. The third outcome is no: the harm risk of this use case exceeds what the agency can manage, and it should not deploy. A risk-screening tool in an agency with no equity-auditing capability and no governance is a clear no, because the history of these tools shows what happens when they are deployed without the posture to catch their harms.

Two final disciplines make the decision durable. First, the readiness assessment is not a one-time gate; it is revisited as the agency changes. A workforce gap closes as training spreads, a systems gap closes as records are digitized, a governance gap closes as a board stands up, and a use case that was not-yet last year may be a go this year. Second, the decision and its reasoning are documented, because the assessment is itself part of the agency's defensibility. When the deputy director declined to sign and recorded why, the agency created evidence that it weighs harm before benefit, which is the posture that protects the people it serves and the agency alike. The strategist's first job is not to say yes to AI. It is to know, honestly and per use case, when the agency is ready to, and to be able to prove the question was asked.

Key Takeaways

  • Readiness is not whether an AI tool works in a demo; it is whether the agency can deploy it without harming the people it serves. The assessment is the diligence that earns the right to deploy, and it precedes the roadmap because you cannot sequence what you have not honestly scored.
  • Readiness is assessed across four dimensions at once: systems and data, workforce capacity and skill, equity and due-process posture, and governance and accountability. A gap in any one is a reason to pause, not a detail to fix later.
  • Readiness is use-case specific. The same agency can be ready for AI-assisted documentation (the goldmine, with grounding and verification) and nowhere near ready for AI risk-screening, which demands a mature equity-auditing program before a single signal reaches a worker.
  • The systems dimension asks whether the tool integrates with the agency's case-management systems (often several, including a state CCWIS), whether records are digital and structured enough to ground the model through RAG, and where sensitive PII lives, who can access it, and whether audit retention matches the case record's lifespan.
  • The workforce dimension asks both skill (can workers verify to a court-record standard) and capacity (do they have the time). An agency that absorbs AI's time savings as more caseload rather than funding verification has planned the harm in from the start, because the human is the error-detection mechanism.
  • The equity and due-process dimension is assessed first and separates documentation readiness from screening readiness. History (Allegheny, the Dutch childcare-benefits scandal, Michigan's MiDAS) shows screening tools can encode inequity, so readiness requires continuous equity auditing and a guarantee that no consequential decision is made by a model.
  • The governance dimension asks whether an accountable owner, a written policy, the legal-equity-practice voices, and an incident-response path exist. Without ownership, the other disciplines erode under caseload pressure the moment they become inconvenient.
  • The assessment produces a go, not-yet, or no decision for the specific use case, is revisited as gaps close over time, and is documented with its reasoning. The documented decision is itself due-process and oversight protection: it proves the agency weighed harm before benefit.