Prioritizing: Documentation, Eligibility, Screening, Navigation
The deputy director had four vendor demos on her calendar in a single week, and every one of them ended the same way: a polished sales engineer telling her that their tool was the obvious first thing the agency should buy. Monday it was an AI documentation assistant that drafted case notes from a recording of a home visit. Tuesday it was an eligibility engine that promised to clear the SNAP (Supplemental Nutrition Assistance Program, the federal food-assistance benefit) backlog. Wednesday it was a predictive risk-screening model that scored incoming child-abuse reports for likelihood of future harm. Thursday it was a resource-navigation chatbot that would answer client questions about housing and food around the clock. Each vendor had a return-on-investment slide. Each one looked compelling in isolation. And the deputy director, who carried responsibility for an agency of three hundred caseworkers, a court that reviewed her staff's reports, and roughly forty thousand families a year, understood that the slides were the wrong way to decide. The question was not which tool had the best demo. The question was which of these four uses should the agency do first, which it should do carefully, and which it should not do at all yet. That question is what this lesson answers, and the tool that answers it is a prioritization matrix that weighs hours returned against the harm a failure could do to a family.
Why Prioritization Is the Strategist's First Real Decision
At Level 1 through Level 3 of this program, the unit of work was a single caseworker and a single document or screening signal. At Level 4 the unit of work changes. A strategist is not deciding whether to use AI on one court report; the strategist is deciding where an entire agency should spend its limited budget, its limited political capital, and its limited capacity for change, knowing that every one of those is scarce and that a wrong sequencing decision can set the program back years.
The four candidate uses the deputy director saw are not interchangeable. They are the four places AI has actually entered human services in 2026, and they sit at radically different points on two axes that matter more than any vendor's feature list. The first axis is benefit: how many hours of caseworker time a use returns, how many families it reaches, how directly it relieves the documentation burden that is the field's defining pain and a top driver of burnout and turnover. The second axis is harm: what happens to a real person if the tool fails. A failed case-note draft that a worker catches in verification costs minutes. A failed eligibility determination leaves a family without food or shelter. A biased risk score, encoded with the inequities of its training data, can contribute to a child being removed from a family that should have been supported instead. These are not the same kind of risk, and a prioritization that treats them as equivalent is malpractice.
Sequencing matters because trust is the agency's scarcest asset and it is spent, not replenished, by early failures. An agency that opens its AI program with the highest-risk use case, a predictive risk-screening model, and suffers a public equity failure in year one will find that caseworkers, advocates, the court, and the community no longer trust anything the agency does with AI, including the safe and beneficial documentation work that could have returned thousands of hours to families. Conversely, an agency that opens with the lowest-risk, highest-benefit use case builds a track record of returned hours and verified accuracy that earns the standing to take on harder problems later. The order is the strategy.
The vendor demo answers "is this tool good." The strategist's matrix answers "should this be the thing we do first, do carefully, or not do yet." Only the second question protects families.
The Two Axes: Benefit and Harm
The matrix has two axes, and getting each one honestly scored is most of the work. Scoring them dishonestly, by letting a vendor's optimistic numbers stand in for the agency's own evidence, produces a matrix that looks rigorous and decides badly.
Scoring the Benefit Axis
Benefit, for a human-services agency, is best measured in hours returned to direct work with families and in the breadth of people a use reaches. The documentation burden is the anchor here. When caseworkers spend half or more of every day documenting instead of being present with families, the use that most directly attacks that half-day is the one with the largest benefit. A documentation assistant that drafts a home-visit note in two minutes that would otherwise take twenty-five, applied across three hundred caseworkers each writing several notes a week, returns hours at a scale that no other use in the four can match. The benefit is large, it is direct, and it accrues to the field's single most painful problem.
Contrast that with a resource-navigation chatbot. Its benefit is real, connecting a person to housing or food faster, but it is narrower: it touches the moments a client is searching for a service, not the daily grind of documentation, and the hours it returns to caseworkers are smaller than the hours a documentation tool returns. Score benefit by asking three concrete questions of each use. How many hours does it return per worker per week, measured against the agency's own time studies rather than the vendor's slide. How many families or staff does it reach. And how directly does it relieve the documentation burden that drives burnout, since a use that lowers burnout also lowers turnover, which lowers caseloads, which compounds the benefit over time.
Scoring the Harm Axis
Harm is scored by asking a single hard question: if this tool fails in the way its technology is known to fail, what happens to the person on the other end, and can it be undone. This question separates the four uses more sharply than benefit does, and it is the axis vendors are least willing to discuss honestly.
A documentation assistant fails by hallucinating, that is, by generating a confident, professionally worded observation, policy citation, or piece of history that is factually false. That failure is serious, because the case note becomes a legal record, but it is catchable. A worker verifying every factual claim against the field notes and the case record before filing catches the invented observation before it enters the record. The harm is real but it sits behind a human review gate that can stop it.
An eligibility tool fails by misapplying policy, producing a denial based on a rule that does not govern the family's situation. That failure reaches a person directly: a wrong denial leaves a family without SNAP, without Medicaid (the joint federal-state health-coverage program for low-income people), without a housing voucher, in a moment of acute need. The harm is more direct than a documentation error because the output of the eligibility tool is closer to the consequential decision. It is still catchable by a worker who verifies the policy against the current manual, but the cost of a miss is higher and lands faster.
A predictive risk-screening tool fails in the most dangerous way of all. It can encode the inequities present in its training data and surface a biased score that, if treated as a verdict rather than as one audited input under mandatory human review, contributes to a decision to investigate or remove a child from a family that should have been supported. The history of these tools, from the contested debate over the Allegheny Family Screening Tool in child welfare to benefits fraud-detection failures such as the Dutch childcare-benefits scandal and Michigan's MiDAS system, shows that the harm is not hypothetical. It has happened, it has fallen disproportionately on the families with the least power to challenge it, and some of it could not be undone. This is the highest-harm use in the four, and the matrix must reflect that.
A resource-navigation tool fails by sending a person in crisis to the wrong door: a closed program, an outdated phone number, a service the person does not qualify for. The harm is real, a person already in distress is sent further from help, but it is generally recoverable and lower in severity than a wrong removal or a wrong benefits denial, provided the agency maintains a freshness check on the resource data and never lets the tool make a promise the agency cannot keep.
Placing the Four Uses on the Matrix
Plot the two axes as a simple grid: benefit on one side, harm on the other, four quadrants. High benefit and low harm is the place to start. High benefit and high harm is the place to proceed carefully, with the heaviest safeguards. Low benefit and high harm is the place to defer or decline. Low benefit and low harm is optional, a candidate for later if capacity allows. Now place the four uses the deputy director was sold.
Documentation assistance sits in the high-benefit, lower-harm quadrant. It returns the most hours, it reaches every caseworker who writes a note, it attacks the documentation burden directly, and its failure mode (hallucination) sits behind a human verification gate that a trained worker can operate. This is the goldmine, and it is the goldmine precisely because it is the safest place for AI to do the most good. It is the agency's first move.
Eligibility support sits in the high-benefit, higher-harm quadrant. It returns substantial hours by applying complex benefits policy at scale and it reaches many families, but its failure lands closer to a consequential decision, because a wrong determination denies someone food or shelter. It is a strong second move, taken with the discipline that the human always makes the determination, that policy is verified against the current source, and that the person retains full due-process rights including notice and a fair hearing.
Predictive risk-screening sits in the high-harm quadrant and must be treated as the most carefully handled use of all. Its potential benefit, surfacing early indicators of abuse or neglect, is real, but its harm potential is the highest in the field and its history of encoding inequity is documented. It is never the agency's first move. When an agency takes it on at all, it does so only after it has built the verification culture, the equity-auditing program, and the mandatory-human-review discipline that the documentation and eligibility work taught the organization. The score on the harm axis is high enough that the matrix should push this use to the back of the sequence, behind an equity gate that can stop deployment before any family is harmed.
Resource navigation sits in the moderate-benefit, lower-harm quadrant. It is a reasonable early or parallel move because its harm is recoverable and its benefit, while narrower than documentation, is genuine and visible to clients. It can run alongside the documentation work without competing for the same political capital, provided the agency keeps the resource data fresh and never lets the tool promise something the agency cannot deliver.
Start where benefit is high and harm is caught by a human gate. Defer where harm is high and history says the tool encodes inequity. The sequence is documentation, then eligibility and navigation, then, only if ever and only behind an equity gate, screening.
A Worked Prioritization for a Real Agency
Return to the deputy director with her three hundred caseworkers and forty thousand families a year. Suppose her own time study, not the vendor's slide, shows that caseworkers spend an average of three hours a day on documentation, roughly half their working time, and that a verified documentation assistant could realistically return one of those three hours per worker per day after accounting for the time verification itself requires. One hour per day across three hundred workers is three hundred hours a day returned to the agency, the equivalent of dozens of additional full-time caseworkers' worth of family-facing time, recovered without a single new hire. That is the benefit number that belongs on the matrix, and it is enormous.
Now weigh the harm. The documentation assistant's worst realistic failure, an invented observation in a court report, is severe, but the agency can build a verification step into the workflow and a supervisory review gate on top of it, and the failure is caught before it reaches the court. The harm is bounded by a control the agency already knows how to operate. High benefit, controllable harm: documentation is the clear first move, and the matrix says so without ambiguity.
The eligibility engine the deputy director saw on Tuesday promised to clear a SNAP backlog of several thousand pending applications. The benefit is real: faster determinations mean families wait days instead of weeks for food assistance, and workers spend less time hunting through nested policy. But the harm axis carries more weight here. If the engine misapplies the categorical-eligibility rule, a family with a child receiving Supplemental Security Income (the federal benefit for people with disabilities and limited income) who should qualify automatically could be wrongly denied. So the matrix places eligibility second, deployed only with the human making every determination, the policy verified against the current manual, and a fair-hearing path preserved. The agency takes the benefit and contains the harm by keeping the determination human.
The predictive risk-screening model from Wednesday is the one the deputy director declines to make a first or second move, no matter how good the demo. Her agency has not yet built an equity-auditing program, has not yet established mandatory-human-review discipline as a habit, and has not yet earned the community trust that would let it deploy a high-harm tool defensibly. The matrix tells her to defer it behind an equity gate, to revisit it only after the documentation and eligibility work has built the agency's verification and audit muscles, and to subject it, if it is ever deployed, to the bias testing that the history of these tools makes non-negotiable. Declining to do the highest-harm thing first is not timidity. It is the strategist protecting the families the agency serves and the program's own future.
The navigation chatbot from Thursday she runs as a parallel, lower-stakes pilot, with a strict freshness check on its resource directory and a hard rule that it routes people to human staff for anything consequential. It builds visible client-facing value while the documentation work builds internal value, and its low harm means a stumble will not cost the program its credibility.
The Traps That Corrupt a Prioritization
A prioritization matrix is only as honest as its inputs, and several predictable forces push agencies toward the wrong sequence. Naming them is part of the discipline.
The vendor-driven sequence. The most aggressive sales motion in 2026 is around the highest-harm tools, because predictive screening and eligibility automation carry the largest contract values. A strategist who lets the loudest vendor or the biggest ROI slide set the order will be steered straight toward the high-harm quadrant first. The matrix exists precisely to put the agency's own benefit-and-harm evidence, not the vendor's, in charge of the sequence.
The efficiency-only score. An agency under budget pressure is tempted to score only the benefit axis, ranking uses purely by hours saved or backlog cleared, and to treat harm as a compliance footnote. This is how the field's documented failures happened. A screening tool scored only on the speed it added to intake, with its equity harm unweighted, will rank far too high. The harm axis is not a footnote; it is the axis that protects people, and it must carry equal weight in the placement.
The reversibility blind spot. Not all harms are equal even at the same severity, because some can be undone and some cannot. A wrong benefits denial is severe but reversible through a fair hearing that restores the benefit and, where the law provides, back-dates it. A child wrongly removed because a biased risk score was treated as a verdict suffers a harm that no later correction makes whole. The matrix should weight irreversible harms more heavily than reversible ones of the same surface severity, which pushes the screening use even further toward the defer-or-decline corner.
The capacity illusion. A use scores well on benefit only if the agency can actually operate the controls that contain its harm. A documentation assistant returns hours only if workers have the capacity to verify every draft; if the agency floods the returned time with more cases instead of protecting time for verification, the benefit evaporates and the harm grows. Scoring a use as high-benefit while ignoring whether the agency can staff its safeguards is the capacity illusion, and it is how a tool that looked safe on the matrix produces harm in the field.
From Matrix to Roadmap
The prioritization matrix is not the roadmap; it is the input that orders the roadmap. Once the four uses are placed, the strategist converts the placement into a sequence with gates between the stages, so that the agency earns its way from the safest, highest-benefit work to the harder problems rather than taking them all on at once.
The sequence that the matrix produces for almost every human-services agency is the same in shape. First, documentation assistance, the goldmine, deployed with a verification step built into the workflow and a supervisory gate, because it returns the most hours at the most controllable risk and it builds the agency's verification culture. Second, in parallel or close behind, resource navigation as a lower-stakes client-facing pilot and eligibility support as a higher-benefit, higher-harm move taken with the determination kept human and due process preserved. Last, and only after the agency has built the equity-auditing program and the mandatory-human-review discipline that the earlier stages teach, the careful consideration of predictive screening behind an equity gate that can stop deployment before harm, if the agency chooses to take it on at all.
Each transition between stages is a gate, not a calendar date. The agency does not move to eligibility because six months passed; it moves because it can demonstrate that its documentation verification is working, that workers have the capacity to operate it, and that the returned hours are real and measured rather than asserted. It does not move to screening because a vendor's contract is up for renewal; it moves, if at all, because it has stood up an equity-auditing program that can test the tool for bias before any family is exposed to it. The matrix decides the order. The gates decide the pace. Together they keep the agency from doing the high-harm thing before it has built the discipline to do it safely, which is the single most important thing a human-services AI strategist does.
Key Takeaways
- Prioritization is the strategist's first real decision: not whether a tool is good in a demo, but which AI use the agency should do first, do carefully, or not do yet, because budget, political capital, and the agency's capacity for change are all scarce and a wrong sequence can set the program back years.
- Score every candidate use on two axes. Benefit is measured in hours returned to family-facing work and the breadth of people reached, anchored on the documentation burden that drives burnout. Harm is measured by what happens to the person if the tool fails in the way its technology is known to fail, and whether that failure can be undone.
- The four real 2026 uses sit at very different points. Documentation assistance is high-benefit and lower-harm (its hallucination failure sits behind a human verification gate), making it the goldmine and the first move. Eligibility support is high-benefit and higher-harm because a wrong denial leaves a family without food or shelter. Resource navigation is moderate-benefit and lower-harm. Predictive risk-screening is the highest-harm use, with a documented history of encoding inequity.
- The recommended sequence for almost every agency is documentation first, then eligibility and navigation, and predictive screening last, only behind an equity gate, only after the agency has built its verification and equity-auditing discipline, and only if the agency chooses to take it on at all.
- Weight irreversible harms more heavily than reversible ones of the same severity. A wrong benefits denial is recoverable through a fair hearing; a child wrongly removed because a biased score was treated as a verdict is a harm no later correction makes whole. This pushes screening toward the defer-or-decline corner.
- Guard against the traps that corrupt a prioritization: letting vendors set the sequence, scoring only efficiency while treating harm as a footnote, ignoring whether a harm is reversible, and the capacity illusion of scoring a use as high-benefit when the agency cannot actually staff the controls that contain its harm.
- The matrix orders the roadmap; gates set the pace. The agency advances from one stage to the next not on a calendar but on demonstrated evidence that its safeguards work, that workers have capacity to operate them, and that the returned hours are real and measured.
- Declining to deploy the highest-harm tool first is not timidity. It is the strategist protecting the families the agency serves and protecting the program's own credibility, which is spent by early failures and earned by a track record of safe, verified, hours-returning work.
Skill.re