โ†
AI for Social Work & Human Services
Visionary ยท M4 ยท lesson 4 of 16 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Identifying Responsible New Use Cases
๐Ÿ“–
now learning

Identifying Responsible New Use Cases

15 min

An innovation manager at a state human-services department opened her inbox on a Monday to find four proposals for new AI use cases, each forwarded by a different program director who had read the same article over the weekend. The first wanted AI to draft the narrative section of adult-protective-services investigation reports. The second wanted AI to predict which families on the child-welfare caseload were most likely to have a placement disruption, so workers could intervene early. The third wanted a chatbot to answer benefit-eligibility questions for the public on the agency's website. The fourth wanted AI to score incoming child-abuse hotline calls by urgency, to help the screening unit triage a backlog of nearly six hundred calls a week. All four were sincere, all four promised relief from a real and crushing workload, and all four landed on her desk as if they were the same kind of request. They were not. Two of them belonged to a documentation pattern her agency had already proved it could run safely. One of them was a screening tool wearing a friendlier name. And one of them was a public-facing eligibility oracle that could deny a person food or shelter with no caseworker in the room. Her job that week was not to say yes or no. It was to tell which was which, and to have a defensible reason for each call. This lesson is about the discipline she used.

Why New Use Cases Are the Hardest Decision

By the time an agency reaches the stage of identifying new use cases, it has usually succeeded at something. The documentation program works. Caseworkers trust the verification discipline. The audit trail satisfies oversight. That success is precisely what makes the next decision dangerous, because success generates appetite, and appetite does not discriminate between use cases that share the safe one's properties and use cases that only share its excitement. A program director who watched AI return four hours a week to home visits naturally assumes the next AI idea will do the same. The innovation manager's job is to hold the uncomfortable truth that the properties that made documentation safe do not automatically transfer to anything else AI can do.

The reason this is the hardest decision in the program is that it is the one most exposed to motivated reasoning. Every new use case arrives attached to a real pain: a backlog, a burnout statistic, a family that fell through a gap. The pain is genuine, and the relief AI promises is genuine, and that combination makes it very easy to wave a use case through on the strength of the problem it solves rather than the liability it imports. The discipline of identifying responsible new use cases is, more than anything, the discipline of separating the problem from the proposed solution, and of asking whether the specific solution brings an equity or due-process liability that the agency cannot defend, no matter how real the problem.

Consider the second proposal, the placement-disruption predictor. The pain is undeniable: a placement disruption, where a child is moved abruptly from one foster home to another, is one of the most damaging things that can happen to a child in care, and a worker who could see it coming might prevent it. But notice what the proposal actually is. It is a predictive risk model that scores families and children, ranks them, and directs scarce intervention resources toward the high scores and away from the low ones. That is the screening pattern, the same pattern behind the Allegheny Family Screening Tool debate and the benefits fraud-detection failures of the Dutch childcare-benefits scandal and Michigan's MiDAS, where automated scoring encoded and amplified inequity against the people it was meant to serve. The pain is real. The proposed solution is the most equity-fraught pattern in the field, and the worthiness of the goal does not change the category of the tool.

A real problem does not make a use case responsible. The question is never how badly you need the relief; it is what liability the specific solution imports.

The Three Questions That Classify a Use Case

The innovation manager did not evaluate the four proposals on their promised benefits. She ran each one through three classifying questions, in order, because the order matters: a use case that fails the first question is disqualified before its benefits are even weighed.

Question One: Does It Decide, or Does It Inform?

The first and most important question is whether the use case, as proposed, would make or directly drive a consequential decision, or whether it would only inform a decision a human still owns. This is the cardinal rule, AI informs and humans decide, applied as a classification test rather than a slogan. A use case that drafts a report a worker then verifies and signs informs. A use case that scores a family and routes agency resources based on that score is much closer to deciding, even if a human technically clicks the final button, because the score shapes the human's judgment before the human ever exercises it.

Apply it to the four proposals. The first, drafting investigation-report narratives, informs: the worker still verifies every claim against the record and owns the report. The third, the public eligibility chatbot, can decide, because a member of the public with no caseworker present may act on the chatbot's answer as if it were a determination, and a wrong answer can cause someone to not apply for a benefit they qualify for. The fourth, the hotline-call urgency scorer, sits in the dangerous middle: it does not formally decide whether to screen a call in, but it powerfully shapes which calls a stretched screening unit looks at first, and a call wrongly scored low can sit while a child is in danger. The first question alone already sorts these four into different risk classes.

Question Two: Whose Data, and Whose Bias?

The second question asks what data the use case learns from or operates on, and whose historical inequities that data carries. Any use case that learns patterns from the agency's own historical case data inherits whatever bias is in that history. If the agency historically investigated certain neighborhoods more, a model trained on that history will treat those neighborhoods as higher risk, not because the families there are higher risk but because the data records past attention as if it were present danger.

This question cleanly separates the documentation proposals from the predictive one. Drafting a narrative from a specific worker's specific case record does not learn a cross-population pattern; it operates on one case at a time, grounded in that case's facts. The placement-disruption predictor, by contrast, learns from the population's history of disruptions, and that history reflects which families got which placements, which were monitored more closely, and which were already disadvantaged in ways the data silently encodes. The chatbot's bias is different again: it inherits the bias of whatever benefit-rules corpus it was trained on and can systematically misstate the rules for the programs that serve the most marginalized applicants. Each use case has a bias profile, and the second question forces it into the open before any deployment.

Question Three: What Is the Cost of Being Wrong?

The third question asks what happens to a real person when the use case is wrong, and how reversible that harm is. This is where the field's stakes make the analysis different from a private-sector one. A wrong product recommendation costs a sale. A wrong eligibility answer costs a family their food benefits. A wrong placement-disruption score directs a scarce caseworker away from a family that needed them. A wrong hotline urgency score leaves a child-abuse report sitting in a queue.

Reversibility is the crucial sub-question. A drafting error caught in verification is fully reversible: it never leaves the draft. A wrong chatbot answer that causes someone not to apply for SNAP (the Supplemental Nutrition Assistance Program, the federal food-assistance benefit) may be nearly irreversible, because the person never enters the system to be corrected. They simply go without, and no fair-hearing right is triggered because no determination was ever made. The most dangerous use cases are the ones where the harm is both severe and invisible, because invisibility removes the due-process mechanisms (notice, a fair hearing, the right to challenge) that exist precisely to catch and correct wrong decisions.

Green, Yellow, Red: A Working Triage

Running the three questions produces a triage the innovation manager could defend to her governance board, to an advocate, and to a court. She sorted use cases into three zones.

Green use cases inform rather than decide, operate on a single case's grounded record rather than a learned cross-population pattern, and produce errors that verification catches before they reach a person. The investigation-report narrative drafter is green. It is the documentation pattern the agency already runs safely: a worker drafts from the record, verifies every claim, and signs. Green does not mean unsupervised; it means the existing verification discipline is sufficient, and the use case can proceed under that discipline without new governance machinery.

Yellow use cases carry real liability that can be managed only with substantial new structure. The hotline-call urgency scorer is yellow. It is genuinely a screening tool, so it imports the screening pattern's equity risk, but its harm is potentially catchable if it is built as an audited input under mandatory human review, with a continuous equity audit, and never as a triage verdict that a stretched unit follows blindly. Yellow means: not now, and only after the equity-auditing and human-review machinery the screening pattern demands is fully standing. Yellow is the answer the placement-disruption predictor receives as well, with an even higher bar, because directing scarce intervention resources by score is precisely the use the history of harm warns against.

Red use cases cannot be made responsible in their proposed form because they remove the human from a consequential decision and place the harm beyond the reach of due process. A public eligibility chatbot that members of the public treat as a determination, with no caseworker present and no fair-hearing right triggered, is red as proposed. Red does not always mean the underlying goal is forbidden; it means the proposed form is, and the goal must be redesigned into a green or yellow shape before it can proceed. The eligibility chatbot can be redesigned: instead of answering "you do not qualify," it can be built to say "here is the program, here is how to apply, and a caseworker will determine your eligibility," which moves it from deciding to informing and restores the human and the due-process path.

Red is rarely a verdict on the goal. It is a verdict on the form. The discipline is to redesign the form until the human and due process are back in the loop.

Redesigning a Red Into a Green

The most valuable skill in identifying responsible new use cases is not rejection. Any cautious manager can say no. The valuable skill is taking a use case that is red as proposed and finding the redesign that captures the underlying benefit while restoring the human and the due-process perimeter. This is where innovation actually lives in this field, not in adopting the boldest tool but in finding the responsible shape of a real need.

Walk through the eligibility chatbot. The underlying need is real: the public has eligibility questions, the agency's phone lines are overwhelmed, and people give up before they apply. The red version answers eligibility questions directly, which means it can deny in effect without ever triggering the protections a denial requires. The redesign keeps the benefit and removes the liability. The chatbot is reframed as a navigation and intake aid: it explains what programs exist, what documents an application needs, and how to start one, and it explicitly does not determine eligibility, routing every actual determination to a caseworker. The person still gets faster help. The agency still relieves its phone lines. But the consequential decision stays with a human, the applicant enters the system rather than being turned away at the door, and the fair-hearing right attaches to a determination that a person actually makes. The benefit survived the redesign; only the liability was removed.

The same redesign logic applies to the placement-disruption predictor, though it lands in yellow rather than green. The red instinct is to let the score direct resources. The responsible redesign treats the model's output as one surfaced signal among many, presented to a supervisor and worker who review it against the actual case, with the equity audit running continuously to detect if the signal disproportionately flags families by race, neighborhood, or disability. Even redesigned, it stays yellow because the equity stakes of a population-trained predictor are high enough that it requires the full screening-support machinery before it can run at all. The redesign does not turn a hard use case into an easy one; it turns an indefensible one into a defensible one that still demands the highest level of oversight.

The Cost of Getting the Classification Wrong

It is worth being concrete about what happens when an agency misclassifies a use case, because the abstraction of green, yellow, and red hides real consequences for real people.

Suppose the innovation manager had waved the eligibility chatbot through as green, reasoning that it was just answering questions and a chatbot cannot make a determination. A single mother working two jobs visits the agency website at eleven at night, the only time she has, and asks whether she qualifies for SNAP. The chatbot, applying a gross-income test that does not account for her child's disability-related deductions, tells her she does not qualify. She never applies. There is no determination, so there is no notice, no fair hearing, and no right to challenge, because the due-process machinery only activates on a decision the agency formally makes. She and her child simply go without food assistance they were entitled to, and the agency never knows it happened. The harm is severe, it is invisible, and it is beyond the reach of every protection the agency built. That is the cost of classifying a red as a green.

Suppose instead the manager had treated the hotline urgency scorer as green and let the stretched screening unit follow its scores to clear the backlog of nearly six hundred calls a week. A call about a child reported by a neighbor gets scored low, perhaps because the model learned that calls from that ZIP code with that pattern of language historically screened out, and it sits in the queue while higher-scored calls are worked first. If that child was in danger, the cost of the misclassification is measured in the worst terms this field knows. The scorer was never a green use case. Treating it as one because it felt like just a triage helper, rather than recognizing it as a screening tool that demanded yellow's full machinery, is the kind of error the three questions exist to prevent.

The mirror-image error is also costly, though less catastrophic. An agency that classifies everything as red, that treats the safe documentation drafter as if it were the dangerous screening predictor, forfeits the genuine benefit AI offers and keeps its caseworkers buried in paperwork they did not need to do. Over-caution is not free in a field where documentation burden drives the burnout and turnover that raise caseloads and let more harm slip through. The discipline is not to fear every use case equally. It is to classify accurately, so that the green ones proceed and return their hours, the yellow ones proceed only with the structure that makes them safe, and the red ones are redesigned rather than either deployed recklessly or rejected reflexively.

How the Manager Answered the Four Proposals

By Friday the innovation manager had four defensible answers, each tied to the three questions and the triage, each one she could read aloud to her governance board and her agency's parent advocates without flinching.

The investigation-report narrative drafter was green and approved to proceed under the existing documentation verification discipline, because it informs rather than decides, operates on a single grounded case record, and produces errors that verification catches before they reach the record. It would return hours to adult-protective-services workers without importing a new liability, the cleanest kind of new use case.

The hotline-call urgency scorer was yellow and held, not because the backlog of six hundred calls a week was not real but because the proposed form was a screening tool, and screening enters only as an audited input under mandatory human review with a continuous equity audit, machinery the agency would have to stand up first. The manager did not reject the goal; she set the conditions under which it could responsibly proceed and put it in the queue behind that structure.

The placement-disruption predictor was yellow with the highest bar, for the same reason at a higher intensity: a population-trained predictor directing scarce resources is the exact pattern the field's history of harm warns against, defensible only under full equity-auditing and human-review structure, and only if the audit could show it did not flag families disproportionately by race, neighborhood, or disability.

The public eligibility chatbot was red as proposed and sent back for redesign, with the redesign specified: reframed as a navigation and intake aid that explains programs and routes every determination to a caseworker, never answering eligibility itself, so the consequential decision and the due-process right both stay with a human. In its red form it could deny food or shelter to a person beyond the reach of any protection; in its redesigned green form it could relieve the phone lines and speed people into the system where a human would decide. The benefit survived; the liability did not.

None of the four answers was a simple yes or no. Each was a classification with a reason, and the reason was the same discipline applied four times: separate the problem from the solution, ask whether it decides or informs, ask whose bias the data carries, ask the cost and reversibility of being wrong, and redesign the form until the human and due process are back in the loop. That discipline is what makes innovation responsible in a field where the cost of an irresponsible use case is not a bad quarter but a harmed child or a family that went without.

Key Takeaways

  • A real problem does not make a use case responsible. New AI ideas always arrive attached to genuine pain (a backlog, a burnout statistic, a family that fell through a gap), and the discipline is to separate the problem from the proposed solution and ask what liability the specific solution imports, no matter how real the problem.
  • Success at the safe use case is what makes the next decision dangerous. The properties that made documentation safe (it informs, it is grounded in one case, its errors are caught in verification) do not automatically transfer to anything else AI can do, and appetite born of success does not discriminate between use cases that share those properties and ones that only share the excitement.
  • Three classifying questions, asked in order, sort any use case: Does it decide or only inform? Whose data and whose historical bias does it carry? What is the cost of being wrong, and how reversible is the harm? A use case that fails the first question is disqualified before its benefits are weighed.
  • Reversibility is decisive. The most dangerous use cases produce harm that is both severe and invisible, like a chatbot that causes someone not to apply for a benefit, because invisibility removes the due-process mechanisms (notice, a fair hearing, the right to challenge) that exist to catch and correct wrong decisions.
  • The triage is green, yellow, red. Green use cases inform, are grounded, and have catchable errors, and proceed under existing verification discipline. Yellow use cases (screening patterns) carry real liability and proceed only after the equity-auditing and mandatory-human-review machinery is fully standing. Red use cases remove the human from a consequential decision and place harm beyond due process, and cannot proceed in their proposed form.
  • Red is a verdict on the form, not usually on the goal. The most valuable skill is redesigning a red use case (an eligibility chatbot that effectively denies) into a green one (a navigation aid that explains programs and routes every determination to a caseworker), capturing the benefit while restoring the human and the due-process perimeter.
  • Misclassification has real costs in both directions: classifying a red as green can leave a family without entitled food assistance with no fair-hearing right triggered, or leave a wrongly low-scored hotline call sitting while a child is in danger; classifying everything as red forfeits the genuine benefit and keeps workers buried in paperwork that drives burnout. The discipline is accurate classification, not uniform fear.
  • The output of the discipline is a defensible answer, not a yes or no. Each use case receives a classification with a reason that a governance board, an advocate, and a court would accept, applying the same discipline every time: separate problem from solution, ask decide-or-inform, ask whose bias, ask the cost of being wrong, and redesign the form until the human and due process are back in the loop.