When Not to Use AI
The supervisor kept a single laminated card taped inside the cabinet above her desk, and new caseworkers always asked about it. It was not a list of how to use the agency's AI tools. It was a list of where not to. At the top, in her own handwriting, was one line: "Do not ask a model to decide whether a child is safe." Below it were nine more, each one a place where, after watching the unit for two years, she had decided the tool must not touch the work at all. She had written the card the week after a worker, exhausted at the end of a long day, had pasted a child's full disclosure of abuse into a public chatbot to "help organize it into a report." Nothing about that worker was careless in the ordinary sense. They were drowning, the tool was right there, and no one had ever told them, in advance and in writing, that this was a line. The card was her answer: a kill-criteria list, named before the pressure arrived, of the decisions and the moments where AI must stay out. This lesson is about why that list exists, what belongs on it, and why naming the no in advance is one of the most protective things a human-services professional can do.
Why Naming the No in Advance Matters
Most AI guidance in this field is about how to use the tools well: how to ground a draft in the record, how to verify a claim, how to read a risk signal as one input. That guidance is essential, and the rest of this program is built on it. But there is a category of decisions and moments where the right answer is not "use the tool carefully." It is "do not use the tool here at all." Naming that category in advance, in writing, is a discipline in its own right, and it is different from verification.
The reason it has to be done in advance is human, not technical. In the moment, under a caseload of thirty families and a court report due tomorrow, the tool is right there, it is fast, and the temptation to reach for it is strongest exactly where the stakes are highest, because that is where the work is hardest. A worker deciding at 9 PM whether to paste a child's disclosure into a chatbot is not in a good position to reason carefully about data privacy, retention, and the dignity of the child. The decision has to have been made already, by the agency, in calm conditions, and communicated as a bright line: not here. A kill-criteria list moves the decision out of the exhausted moment and into the policy, where it belongs.
The second reason is that a bright line is enforceable and a judgment call is not. "Use AI carefully on sensitive cases" gives a tired worker nothing to hold onto and an agency nothing to audit. "Never paste a client's unredacted disclosure into a tool the agency has not approved" is a rule a worker can follow, a supervisor can check, and an oversight body can verify. The kill-criteria list converts the soft, erodible guidance of "be careful" into hard lines that survive contact with a hard day.
The time to decide where AI must not go is not at 9 PM under a deadline. It is in advance, in writing, by the agency, before the pressure arrives.
The Decisions That Must Stay Human
The first and largest category on any kill-criteria list is the set of consequential decisions an AI must never make. This is the cardinal rule of the field, AI informs and humans decide, expressed as a list of specific calls. The point of writing them out is that "humans decide" can quietly erode into "the human approved what the tool said," and naming the specific decisions keeps the boundary visible.
The decision to remove a child from a home must stay human. This is among the most consequential acts a government performs, bound by due process and reviewed by a court, and the judgment it requires, weighing safety, family integrity, and the harm of removal itself, is not a judgment a next-token prediction system can make. An AI tool may help organize the documentation that informs the decision, but the call to remove belongs to the caseworker, the supervisor, and the court, and the tool must never be the thing that decides it.
The decision to substantiate a report of abuse or neglect must stay human. Substantiation is a formal finding that abuse or neglect occurred, and it carries lasting consequences for the family, including entry into a registry in many jurisdictions. The standard is legal and the evidence is human: interviews, observations, records, professional judgment. A model that "scores" a case or suggests a finding is on the wrong side of the line if its output becomes the substantiation rather than one input a human weighs and owns.
The decision to deny, reduce, or terminate a benefit must stay human, for the reasons due process demands: the person is owed a reason, a notice, a hearing, and an accountable human. A determination that is in substance the tool's verdict, accepted without independent human judgment, fails due process even if it happens to be correct. The eligibility worker owns the determination.
Decisions about a child's permanency and placement, whether to reunify, whether to move toward adoption, where a child lives, must stay human. These decisions shape a child's entire life, they are reviewed by courts, and they require weighing relationships and wellbeing that no model can assess. The tool may help assemble the record; it must not be the thing that decides where a child belongs.
Across all of these, the test is the same: if the model's output would become the decision rather than inform a decision a human makes and can defend, the tool is being used where it must not be. The kill-criteria list names these decisions so that no one, under pressure, mistakes "the tool said so" for a sufficient reason.
It helps to see how the line holds in a single case. A worker is preparing for a safety decision on a family where a report has come in. The agency's documentation tool can organize her notes, draft a chronology of contacts, and pull the relevant policy language into one place, and using it there saves her perhaps two hours she would otherwise spend assembling the file by hand. All of that is on the right side of the line: the tool is informing the decision by organizing what the worker already knows and has gathered. The line is crossed the moment the question shifts from "help me assemble the record" to "tell me whether this child is safe." The first is documentation support. The second asks a next-token prediction system to make the call that the worker, the supervisor, and the court are accountable for, weighing a child's safety against the serious harm of removal itself. The same tool, the same case, the same afternoon: helpful on one side of the line, forbidden on the other. The kill-criteria list is what makes that boundary legible to a tired worker who would otherwise see only a single helpful tool.
The Moments and Data That Must Stay Out
The second category is not about decisions at all. It is about specific moments, settings, and kinds of data where AI must stay out even when no consequential decision is being made, because the act of involving the tool itself causes harm or risk. The worker pasting a disclosure into a public chatbot was not making a removal decision. They were creating a data and dignity harm by the act of using the tool.
Sensitive Data in Unapproved Tools
The clearest line is this: a client's sensitive information must never be entered into an AI tool the agency has not approved and contracted for. The people in the human-services system carry the most sensitive data the government holds: disclosures of abuse, medical and behavioral-health histories, immigration status, substance use, addresses of people fleeing violence. A public chatbot is not a secure, contracted environment. Pasting that data into it can mean the information is retained, used to train a model, exposed in a breach, or simply placed outside the agency's control and outside the protections the law requires for it. This is true even when the worker's intent is entirely benign, even when they are only trying to organize a note. The harm is in the transfer of the data, not in the intent. A kill-criteria list names the approved tools and makes everything else off limits for client data.
The Direct Crisis and Therapeutic Moment
AI must not stand in for the human presence in a crisis or a clinical-relational moment. A person in acute crisis, a child making a first disclosure, a family in the worst hour of their life, needs a human being, not a generated response. Using a tool to produce what should be a worker's own compassionate, present engagement crosses a line that is about the dignity and safety of the person, not about documentation accuracy. This does not forbid using AI to draft a routine letter or organize notes after the fact; it forbids letting the tool occupy the relational space that is the actual help the person needs. The moment of human connection is the mission, and it is precisely the thing the field is trying to protect by taking documentation off the worker's plate. Spending the returned time letting a tool mediate the human moment would invert the entire purpose.
Novel, High-Stakes, or Unverifiable Situations
AI must stay out where its output cannot be verified to the standard the situation requires, and where the stakes make an unverified output dangerous. If a worker cannot independently check what the tool produced, against the record, the policy, or an authoritative source, then in a high-stakes context the tool's output cannot be safely used, because the verification step that makes AI use defensible is unavailable. A novel legal question with no clear policy source, a clinical judgment outside the worker's competence, a situation where the facts cannot be confirmed: these are places where a confident, plausible, unverifiable answer is more dangerous than no answer, because it invites reliance the situation cannot support.
The reason this belongs on the kill-criteria list, rather than being left to the verification discipline taught elsewhere in the program, is that verification assumes a source exists to verify against. The rest of this program teaches workers to trace every AI-touched claim to the field notes, the policy manual, or the case record. That discipline is powerful precisely where those sources exist. But there are situations where they do not: a question the policy manual does not address, a clinical determination no record can confirm, a fact pattern with no authoritative source to check. In those situations the verification step is not just skipped, it is impossible, and a confident model answer fills the vacuum with something that looks authoritative and is not. The kill-criteria list names these as off limits so a worker does not mistake the absence of a verification source for permission to rely on the tool. When the situation is high-stakes and the answer cannot be verified, the safe move is to escalate to a human who can, a supervisor, agency counsel, a clinician, not to accept what the model offered.
How to Write an Agency's Kill-Criteria List
A kill-criteria list is only protective if it is concrete, owned, and reachable in the moment. A vague list is no better than no list. Writing a good one takes a few deliberate properties.
It names specific decisions and specific moments, not categories. "Do not use AI for sensitive decisions" is unenforceable because every worker draws the line of "sensitive" differently. "Do not let an AI tool's output be the basis for a removal, a substantiation, a benefit denial, or a placement decision" is a line a worker can apply. "Do not enter a client's unredacted disclosure into any tool not on the approved list" is a line a worker can follow at 9 PM. Specificity is what makes the list survive the moment.
It is owned by the agency, not the individual worker. The whole point is to move the decision out of the exhausted worker's hands and into policy made in calm conditions. The list should be written, adopted, and communicated by the agency, with legal, practice, equity, and frontline voices at the table, so that a worker following it is backed by the agency rather than improvising alone. A worker who declines to use a tool because the agency's list forbids it should never be second-guessed for slowing down; they are doing exactly what the policy requires.
It is reachable in the moment the temptation arrives. A kill-criteria list buried in a policy binder no one opens is not a control. The list has to live where the work lives: at the point of use, in training that workers actually remember, on the supervisor's cabinet door if that is what it takes. The laminated card in the opening story worked because it was visible at the moment a worker might reach for a tool, not filed away where it could not help.
It is paired with a clear escalation path. A no is more followable when it comes with a yes-instead. "Do not paste this into a chatbot" should be paired with "here is the approved tool, here is who to ask, here is the supervisor to escalate to." A kill-criteria list that only forbids, without offering the supported alternative, pushes a desperate worker toward the workaround. The list should close the door and open the right one in the same breath.
It is reviewed and updated. The tools change, the data flows change, and new failure modes appear. A list written once and never revisited will miss the next pasted-disclosure incident waiting to happen. The list is a living document, reviewed on a schedule and after any incident, so it grows with the work rather than ossifying.
The Worker Under Pressure Is the Real Test
It is worth returning to the worker in the opening story, because the kill-criteria list exists for them, not for the careful worker on a calm day. The careful worker on a calm day does not need a list; they would have caught the problem themselves. The list exists for the worker at the end of a brutal day, behind on documentation, carrying a caseload that the agency knows is too high, with a tool sitting right there that promises to make the next hour easier. That worker is not the exception. Under the caseload and burnout conditions that define the field, that worker is the norm, and any AI policy that assumes the careful-worker-on-a-calm-day is a policy that will fail exactly when it matters.
This is why the kill-criteria list is framed as protection rather than restriction. It protects the family whose disclosure stays inside the agency's secure systems. It protects the child whose safety decision stays with the people and the court accountable for it. It protects the person whose benefit determination carries a real reason and a real human behind it. And it protects the worker, by taking the hardest judgment calls off their exhausted shoulders and putting them where they belong, in policy decided in advance. A worker who can point to the list and say "I did not use the tool there because the agency told me not to" is a worker the agency has set up to do the right thing under pressure rather than left to improvise it alone.
Consider the arithmetic of the alternative. An agency that deploys AI tools widely, returns hours to its workers, and never writes the kill-criteria list has built a system where the highest-stakes misuse, the pasted disclosure, the tool-driven removal recommendation, the chatbot standing in for a crisis response, is most likely to happen at the moment of greatest pressure, by a worker who was never told it was a line. The hours returned by the documentation tools are real and valuable, but they do not offset a single family harmed because no one named the no. The kill-criteria list is what lets an agency capture the genuine benefit of AI while holding the lines that protect the people the agency exists to serve. It is the discipline that says: here is everywhere the tool helps, and here, named in advance and in writing, is everywhere it must not go.
Key Takeaways
- A kill-criteria list is a written, agency-owned list of the decisions and moments where AI must not touch the work at all. It is a distinct discipline from verification: verification governs how to use the tool well, while the kill-criteria list governs where not to use it.
- The no must be named in advance, in calm conditions, because the temptation to reach for the tool is strongest exactly where the stakes are highest and the worker is most exhausted. A bright line decided in advance is enforceable and auditable; "be careful" is neither.
- The decisions that must stay human include removing a child, substantiating a report, denying or terminating a benefit, and deciding a child's permanency and placement. The test is whether the model's output would become the decision rather than inform a decision a human makes and can defend.
- Some moments and data must stay out regardless of any decision: a client's sensitive information must never go into a tool the agency has not approved, AI must not stand in for human presence in a crisis or first-disclosure moment, and AI must stay out where its output cannot be verified to the standard the situation requires.
- The harm of pasting a disclosure into a public chatbot is in the transfer of the data, not in the worker's intent. Benign intent does not make the line safe to cross, which is why the line must be named for the data and the tool, not left to judgment.
- A good kill-criteria list names specific decisions and moments rather than vague categories, is owned by the agency rather than the individual worker, is reachable at the point of use, is paired with a supported alternative and escalation path, and is reviewed and updated as tools and data flows change.
- The list exists for the worker under pressure, not the careful worker on a calm day. Any AI policy that assumes the calm-day worker will fail exactly when the caseload and burnout conditions of the field make misuse most likely.
- The kill-criteria list is protection, not restriction. It protects the family, the child, and the person served, and it protects the worker by moving the hardest judgment calls off their exhausted shoulders and into policy decided in advance.
Skill.re