Incident Response for AI-Related Case Problems
The call came in on a Thursday afternoon, and the agency had no plan for it. A defense attorney representing a mother in a dependency case had been preparing for a hearing and noticed something wrong in the caseworker's court report: it described a prior substance-use evaluation that the mother had never undergone. The attorney pulled the underlying case file. The evaluation was not there, because it had never happened. The report had been drafted with the agency's AI documentation tool, and the fabricated history had survived verification and supervisory review and entered the legal record. The attorney's next call would be to the judge. By the time the deputy director heard about it, two questions were already urgent and the agency could answer neither. First, was this one report or were there others, because the same tool and the same gap in verification might have produced fabricated history in dozens of other reports across the unit. Second, who needed to be told, and how fast, because a family was about to walk into a hearing built partly on a record that contained a false statement. An agency that has rehearsed an incident-response process answers both questions in hours. An agency that has not loses days it does not have, and a family pays for the delay. This lesson is about building the process before the Thursday call.
Why AI Incidents Need Their Own Response
Human-services agencies already have incident-response procedures for many things: a data breach, a child fatality review, a complaint of worker misconduct. An AI-related case problem is different enough from all of these that an agency cannot simply fold it into an existing checklist and assume it is covered. The differences are what make a dedicated process necessary.
The first difference is scale and speed of propagation. A single human error usually stays in a single case. An AI error can be systemic, because the same tool, the same prompt, the same gap in a verification step, and the same model behavior repeat across every case the tool touched. When a caseworker discovers that the documentation tool fabricated a piece of history in one court report, the agency cannot treat it as an isolated mistake. The tool may have done the same thing in a hundred reports, because the failure mode is structural to how the model works, not particular to one tired worker on one bad day. An AI incident is presumptively a population problem until the agency proves otherwise, and that presumption changes everything about how fast and how broadly the agency must respond.
The second difference is the due-process stakes. The records AI touches in this field are not internal memos. They are court reports that influence whether a child stays with a family, eligibility determinations that decide whether a person eats this month, safety assessments that drive removal decisions. When an AI error reaches one of these, it does not merely embarrass the agency; it potentially corrupts a decision that a court made or is about to make about a person's life, and it implicates that person's right to notice, to a fair hearing, and to challenge a determination on the basis of accurate information. An AI incident is therefore often a due-process incident, and the response has to be built to protect rights, not just to fix data.
The third difference is the question of disclosure to people outside the agency. A fabricated observation in a court record is a false statement that the court relied on and that the affected family had a right to know about and contest. Deciding whether, when, and how to disclose an AI error to a court, to an attorney, to an advocate, and to the affected family is one of the hardest and most consequential parts of the response, and it is one that has no parallel in routine internal incident handling. Transparency and disclosure are what keep AI-assisted work defensible; concealment is what turns a contained error into a scandal that costs the program its license to operate.
An AI error is a population problem until proven otherwise, it is usually a due-process problem and not just a data problem, and it almost always raises a disclosure obligation to people outside the agency. None of those is true of an ordinary mistake.
The Four Phases: Detect, Contain, Remediate, Disclose
A workable incident-response process for AI-related case problems moves through four phases. They overlap in practice and the order is not rigid, but naming them separately keeps the response from collapsing into panic, where an agency does the easy parts and skips the hard ones.
Detect
Detection is the phase most agencies fail at, because an AI error does not announce itself. A fabricated observation reads exactly like a real one, in the same confident professional tone, so the agency that waits for errors to surface on their own will mostly learn about them the way the Thursday agency did: from an opposing attorney, an advocate, or a journalist, which is the most damaging possible source. A real incident-response capability begins long before an incident, with detection channels the agency builds deliberately.
Those channels include a clear, blame-light reporting path that a caseworker who catches a hallucination in verification can use without fear, because a worker who fears punishment will quietly fix the one error and never report the pattern. They include supervisory review that is structured to surface AI errors specifically, not just general quality issues. They include periodic audits that sample AI-touched records and check them against the source, so the agency finds its own errors before an outsider does. And they include a way for the people served, and their advocates, to raise a concern that a record may be wrong, because the affected family is sometimes the first to know that a documented fact never happened. An agency that builds these channels turns detection from luck into a system.
Contain
Containment is the phase where speed matters most, because an AI error propagates. The instant an error is confirmed, the agency must answer the systemic question: is this one record or many. Because an AI failure mode is structural, the safe assumption is many, and containment means quickly identifying the population of records the same tool, prompt, and workflow gap could have affected, then freezing the relevant use until the scope is understood. If a documentation tool fabricated history because of a specific weakness in how it handles partial records, every report drafted from a partial record is suspect, and the agency may need to pause that workflow, flag those records, and prevent any of the affected reports from being filed or relied upon while the review runs.
Containment also means stopping the immediate harm in the triggering case. In the Thursday example, containment is not waiting for the full audit; it is immediately notifying the court and the attorney that the report contains an unverified statement that the agency is investigating, so the hearing does not proceed on a corrupted record. Containment protects the person in front of you while remediation protects everyone else the tool may have touched. An agency that does only the audit and forgets the family at Thursday's hearing has contained the systemic problem while letting the acute one cause its harm.
Remediate
Remediation has two layers, and an agency that does only the first will see the same incident again. The first layer is fixing the affected records and the affected decisions: correcting the false statements in the case record, reopening or re-reviewing any determination that relied on the corrupted information, and, where a consequential decision was made on a corrupted record, doing whatever the law and the agency's duty require to revisit that decision. A correction to the record is not enough if a child was removed or a benefit was denied on the basis of the error; the decision itself has to be revisited.
The second layer is fixing the cause so it does not recur. That means asking why the error survived: was the verification step skipped under caseload pressure, was the supervisory gate not actually checking for this failure mode, was the tool deployed into a workflow it was not safe for, were workers not trained to catch this specific kind of hallucination. The cause is usually not one careless worker; it is a systemic gap in the workflow, the training, the tool configuration, or the capacity workers had to verify. Remediation that ends with disciplining the worker who filed the report and changes nothing about the system guarantees the next incident, and it teaches every other worker that reporting an error is dangerous, which destroys the detection channel the agency depends on.
Disclose
Disclosure is the phase agencies most want to avoid and most need a pre-decided policy for, because the pressure to minimize and conceal is highest exactly when transparency matters most. The question is who has a right to know, and in this field the answer usually includes people outside the agency. When an AI error corrupted a court record, the court and the affected family's attorney have a right to know, because the family has a due-process right to challenge a determination made on accurate information. When an error caused a wrong benefits denial, the affected person has a right to know so they can exercise their fair-hearing rights. When an error is systemic across many cases, the agency's governance board, and depending on the jurisdiction its oversight bodies and the public, may have a right to know.
Deciding disclosure case by case, in the heat of an incident, under the influence of the people whose reputations are exposed, produces concealment. The defense against that is a disclosure policy decided in advance, in calm, that names who must be told for each category of incident and on what timeline, so that when the Thursday call comes the agency executes a pre-committed obligation rather than negotiating its own honesty under pressure. The agencies whose AI failures became scandals were rarely undone by the original error. They were undone by the concealment that followed.
A Worked Incident from the Thursday Call
Walk the Thursday incident through the four phases to see how a prepared agency would have handled it, and how much the preparation is worth.
Detect. A prepared agency would not have learned about the fabricated evaluation from the defense attorney. Its periodic audit of AI-touched court reports, sampling even ten percent of reports drafted from partial records, would likely have caught the pattern of fabricated history before this report reached a hearing, because the same structural weakness that produced this fabrication would have produced others the audit would sample. Even if this particular report slipped through to the attorney, a prepared agency would have an open, blame-light channel so that the moment the caseworker or supervisor learned of the attorney's finding, the incident was logged and escalated in minutes rather than circulating as an anxious rumor for days.
Contain. Within hours of confirming the fabrication, the prepared agency does two things at once. It notifies the court and the attorney that the report contains an unverified statement the agency is actively investigating and asks that the hearing not proceed on the uncorrected record, protecting the mother in front of them. And it identifies the population of at-risk records: every court report the same tool drafted from a partial record in the relevant period, freezing reliance on them and flagging them for review. The agency does not yet know how many are affected, so it treats all of them as suspect until the review proves otherwise.
Remediate. The agency corrects the false statement in the mother's record and, because the report was about to inform a consequential hearing, ensures the court has an accurate record before any decision is made. It then reviews the flagged population, correcting every report where the audit finds fabricated history and revisiting any decision already made on a corrupted one. In parallel it fixes the cause: it discovers that the verification step was being skipped on history sections under caseload pressure because workers had not been given time or a specific procedure to trace each historical claim to the case-management record, and it changes the workflow so that history verification is a required, staffed, supervised step rather than an aspiration. The fix is to the system, not to the one worker.
Disclose. Following its pre-decided policy, the agency discloses to the court and the affected attorney in the triggering case immediately, discloses to every other affected family and their representatives as the review identifies corrupted records, and reports the systemic incident to its governance board with the scope, the cause, and the remediation. It documents the entire response so that an oversight reviewer, an advocate, or the court can reconstruct exactly what happened and what the agency did about it. The disclosure is uncomfortable, but it is the thing that keeps the program defensible and keeps the agency's word worth trusting the next time it says a record is accurate.
The difference between the prepared agency and the Thursday agency is not the original error, which either could have suffered. It is the days saved, the families protected, the systemic scope caught early, and the trust preserved by disclosing rather than concealing. That difference is what the incident-response process buys, and it is only available to an agency that built the process before it needed it.
Building the Process Before You Need It
An incident-response process is worthless if it is invented during the incident, because the incident is exactly when judgment is worst and pressure is highest. The strategist's job is to build, rehearse, and resource the process in calm so that it executes under stress. Several elements make the difference between a plan on a shelf and a capability that works.
A named owner and a standing team. An AI incident touches legal, practice, the affected unit, data and IT, and often communications. Someone must own the response with the authority to freeze a tool's use, and the standing team must be identified in advance so it convenes in hours, not after a week of figuring out who is responsible. Diffuse ownership is how the first critical days are lost.
A severity classification decided in advance. Not every AI error is a five-alarm incident, and treating a caught-in-verification hallucination the same as a fabrication that reached a court will exhaust the team and dull its response. A simple tiering, by whether the error reached a consequential decision, whether it is systemic, and whether it touched a person's due-process rights, lets the agency match its response to the stakes and reserve the full mobilization for the incidents that warrant it.
A pre-committed disclosure policy. As established, disclosure decided under pressure becomes concealment. The policy that names who must be told, for which tier, on what timeline, must exist before the incident, and it must be owned at a level high enough that no individual whose reputation is exposed can quietly override it.
Rehearsal. A tabletop exercise that walks the team through a realistic scenario, the Thursday call, a wrong benefits denial discovered by an advocate, a biased screening score surfaced by an equity audit, reveals the gaps in the plan while they are cheap to fix. An agency that has rehearsed the response moves with practiced calm when the real call comes; an agency that has only written the plan discovers its holes in public.
Integration with governance and the audit trail. The incident-response process does not stand alone. It draws on the audit trail that logs every AI-touched record and every verification, because that trail is what makes detection and scope-identification possible, and it reports into the governance board that owns the agency's AI program, because the board needs the pattern of incidents to decide whether a tool should keep operating at all. An agency whose AI work is logged and governed can run this process; an agency whose AI use is untracked cannot even determine the scope of its own incident.
Key Takeaways
- AI-related case problems need their own incident-response process because they differ from ordinary errors in three ways: they propagate systemically (the same tool and workflow gap repeat across every case), they usually implicate due process (the records AI touches drive court, eligibility, and removal decisions), and they almost always raise a disclosure obligation to people outside the agency.
- Treat an AI error as a population problem until proven otherwise. Because the failure mode is structural to how the model works, one discovered error means the same tool may have produced the same error across many cases, which changes how fast and how broadly the agency must respond.
- The response moves through four phases. Detect: build deliberate channels (blame-light reporting, AI-specific supervisory review, periodic audits, a path for families and advocates) so the agency finds its own errors before an outsider does. Contain: freeze the affected use, identify the at-risk population, and stop the acute harm in the triggering case. Remediate: fix the records and revisit any decisions made on them, then fix the systemic cause. Disclose: tell the people who have a right to know.
- Remediation has two layers, and fixing only the records guarantees recurrence. The second layer asks why the error survived, which is almost always a systemic gap in workflow, training, tool configuration, or verification capacity, not one careless worker. Disciplining the worker and changing nothing repeats the incident and destroys the detection channel.
- Disclosure must be governed by a policy decided in advance, in calm, because disclosure decided under pressure becomes concealment. In this field the people with a right to know usually include the court, the affected family and their attorney, advocates, and the governance board. The original error rarely ends a program; the concealment that follows does.
- Containment protects the person in front of you (notify the court so the hearing does not proceed on a corrupted record) while remediation protects everyone else the tool may have touched (audit and correct the flagged population). A response that does only one leaves real harm in place.
- Build the process before you need it: a named owner with authority to freeze a tool, a standing cross-functional team, a severity classification matched to stakes, a pre-committed disclosure policy, and rehearsal through tabletop exercises that find the gaps while they are cheap.
- Incident response depends on governance and the audit trail. The trail that logs every AI-touched record and verification is what makes detection and scope-identification possible; an agency whose AI use is untracked cannot even determine the scope of its own incident.
Skill.re