Incident Response: When the AI Got It Wrong
At 8:15 on a Wednesday morning, Jordan's clinical director forwards an email with the subject line "Did you see this?" An associate signed a note last Thursday in which the AI scribe rendered a Columbia Suicide Severity Rating Scale result the clinician never administered: a hallucinated CSSRS score, sitting in the legal record, in a chart that a custody evaluator subpoenaed Monday. The practice has a risk register that scored exactly this scenario. What it does not have is the answer to the only question that matters now: what do we do in the next four hours? This lesson builds that answer, a 5-step incident response runbook (detect, contain, document, notify, remediate) tuned to the three failures behavioral health AI actually produces: the hallucinated risk score, the 42 CFR Part 2 leak, and the wrong-client note. It covers the HIPAA breach-notification 60-day clock, the OCR notification thresholds, and the FTC Health Breach Notification Rule for non-HIPAA digital health components. By the end you will have a Five-Step AI Incident Response Runbook with Notification Clocks, the document you write on a calm Tuesday so you never have to improvise on a bad Wednesday.
The Code Blue Principle: You Do Not Design Resuscitation During the Arrest
Every hospital floor has a crash cart, and every nurse on that floor knows where it is, what is in each drawer, and who calls the code. None of that knowledge was acquired during a cardiac arrest. It was drilled in advance, precisely because the moment of crisis is the worst possible moment to make design decisions. Adrenaline narrows thinking, hierarchy gets fuzzy, and the clock punishes every minute of deliberation. That is the controlling analogy for this lesson: an AI incident in a behavioral health practice is a code, and the runbook is your crash cart. The practice that writes the runbook on a calm Tuesday executes five rehearsed steps on the bad Wednesday. The practice without one holds a panicked meeting, makes the containment decision late, makes the notification decision wrong, and discovers afterward that the improvised choices are now part of the record a regulator will read.
The runbook earns its keep in behavioral health specifically because our incidents are not generic IT events. A leaked spreadsheet of orthopedic appointment times is a breach; a leaked therapy transcript naming a client's affair, their substance use, and their suicidal ideation is a different category of harm, and if the record carries 42 CFR Part 2 protection, it is also a different category of federal exposure. The three scenarios this lesson drills, the hallucinated CSSRS score, the Part 2 leak, the wrong-client note, were chosen because they map to the three distinct harms an AI incident can cause: harm to clinical decision-making, harm to confidentiality with heightened statutory protection, and harm to record integrity. Each scenario runs the same five steps; what changes is the content of each step.
Step 1, Detect: Build the Tripwires Before You Need Them
Incidents do not announce themselves; they are noticed, and the question is by whom and how late. Detection in a behavioral health AI program rests on four tripwires. First, the clinician's own verification pass: the read-every-word review before signing is not only a prevention control, it is the single most productive detection mechanism the practice has, because the clinician who catches a hallucinated assessment before signing has converted an incident into a near miss. Second, the spot audit: the clinical director's quarterly review of signed notes per clinician catches what signing missed. Third, vendor notifications: the BAA must obligate the AI vendor to report security incidents to the practice on a defined timeline, because under the Breach Notification Rule a business associate must notify the covered entity without unreasonable delay and within 60 days of discovering a breach, and the practice's own clocks start running on discovery. Fourth, the human channel: a no-blame reporting culture in which the associate who notices a wrong-client paragraph at 9 PM reports it at 9:05 rather than quietly deleting it and hoping.
That last tripwire deserves a sentence in the runbook itself: anyone who reports an AI incident in good faith will not be disciplined for the reporting. The reason is arithmetic. A wrong-client note caught in two hours is an amendment and an apology; the same note discovered eight months later by the wrong client's attorney is a breach with an aged clock, an integrity problem in two charts, and a credibility problem for everyone who touched it. The practice's exposure scales with time-to-detection more than with the underlying error, which means the culture that surfaces errors fast is itself a risk control. Log every report, including the near misses, in an incident log with date, reporter, tool involved, and disposition. Near misses are free intelligence; they are the same failure modes as incidents, delivered without the harm.
Step 2, Contain: Stop the Spread Without Destroying the Evidence
Containment answers one question: is the failure still propagating? For the hallucinated CSSRS score, containment means freezing reliance on the contaminated record: flag the note immediately, notify every clinician treating that client that the assessment result in the record was never administered, and check whether any downstream decision, a level-of-care change, a safety plan, a referral, relied on the phantom score. For the Part 2 leak, containment means severing the flow: suspend the implicated tool or integration, confirm with the vendor what was disclosed, to whom, and whether it can be clawed back, and segregate any further Part 2 content away from the implicated pipeline. For the wrong-client note, containment means correcting both charts through proper amendment procedures, never deletion: the erroneous content is removed from the wrong chart by amendment with an audit trail, and the right client's chart receives the note it was missing.
The non-negotiable rule of containment: preserve evidence. The panicked instinct is to delete the bad note, purge the transcript, and pretend the system never produced the error. Resist it completely. Spoliation converts a defensible incident into an indefensible cover-up, and EHR audit trails make deletion both discoverable and damning. Amend, flag, version, and quarantine; never erase. The second rule: contain at the system level, not just the instance level. If the scribe hallucinated one assessment score, the same model is drafting notes for every clinician in the practice today; the containment decision includes whether to suspend the tool practice-wide pending the vendor's explanation, and the runbook should name who has authority to make that call within the hour, without convening a committee.
Step 3, Document: The Incident File You Will Hand to Someone
Assume from the first minute that the incident file will eventually be read by someone adversarial: an OCR investigator, a board reviewer, a plaintiff's attorney, a carrier's claims adjuster. Document accordingly, factually, contemporaneously, and without speculation about fault. The incident file contains: a timeline (when the error occurred, when it was detected, by whom, every action with a timestamp), the scope (which clients, which records, which data elements, whether any record carries Part 2 protection), the tool and version involved, the containment actions taken, the clinical impact assessment, and the notification analysis that step four will produce. For the hallucinated CSSRS scenario, the clinical impact assessment is the heart of the file: the treating clinician re-performs and documents the actual risk assessment, the clinician's own determination, because AI never scores the CSSRS and never assigns a risk level, and the corrected record must show the licensed human's judgment replacing the fabricated number.
Two documentation disciplines matter here. First, separate facts from conclusions: "the note dated March 12 contained a CSSRS result; no CSSRS was administered on March 12" is a fact; "the vendor's model is dangerous" is a conclusion that does not belong in the file. Second, loop the incident back into the risk register from the previous lesson: an incident is an automatic re-score trigger, and the file should record which register category fired, what the pre-incident score was, and what the post-incident re-score concludes. This is what mature governance looks like from the outside: the incident file and the risk register cite each other, and a reviewer can trace the practice's reasoning from anticipation through event through correction.
The breach is rarely what ends a practice. The improvised, undocumented, late response is.
Step 4, Notify: The Clocks, the Thresholds, and the Regulators
Notification is where incident response becomes law rather than management, and the clocks are unforgiving. Under the HIPAA Breach Notification Rule, when a breach of unsecured PHI occurs, the covered entity must notify affected individuals without unreasonable delay and in no case later than 60 days after discovery. The 60-day clock is a ceiling, not a target; "without unreasonable delay" is the operative standard, and a practice that sits on a known breach for seven weeks invites exactly the scrutiny it hoped to avoid. OCR notification runs on thresholds: breaches affecting 500 or more individuals must be reported to OCR contemporaneously with individual notice and trigger media notification obligations, while breaches affecting fewer than 500 individuals are logged and reported to OCR annually. A solo practice's wrong-client note is almost always an under-500 event; a vendor-side leak across a tool's whole customer base can put a 25-clinician practice over the line faster than anyone expects, which is why the scope analysis in step three feeds directly into this step.
Behavioral health adds two layers. If the leaked records carry 42 CFR Part 2 protection, the practice faces the Part 2 regime's own consequences on top of HIPAA, and the notification analysis must address the redisclosure exposure separately; the 2024 Part 2 final rule's alignment with HIPAA breach notification makes the HIPAA clock the operational backbone, but the incident file must explicitly analyze the Part 2 status of every affected record. And if the failing component sits outside HIPAA entirely, a wellness app, a client-facing digital tool, a direct-to-consumer service with no covered-entity relationship, the FTC Health Breach Notification Rule governs: it requires vendors of personal health records and related entities not covered by HIPAA to notify individuals and the FTC when identifiable health data is breached, and the FTC's 2024 update made clear it reaches modern health apps. The runbook's notification section is therefore a decision tree: HIPAA-covered PHI runs the 60-day individual-notice clock with OCR thresholds; Part 2 records add the Part 2 analysis; non-HIPAA digital components run the FTC rule. One more notification belongs on the list and is habitually forgotten: the malpractice carrier. Carriers expect prompt notice of incidents that could mature into claims, and late notice can jeopardize coverage exactly when the practice needs it most.
Who decides whether an event is a reportable breach at all? Not the clinician who found it, and not by gut. The runbook names a breach-decision owner (in a group practice, the compliance officer with counsel on call; in a solo practice, the clinician with a health care attorney's number already in the runbook) who performs the breach risk assessment, documents it even when the conclusion is "not a reportable breach," and starts the clocks when it is. An unreported breach is a gamble; an undocumented non-breach determination is nearly as bad, because it looks like the question was never asked.
Step 5, Remediate: Fix the System, Not Just the Note
Remediation has three rings. The innermost ring is the client: the corrected record, the clinician's re-performed assessment where clinical content was contaminated, and where notification occurred, the conversation that explains what happened, what was done, and what changes, handled with the same clinical care as any rupture-and-repair work. The middle ring is the workflow: what control would have caught this earlier, and is it now installed? A hallucinated assessment score argues for a hard rule that no instrument result enters a note unless the clinician administered the instrument, verified at signing. A wrong-client note argues for client-identity verification at the top of every draft review. A Part 2 leak argues for tool segmentation so SUD records never transit a vendor whose BAA does not explicitly cover Part 2 data. The outermost ring is the vendor and the register: a root-cause explanation demanded from the vendor in writing, the contract or the tool reconsidered if the answer is unsatisfying, and the risk register re-scored with the incident logged as likelihood evidence.
Close every incident with a written debrief, scheduled within two weeks, no longer than a page: what happened, what worked in the response, what was slow, what changes in the runbook itself. The runbook is a living document on the same cadence logic as the register; an incident that does not update the runbook was only half-processed. And resist the most human remediation error in supervisory practice: making the incident about the individual. If the associate who signed the hallucinated score was following the workflow the practice trained, the failure is the workflow's. Discipline-first responses teach clinicians to hide the next error, which dismantles the detection tripwire you built in step one. The fix that holds is almost always structural.
Running the Three Scenarios Through the Five Steps
Drill the hallucinated CSSRS score: detected at the spot audit or by subpoena-driven review; contained by flagging the note, alerting treating clinicians, and checking downstream reliance; documented with the timeline and the clinician's re-performed risk assessment; notification analysis usually concludes no PHI breach occurred (the data went nowhere) but the record-integrity correction is documented and, where the chart is under subpoena, counsel is engaged immediately; remediation installs the no-instrument-without-administration rule and interrogates why the model fabricated a score. Drill the Part 2 leak: detected by vendor notice or client report; contained by suspending the pipeline and segregating SUD content; documented with explicit Part 2 status analysis per record; notified on the HIPAA clock with the Part 2 exposure analyzed separately and the carrier informed; remediated with tool segmentation and a BAA that names Part 2 data or a vendor that exits the SUD workflow entirely.
Drill the wrong-client note: Maria's version of this, solo practice, 9:54 PM, is detecting that her scribe merged content from her 4 PM and 5 PM sessions into one draft. Caught before signing, it is a near miss for the log. Caught after signing, it is both an integrity incident (two charts amended, audit trail preserved) and a potential breach (client A's session content resided in client B's record; the breach-decision owner assesses whether client B or anyone else accessed it, documents the determination, and starts the individual-notice clock if the answer supports it). The under-500 OCR pathway means annual reporting rather than immediate, but the individual-notice obligation runs on its own 60-day ceiling. The remediation is the identity-verification habit: client name, session date, and one fact only the right session contains, checked at the top of every draft before any other reading begins.
The Applied Problem: Build Your Five-Step AI Incident Response Runbook with Notification Clocks
Your artifact is the Five-Step AI Incident Response Runbook with Notification Clocks, and it should fit on four pages, because a runbook nobody can execute under stress is a binder, not a crash cart. Page one: the five steps as a one-glance flowchart (detect, contain, document, notify, remediate), the incident-log location, the no-blame reporting sentence, and the three names that matter with phone numbers: the containment authority who can suspend a tool within the hour, the breach-decision owner, and the health care attorney. Page two: containment playbooks for the three drilled scenarios, hallucinated risk score, Part 2 leak, wrong-client note, each as five or six imperative sentences, including the evidence-preservation rule: amend and quarantine, never delete.
Page three: the notification decision tree with the clocks stated plainly. Is the affected data HIPAA-covered PHI? Then individual notice without unreasonable delay, 60 days at the outside, from discovery; 500 or more individuals means contemporaneous OCR notice and media obligations; under 500 means the annual OCR log. Does any record carry 42 CFR Part 2 protection? Then the Part 2 analysis is documented separately. Is the component outside HIPAA, a consumer-facing app or non-covered digital tool? Then the FTC Health Breach Notification Rule path. In every branch: notify the malpractice carrier promptly, and document the breach-or-not determination even when the answer is no. Page four: the documentation template (timeline, scope, tool, containment, clinical impact, notification analysis) and the two-week debrief form that loops findings into the risk register and back into this runbook.
You may use AI to format the runbook from your notes; here is the prompt: "Lay out the following incident-response content as a four-page runbook with a one-page flowchart, three scenario playbooks, a notification decision tree, and a documentation template. Do not add legal requirements, deadlines, or thresholds I have not provided, and flag any step that lacks a named owner." Then run the verification pass: every step has an owner with a phone number; the 60-day clock, the 500-individual OCR threshold, and the FTC rule branch appear on the decision tree; the CSSRS playbook states that the clinician re-performs the risk assessment because AI never scores risk instruments; the evidence-preservation rule appears in every containment playbook. "Done" is a dated runbook, acknowledged by every clinician, drilled once as a thirty-minute tabletop exercise using one of the three scenarios, and filed beside the risk register it serves.
Key Takeaways
- An AI incident is a code blue: you do not design resuscitation during the arrest. The five-step runbook (detect, contain, document, notify, remediate) is written on a calm Tuesday, names its owners with phone numbers, and is drilled before it is needed, because improvised responses become part of the record a regulator reads.
- Detection is the highest-leverage step, and exposure scales with time-to-detection more than with the underlying error. The four tripwires are the clinician's read-every-word pass before signing, the quarterly spot audit, BAA-obligated vendor incident reporting, and a written no-blame reporting culture that logs near misses as free intelligence.
- Containment stops propagation without destroying evidence: flag and amend, never delete, because spoliation converts a defensible incident into an indefensible cover-up and EHR audit trails make deletion discoverable. Contain at the system level too; the model that hallucinated one score is drafting everyone's notes today.
- The hallucinated CSSRS scenario carries the chapter's hard rule: AI never scores the CSSRS, never assigns a risk level, and the remediation always includes the clinician re-performing and documenting the actual risk assessment, the licensed human judgment replacing the fabricated number in the record.
- The notification clocks are law: individual notice of a breach of unsecured PHI without unreasonable delay and no later than 60 days from discovery; 500 or more affected individuals triggers contemporaneous OCR notice and media obligations, under 500 goes on the annual OCR log; 42 CFR Part 2 records get a separately documented analysis; non-HIPAA digital health components run the FTC Health Breach Notification Rule; and the malpractice carrier gets prompt notice in every branch.
- A named breach-decision owner performs and documents the breach risk assessment even when the conclusion is "not a breach," because an undocumented non-breach determination looks like the question was never asked.
- Remediation fixes the system, not the person: install the structural control (no instrument result without administration, identity verification atop every draft, Part 2 tool segmentation), demand the vendor's written root cause, re-score the risk register with the incident as likelihood evidence, and debrief within two weeks so the runbook itself learns.
Skill.re