Behavioral Health AI Risk Assessment
Jordan's compliance officer asked a question in Friday's leadership meeting that stopped the room: "If our AI scribe leaked a client file tomorrow, who would find out first, us or the client?" Nobody knew. Twelve of the practice's twenty-five clinicians use AI tools today, and still nobody could answer the most basic risk question a regulator, a malpractice carrier, or a plaintiff's attorney would ask. This lesson builds the answer: a formal behavioral health AI risk assessment, scored across six categories on a likelihood times severity grid, with a written mitigation plan for the top five risks. This is the grown-up risk function. Many practices skip it, and the ones that skip it do not survive their first complaint, because the complaint arrives at a practice that cannot show it ever looked. By the end you will have a Six-Category AI Risk Register with Top-Five Mitigation Plan, the document that turns "we use AI carefully" from a feeling into evidence.
Why a Risk Assessment Is the Grown-Up Move, Not Paperwork Theater
Think of the risk assessment the way you think of a fire marshal's inspection of a building. The building does not become fireproof because the marshal walked through it. The inspection does three things instead: it finds the hazards that actually matter before they ignite, it documents that a competent person looked, and it creates a dated record that the owner acted on what was found. When a fire eventually happens, and in any building occupied long enough something eventually happens, the difference between negligence and misfortune is the inspection trail. That is the controlling analogy for this lesson. Your practice is the building. The AI tools your clinicians use are new wiring run through old walls. The risk assessment is the marshal's walk-through, and you are the marshal, because nobody else is coming to do it for you.
Practices resist this work because it feels like inviting trouble to write down everything that could go wrong. The legal reality is the opposite. When OCR investigates a breach, when a state board (the California BBS, New York's OPP, the Texas BHEC, Florida's DOH for Chapter 491 professions, the Illinois DPR) processes a complaint, when a malpractice carrier evaluates a claim, the first document request is some version of "show us your risk analysis." HIPAA's Security Rule has required a risk analysis for two decades; the AI risk assessment extends the same discipline to a new class of vendor and failure. A practice with the document took the duty seriously and got unlucky. A practice without it never looked, and regulators treat the two very differently.
There is also a managerial reason. Without a scored register, AI risk discussions run on anxiety and anecdote: whoever read the scariest headline this week sets the agenda. With a register, the conversation becomes boring in the best way: the breach risk scores sixteen, the biased-output risk scores six, so the next dollar of mitigation goes to the breach exposure, and everyone can see why. Scoring converts fear into a queue.
The Six Risk Categories Every Behavioral Health Practice Must Inventory
The register covers six categories, and the discipline is to inventory all six even when some feel remote. First, PHI breach: the AI vendor, a subprocessor in its chain, or a clinician's personal account exposes protected health information. This includes the quiet versions, a clinician pasting session content into a free consumer chatbot with no BAA, a vendor whose subprocessor list includes a model provider that does not sign a BAA at the tier the practice pays for, a recording stored longer than the retention policy says. For any practice with a substance use disorder component, the breach category must separately flag 42 CFR Part 2 records, because a Part 2 disclosure failure carries its own federal exposure on top of HIPAA.
Second, mis-diagnosis influence: an AI-drafted intake summary or case conceptualization frames a presentation in a way that nudges the clinician toward a wrong diagnostic conclusion, and the clinician adopts the frame without independent verification. The AI never diagnoses, but a confidently worded draft can anchor a tired clinician at 9:54 PM, and the anchoring is the risk. Third, and most serious in this field, missed risk: suicidal ideation, homicidal ideation, abuse disclosure, or intimate partner violence that was present in session but absent or softened in the AI draft, and the clinician signs the note without catching the gap. State this plainly in the register and in every training built from it: AI never scores the CSSRS, never assigns a risk level, never makes the duty-to-protect determination, never makes the mandated-report call. The clinician makes every risk determination; AI structures and formats only after that determination. The missed-risk category exists precisely because a workflow that forgets this rule produces the worst outcomes a practice can produce.
Fourth, biased output: AI drafts that systematically describe some clients differently, more pathologizing language for some populations, less crisis salience for others, culturally tone-deaf framing that ends up in the legal record. Fifth, payer audit exposure: AI-drafted notes that read templated, repeat phrasing across clients, or lack the verifiable specifics (time in session for a 90837, a specific PHQ-9 delta, the modality actually used) that medical-necessity review demands. A pattern of cloned notes is an audit flag and a recoupment engine. Sixth, board complaint: a client, a former client, a colleague, or a supervisee files with the licensing board over undisclosed AI use, a confidentiality concern, or an AI-tainted note, and the board asks for the consent addendum, the BAA, and the policy the practice cannot produce.
Scoring the Grid: Likelihood Times Severity
Score each risk on two five-point scales and multiply. Likelihood: 1 means remote in the next two years, 3 means plausible within a year given current practices, 5 means it is probably already happening somewhere in the practice. Severity: 1 means an internal correction with no client impact, 3 means reportable harm, regulatory notification, or a material financial hit, 5 means client safety harm, license action, or practice-threatening liability. The product runs from 1 to 25. Anything at 15 or above is a red risk demanding mitigation this quarter; 8 to 14 is amber, mitigated this year; below 8 is green, monitored and revisited at the next review.
Be ruthless with the likelihood column, because this is where practices lie to themselves. If twelve clinicians use AI tools and the practice has never audited which accounts they use, the likelihood that at least one is on a consumer tier with no BAA is not a 2; it is a 4 or 5, and Jordan's practice discovered exactly that: a vendor whose subprocessor list included a model provider that does not sign a BAA at the paid tier. Severity is where behavioral health differs from the rest of health care. A missed-risk event, an unflagged suicidal ideation that precedes an attempt, is a 5 on any honest scale, so even a low likelihood demands a workflow control. That is why the missed-risk category almost always lands in the top five even in careful practices: the severity ceiling is absolute.
Run the scoring as a group exercise: the clinical director, the compliance officer, a line clinician who actually uses the tools, and, where the practice supervises associates, a supervisor. The line clinician will correct the likelihood estimates ("everyone exports the transcript, the policy just doesn't know it"), and the supervisor will surface the supervision-chain exposures the owner never sees. Date the meeting, record attendees and dissents. The register's credibility in front of a board or carrier comes from evidence that people with real knowledge scored it.
A risk you scored and mitigated is a defense exhibit. A risk you never wrote down is the plaintiff's opening slide.
The Parity Wrinkle: Scoring Payer Risk in the 2026 MHPAEA Landscape
The payer-audit category needs a 2026-specific calibration, because the federal parity picture moved under everyone's feet. As of 2026, the federal Departments have signaled reduced enforcement of significant portions of the 2024 MHPAEA final rule on non-quantitative treatment limitations: a May 2025 non-enforcement statement, then a March 2026 court-filing disclosure that the Departments intend to propose replacement regulations. Rule-specific framing in practice documentation should be reviewed quarterly until the regulatory picture stabilizes.
What did not move: MHPAEA's statutory parity rights and the CAA 2021 requirement that plans perform NQTL comparative analyses both persist, and the state layer remains fully live. New York's Timothy's Law, the Illinois parity statute, and California's SB 855 enforced through the DMHC remain enforceable and often exceed the federal floor. The consequence for your register is twofold. First, the payer-audit risk does not shrink because federal NQTL enforcement softened; plans emboldened by the federal posture may press utilization review harder, raising the likelihood score for high-frequency 90837 practices. Second, the mitigation shifts: the practice's parity dossier, the asset that documents wrongful-denial patterns and supports appeals, is still worth maintaining, but it should pivot toward state-statute citations where the federal rule is unsettled. A dossier citing CA SB 855 and DMHC authority is stable ground; one resting solely on the 2024 federal NQTL rule is built on a regulation the Departments intend to replace.
Put a literal line item in the register for this: "parity-citation drift," the risk that practice templates, appeal letters, and AI prompt libraries cite a federal rule provision that is no longer being enforced as written. Likelihood is high because templates fossilize; severity is moderate, a weakened appeal rather than a safety event. The mitigation is a quarterly review of every parity citation in the practice's document library, which is cheap and calendar-able.
Mitigating the Top Five: Controls That Actually Hold
Rank the six categories by score and write a mitigation plan for the top five. Resist the urge to mitigate everything; a plan with nineteen action items is a plan nobody executes. Each mitigation gets four fields: control, owner by name, deadline, and the evidence that will prove the control exists. "Improve note review" is not a control. "Every clinician completes the read-every-word verification pass before signing any AI-drafted note; the clinical director spot-audits five signed notes per clinician per quarter; audit log retained" is a control.
For PHI breach, the controls are vendor-side (a current BAA for every AI tool touching PHI, a reviewed subprocessor list, confirmed retention and deletion settings, and for any Part 2 records, explicit BAA coverage of Part 2 data or strict tool segmentation) and human-side (an approved-tool list, a written prohibition on consumer-tier accounts for clinical content, and an annual attestation from every clinician naming the tools they use). For missed risk, the control is structural: the clinician performs and documents their own risk assessment before consulting any AI draft, the signing workflow requires the clinician to verify risk content against their own session memory, and any session containing SI, HI, abuse, or IPV content triggers a no-shortcut review of the full draft. For mis-diagnosis influence, train the anchoring problem explicitly and require that diagnostic impressions be recorded by the clinician before the AI summary is opened.
For payer audit, the control is the verifiable-detail discipline: every 90837 note carries actual time in session, every progress claim carries a measurable anchor like a PHQ-9 score change, and a monthly template-similarity spot check catches phrase cloning before a payer's algorithm does. For board complaint, the control is the disclosure stack: signed AI consent addendum in every active chart, the AI policy distributed and acknowledged, and supervision agreements that address associate AI use explicitly, because the supervisor's license is on the line for every AI-drafted note an associate signs under their number, and the agreement must say so. Whichever category falls sixth still gets a named owner and a review date; unmitigated is acceptable, unwatched is not.
The Register as a Living Document: Cadence, Triggers, and Ownership
A risk register dated eighteen months ago is almost worse than none, because it proves the practice knew how to look and stopped looking. Set a review cadence on the compliance calendar: a full re-score annually, a light review quarterly, and an immediate re-score on any of four triggers: a new AI tool enters the practice, a vendor changes its model provider or subprocessor list, an incident occurs (the subject of the next lesson), or the regulatory ground shifts, as it did with the MHPAEA enforcement posture and continues to do with state AI statutes like the Illinois WOPR Act and Nevada AB 406.
Ownership must be singular. In a solo practice the owner is the clinician and the register can be two pages; the discipline matters more than the formatting. In a group practice like Jordan's, name one accountable owner, usually the compliance officer or clinical director, who maintains the document, schedules reviews, and reports red and amber items to leadership. The owner does not own the risks; clinical risks belong to clinicians, vendor risks to whoever signs contracts. The owner owns the register, the cadence, and the escalation when a mitigation deadline slips.
Finally, connect the register to the rest of your governance stack. The consent addendum mitigates board-complaint risk; the BAA file mitigates breach risk; the pilot log shows tools were evaluated before deployment; the incident-response runbook (next lesson) answers "what happens when a red risk fires." The register is the index page of that binder. When the malpractice carrier's renewal questionnaire asks whether the practice uses AI scribes and what controls exist, the register is the one-document answer, and carriers like CPH & Associates have already started asking.
A Worked Example: Jordan's First Register
Here is how the scoring fell out when Jordan's leadership team ran the exercise. PHI breach: likelihood 4 (twelve clinicians on unaudited tools, one vendor with a non-BAA subprocessor), severity 4 (OCR exposure, notification costs, reputational damage), score 16, red. Missed risk: likelihood 2 (the practice already requires clinician-first risk documentation), severity 5 (client safety, license action), score 10, amber but treated as red by policy: Jordan's team adopted a standing rule that any severity-5 risk gets a mandatory workflow control regardless of likelihood. Payer audit: likelihood 4 (heavy 90837 utilization, similarity checks only recently begun), severity 3 (recoupment, a peer's $14,200 letter fresh in memory), score 12, amber. Board complaint: likelihood 3 (consent addenda rolled out, but two associates' supervision agreements are silent on AI), severity 4, score 12, amber. Mis-diagnosis influence: likelihood 2, severity 4, score 8. Biased output: likelihood 2, severity 3, score 6, green, monitored through the quarterly note-quality audit with a bias-language checklist.
The top five wrote themselves. The biggest red item, the breach exposure, produced three mitigations inside thirty days: a tool amnesty week in which every clinician declared every tool with no discipline attached, a BAA audit that forced one vendor substitution, and a written consumer-tier prohibition. The missed-risk control cost nothing but workflow discipline. The supervision-agreement gap got a deadline tied to the next quarterly supervision cycle. The scored register took a practice that felt vaguely anxious about AI and gave it five concrete, owned, dated tasks. Risk management for behavioral health AI is not a mood. It is a table.
The Applied Problem: Build Your Six-Category AI Risk Register with Top-Five Mitigation Plan
Your artifact is the Six-Category AI Risk Register with Top-Five Mitigation Plan, and you can draft it in ninety minutes. Step one: open a table with seven columns: risk category, specific scenario in your practice, likelihood (1-5), severity (1-5), score, mitigation, owner and deadline. Enter the six categories: PHI breach, mis-diagnosis influence, missed risk, biased output, payer audit exposure, board complaint. Under each, write the scenario in your practice's own facts, not generic language: name the actual tools, the actual payer mix, the actual supervision arrangements. If you cannot name the tools your clinicians use, that discovery is itself a likelihood-5 finding; run the tool amnesty first.
Step two: score with the right people in the room and record who scored. You may use AI to help structure the document; here is a prompt that stays on the right side of the line: "Format the following risk-assessment notes into a six-row register table with columns for category, scenario, likelihood, severity, score, mitigation, owner, deadline. Do not change any scores or add any risks I have not listed. Flag any row where the mitigation field is vague." The scores are yours; the AI formats after your determination, the same division of labor as a clinical note. Never ask AI to score your risks, for the same reason it never scores a CSSRS: the judgment is the licensed work.
Step three: rank by score, apply the severity-5 override rule (any severity-5 risk gets a workflow control regardless of likelihood), and write the top-five mitigation plan with the four fields per item: control, named owner, deadline, evidence. Step four: the verification pass. Read the register asking three adversarial questions: Could I hand this to an OCR investigator tomorrow without embarrassment? Does every mitigation name a person and a date? Does the missed-risk row state explicitly that AI never scores risk instruments and never makes the duty-to-protect or mandated-report determination, with the clinician's verification step spelled out? Fix what fails. "Done" looks like this: a dated, attendee-listed, scored register; five mitigations with owners and deadlines; a review cadence on the calendar; and the document filed alongside the BAAs, the consent addenda, and the AI policy, where the carrier questionnaire and the next lesson's incident runbook can point to it.
Key Takeaways
- The AI risk assessment is the fire marshal's inspection of your practice: it does not make the building fireproof, but it finds the hazards that matter, proves a competent person looked, and converts "we are careful" into a dated, defensible document. Practices that skip it do not survive their first complaint, because the complaint arrives at a practice that cannot show it ever looked.
- Inventory exactly six categories: PHI breach, mis-diagnosis influence, missed risk, biased output, payer audit exposure, and board complaint. Score each as likelihood times severity on five-point scales; 15 or above is red and gets mitigated this quarter, 8 to 14 is amber, below 8 is green and monitored.
- Missed risk carries an absolute severity ceiling, so it earns a workflow control even at low likelihood. The non-negotiable rule belongs in the register verbatim: AI never scores the CSSRS, never assigns a risk level, never makes the duty-to-protect determination, never makes the mandated-report call; the clinician makes every risk determination and AI formats only afterward.
- Calibrate the payer-audit category for 2026: the federal Departments signaled reduced enforcement of the 2024 MHPAEA final rule's NQTL provisions (May 2025 non-enforcement statement, March 2026 court-filing disclosure of intended replacement regulations), but MHPAEA's statutory rights and the CAA 2021 comparative-analysis requirement persist, and state parity laws (NY Timothy's Law, the IL parity statute, CA SB 855 via DMHC) remain enforceable and often exceed the federal floor. Review parity citations in practice documents quarterly and pivot the parity dossier toward state-statute citations.
- Mitigate only the top five, with four fields per mitigation: control, named owner, deadline, and the evidence that proves the control exists. A nineteen-item plan is a plan nobody executes; the sixth category still gets an owner and a review date.
- Make the register a living document: full re-score annually, light review quarterly, immediate re-score when a new tool arrives, a vendor changes subprocessors, an incident occurs, or the regulatory ground shifts. One named owner maintains the register; the risks themselves stay with the clinicians and contract signers who can actually control them.
- The register is the index page of your governance binder, connecting the consent addendum, the BAA file, the pilot log, and the incident-response runbook, and it is the one-document answer when a malpractice carrier's renewal questionnaire asks what AI controls the practice has.
Skill.re