Incident Response for AI Failures - Adverse Decisions, Drift Events, Vendor Outages, MHPAEA NQTL Violations
AI incidents in insurance operations are not edge cases - they are the predictable failure modes of complex models running on shifting consumer populations across changing regulatory regimes. The five most common 2026 incidents are: (1) adverse-decision errors traced to a model output (premium impact, claim denial, accelerated UW knockout, fraud referral); (2) drift events where a model's behavior changes faster than the monitoring cadence catches; (3) vendor outages (Hi Marley, Five Sigma, Tractable, Cytora, Federato, Akur8 down for hours to days); (4) data exposure where AI workflow inadvertently routes NPI or PHI to an environment without appropriate safeguards; (5) MHPAEA Non-Quantitative Treatment Limitation (NQTL) violations where behavioral-claims AI produces disparate denial rates. Each has a documented runbook; each runbook integrates with NAIC §4 incident response, Colorado Reg 10-1-1 notification expectations, state DOI bulletins, FCRA workflow, and HIPAA breach-notification timing. Carriers without runbooks improvise under pressure and produce inconsistent regulatory disclosure, missed notification windows, and inadequate consumer remediation. Carriers with runbooks execute mechanically - incident classification, severity scoring, notification SLA, remediation workflow, post-incident review - and produce defensible documentation. This lesson is the five runbooks: Hi Marley outage, Akur8 model-drift, Shift false-positive cluster, vendor data-exposure event, and MHPAEA NQTL discovery with 60-day quantitative parity test.
The Incident Response Program - Classification, Severity, Notification
NAIC §4 governance and risk management element requires documented incident response. The program operationalizes §4 through a classification scheme, severity scoring, and notification SLA matrix. Incident classification (six categories): (1) adverse-decision error (consumer-impacting model error), (2) drift event (model behavior changed materially), (3) vendor outage (third-party AI down or degraded), (4) data exposure (NPI/PHI routed to unauthorized environment or party), (5) bias or fairness failure (disparate-impact threshold breach), (6) MHPAEA NQTL violation (behavioral-claims AI disparate treatment).
Severity scoring (four levels): Severity 1 - material consumer harm or regulatory exposure (Hispanic-surname false-positive cluster, accelerated UW knockout pattern with FCRA violation, multi-claim NPI exposure). Severity 2 - single-claim or limited consumer harm with remediation potential, vendor outage exceeding 4 hours during business operations, drift exceedance triggering enhanced monitoring. Severity 3 - operational incident without consumer harm (vendor degraded performance, single-user data-handling error, drift within yellow band). Severity 4 - minor operational anomaly (logged for trend analysis, no operational response required).
Notification SLA matrix: Severity 1 triggers CRO + CCO + General Counsel + CEO notification within 4 hours; AI committee emergency session within 24 hours; board risk committee chair within 48 hours; state DOI notification per applicable bulletin timing (Colorado Reg 10-1-1 expects timely notification of material adverse-impact incidents). Severity 2 triggers algorithm owner + accountable executive + CRO within 24 hours; AI committee at next standing meeting. Severity 3 logged in incident system with 5-day review by accountable executive. Severity 4 trended only.
Runbook 1 - Hi Marley Outage
Hi Marley provides SMS-first claims communication used by most major P&C carriers in 2026. Service-affecting outages are infrequent but consequential because the carrier's claims comms workflow depends on it. The runbook activates when Hi Marley reports a Severity 1 or Severity 2 incident on its status page, or when carrier's monitoring detects degraded API response for more than 30 minutes.
T+0 (detection): Hi Marley status page or internal monitoring alert. Algorithm owner (Director Claims Operations) notified by alert system. Initial assessment within 15 minutes - confirm scope (full outage vs. degraded), estimated duration if Hi Marley communicates.
T+30 minutes: Activate claims comms continuity protocol - claims floor falls back to direct SMS via the carrier's MMS gateway, with templated messaging matching Hi Marley's standard openings; pending Hi Marley conversations queued for catch-up post-restore; FNOL intake re-routes to phone with extended hold capacity. Communication to claims floor: 5-minute standup with Director Claims and Claims VP; runbook script ready.
T+1 hour: Customer-facing communication if outage exceeds 1 hour - proactive SMS via fallback channel: "We're experiencing temporary delays in messaging. Your claim is being handled - adjuster will call within 4 hours." Cycle-time impact tracked.
T+4 hours: Severity 2 notification to CRO; AI committee chair informed; cycle-time impact assessed and projected; if outage continues, evaluate prolonged-outage protocol (full phone fallback, contractor surge from claims TPA partners, customer remediation messaging).
T+8 hours: Severity 1 reclassification if outage continues with material claim-cycle impact. Notification expands to CCO and General Counsel. Consider state DOI notification depending on jurisdiction and exposure (Colorado, New York for material consumer impact).
Restoration and post-incident: Hi Marley restores; queued conversations replayed; adjusters verify customer-facing impact; carrier produces post-incident review within 5 business days documenting: outage timeline, customer impact (delayed conversations, cycle-time delta, complaint volume), vendor communication and remediation, runbook execution gaps, registry entry update for Hi Marley vendor scorecard, lessons learned. Vendor scorecard notes the incident; concentration analysis re-checked.
Runbook 2 - Akur8 Model-Drift Response
Drift in a personal-auto pricing model produces premium-impact risk: customers paying rates inconsistent with the current loss experience embedded in the model. Drift is detected by registry-integrated monitoring against PSI, calibration, AUC, and segment-level metrics. PSI exceeding 0.10 (yellow) triggers enhanced monitoring; exceeding 0.25 (red) triggers Severity 2 incident response.
T+0 (detection): Registry dashboard alerts on PSI yellow or red threshold. Algorithm owner (AVP Personal Lines Pricing) notified within 1 hour. Initial assessment within 4 hours - segment-level analysis (which features drifted, which states, which book segments), calibration check (is the drift directional bias or pure shift), business-impact estimate (premium impact at current rate plan).
T+24 hours: Severity 2 notification; AI committee chair informed; Chief Actuary briefed (accountable executive). Triage decision: (a) drift within acceptable range - enhanced monitoring only, no model change; (b) drift requiring rate-filing response - Akur8 platform refresh, new rate filing under accelerated SERFF timeline; (c) drift requiring immediate model intervention - rollback to prior version, recalibration, or emergency rate adjustment.
T+72 hours: Triage decision made and documented. If (b) or (c): AI committee emergency session if Severity 1 reclassification (material adverse impact). Notification to state DOIs in 14 deployment states if material - Colorado Reg 10-1-1 expects timely notification; NY DFS expects notification under Circular Letter 2024-7 if material impact on NY policyholders.
T+5 days: Remediation plan documented and approved. Communication to producer force if rate adjustment imminent. Registry entry updated with drift incident, decision, remediation. Akur8 vendor scorecard updated with incident reference.
T+30 days: Post-incident review documenting drift cause (population shift, regulatory change in deployment states, external data source change), monitoring adequacy, response timeliness, business impact realized, registry refresh.
Runbook 3 - Shift Fraud False-Positive Cluster Response
Fraud-detection models produce false positives - legitimate claims flagged for SIU referral. A cluster - disproportionate false-positive rate concentrated in a demographic class, geography, or claim type - is a Severity 1 incident under §4 because it implies disparate treatment. The October 2025 Hispanic-surname false-positive cluster at the carrier in the registry example produced exactly this response.
T+0 (detection): Quarterly fairness test or operational anomaly surface the cluster. Detection signals: false-positive rate by demographic class ratio exceeding 1.15, surge in SIU referrals from a specific geography or surname pattern, customer complaints with discrimination allegations. Algorithm owner (Director SIU Operations) notified within 2 hours.
T+4 hours: Severity 1 classification. CRO + CCO + General Counsel + CEO notified. AI committee emergency session scheduled within 24 hours. Initial scope assessment - affected claim count, demographic concentration, model version implicated, SIU referrals already actioned.
T+24 hours: AI committee emergency session. Decisions: (1) immediate model action - rollback to prior version if available, threshold adjustment to reduce false-positive rate, or model offline pending v-next; (2) consumer remediation - affected claims re-reviewed by human SIU specialists without model input, denied claims re-opened where appropriate, customer communication with apology and review status; (3) regulatory notification - state DOI notification under applicable bulletins (Colorado Reg 10-1-1, NY DFS Circular Letter 2024-7 for affected states); (4) vendor communication - Shift notified, root-cause analysis requested, model retraining commitment.
T+72 hours: Notification executed to Colorado DOI, NY DFS, other implicated states. Customer remediation begun on impacted claims. Vendor working on root-cause analysis with carrier oversight. Press preparation in case of media inquiry. Registry incident entry created with severity 1 classification.
T+30 days: Remediated model version deployed (Shift v4.0 in the October 2025 example) with documented feature removal (surname-derived feature) and re-tested fairness ratios. Customer remediation completed across affected claims. Post-incident review filed with AI committee. Final state DOI updates with remediation evidence.
Runbook 4 - Data Exposure Event (GLBA NPI / HIPAA PHI)
The most common AI data-exposure event in 2026: an employee pastes NPI or PHI into a consumer-grade LLM (ChatGPT, Claude.ai, Gemini, Copilot in consumer mode rather than M365 contracted environment) for a quick task. The exposure is technically a §501(b) GLBA Safeguards breach and, for L&H, potentially a HIPAA breach. Less common but more consequential: vendor's AI workflow inadvertently routes data outside the contracted environment (e.g., a fine-tuning run that included carrier NPI in training data exposed to third parties).
T+0 (detection): Detection via DLP (data loss prevention) alert, employee self-report, vendor disclosure, or external researcher report. Algorithm owner and CISO notified within 2 hours. General Counsel notified for legal-privilege review.
T+4 hours: Severity classification - typically Severity 1 if PHI exposure to public LLM (HIPAA breach analysis), NPI exposure to public LLM (GLBA Safeguards breach), or vendor data outside contracted environment. CRO + CCO + General Counsel + CEO notified. Initial scope - record count, data classification, exposure duration, recipient (public LLM provider's training data, third party, etc.).
T+24 hours: Forensic investigation. For consumer-LLM pastes: ChatGPT / Claude / Gemini policies on training-data inclusion reviewed; carrier requests data deletion under vendor terms; assess whether data was used for model training (typically retained for limited period unless training-opt-out activated). For vendor data exposure: vendor incident response coordination, root-cause analysis, data return or deletion verification.
T+72 hours: Notification analysis. HIPAA Breach Notification Rule requires individual notification within 60 days of discovery for breaches affecting 500+ individuals; HHS notification same timing; media notification if affecting 500+ residents of a state. GLBA Safeguards expects timely consumer notification (state-specific timing under state breach-notification laws). State breach-notification statutes typically 30-90 days. State DOI notification under applicable bulletins.
T+30-60 days: Notification executed per timing requirements. Employee training and DLP enhancement implemented. Vendor governance review - additional controls, contractual remediation, or vendor exit decision. AI committee post-incident review with §4 program enhancement. Registry entries updated for affected AI Systems with data-handling controls strengthened.
Runbook 5 - MHPAEA NQTL Violation Discovery
The Mental Health Parity and Addiction Equity Act prohibits health plans from imposing more restrictive Non-Quantitative Treatment Limitations (NQTLs) on mental-health and substance-use-disorder benefits than on medical/surgical benefits. AI-powered utilization review applied to behavioral health claims is an NQTL - its design, application, and outcomes must meet parity standards. Discovery of disparate denial rates between behavioral and medical/surgical claims triggers a 60-day quantitative parity test requirement.
T+0 (detection): Routine NQTL monitoring, complaint pattern analysis, or quarterly bias review surfaces disparate denial rates. Threshold for Severity 1: behavioral denial rate exceeds medical/surgical denial rate by material margin (typical operational threshold - 1.5x or 50% higher) sustained across review period. L&H Claims Supervisor and Chief Claims Officer notified within 4 hours.
T+24 hours: Severity 1 classification. CRO + CCO + General Counsel + CEO notified. AI committee emergency session within 48 hours. Initial scope - affected claim count, model version, behavioral vs. medical/surgical denial ratio over relevant period, complaints received.
T+48-72 hours: AI committee session. Decisions: (1) immediate model action - utilization review AI offline for behavioral claims pending review, human-only review for behavioral claims in interim; (2) 60-day quantitative parity test - design and execute formal NQTL comparative analysis covering processes, strategies, evidentiary standards, and other factors used in design and application of the NQTL, in writing, comparing behavioral and medical/surgical NQTLs; (3) DOL / HHS notification preparation - federal MHPAEA enforcement bodies; (4) state insurance department notification per state mental-health parity overlays.
T+60 days: NQTL comparative analysis completed. Findings documented. If disparate treatment confirmed, remediation: model retraining with parity controls, new utilization review protocols ensuring parity, customer remediation on impacted denials (re-review, reopening where appropriate, payment with interest). DOL / HHS / state DOI notifications with remediation evidence.
T+90 days: Remediated AI deployed with documented parity controls. Ongoing parity monitoring established with quarterly NQTL comparative analysis. Registry entry updated with parity-test methodology and cadence. Customer remediation continued. Post-incident review filed.
Severity 1 Notification Matrix and the State DOI Coordination
Severity 1 incidents trigger multi-party notification coordinated through CCO. The matrix:
- CEO + Board Risk Committee chair - 48 hours.
- CRO + CCO + General Counsel - 4 hours.
- AI committee emergency session - 24 hours.
- State DOI of domicile - typically 5 business days under Colorado Reg 10-1-1 for material adverse-impact incidents; varies by state.
- State DOIs in deployment states with material consumer impact - 5-10 business days depending on bulletin.
- NY DFS for material NY-policyholder impact - within applicable Circular Letter 2024-7 timing.
- Federal - HHS for HIPAA breach (60 days for 500+ individuals; same day media if 500+ in a state); DOL for MHPAEA enforcement; FTC for FCRA-related events under §615(d).
- Reinsurer - under treaty wording's AI-event reporting clauses if present; bordereau entry if material.
- AM Best - typically not direct incident notification; surfaces at next analyst meeting as mid-cycle update if material development.
The coordination discipline is documented: CCO leads notification; General Counsel reviews legal and privilege posture; CRO coordinates enterprise risk implications; algorithm owner and accountable executive provide operational detail. Without documented coordination, notifications miss windows, contradict each other, or fail to invoke privilege appropriately.
Post-Incident Review and Program Evolution
Every Severity 1 and Severity 2 incident triggers post-incident review within 30 days. The review documents: incident timeline, severity assessment, response execution against runbook, consumer impact, regulatory notifications, remediation, root-cause analysis, runbook gaps, registry updates, program-level lessons. The post-incident review feeds the AI committee, the registry's incident-history field for affected entries, the §4 program annual refresh, and the AISET Exhibit B narrative on incident response posture.
Program evolution from incidents is the §4 continuous improvement obligation. The Hispanic-surname false-positive cluster produced not just Shift v4.0 deployment but also: (1) enhanced fairness-test design adding surname-pattern analysis to standard methods; (2) quarterly cluster-detection analysis across all Tier 1 models, not just fraud; (3) vendor §4 attestation refresh to confirm vendor governance posture; (4) registry field addition for "surname-pattern test status" on consumer-facing models; (5) board risk-committee briefing on incident learnings. The carrier that treats post-incident review as compliance theater misses the program evolution; the carrier that treats it as program input absorbs the learning and reduces future incident exposure.
Key Takeaways
- Five most common 2026 incidents: adverse-decision error, drift event, vendor outage, data exposure, MHPAEA NQTL violation. Each has documented runbook integrated with §4, Colorado Reg 10-1-1, state DOI bulletins, FCRA, HIPAA.
- Severity scoring is four levels: S1 (material consumer harm or regulatory exposure), S2 (limited harm with remediation), S3 (operational without consumer harm), S4 (minor anomaly trended). Notification SLA matrix maps severity to notification scope and timing.
- Hi Marley outage runbook: T+30 min activate continuity protocol with phone and templated MMS fallback; T+1 hr customer-facing communication; T+4 hr CRO notification; T+8 hr S1 reclassification if continuing; post-incident review within 5 business days.
- Akur8 model-drift runbook: PSI yellow (0.10) enhanced monitoring; red (0.25) S2 incident; T+24 hr AI committee chair and Chief Actuary; T+72 hr triage decision; T+30 days post-incident review. State DOI notification if material under Colorado Reg 10-1-1 or NY DFS.
- Shift false-positive cluster runbook: cluster detection at false-positive ratio >1.15 triggers S1; T+4 hr CRO/CCO/GC/CEO; T+24 hr AI committee emergency session with model action + consumer remediation + regulatory notification + vendor communication; T+30 days remediated version deployed. October 2025 Hispanic-surname cluster ran this exact playbook.
- Data exposure runbook: T+4 hr severity classification; T+24 hr forensic investigation; T+72 hr notification analysis under HIPAA Breach Notification (60 days for 500+; HHS + media if 500+ in state), GLBA, state breach laws (30-90 days). Most common 2026 source: employee paste into consumer LLM.
- MHPAEA NQTL violation runbook: behavioral denial rate >1.5x medical/surgical triggers S1; T+24 hr CRO/CCO/GC/CEO; T+48-72 hr AI committee with model offline for behavioral + 60-day quantitative parity test; T+60 days NQTL comparative analysis with remediation; T+90 days remediated AI with ongoing parity monitoring. DOL + HHS + state DOI notification chain.
- Severity 1 notification matrix: CEO + Board chair 48 hr; CRO/CCO/GC 4 hr; AI committee 24 hr; state DOI of domicile 5 business days (Colorado Reg 10-1-1); deployment states 5-10 days; NY DFS Circular 2024-7 timing; HHS for HIPAA 60 days; DOL for MHPAEA; FTC for FCRA §615(d); reinsurer per treaty.
- Post-incident review is §4 continuous improvement obligation, not compliance theater. Hispanic-surname cluster produced enhanced fairness-test design, quarterly cluster detection across Tier 1, vendor attestation refresh, registry field addition, board briefing. Carrier that treats review as program input reduces future incident exposure.
Skill.re