Catch Hallucinations in Coverage Analyses, Reserve Memos, and Filings
Hallucinations in insurance artifacts come in patterns. Fake case citations, phantom endorsements, invented case-law holdings, made-up NAIC bulletin numbers, fictional ISO and AAIS form editions, fabricated NAIC IRIS ratios, invented ASOP section numbers. Each one is detectable with a curated regex or lookup pattern and a reference database; each one is a real LLM failure case the 2026 insurance industry has logged dozens of times across coverage opinions, reserve memos, SERFF filings, ORSA narratives, and Statement of Actuarial Opinion documentation. This lesson catalogs the seven highest-frequency hallucination categories on coverage analyses, reserve memos, and SERFF filings; the detection patterns the verification layer runs to catch them; the audit-defensible documentation of each catch; and the discipline that closes the loop between the verification layer, the credentialed reviewer, and the model risk management governance cycle. The headline rule: every credentialed reviewer in insurance must know what AI hallucinations look like in the artifacts they sign, because the AI is generating fluent prose that reads correct and the credentialed reviewer is the last line of defense before the artifact ships to an insured, a DOI, a treaty reinsurer, or a court.
Hallucination Category 1 - Fake Case Citations
LLMs frequently invent legal case citations. The pattern: a plausible-sounding case name, a plausible reporter and citation format, a plausible holding - but the case does not exist. The hallucination passes a casual read because the syntax is correct (court abbreviation in the right format; reporter volume and page number look real; jurisdiction reference matches).
Real 2026 example caught by a verification layer at a Texas WC carrier. AI-drafted coverage opinion cited "Smith v. Acme Industries, 247 S.W.3d 412 (Tex. App.-Houston [1st Dist.] 2024), holding that workers' compensation exclusivity bars dual-capacity claims where the employer also acted as product manufacturer." The case does not exist. The closest real case is Tigner v. First National Bank, 264 S.W.2d 85 (Tex. 1954), which addresses an entirely different issue. The AI fabricated the citation because the reasoning chain needed a Texas authority on WC exclusivity and the AI's training did not surface a real case at that point in the reasoning.
Detection pattern. Cross-check every case citation against the Westlaw API or the Lexis API. The verification engine extracts citations via regex: typical formats include \d+ [A-Z]\.\d{1,3} \d+ for federal reporters (e.g., 247 F.3d 412); \d+ S\.W\.\d+ \d+ for state regional reporters (e.g., 247 S.W.3d 412); jurisdiction-and-year pattern in parentheses for state appellate cases. Each citation is queried against the legal database; if no match, the citation is flagged as suspicious and routed to the analyst's queue. Detection accuracy in 2026 carriers running this discipline: 92-97% of fabricated citations are caught at verification before the artifact reaches the credentialed reviewer.
The remaining 3-8%. The verification layer misses citations that are syntactically valid and partially match a real case but with subtly fabricated details (a real case in a real reporter at a real volume but with the wrong page number; a real case with a slightly mis-paraphrased name). These reach the credentialed reviewer, who must run the citation lookup as part of the skeptic's checklist.
Hallucination Category 2 - Phantom Endorsements
AI invents endorsement numbers that do not exist or attributes content to endorsements that exist but cover entirely different topics. The pattern: the form number reads as a plausible ISO or AAIS endorsement; the edition date reads plausibly; the AI's paraphrased description reads as plausible coverage language. But the actual endorsement either does not exist or covers a different subject.
Real example. AI-generated coverage analysis cites "CG 21 47 Employment-Related Practices Exclusion (12 07) excludes claims for wrongful termination, sexual harassment, and discrimination." The form number CG 21 47 exists and the edition exists, but the AI has paraphrased the actual exclusion language inaccurately - the real CG 21 47 has different specific language, and the paraphrase misrepresents the scope of the exclusion. Another pattern: AI cites "ISO CG 24 26 Limited Pollution Liability Endorsement." The form number CG 24 26 exists but covers Amendment of Insured Contract Definition, not pollution. The AI conflated two unrelated endorsements.
Detection pattern. Form-edition database with the full text of every ISO and AAIS endorsement, indexed by form number and edition date. Validator checks: (a) endorsement number exists; (b) edition date exists for that number; (c) the AI's paraphrased description matches the endorsement's actual topic via semantic similarity check, not strict text match - the AI's summary is allowed; the AI's invention is not. Threshold: cosine similarity above 0.78 between the AI description and the form's actual purpose statement. Below threshold flags for analyst review.
The maintenance discipline. The form-edition database refreshes quarterly from ISO's published form releases and AAIS's bulletin updates. New endorsements are added; deprecated editions are flagged as historical-only. The MRM team owns the database; the verification engine queries it on every coverage analysis.
Hallucination Category 3 - Invented Case-Law Holdings
Even when the case exists, the AI sometimes attributes holdings to the case that the case did not produce. The pattern: real case name plus real citation plus fabricated holding. The verification layer's citation-existence check passes (the case is real); but the holding is wrong.
Real example. AI-generated property coverage opinion cited "Allstate v. Watts, 811 F.3d 1093 (10th Cir. 2016), holding that anti-concurrent-cause language is unenforceable in Colorado because it violates public policy." The case exists at the cited reporter and volume, but its actual holding addresses an entirely different question (specifically, the duty to defend under a different policy provision, not ACC enforceability). The AI conflated two unrelated cases that share parties or jurisdictions or topical themes.
Detection pattern. Beyond citation verification (Category 1), pull the case's actual headnote and syllabus from Westlaw or Lexis and compare to the AI's stated holding. Compare via semantic similarity at the holding level. Below threshold (0.72) flags. The 2026 best-practice carriers also maintain a curated coverage-litigation database of the 200-400 cases most commonly cited in coverage opinions for their lines of business, with verified holdings for each. AI-cited holdings on these cases are matched against the curated database; mismatches flag immediately at higher confidence than the general Westlaw / Lexis comparison.
The credentialed reviewer's role. The reviewer reads the AI's holding statement and compares to the actual case. If the AI cites Allstate v. Watts for an ACC public-policy holding and the reviewer reads the actual case, the reviewer immediately catches the conflation. Training on common case-law authorities is part of the reviewer's annual continuing education.
Hallucination Category 4 - Made-Up NAIC Bulletin Numbers
NAIC publishes model laws, bulletins, and circular letters with specific identifiers. State DOIs publish their own circulars and bulletins. The AI invents identifiers that look right but do not exist.
Real example. AI-generated SERFF filing memo cited "NAIC Bulletin 2024-09 Use of External Consumer Data in Rating." No such bulletin exists. The actual reference the AI should have produced is NAIC Model Bulletin on the Use of Algorithms, Predictive Models, and AI Systems by Insurers (adopted December 2023). The AI confused a real document with a fabricated identifier or conflated multiple NAIC documents.
Detection pattern. Reference database of NAIC model laws, model bulletins, white papers, and circular letters - plus each state DOI's bulletins, circular letters, and rulemaking documents. The validator parses citations and verifies existence. Citations to "NAIC Bulletin 2024-09" return "not found." The carrier maintains a curated reference list refreshed quarterly from the NAIC website, the state DOI websites, and the NCOIL (National Council of Insurance Legislators) publications. The validator queries the curated list on every artifact that cites NAIC or state DOI authority.
Aliases and informal references. The AI may use informal references ("the AI bulletin" for NAIC Model Bulletin on AI; "the proxy-test letter" for NY DFS Circular Letter 2024-7). The validator maintains an alias dictionary mapping informal references to formal identifiers; if the AI uses an alias, the validator confirms the alias maps to a real document.
Hallucination Category 5 - Fictional ISO/AAIS Form Editions
Covered in the verification-workflow lesson at higher level; here the focus is the specific failure pattern. The AI invents edition dates that do not exist or attributes provisions to an edition that has been updated or deprecated.
Real example. AI cites "ISO CG 00 01 (04 25) Section I Coverage A excluding any claim arising from the use of AI systems by the named insured." Neither the edition (04 25) nor the AI-exclusion language exists. The actual current edition of CG 00 01 has standard exclusions that do not include an AI-systems exclusion. The AI fabricated both the edition date and the exclusion content because the reasoning chain needed a 2025-era CGL form with an AI exclusion and the AI's training did not surface a real form matching that description.
Detection pattern. Form-edition database with full text per edition (per the phantom-endorsements pattern). Validator confirms edition exists; pulls the actual edition text; compares the AI's quoted language to the actual language via fuzzy match at 0.95 threshold. Mismatch flags. The fuzzy-match threshold is high because direct quotes from policy forms should match almost exactly; even slight variation suggests fabrication.
The 2026 AI-endorsement reality. Coalition Control 2.0 introduced the affirmative AI endorsement in 2025-2026 for cyber coverage; some cyber carriers have followed with their own AI-related endorsements. The form-edition database includes these as they emerge; the validator recognizes legitimate AI endorsements and flags AI-related citations that do not match any real form.
Hallucination Category 6 - Fabricated NAIC IRIS Ratios
NAIC IRIS (Insurance Regulatory Information System) computes 13 standard ratios annually per carrier, with established "usual" ranges. The AI sometimes invents IRIS ratio values or attributes wrong usual-range thresholds.
Real example. AI-drafted ORSA narrative cited "Net Premium-to-Surplus Ratio 2.4 (within usual range of -1.5 to 4.0)." The carrier's actual ratio for the period was 1.8, not 2.4; the AI value flowed forward into the narrative without check. The usual range citation was accurate (the actual NAIC IRIS Net Premium-to-Surplus usual range is -1.5 to 4.0), but the carrier's specific ratio was fabricated.
Detection pattern. Cross-check every IRIS ratio in AI-generated artifacts against the carrier's actual annual statement data. The annual-statement data is structured (the NAIC's standardized annual-statement format); the validator extracts IRIS ratios computed from the statement and compares to the AI-cited values. Mismatch above 5% relative tolerance flags. The 5% threshold accommodates minor rounding differences while catching material fabrications. Usual ranges are also validated against the NAIC IRIS Manual, which the carrier's MRM team maintains as a reference document.
The ORSA-narrative-and-SAO concentration. IRIS-ratio hallucinations concentrate in ORSA narratives and in the Statement of Actuarial Opinion documentation. Both artifacts cite financial-position metrics and both reach the DOI as part of annual filings. A fabricated IRIS ratio in an ORSA narrative that the DOI reads is a regulatory exposure; the verification layer catches before submission.
Hallucination Category 7 - Invented ASOP Section Numbers
Actuarial Standards of Practice (ASOPs) have specific section structures published by the Actuarial Standards Board. The AI sometimes invents section numbers or attributes content to a section that exists but addresses a different topic.
Real example. AI-drafted opinion memo cited "ASOP No. 43 Section 3.4.6 requires the actuary to disclose specific data limitations in the actuarial communication." ASOP No. 43 does not have a Section 3.4.6 - the standard's structure does not extend to that subsection level. The data-limitation disclosure requirement the AI was reaching for is in ASOP No. 23 (Data Quality), not ASOP No. 43 (Reserving). The AI conflated two related but distinct standards.
Detection pattern. ASOP database with section structure per standard, sourced from the Actuarial Standards Board's published documents. Validator checks: (a) the ASOP number exists (ASOPs are numbered 1 through 56+ with gaps from retired standards); (b) the section number exists in that ASOP; (c) the cited content matches the section's actual subject matter via semantic similarity check. The 2026 best-practice carriers maintain the ASB's published ASOPs in structured form for verification queries; the structure includes the standard's table of contents, section headings, and key obligations per section.
The cross-ASOP confusion pattern. The AI often conflates related ASOPs - No. 23 (Data Quality) with No. 43 (Reserving); No. 38 (Catastrophe Models) with No. 56 (Modeling); No. 41 (Communications) with No. 36 (Loss-Reserve Opinion). The validator's semantic check catches the conflation; the credentialed actuary's training on ASOP boundaries catches what the validator misses.
The Curated Regex and Lookup Discipline
The verification layer's hallucination-detection patterns are curated, versioned, and continuously updated. The MRM team owns the patterns. Each pattern has: (a) regex or extraction rule to identify citation candidates in AI output; (b) reference database lookup; (c) similarity threshold for content match; (d) failure action (block the artifact, flag for analyst review, log for monthly review).
Pattern updates. Monthly review of caught hallucinations and missed hallucinations. Pattern refinement based on observed AI failure modes - new model versions may produce novel hallucination patterns that require new detection rules. The carrier's "hallucination catch register" logs every caught hallucination with details: AI model version, artifact type, citation cited, failure type, detection pattern that caught it, and analyst's confirmation.
Real LLM failure cases by category 2026. Category 1 (fake citations) is most frequent in coverage opinions and reserve memos. Category 2 (phantom endorsements) is most frequent in coverage opinions and quote-with-restriction memos. Category 3 (invented holdings) is most frequent in coverage opinions and SIU referral memos. Category 4 (NAIC bulletins) is most frequent in SERFF filing memos and ORSA narratives. Category 5 (form editions) is most frequent in coverage opinions and renewal-stewardship narratives. Category 6 (IRIS ratios) is most frequent in ORSA narratives and SAO documentation. Category 7 (ASOP sections) is most frequent in opinion memos, actuarial certifications, and Schedule P narratives.
The pattern-library evolution. 2024 patterns focused on Categories 1-3 (legal citations and holdings). 2025 added Categories 4-5 (regulatory references). 2026 added Categories 6-7 (financial and actuarial reference). The library evolves as the AI failure modes evolve; the 2027 horizon includes new categories for cyber-coverage hallucinations, ESG-disclosure hallucinations, and treaty-clause hallucinations.
The Credentialed Reviewer Discipline
The verification layer catches 70-95% of hallucinations depending on category. The remaining 5-30% reach the credentialed reviewer's desk. The reviewer's job is to read with the right paranoia: every citation is a candidate for verification; every quoted holding is a candidate for cross-check; every reference number is a candidate for lookup; every IRIS ratio is a candidate for reconciliation to the annual statement.
The 2026 best-practice carriers train reviewers explicitly on the seven hallucination categories with example failure cases. Annual training cadence; documented training completion in the MRM registry; the DOI examiner asks for training records in financial-condition examinations and market-conduct exams.
The skeptic's checklist. Before signing any AI-generated coverage opinion, reserve memo, SERFF filing, ORSA narrative, SAO documentation, or actuarial certification:
(1) Verify every case citation via Westlaw or Lexis quick-cite lookup.
(2) Check every endorsement number against the ISO and AAIS form-edition database.
(3) Cross-check every NAIC bulletin and circular letter against the curated reference list.
(4) Confirm every form edition exists and pull the actual text to verify quoted language.
(5) Reconcile every IRIS ratio to the annual statement data.
(6) Verify every ASOP section reference against the ASB's published structure.
(7) Sanity-check magnitude (rate change, reserve, sublimit, premium-to-surplus ratio) against the carrier's documented benchmarks.
Each step takes 30-60 seconds with the right tooling. The cumulative time on a multi-page artifact is 5-15 minutes. The cost of skipping the discipline is the E&O claim, the DOI finding, the treaty-renewal exposure, the bad-faith judgment, or the AM Best survey response that fails - each one orders of magnitude more expensive than the disciplined review.
The Governance Loop and the 2026 Reinsurer Perspective
The hallucination catch register feeds the carrier's AI governance committee monthly review. The committee tracks: caught-hallucination volume by category; missed-hallucination examples surfaced by credentialed reviewer or post-hoc audit; pattern-library updates required; reviewer training gaps. The committee reports to the chief risk officer and to the board's risk committee.
The reinsurer's view. The 2026 reinsurer treaty wordings (covered in the treaty wording markup lesson) include representations about the cedent's AI governance. A cedent with a documented hallucination-catch register, regular pattern updates, and trained credentialed reviewers can demonstrate AI-governance discipline to the reinsurer's treaty broker during renewal due diligence. A cedent without such documentation faces tighter representations, potential rate concessions, or capacity constraints.
The AM Best 2026 readiness survey. The survey asks specifically about AI output verification and hallucination management. Carriers that can answer with documented discipline (verification layer architecture, hallucination categories tracked, credentialed reviewer training) receive favorable survey treatment; carriers that cannot face follow-up questions that pressure capital and rating elements.
The DOI examiner's view. Market-conduct examinations touching AI workflow routinely sample AI-generated artifacts for hallucinations. The examiner reads coverage opinions for fabricated citations; reads SERFF filings for invented bulletin references; reads ORSA narratives for fabricated IRIS ratios. The carrier that has caught and documented these patterns produces clean examination findings; the carrier that has not faces findings letters and remediation cycles.
Key Takeaways
- Seven hallucination categories dominate 2026 insurance artifacts. Fake case citations; phantom endorsements; invented case-law holdings; made-up NAIC bulletins; fictional form editions; fabricated IRIS ratios; invented ASOP sections. Each has a named detection pattern and a named credentialed-reviewer skeptic's-checklist step.
- Case-citation verification via Westlaw and Lexis API catches 92-97% of fabricated citations. Regex extraction plus database lookup. Real failure example: Smith v. Acme Industries, 247 S.W.3d 412 (Tex. App. 2024) - the case does not exist. The remaining 3-8% (syntactically valid but subtly fabricated) reach the reviewer.
- Phantom endorsements use real form numbers with paraphrased-inaccurate content. Detection: semantic similarity above 0.78 between AI description and form's actual purpose statement. Form-edition database refreshes quarterly from ISO and AAIS releases.
- Invented holdings attribute fabricated rulings to real cases. Beyond citation verification, pull actual case headnote and compare via semantic similarity above 0.72. The 2026 best-practice carriers maintain a curated coverage-litigation database of 200-400 frequently-cited cases for higher-confidence catch on common authorities.
- NAIC bulletin numbers are made-up frequently. Real example: "NAIC Bulletin 2024-09" does not exist; the actual reference is NAIC Model Bulletin on AI Systems (December 2023). Reference database refreshed quarterly from NAIC, state DOI, and NCOIL websites; alias dictionary maps informal references to formal identifiers.
- IRIS ratios get fabricated in ORSA narratives and SAO documentation. Cross-check every IRIS ratio in AI artifacts against the carrier's actual annual statement data; mismatch above 5% flags. Usual ranges validated against the NAIC IRIS Manual.
- ASOP section numbers are invented or mis-attributed. Real example: "ASOP No. 43 Section 3.4.6" - that section does not exist; the data-limitation disclosure requirement the AI was reaching for is in ASOP No. 23. Validator maintains structured ASOP database from the Actuarial Standards Board's published documents.
- The verification layer catches 70-95% of hallucinations depending on category. The remaining 5-30% reach the credentialed reviewer's desk. The reviewer is the last line of defense; training on the seven categories plus the skeptic's-checklist discipline are the operational standard.
- The hallucination catch register logs every caught hallucination with AI model version, artifact type, citation, failure type, detection pattern, and analyst confirmation. Monthly AI governance committee review surfaces new failure modes; pattern library evolves; the 2026 reinsurer treaty wordings, the AM Best 2026 readiness survey, and the DOI market-conduct examiner all read the carrier's verification posture as a quality signal.
Skill.re