Recognizing Bad Insurance Output - Phantom Endorsements, Wrong Form Editions, Made-Up Case Law
An LLM that confidently cites ISO CG 00 03 04 13 on a CGL coverage analysis is producing a phantom endorsement; the ISO CGL coverage form is CG 00 01 (occurrence) or CG 00 02 (claims-made), and there is no CG 00 03. An LLM that drafts an ROR citing the 1985 Carrothers v. Carolina Casualty venue case has invented the case; the Carrothers reasoning the model is likely paraphrasing comes from a real but differently-named decision in a different jurisdiction. An LLM that produces a WC class-code analysis citing NCCI code 8810 for a "general office clerical" exposure in Pennsylvania is conflating the federal NCCI taxonomy with the Pennsylvania Compensation Rating Bureau (PCRB) state-specific classification system - PA is one of the monopolistic-equivalent states with its own bureau and different class codes. An LLM that drafts a producer-side compliance memo citing "California Insurance Code §10133.95" on AI disclosure is producing a section that does not exist; the actual operative authority on California broker AI disclosure as of mid-2026 is CDI guidance under §790.03 unfair-practices and pending §1749-series CE expansion. Each of these failures is a documented hallucination surface. Each surfaces in real carrier and broker AI deployments in 2025-2026. Each has a 60-second verification habit that catches it before the artifact lands on the next reviewer's desk. This lesson is the diagnostic checklist - by category, by failure pattern, by verification source - that every L2 insurance professional builds into the muscle memory of daily LLM use.
The Six Categories of Bad Insurance Output
Hallucination in insurance LLM output clusters into six categories. Each category has a mechanism (why the model produces it), a signal (how to spot it in output), and a verification source (the authority that confirms or denies). Memorize the categories; the daily verification habit operates against them.
Category 1 - Phantom ISO form numbers and editions. The model cites a form number that doesn't exist (CG 00 03), or cites a valid form with the wrong edition (CG 00 01 11 88 when the dec page is 04 13), or invents endorsement code numbers (CG 24 35 04 13 cited but the actual endorsement is something else). Mechanism: ISO code patterns are dense in training corpus; the model interpolates plausible-looking codes. Signal: any form citation without a source confirmation. Verification: ISO Mercury portal, the AAIS form library equivalent, or the policy's bound endorsement schedule.
Category 2 - Fake or misattributed case-law citations. The model cites cases that don't exist (made-up case names with plausible court abbreviations), cites real cases with wrong holdings, cites cases from the wrong venue (a California-only Anti-Concurrent-Cause holding cited as if Texas-applicable), or cites unpublished opinions as binding authority. Mechanism: case-law citations follow predictable name + court + year patterns; the model interpolates. Signal: any case citation that doesn't include the Westlaw or LexisNexis identifier and the holding language. Verification: Westlaw, Lexis, or the carrier's controlled case-law database.
Category 3 - Invented bulletin citations and circular numbers. The model cites NAIC bulletins, state DOI circulars, or NAIC working-group documents by name and number when the citation is wrong. Mechanism: NAIC and DOI citation patterns follow stable formats (NY DFS Circular Letter 2024-X, Colorado Reg 10-1-1, Connecticut MC-X-X, Nevada Bulletin X-XXX); the model interpolates. Signal: any bulletin citation not traceable to NAIC content.naic.org, the state DOI website, or the carrier's regulatory-tracking system. Verification: NAIC website, state DOI publications, regulatory-tracking subscriptions (RegEd, S&P Global Regulatory, OneTrust DOI module).
Category 4 - Made-up Schedule P, Schedule F, or statutory-filing references. The model cites Schedule P loss-development triangles, Schedule F assumed/ceded breakdowns, or other statutory annual-statement schedules with fabricated numbers or non-existent line references. Mechanism: statutory-filing structures are dense in training corpus; line numbers and aggregate figures interpolate. Signal: any statutory citation not traceable to the carrier's own filed annual statement or NAIC IRIS database. Verification: carrier's statutory filing PDFs, NAIC IRIS, S&P Global Market Intelligence carrier-data.
Category 5 - Fictional WC class codes and state-specific code conflations. The model cites NCCI class codes that don't exist, applies NCCI codes in monopolistic-equivalent states with their own bureaus (PCRB in PA, WCIRB in CA, NYCIRB in NY, NJCRIB in NJ), or invents state-specific exception codes. Mechanism: NCCI and state-bureau codes follow predictable 4-digit patterns; the model interpolates. Signal: any WC class code without state-specific bureau confirmation. Verification: NCCI Scopes Manual, PCRB Manual (PA), WCIRB California Workers' Compensation Uniform Statistical Reporting Plan, NYCIRB Manual, NJCRIB Manual.
Category 6 - Hallucinated state DOI guidance, statutory citations, and producer-license requirements. The model cites state insurance codes by section number with invented content, cites DOI guidance documents that don't exist, fabricates producer-licensing CE requirements, or invents stamping-office requirements for surplus-lines transactions. Mechanism: state insurance codes follow numeric patterns; the model interpolates. Signal: any state insurance code citation, DOI guidance reference, or producer-license requirement not traceable to the state's official statute, DOI publication, or NIPR. Verification: state statute database (state-specific), DOI website, NIPR for licensing, AAMGA or NAPSLO for surplus-lines stamping requirements.
Phantom ISO Form Numbers - The Canonical Failure
The CG 00 01 vs. CG 00 02 vs. non-existent CG 00 03 example is the canonical phantom-form failure. ISO's CGL coverage form library has exactly two base coverage forms: CG 00 01 (Commercial General Liability Coverage Form - Occurrence) and CG 00 02 (Commercial General Liability Coverage Form - Claims-Made). There is no CG 00 03. When an LLM cites CG 00 03 in a coverage analysis, the model has interpolated a plausible-looking number from the dense ISO code pattern; the citation is structurally impossible.
The edition trap. Even when the form number is correct (CG 00 01 is real), the edition can be wrong. ISO updates the CGL form periodically; the canonical editions in active use as of mid-2026 are CG 00 01 04 13 (April 2013) and CG 00 01 12 07 (December 2007) for legacy policies, with state-amendatory variations layered on top. When an LLM cites "CG 00 01 11 88" - a real ISO edition pattern format with a plausible date - the citation is plausible but probably wrong for the bound policy. Verification: the dec page carries the form edition. If the dec page is not in context, mark [verify against dec page].
The endorsement trap. ISO endorsement codes follow the pattern CG XX XX MM YY (two letters for line, four digits for endorsement, two digits for edition month, two digits for edition year). Real ISO endorsements: CG 21 47 (Employment-Related Practices Exclusion), CG 24 26 (Amendment of Insured Contract Definition), CG 04 35 (Employee Benefits Liability Coverage). The pattern is dense; the model interpolates non-existent codes like CG 24 35 (real) and CG 24 36 (not real, may be confused with CG 24 35), or cites real codes with wrong editions. Verification: ISO Mercury portal subscription, or the carrier's controlled endorsement schedule against the bound policy.
The AAIS confusion. Carriers writing AAIS forms instead of ISO forms have a parallel form library with different codes. AAIS CL 0001 (Commercial Lines General Liability) is the AAIS equivalent of ISO CG 00 01; the AAIS Aggregate Endorsements series follows AAIS-specific codes. LLMs trained on insurance corpora confuse the two systems freely. Verification: identify whether the bound policy is on ISO or AAIS forms; load the right form library into RAG retrieval.
Fake Case Law and the Mata v. Avianca Failure Mode
Fabricated case law is the most-publicized LLM failure mode in legal and quasi-legal work since the May 2023 Mata v. Avianca sanctioned-attorney case, where an attorney submitted a brief citing six fabricated cases produced by ChatGPT. The mechanism is durable: case-law citations follow stable patterns (Plaintiff v. Defendant, Court, Year, with optional Reporter citation), and models interpolate plausible-looking citations on demand. In insurance coverage analysis, the failure surfaces on coverage-opinion drafts citing Anti-Concurrent-Cause holdings, bad-faith standards by venue, Cumis-counsel application, and Brandt-fee awards.
The made-up case. An LLM cites "Carrothers v. Carolina Casualty Insurance Co., 421 F.2d 814 (4th Cir. 1985)" in an ROR draft. The carrier's research department checks: no Carrothers v. Carolina Casualty exists at that citation in any reporter. The model has produced a plausible name (Carrothers reads as a real surname), a plausible court (4th Cir.), a plausible year (1985), and a plausible reporter (F.2d 814 is in the range of real Fourth Circuit 1985 decisions). The citation is structurally fabricated.
The real-case-wrong-holding failure. An LLM cites a real case but attributes a holding the case doesn't carry. The classic failure pattern is the Anti-Concurrent-Cause analysis: an LLM cites Wallis v. United Services Automobile Assn. (the actual California ACC line of cases is real) but applies California's ACC posture as if it controlled in Texas, when Texas Anti-Concurrent-Cause analysis follows a different framework rooted in different case law. Verification: every case citation in a coverage opinion verified by venue, by holding text, and by the Westlaw/Lexis identifier.
The unpublished-as-binding failure. An LLM cites an unpublished or non-precedential opinion as binding authority. Unpublished opinions in federal courts and many state courts cannot be cited as binding precedent under court rules; using them as the basis for a coverage opinion creates §4 reason-chain failure and bad-faith exposure. Verification: status flag on every case citation (published / unpublished, precedential / non-precedential, current good standing / overruled / superseded).
Invented NAIC Bulletin and State DOI Circular Citations
The NAIC Model Bulletin and state DOI circulars are dense citation targets for LLMs working on insurance compliance and AI-governance topics. Real citation patterns: NY DFS Circular Letter 2024-7 (the proxy-test bulletin for NY AI use); Colorado Reg 10-1-1 (the algorithm-inventory and bias-testing requirement); Connecticut MC-25-8 (the Connecticut AI bulletin); Nevada Bulletin 24-006 (the Nevada AI bulletin); NAIC Model Bulletin on Use of AI Systems by Insurers (December 2023 draft adoption).
The plausible non-existent circular. An LLM cites "NY DFS Circular Letter 2024-12" on a topic where no such circular exists (the actual 2024 series has specific numbered letters, and the model invents an additional one). Or cites "Colorado Reg 10-1-2" as if it were a real regulation expanding Reg 10-1-1 - when Reg 10-1-1's expansion was via amendment rather than a new regulation number. Verification: NAIC content.naic.org for NAIC bulletins; NY DFS website for DFS letters; Colorado DOI website for Colorado regulations; state DOI websites for state-specific circulars. Carrier-side: regulatory-tracking system (RegEd, S&P Global Regulatory, OneTrust DOI module).
The wrong-year-pattern citation. An LLM cites "NY DFS Circular Letter 2025-3" when the actual operative circular on the topic is from 2024. The 2024 numbering carries through; 2025 has its own numbering. Verification: temporal check against the citation's claimed year against the current circular index.
The misattributed working-group document. An LLM cites "NAIC Big Data and Artificial Intelligence (H) Working Group Issue Brief, January 2026" - when the actual H Working Group's January 2026 documentation is a calendar-call attendance record, not an issue brief. Verification: NAIC working-group calendars and document repositories on content.naic.org.
Made-Up Schedule P, Schedule F, and Statutory-Filing References
Schedule P (Loss Development) and Schedule F (Reinsurance) are the dense data-citation surfaces in statutory annual statements. An LLM working on a reserve memo, a reinsurance review, or a Statement-of-Actuarial-Opinion narrative routinely cites Schedule P loss-development triangles and Schedule F assumed/ceded breakdowns. The failure mode: invented line numbers, fabricated aggregate figures, or non-existent column references.
The fabricated triangle figure. An LLM cites "Schedule P Part 2A, Section 1, Column 11, Line 8 showing $14.2M of net incurred losses for accident year 2021" - when the carrier's actual filed Schedule P Part 2A doesn't carry that figure. Verification: the carrier's filed annual statement PDF, the NAIC IRIS database, or S&P Global Market Intelligence carrier-data extract.
The wrong-schedule citation. An LLM cites "Schedule F Part 3" for assumed reinsurance when assumed is Schedule F Part 1 (carrier as assuming party) and ceded is Schedule F Part 3 (carrier as ceding party). The conflation is plausible; the citation is wrong. Verification: NAIC annual-statement instructions, or the carrier's actuarial team's Schedule F reference.
The IRIS-ratio fabrication. An LLM cites IRIS ratios with invented figures or interprets them against wrong thresholds. NAIC IRIS ratios have specific definitions and unusual-value thresholds published in NAIC's Annual Statement Instructions. Verification: NAIC IRIS publication, carrier's IRIS-ratio extract from the chief actuary's annual review.
Fictional WC Class Codes and State-Bureau Conflations
Workers' Compensation classification is structured around the NCCI Scopes Manual in 38 states and DC. Twelve states - California (WCIRB), Delaware (DCRB), Indiana (ICRB), Massachusetts (WCRIBMA), Michigan (CAOM), Minnesota (MWCIA), New Jersey (NJCRIB), New York (NYCIRB), North Carolina (NCRB), Pennsylvania (PCRB), Texas (Texas DOI), Wisconsin (WCRB) - operate their own rating bureaus with state-specific class codes that differ from NCCI codes.
The state-bureau conflation. An LLM applies NCCI class code 8810 (Clerical Office Employees) to a Pennsylvania payroll classification - when PA uses PCRB code 953 for the equivalent exposure. The application produces wrong premium calculation, wrong rate, and a misclassified ACORD 130 that the producer signature certifies under license. Verification: NCCI Scopes Manual for NCCI states; PCRB Manual for PA; equivalent for the eleven other state-bureau jurisdictions.
The made-up exception code. An LLM cites a "PCRB code 8810-PA" or "WCIRB code 8810-CA" - applying NCCI numbering as if it were the state-bureau code. Verification: state-bureau classification manual; the state's WC rate filing process at the state DOI.
The wrong dual-state application. An LLM applies NCCI 5645 (Carpentry - Residential) for a contractor working both Texas (Texas DOI rates) and Pennsylvania (PCRB code 951). Each state requires its own classification on the ACORD 130; the LLM conflates. Verification: per-state classification manuals; the producer's state-by-state WC compliance reference.
The 60-Second Verification Habit
The L2 insurance professional's daily LLM use builds a 60-second verification habit into the artifact-review cycle. The habit operates at three checkpoints: pre-prompt (RAG architecture loads authority sources), in-output (model marks uncertain assertions), post-output (reviewer runs the verification scan).
Pre-prompt checkpoint. Five-line preamble (Lesson 1) loads the role + jurisdiction + form edition + reason-code policy + uncertainty handling. L3 RAG architecture loads ISO/AAIS form library, NAIC bulletin text, state DOI circular index, case-law corpus from Westlaw/Lexis, NCCI/state-bureau classification manual, NAIC IRIS data. The prompt structurally cannot draw on hallucinated authority because retrieval bounds generation.
In-output checkpoint. Model marks uncertain assertions [verify against dec page], [verify against ISO form library], [verify against Westlaw/Lexis], [verify against state DOI website], [verify against NCCI / state bureau], [verify against NAIC IRIS]. The L2 prompt constraint demands these markers rather than fabrication.
Post-output checkpoint - the 60-second scan. Reviewer scans every form citation against ISO/AAIS reality (10-15 seconds via the carrier's form-library reference); every case citation against Westlaw/Lexis (15-20 seconds via citation lookup); every bulletin citation against NAIC/DOI websites (10-15 seconds via search); every statutory schedule figure against the carrier's filed annual statement (10-15 seconds via cross-reference); every WC class code against the state bureau manual (10-15 seconds via lookup). Total: roughly 60 seconds for a typical insurance artifact carrying 3-6 such citations.
Real Hallucinations Anthropic / OpenAI Models Have Produced in Insurance Contexts
Empirical examples drawn from published reports, internal carrier evaluation logs, and the public Mata v. Avianca legal-corpus literature. Each is a documented hallucination surface; each illustrates one of the six categories.
The CGL coverage analysis with phantom CG 00 03. A mid-size P&C carrier's enterprise LLM running on Claude or GPT for ROR drafting cited CG 00 03 04 13 as the coverage form on a faulty-workmanship coverage analysis. The form does not exist. The carrier's quality-assurance reviewer caught it via the post-output checkpoint; the file note logged the hallucination and the corrective action (RAG architecture adjustment to load the actual policy form by edition into context, plus the L2 prompt constraint demanding [verify against dec page] on uncited form numbers).
The fabricated Anti-Concurrent-Cause case. A defense-panel attorney working with carrier coverage counsel on a wind-and-flood property loss in Texas used an LLM to draft an initial coverage memo. The memo cited "Carrothers v. Carolina Casualty Insurance Co., 421 F.2d 814 (4th Cir. 1985)" as authority on the Anti-Concurrent-Cause posture; no such case exists at that citation. Caught by the senior attorney's verification scan. The corrective discipline: RAG architecture loads the Westlaw/Lexis case-law corpus; the L2 prompt constraint demands the Westlaw/Lexis identifier on every cited case.
The wrong PCRB class code application. A producer's agency-overlay LLM working on a Pennsylvania contractor's ACORD 130 classified clerical staff under NCCI 8810 - when PA uses PCRB code 953. The producer caught it before submission via post-output scan. The corrective discipline: RAG architecture loads state-specific bureau manuals by state of risk; the L2 prompt constraint flags any WC class code without state-bureau confirmation.
The made-up Colorado Reg 10-1-1 amendment number. A carrier compliance team's LLM produced a memo on the October 2025 expansion of Reg 10-1-1 (which actually was an amendment to the existing regulation, not a new regulation), citing "Colorado Reg 10-1-2" as the expansion. The regulation does not exist by that name. Caught by the chief compliance officer's verification scan. The corrective discipline: RAG architecture loads state DOI publication index; the L2 prompt constraint demands [verify against state DOI website] on any cited regulation number.
The fabricated Schedule P aggregate. A reserving-actuary's LLM produced a draft commentary citing Schedule P Part 2A figures that didn't reconcile to the carrier's actually filed annual statement. The chief actuary caught it via cross-reference to the filed PDF. The corrective discipline: RAG architecture loads the carrier's filed annual-statement data as authority; the L2 prompt constraint demands any statutory schedule figure to match the carrier's filed source.
The L3 RAG Architecture and the L4 Incident Runbook
The L3 RAG architecture loads the authority surfaces that prevent each category of hallucination structurally: ISO Mercury portal subscription for ISO form library; AAIS form library where applicable; Westlaw / Lexis case-law corpus; NAIC content.naic.org bulletin index; state DOI publication indexes (CA CDI, NY DFS, CO DOI, CT DOI, NV DOI, etc.); NCCI Scopes Manual; state-bureau classification manuals (PCRB, WCIRB, NYCIRB, NJCRIB, others); NAIC IRIS database; the carrier's filed annual statement; the carrier's controlled appetite guide; the carrier's bound endorsement schedule. Retrieval against these sources bounds generation; the model cites what retrieval returns rather than what training prior interpolates.
The L4 incident runbook on a caught hallucination. (1) Reviewer logs the hallucination with category (phantom form / fake case / invented bulletin / made-up schedule / fictional class code / hallucinated DOI guidance), the artifact, the model version, the prompt version, the retrieved context, the actual output, and the corrected output. (2) Production pause on the template if material defect. (3) Root-cause analysis - was retrieval missing the authority source? Was the prompt constraint missing the verification marker? Did the reviewer's verification scan operate? (4) Corrective action - RAG source addition or refresh, prompt constraint update, reviewer training cadence update. (5) §4.4 documentation update. (6) State-overlay regulatory notification where material adverse impact. (7) Post-mortem to L4 governance committee with corrective action signed off. (8) Template version-controlled and redeployed.
Key Takeaways
- Six categories of bad insurance LLM output: phantom ISO form numbers and editions; fake or misattributed case-law citations; invented NAIC bulletin and state DOI circular citations; made-up Schedule P, Schedule F, statutory-filing references; fictional WC class codes and state-bureau conflations; hallucinated state DOI guidance, statutory citations, producer-license requirements.
- The canonical phantom-form failure is CG 00 03 - a form number that does not exist. ISO's CGL coverage form library has exactly two base forms (CG 00 01 occurrence and CG 00 02 claims-made). The edition trap (CG 00 01 11 88 cited when dec page shows 04 13) and the AAIS confusion (mixing ISO with AAIS form codes) are the related failure surfaces.
- The Mata v. Avianca fabricated-case failure mode is durable across insurance contexts. LLMs interpolate plausible case-law citations from training-corpus patterns; coverage opinions citing Anti-Concurrent-Cause, bad-faith standards, Cumis-counsel application, and Brandt-fee awards all carry the failure surface. Real-case-wrong-holding and unpublished-as-binding are the related failure modes.
- NAIC bulletin and state DOI circular citations are dense LLM citation targets. NY DFS Circular Letter 2024-7, Colorado Reg 10-1-1, Connecticut MC-25-8, Nevada Bulletin 24-006 are real; the model interpolates non-existent series. Wrong-year-pattern, made-up amendment, and misattributed working-group document are the failure surfaces.
- WC class codes have a federal NCCI baseline and 12 state-bureau exceptions. PCRB (PA), WCIRB (CA), NYCIRB (NY), NJCRIB (NJ), and 8 other bureaus operate state-specific classifications. LLMs conflate NCCI codes with state-bureau equivalents; a Pennsylvania payroll classified as NCCI 8810 should be PCRB 953.
- Schedule P, Schedule F, and IRIS-ratio citations are statutory-filing failure surfaces. Fabricated triangle figures, wrong-schedule conflations (Schedule F Part 3 for assumed when it's Part 1), and invented IRIS-ratio interpretations all surface in actuarial draft memos. Verification against the carrier's filed annual statement PDF + NAIC IRIS database closes them.
- The 60-second verification habit operates at three checkpoints. Pre-prompt (RAG architecture loads authority sources); in-output (model marks [verify against ...]); post-output (reviewer scans every citation in roughly 60 seconds total across the typical 3-6 citations per artifact).
- The L3 RAG architecture + L4 incident runbook produces the structural mitigation. RAG loads ISO Mercury, AAIS, Westlaw/Lexis, NAIC content, state DOI indexes, NCCI/state-bureau manuals, NAIC IRIS, carrier-filed annual statements; incident runbook captures category, root cause, corrective action, and §4.4 documentation update on every caught hallucination.
- Real hallucinations are documented in 2025-2026 insurance contexts. Phantom CG 00 03 in ROR drafts; fabricated Carrothers v. Carolina Casualty Anti-Concurrent-Cause cite; wrong NCCI 8810 on PA payroll when PCRB 953 applies; made-up Colorado Reg 10-1-2; fabricated Schedule P Part 2A figures. Each is a real surface; each has a documented fix.
Skill.re