The NAIC Model Bulletin on AI Systems, Model 880, and the AI Systems Evaluation Tool
The NAIC Model Bulletin on the Use of Artificial Intelligence Systems by Insurers, adopted in draft form in December 2023 and now adopted in some form by 25+ U.S. insurance jurisdictions as of mid-2026, is the single document every insurance professional must be able to navigate from memory. It sits on top of NAIC Model 880 (Unfair Trade Practices Act) as the operational framework that translates §4 of Model 880 into a four-pillar governance program - and it is the document that produces the NAIC AI Systems Evaluation Tool, the four-exhibit (A, B, C, D) information request that the 12-state pilot (California, Colorado, Connecticut, Florida, Iowa, Louisiana, Maryland, Pennsylvania, Rhode Island, Vermont, Virginia, Wisconsin) is sending to carriers from Q1 2026 kickoff through the September 2026 close, with re-exposure for public comment in September–October 2026 and adoption expected at the NAIC Fall National Meeting in November 2026. If you are an underwriter, an adjuster, a producer, or an actuary in 2026, the Model Bulletin and the AI Systems Evaluation Tool are the regulatory rails you will be operating on for the next decade - and this lesson is the map.
Why the Bulletin Exists and What Model 880 Actually Says
The NAIC's Unfair Trade Practices Act - Model 880 - has been the statutory backbone of state insurance market-conduct enforcement since the 1940s. §4 of Model 880 prohibits unfair discrimination, false advertising, misrepresentation, twisting, and a list of other consumer harms. The clause that matters most for AI in 2026 is the unfair-discrimination prohibition, which makes it a market-conduct violation for an insurer to discriminate against an applicant or insured on the basis of a protected class (race, color, religion, national origin, ancestry, sex/gender, marital status, sexual orientation, gender identity, age, disability - the specific list varies by adopting state). The bulletin's pivotal contribution is that it extends the §4 analysis to algorithmic and AI-driven outcomes: a fraud score, a triage decision, an accelerated-UW knockout, or a rate-impact recommendation produced by a machine-learning model can be a §4 violation if the model's outputs disparately affect a protected class - even if no human ever consciously discriminated and even if the model never saw a protected-class variable directly.
The bulletin is not itself binding law. It is a model bulletin - a template that each state's Department of Insurance (DOI) chooses to adopt, adapt, or reject. By mid-2026, 25+ jurisdictions have adopted some form of it (the count varies by tracker; Holland & Knight, Faegre Drinker, and the NAIC's own Issue Brief from March 2026 are the working references), with the holdouts mostly in the smaller-population states where DOI rule-making bandwidth is constrained. The states that have adopted it most aggressively - Colorado (which goes further with Reg 10-1-1), Connecticut (MC-25-8), Nevada (Bulletin 24-006), New York (DFS Circular Letter 2024-7, technically pre-bulletin but operationally aligned), Maryland, Vermont, Rhode Island, Virginia, Washington, Oregon - are the ones now participating in the 12-state Evaluation Tool pilot, and the bulletin's text is what their examiners will cite when they walk into your shop.
Model 880 §4(a). The unfair-discrimination prohibition that the bulletin extends to algorithmic outcomes. Every adverse AI-influenced decision - a decline, a quote-with-restriction, a knockout, a fraud referral, a triage assignment that materially affects the consumer - needs a documented, human-reviewable reason that does not reduce to "the model said so." This is the rule that all four governance pillars below ultimately serve.
The Four Governance Pillars
The bulletin organizes the insurer's obligations into four explicit governance pillars. Memorize the names. Every DOI examiner who has read the bulletin once will use these four words.
Pillar 1 - AI Program (Governance Framework)
The first pillar requires the insurer to maintain a written AI program: a policy, an accountable senior executive (the bulletin language is "senior management"; in practice this is the Chief Risk Officer, Chief Underwriting Officer, Chief Claims Officer, Chief Actuary, or the Chief Compliance Officer, depending on the insurer's structure), a committee or committee-equivalent that oversees AI use across the enterprise, written escalation paths, and risk-tiering. The §4.1 language asks for a "risk-based framework" - meaning the insurer must classify AI systems by the consequence of error and apply heavier controls to higher-risk systems. A photo-classification model that auto-routes auto FNOL to fast-track is low-risk; a model that produces an accelerated-UW knockout reason code on a life application is high-risk. The framework documentation is the first artifact a market-conduct examiner asks for.
Pillar 2 - Third-Party AI Systems
The second pillar - §4.2 - extends the program to third-party AI: the Cytora submission triage, the Federato RiskOps workbench, the Akur8 GLM, the Earnix dynamic-pricing engine, the Tractable photo estimator, the CCC Intelligent Solutions auto pipeline, the Shift Technology fraud model, the Hi Marley SMS triage, the Five Sigma auto-coverage summary, the Coalition cyber-scan engine, the Munich Re / Swiss Re Magnum / RGA AURA NEXT / SCOR Velogica L&H underwriting engines. The bulletin makes clear that delegation does not relieve the insurer. If a third party's model produces a §4 violation, the insurer is on the hook - which is why the bulletin specifies that contractual due diligence, model documentation, audit rights, sub-processor disclosure, and exit clauses are required components of the third-party relationship. The third-party AI inventory schedule is the second artifact an examiner asks for, and it is the artifact your contract addendum has to satisfy.
Pillar 3 - Testing and Validation
The third pillar - §4.3 - requires ongoing testing and validation. The bulletin does not prescribe specific metrics, but the working consensus in 2026 - informed by Colorado SB 21-169, NY DFS Circular Letter 2024-7, and the AI Systems Evaluation Tool's Exhibit C - is that high-risk AI systems need (1) a bias test on a defined cadence (annual at minimum; quarterly for the most consumer-impacting models), (2) a drift-monitoring plan with PSI / KS / AUC tracking and segment-level monitoring, (3) a calibration check, (4) explainability artifacts (SHAP, PDP, ALE plots) for high-risk pricing and underwriting models, and (5) a champion-challenger testing protocol for major model updates. The disparate-impact ratio calculation, the marginal-effects analysis on protected classes, and the proxy-test memo (NY DFS) are the named outputs of this pillar. Every insurer running an Akur8 GLM, a Munich Re accelerated-UW engine, or a Shift fraud model is expected to be able to show the testing log on demand.
Pillar 4 - Documentation
The fourth pillar - §4.4 - is the documentation pillar, and in practice it is the pillar that decides whether the previous three pillars survive a market-conduct exam. The bulletin requires documentation of (a) the AI inventory itself - the algorithm registry with model name, version, vendor, owner, last-tested date, bias-test status, drift-monitoring status, risk tier, deployment scope; (b) the reason-code framework for every AI-influenced adverse decision; (c) the policies, procedures, and training records; (d) the complaint-handling integration - meaning a consumer complaint involving an AI-driven adverse decision has to be routable to a human who can re-examine the model output and produce a human-reviewable reason. The model card, the algorithm-inventory entry, the reason-code memo, and the complaint-routing log are the documentation artifacts the bulletin contemplates. If they do not exist, the carrier is exposed.
The AI Systems Evaluation Tool - The Four Exhibits
The AI Systems Evaluation Tool is the NAIC's structured information request developed under the Big Data and Artificial Intelligence (H) Working Group. It is the operational instrument that translates the four bulletin pillars into a packet a regulator can actually send and an insurer can actually answer. The Tool has four exhibits - and every L1-trained insurance professional in 2026 should know the exhibits by name.
Exhibit A - AI Quantification
Exhibit A is the inventory exhibit. It asks the insurer to enumerate every AI system in use across the operation, by line of business, by function (underwriting, pricing, claims, fraud, distribution, customer service), by vendor or in-house build, by deployment scope, and by risk tier. For a mid-size P&C carrier this is typically 25–60 systems by the time you count the Cytora triage, the Federato workbench, the Akur8 pricing model on each line, the Earnix dynamic engine, the Tractable photo flow, the CCC pipeline, the Shift fraud model, the Hi Marley SMS, the Five Sigma coverage summarization, the Coalition cyber engine, the agency-built chatbot, the Convr enrichment pull from FEMA NRI, the Hyperscience or Indico IDP, and the Roots Automation overlays. For an L&H carrier, add Munich Re, Swiss Re Magnum, RGA AURA NEXT, SCOR Velogica, the MIB / Rx triage, and the ECDIS-driven simplified-issue routing. The Exhibit A inventory is the artifact a Colorado examiner has been asking for in some form since July 2026 under Reg 10-1-1 and that the NAIC pilot now formalizes across the 12 states.
Exhibit B - Governance
Exhibit B is the governance exhibit - the bulletin's §4.1 made concrete. It asks for the AI program charter, the senior-management accountability statement, the committee structure and cadence, the policies (acceptable use, third-party AI, data, fairness, monitoring, incident response), the training records, the escalation paths, the board reporting cadence, and the incident-response runbook. A complete Exhibit B response is 30–80 pages for a mid-size carrier. The model documentation here is the AI Committee charter (membership, voting, cadence, authority), the AI policy stack (typically 5–9 distinct policies), and the senior-management report template that goes to the board's risk/audit/technology committee.
Exhibit C - High-Risk Systems
Exhibit C is the high-risk-systems deep-dive - the bulletin's §4.3 made concrete. For each AI system the insurer flags as "high-risk" in Exhibit A (typically 4–12 systems for a mid-size carrier; the threshold is usually consumer-adverse decisions and pricing impact), Exhibit C asks for a full model card: data sources, feature list with protected-class proxy analysis, training-data characterization, validation methodology, bias-test results (disparate-impact ratio, marginal effects on protected classes, segment-level performance), drift-monitoring plan with PSI / KS / AUC thresholds, explainability artifacts (SHAP, PDP, ALE), incident-response history, model version, retirement criteria. Exhibit C is where the rate-filing actuary, the L&H accelerated-UW owner, and the SIU fraud owner each spend serious weeks of work producing artifacts.
Exhibit D - AI Data
Exhibit D is the data exhibit - focused on the data feeding the AI systems. It asks for the data sources by category (first-party policy/claims data; third-party data such as LexisNexis MVR / Rx / public records; behavioral and telematics data; external geospatial data such as FEMA NRI / EagleView / RMS / Verisk AIR; consumer-report data subject to FCRA; ECDIS attributes under Colorado Reg 10-1-1); the lineage and refresh cadence; the data-quality controls; the privacy and security controls under GLBA Safeguards and HIPAA where applicable; the consent and notice posture under state privacy laws (CCPA/CPRA, CPA, VCDPA, CTDPA, UCPA, TIPA); and the sub-processor disclosure when third-party AI vendors access the data. Exhibit D is where GLBA, HIPAA, and the state privacy patchwork intersect with the bulletin - and where Lesson 3 of this chapter picks up the thread.
The 12-State Pilot Timeline and Who Is In It
The Evaluation Tool is being run as a pilot before broad adoption. The 12 pilot states are California, Colorado, Connecticut, Florida, Iowa, Louisiana, Maryland, Pennsylvania, Rhode Island, Vermont, Virginia, and Wisconsin. The kickoff window is Q1 2026 - sources conflict on the precise month. Fenwick and Swept AI cite a March 2, 2026 launch tied to a public NAIC announcement; NAIC's own Big Data and AI (H) Working Group calendar references January 2026 monthly state-coordination working calls under which the pilot operationally started. Both are consistent. The authoritative reference is the H Working Group's public calendar at content.naic.org. The pilot runs through September 2026, with re-exposure for public comment in September–October 2026 and adoption expected at the NAIC Fall National Meeting in November 2026.
Operationally, the 12 pilot states are sending Exhibit A / B / C / D information requests to carriers writing in their states across the pilot window. A national carrier with a writing footprint across all 12 states could receive 12 separate requests in 2026 - or, more commonly, a coordinated request through the lead-state regulator under the NAIC's lead-state framework, with the other 11 states observing. The 8:14 a.m. Dallas underwriter from the program's Insider Brief now has another deadline on her week: her chief actuary's team is producing the Exhibit C model card on the Akur8 commercial-property pricing GLM, the Federato appetite-scoring model, and the Cytora triage model - and her UW comments feed the variable-attribution narrative for each.
The non-pilot states are watching. Texas, Florida (Florida is a pilot state, but TDI / OIR is watching adjacent), Arizona, Georgia, North Carolina, Ohio, New Jersey, Massachusetts - none in the pilot - are tracking the pilot and are expected to formally adopt the bulletin (or their own variant) after the November 2026 NAIC Fall National Meeting adoption vote. The realistic 2027 forecast is convergence: the Evaluation Tool becomes the de facto national information request, with state-specific overlays from Colorado, NY DFS, Connecticut, and a handful of others.
Mapping the Bulletin to the Artifacts You Already Touch
The bulletin is not abstract. It maps onto artifacts that already exist on your desk - and the bulletin's effect is to require that those artifacts be produced in a particular form and stored in a particular place. The L1 learner should be able to point at each of these.
The algorithm inventory. A spreadsheet or registry tool listing every AI system in use. Columns include model name, version, vendor, internal owner, function, line of business, risk tier, last-tested date, last bias-test result, drift-monitoring status, deployment scope, NAIC §4 reason-code reference, and decommission date if applicable. Colorado Reg 10-1-1 requires this for life insurance carriers and as of October 15, 2025, also for private-passenger auto and health-benefit plans. The Evaluation Tool Exhibit A is the NAIC-aligned version. Monitaur, Credo AI, and Trustible publish working templates; most large carriers maintain the inventory in a GRC tool (ServiceNow, Archer, OneTrust) with a custom AI module.
The model card. A 4–12 page document per high-risk model with the structure: data sources and lineage, feature list with proxy-class analysis, training-data characterization, model methodology (GLM, GBM, deep-learning), validation methodology (out-of-time, out-of-sample, cross-validation), bias test (disparate-impact ratio across protected classes, marginal effects, segment-level performance), drift-monitoring plan, explainability artifacts (SHAP for global feature importance, PDP for marginal effects, ALE for interactions), incident history, version control, retirement criteria. The model card is the Exhibit C input. Akur8's transparent-GLM workflow auto-produces most of it; Earnix's documentation module produces a comparable artifact; in-house models built in R or Python need a custom template.
The third-party AI contract addendum. A standardized addendum the carrier requires from every third-party AI vendor. Required components: data-handling and sub-processor schedule, model documentation rights, audit rights (annual onsite or remote), model-update notification SLA, fairness-testing attestation, security attestation (SOC 2 Type II, ISO 27001), incident-response notification SLA, BAA chain where HIPAA-eligible, exit clauses with data return / destruction, and indemnity language. The bulletin's §4.2 is what this addendum is satisfying. Vendor procurement at a carrier of any size has been rewriting these addenda through 2025–2026; the addendum is now a 12–30 page document at most large carriers.
The testing schedule. A calendar of bias tests, drift checks, calibration reviews, and champion-challenger refreshes per AI system. High-risk systems get tested quarterly; medium-risk semiannually; low-risk annually. The testing schedule lives in the GRC tool and is the working artifact behind Exhibit C. The cadence for a Colorado-writing life carrier is set by Reg 10-1-1; for an auto carrier writing in Colorado, by the October 2025 expansion of Reg 10-1-1; for everyone else, by the carrier's risk-based judgment under bulletin §4.1.
The complaint-routing log. A log of consumer complaints that touch an AI-driven decision, with the resolution path documented, the human reviewer named, and the reason code reaffirmed (or, if the original AI output was wrong, the corrective action taken). Bulletin §4.4 contemplates this. Most carriers route AI-touched complaints through the same DOI-complaint-handling team they already use, with an AI flag added to the case file.
What the Bulletin Does Not Say - And Why That Matters
The bulletin is deliberately principles-based. It does not specify a disparate-impact-ratio threshold (the "80% rule" from federal EEOC employment-discrimination law is sometimes invoked as a benchmark but is not in the bulletin). It does not name specific bias-testing methodologies. It does not list which models are "high-risk" - that classification is left to the insurer subject to examiner challenge. It does not name specific vendors. It does not specify the frequency of testing in absolute terms. This is intentional: the NAIC wanted a framework durable enough to survive five years of AI evolution. The cost is that examiners apply judgment, and judgment varies state-to-state and examiner-to-examiner.
What fills the gap is a combination of (a) state-specific overlays that do specify thresholds and cadences - Colorado Reg 10-1-1 being the most prescriptive, NY DFS Circular Letter 2024-7 the second most - and (b) emerging carrier-industry working consensus on what "good" looks like for a §4.3 testing program. The next lesson walks through the state DOI bulletins that fill the prescriptive gap. Lesson 3 of this chapter walks through the other federal and state frameworks (GLBA, HIPAA, FCRA, MHPAEA, state privacy) that overlap with the bulletin.
What This Means for the People on the Desk
For the underwriter: every adverse decision - a decline, a knockout, a quote-with-restriction - needs a §4-compliant reason code that survives examiner review. The Cytora or Federato workbench should be configured to produce reason codes that map to the carrier's appetite guide, not to opaque scores. The variable-attribution narrative is part of the file.
For the adjuster: every AI-influenced claim decision - a fraud referral, a triage assignment, a Tractable estimate that drives a reserve recommendation, a Five Sigma coverage summary that drives an ROR - needs a human-reviewable reason that does not reduce to the score. The SIU referral memo names the human-reviewable facts separate from the Shift score. The ROR cites the policy form by edition and the controlling exclusion, not the AI's summary.
For the producer: every AI-generated client communication needs producer review under the agency's written supervisory procedure. The agency-built chatbot needs a disclaimer that survives a misrepresentation claim. The four-carrier comparison drafted by an LLM needs producer verification before the client sees it.
For the actuary: every rate filing built on an AI model needs a SERFF-ready package that includes the model card, the bias-test exhibit, the variable-selection narrative, and the proxy-test memo where the carrier writes in NY. The actuarial certification under ASOP 56 (Modeling), ASOP 41 (Communications), and ASOP 23 (Data Quality) names the AI involvement explicitly.
Key Takeaways
- The NAIC Model Bulletin (December 2023) sits on top of Model 880 §4 and extends unfair-discrimination analysis to algorithmic outcomes. 25+ states adopted some form by mid-2026. The four pillars - AI Program, Third-Party AI, Testing and Validation, Documentation - are §4.1, §4.2, §4.3, §4.4 respectively, and every examiner uses these four names.
- The AI Systems Evaluation Tool has four exhibits: A (Quantification / inventory), B (Governance), C (High-Risk Systems / model cards), D (AI Data). A mid-size P&C carrier inventories 25–60 systems; an L&H carrier adds Munich Re, Swiss Re Magnum, RGA AURA NEXT, SCOR Velogica, MIB/Rx triage, and ECDIS routing.
- The 12 pilot states are CA, CO, CT, FL, IA, LA, MD, PA, RI, VT, VA, WI. Q1 2026 kickoff (Jan or Mar per source - both consistent with the H Working Group calendar), September 2026 close, September–October 2026 re-exposure, November 2026 NAIC Fall National Meeting adoption.
- The bulletin is principles-based - no fixed disparate-impact-ratio threshold, no prescriptive testing cadence. State-specific overlays (Colorado Reg 10-1-1, NY DFS Circular Letter 2024-7, Connecticut MC-25-8, Nevada Bulletin 24-006) fill the prescriptive gap. Federal frameworks (GLBA, HIPAA, FCRA, MHPAEA) overlay on top.
- Five working artifacts implement the bulletin on the desk: the algorithm inventory (Exhibit A), the model card (Exhibit C), the third-party AI contract addendum (§4.2), the testing schedule (§4.3), and the complaint-routing log (§4.4). If they do not exist when the examiner asks, the carrier is exposed.
- Delegation to a third-party vendor does not transfer §4 liability. Cytora, Federato, Akur8, Earnix, Tractable, CCC, Shift, Hi Marley, Five Sigma, Coalition, Munich Re, Swiss Re Magnum, RGA AURA NEXT, SCOR Velogica - if their model produces a §4 violation in your book, your carrier owns it. The contract addendum has to satisfy §4.2.
- "The model said so" is not a reason code under §4. Every adverse AI-influenced decision - UW decline, knockout, claim referral, rate impact - needs a human-reviewable reason that maps to documented criteria. The variable-attribution narrative, the proxy-test memo, and the reason-code framework are the file.
- Adoption expected at the NAIC Fall National Meeting in November 2026. Non-pilot states (Texas, Arizona, Georgia, New Jersey, Massachusetts and others) are expected to follow in 2027. The Evaluation Tool becomes the de facto national information request, with state overlays from Colorado, NY DFS, Connecticut, and others on top.
Skill.re