Bias, Proxy Discrimination, and Disparate Impact in Insurance AI
A fraud-score cluster correlated with a protected class is a §4 violation under the NAIC Model Bulletin even if no human ever consciously saw the cluster and even if the model never saw a protected-class variable directly. That single sentence is the operational reality of insurance AI bias in 2026 - and the reason every underwriter, adjuster, producer, and actuary on the desk needs to be able to identify proxy discrimination, calculate disparate impact, and produce the proxy-test memo and the bias-test exhibit that survive NY DFS Circular Letter 2024-7, Colorado SB 21-169, and a market-conduct exam. ZIP code is a problem because it correlates with race in U.S. census geography. Credit-based insurance scoring is fraught because it correlates with race-and-income indirectly. Vehicle make and model correlates with both income and ethnicity. Occupation correlates with race, gender, and age. Bayesian Improved Surname Geocoding (BISG) is the defensible-but-careful proxy-attribution technique. The disparate-impact ratio under the EEOC 80% rule is the informal benchmark. The marginal-effects analysis is the variable-attribution narrative. This lesson is the map.
Why ZIP Code Is the Canonical Proxy Problem
U.S. residential geography is highly segregated by race, ethnicity, and income - a legacy of redlining, racially restrictive covenants, the Federal Housing Administration's historical mortgage policies, and decades of segregation-perpetuating zoning. Census-block-group and ZCTA (ZIP Code Tabulation Area) statistics confirm that American ZIP codes are heavily race-correlated; the Home Mortgage Disclosure Act (HMDA) data and the Census Bureau's American Community Survey are the canonical references showing the correlation directly. The consequence for insurance AI: any model feature that ingests ZIP code, ZIP-derived geography, territory codes, census-block geography, or any geographic proxy is implicitly ingesting race-correlated signal - even if the model is "just" pricing wind exposure or evaluating cat aggregation.
The defensive position carriers have historically taken is "ZIP code is actuarially necessary for cat-exposure pricing and accident-frequency rating." That defense survives in courts and rate filings in 2026, but it survives only when accompanied by quantitative documentation: the disparate-impact ratio analysis, the marginal-effects narrative showing the geographic feature's predictive lift, the substitution analysis showing no equivalent non-correlated variable produces comparable lift, and the proxy-test memo documenting the carrier's methodology. NY DFS Circular Letter 2024-7 essentially requires this documentation explicitly. Colorado Reg 10-1-1's compliance report under SB 21-169 quantifies the disparate-impact ratio. Carriers that file SERFF rate increases citing ZIP-driven territory factors without the documentation face examiner challenge.
The operational test is simple: hold every other variable constant, change the ZIP code from a majority-white ZCTA to a similarly situated majority-Black ZCTA with comparable cat exposure and accident frequency, and observe whether the model's output (the rate, the underwriting decision, the fraud score) materially changes. If the output changes meaningfully without a corresponding change in actuarial risk, the ZIP feature is producing proxy-class outcome. If the change is fully explained by actuarial risk, the variable is necessary. The carrier's job is to document the analysis and the conclusion.
Credit-Based Insurance Scoring - The Fraught Variable
Credit-based insurance scoring (CBIS) - using elements of a consumer's credit history to predict insurance loss frequency - has been one of the most-debated variables in personal-lines pricing since the late 1990s. The actuarial defense is well-documented: credit-based scores have meaningful predictive lift on personal-auto loss frequency, supported by extensive industry studies (the 2007 FTC Report to Congress; multiple academic and consulting analyses since). The proxy-class problem is also well-documented: credit attributes correlate with race and income through structural causes (historical employment discrimination, wealth gaps, geographic and educational disparities), and the CBIS-driven rate disparity falls disproportionately on minority and low-income consumers.
The state regulatory landscape on CBIS has fractured. California prohibits CBIS in personal-auto rating under the Proposition 103 framework and CDI rulemaking. Washington banned CBIS for personal auto and several other lines (effective 2021); the state has been an enforcement leader. Maryland restricts CBIS for personal-lines rating with specific limits on use. Massachusetts prohibits CBIS in personal auto rating. Hawaii prohibits CBIS in personal-auto rating. Other states (most of the country) permit CBIS subject to disparate-impact-testing expectations. New York under DFS 2024-7 applies the proxy-test framework to CBIS; Colorado Reg 10-1-1 includes CBIS in the quantitative bias-testing scope; Connecticut MC-25-8 treats CBIS as a high-attention variable.
The carrier-side operational posture in 2026 is to (a) maintain the actuarial defense for CBIS where it's used, including the predictive-lift documentation and the substitution-analysis narrative; (b) run the disparate-impact ratio analysis across BISG-attributed cohorts as part of the model card; (c) be prepared to drop or modify CBIS in jurisdictions that restrict or prohibit; and (d) develop substitute variables where CBIS is restricted, while documenting that the substitute does not itself produce equivalent proxy-class outcomes. The compliance overhead on CBIS is one of the highest of any single variable, and many carriers are reassessing whether the predictive lift justifies the operational burden.
Colorado SB 21-169 Quantitative Bias Testing - The Most Prescriptive Protocol
Colorado SB 21-169, passed in 2021 and implemented through Reg 10-1-1, is the most prescriptive state framework for quantitative bias testing in insurance. The statute prohibits unfair discrimination in any insurance practice through the use of ECDIS or AI/predictive models. The implementing regulation requires a testing protocol with these elements.
The Testing Protocol
Step 1 - Identify the protected classes in scope: race, color, national origin, ancestry, religion, sex (including pregnancy and gender), sexual orientation, gender identity, disability, age, marital status, and other classes per Colorado statute. Step 2 - Construct cohort attributions for each protected class where direct data is unavailable. For race and ethnicity, BISG is the standard methodology; the carrier documents the BISG bucket assignments. Step 3 - Compute the disparate-impact ratio of the model's adverse decisions across cohorts. The disparate-impact ratio is calculated as (adverse rate for minority cohort) / (adverse rate for reference cohort); the federal EEOC 80% rule (the "4/5ths rule") is the informal threshold - a ratio below 0.80 indicates potentially actionable disparate impact requiring justification or remediation. Step 4 - Compute the marginal effects of each model variable across cohorts. The partial-dependence plot and Shapley-value decomposition are the standard tools. Step 5 - For variables producing meaningful disparate impact, justify the variable as actuarially necessary and non-substitutable, or remove. Step 6 - Document the methodology, the results, the justification, and the remediation in the proxy-test memo and the SB 21-169 compliance exhibit. Step 7 - Refresh the analysis on a defined cadence (quarterly for high-risk models). Step 8 - File the relevant artifacts in the annual Colorado compliance report (first due July 1, 2026).
The Disparate-Impact Ratio Calculation Worked
Worked example. A personal-auto pricing GLM produces an "underwriting decline" outcome for 12.4% of applicants across a Colorado book. BISG attribution buckets applicants into reference (majority non-Hispanic white) and minority cohorts. The decline rate for the reference cohort is 11.1%; the decline rate for the Black cohort is 16.3%; the decline rate for the Hispanic cohort is 14.7%. The disparate-impact ratio Black-vs-reference is 11.1 / 16.3 = 0.68 (below the 0.80 informal threshold). The Hispanic-vs-reference ratio is 11.1 / 14.7 = 0.76 (also below 0.80). Both ratios indicate disparate impact requiring justification or remediation. The carrier runs the marginal-effects analysis identifying which features drive the disparity (typically ZIP, vehicle make/model, occupation, in some states CBIS). The carrier then documents the actuarial necessity of each contributing feature, runs substitution analysis to see if non-correlated alternatives exist, and either justifies retention with documentation or removes/replaces the feature. The compliance exhibit captures every step.
BISG - Bayesian Improved Surname Geocoding - The Defensible-But-Careful Technique
BISG was developed by RAND for the federal government in 2009 and is the most statistically rigorous publicly available method for attributing race and ethnicity from surname and geography. The method combines a surname-based prior probability (from Census Bureau surname-frequency data) with a geographic likelihood (from Census block-group race composition) to produce a posterior probability distribution of race/ethnicity for each individual. The method does not assert that any individual is of a particular race; it produces a probabilistic attribution that, aggregated across a population, supports statistical bias testing.
BISG is the standard because (a) it's the most widely accepted method in fair-lending compliance under the CFPB and federal banking regulators, (b) it's publicly documented and reproducible, (c) it's the method NY DFS, Colorado, and most state DOIs have implicitly endorsed as the proxy-test methodology, and (d) it produces results that survive examiner review. BISG is also "careful" because it is probabilistic and can mis-attribute individuals - meaning the disparate-impact analysis is on aggregated cohort behavior, not individual decisions. Carriers using BISG document the methodology, the bucket-attribution thresholds, and the aggregation framework.
Alternatives to BISG exist (deterministic surname matching, geocoded census attribution alone, self-reported race where available) but none have BISG's combination of accuracy, defensibility, and regulatory acceptance. Carriers that depart from BISG for proxy attribution face additional examiner challenge and must document why the alternative is comparable or superior. The 2026 working consensus is that BISG is the proxy-test default; departures are exceptions.
Why a Fraud-Score Cluster Correlated With a Protected Class Is a §4 Violation
The most counterintuitive concept in insurance AI bias is that a model can violate §4 without anyone - neither the carrier, nor the model designer, nor any human reviewer - ever consciously discriminating, and even when the model has no protected-class variable in its training set. The Shift Technology fraud-score example from the Atlanta playbook is the canonical case.
The scenario: Shift's fraud model identifies a soft-tissue-injury claim cluster involving a third-party medical clinic that appears on multiple plaintiff-attorney-represented claims. The cluster is a fraud-investigation signal. The model has no race variable; the clinic location is in a specific ZIP cluster; the plaintiff-attorney representation pattern correlates with the clinic's neighborhood; the cluster's claimants are disproportionately of a protected class because of the ZIP-and-neighborhood correlation. The model produces SIU referrals at a disparate rate across protected classes. No human SIU investigator ever saw the cluster as a protected-class pattern; no SIU referral memo cites race. But the disparate referral rate is itself the §4 unfair-discrimination signal - and the carrier is exposed to a §4 finding regardless of intent.
The remediation discipline is the bias-test workflow applied to fraud models. The carrier runs the disparate-impact ratio on SIU referrals across BISG-attributed cohorts; documents the marginal effects of each fraud-score component; justifies or remediates features producing meaningful disparate impact; produces the proxy-test memo for the fraud model; integrates the analysis into the algorithm-inventory entry and the §4.3 testing schedule. Every adverse output of the fraud model - SIU referral, claim hold, additional investigation requirement - gets the same disparate-impact discipline as a pricing GLM. The Shift documentation, the carrier's contract addendum, and the bias-test exhibit all reference the analysis.
The NY DFS Proxy Test Operationally - A Worked Walkthrough
The NY DFS proxy test under Circular Letter 2024-7 is the de facto national methodology in 2026. Walked through operationally for a homeowners pricing GLM with 47 features.
Step 1 - Identify candidate proxy variables. From the feature set, flag ZIP code, ZIP-derived territory factor, occupation category, education-attainment level (if used), credit-based insurance score (if used), prior-carrier indicator, and any variable with documented or suspected protected-class correlation. The candidate list is typically 5–12 variables out of a 30–50 feature model.
Step 2 - Construct race/ethnicity attributions. Apply BISG using the policyholder's surname plus the property ZIP. Bucket the book into BISG cohort categories (Non-Hispanic White, Black, Hispanic, Asian, Other). The attribution is probabilistic; the bucket assignment uses a threshold (typically the highest-probability cohort if probability exceeds a defined floor).
Step 3 - Compute marginal effects. For each candidate variable, generate partial-dependence plots showing the variable's effect on the model output across the feature's range, separately for each BISG cohort. Identify where the marginal effects diverge across cohorts. The divergence is the proxy-class signal.
Step 4 - Compute disparate-impact ratios. For each adverse model outcome (decline, rate increase above threshold, restriction), compute the ratio of adverse rates between minority and reference cohorts. Apply the 0.80 informal threshold; investigate ratios below 0.80.
Step 5 - Justify or remediate. For each variable producing meaningful disparate impact, run substitution analysis: is there a non-correlated alternative variable producing comparable predictive lift? If yes, replace. If no, document the actuarial necessity - citing the loss-experience data, the actuarial study, the substitution-analysis results, and the actuarial certification.
Step 6 - Document. The proxy-test memo captures every step: the methodology, the data, the BISG bucketing, the marginal effects, the disparate-impact ratios, the substitution analysis, the justifications, the remediations. The memo is filed with the model card, referenced in the SERFF rate filing, and produced on examiner request.
Step 7 - Refresh quarterly. The proxy test is not a one-time exercise. The book composition shifts, model features drift, and the analysis must be refreshed on a defined cadence - quarterly for high-risk models per the working 2026 consensus.
Step 8 - Reference in SERFF. The rate-filing memorandum references the proxy-test methodology and the conclusion. The bias-testing exhibit attached to the filing reproduces key elements. The actuarial certification under ASOP 41 names the proxy-test work.
Why the Fraud Cluster, the Pricing GLM, and the Triage Score All Share the Same Discipline
The discipline is identical across model types because §4 applies to all adverse insurance outcomes regardless of which model produces them. A Cytora triage score that routes submissions disparately is a §4 risk; an Akur8 pricing GLM that produces disparate rate impacts is a §4 risk; a Shift fraud score that produces disparate SIU referrals is a §4 risk; a Munich Re accelerated-UW knockout model that produces disparate decline rates is a §4 risk; a Tractable photo-estimator that produces disparate severity outcomes is a §4 risk. Every adverse model output needs the disparate-impact discipline; every model needs the BISG-cohort analysis; every model needs the marginal-effects narrative; every model needs the substitution analysis; every model needs the proxy-test memo. The discipline is the bulletin's §4.3 testing pillar applied uniformly.
What This Means for the People on the Desk
For the underwriter: every adverse decision that touches a model output needs a §4 reason code documenting the variables driving the decision, separating protected-class proxies, and reproducing the marginal-effects narrative. The Cytora or Federato workbench's reason-code output should be cross-checked against the proxy-test memo before the decision letter goes out.
For the adjuster: every fraud referral, every claim hold, every coverage decision that touches a model output needs the disparate-impact discipline. The SIU referral memo cites the human-reviewable facts separate from the Shift score; the file note documents the AI involvement and the adjuster's independent verification.
For the producer: every AI-generated recommendation or product-suitability output needs review against the carrier's bias-testing posture. Agency-built chatbots making product recommendations need to be tested against disparate-impact across the producer's book.
For the actuary: every model produces a model card with the bias-test exhibit. The SERFF rate filing references the proxy-test methodology. The actuarial certification under ASOP 41/56 names the AI involvement and the bias-testing program. The Colorado SB 21-169 compliance exhibit, the NY DFS proxy-test memo, and the NAIC Exhibit C model card all reference the same underlying analysis.
Key Takeaways
- A fraud-score cluster correlated with a protected class is a §4 violation even if no human ever saw the cluster and the model never saw a protected-class variable. The disparate adverse-outcome rate is the violation regardless of intent. Shift Technology fraud-model SIU referrals, Cytora triage routing, Akur8 pricing GLMs, Tractable photo-estimating outputs - all subject to the same discipline.
- ZIP code is the canonical proxy problem. U.S. residential geography is heavily race-correlated per HMDA, Census ACS, and ZCTA-level analyses. Models using ZIP need the disparate-impact ratio analysis, the marginal-effects narrative, and the substitution analysis documented.
- Credit-based insurance scoring (CBIS) is restricted in CA (Prop 103), WA (2021), MA, HI, and limited in MD. Most other states permit subject to disparate-impact-testing expectations. NY DFS applies the proxy-test framework; Colorado requires quantitative testing under SB 21-169.
- Colorado SB 21-169 quantitative bias testing is the most prescriptive state protocol. Eight steps: protected-class scoping, BISG cohort construction, disparate-impact ratio computation, marginal-effects analysis, actuarial-necessity justification or remediation, documentation, quarterly refresh, annual compliance-report filing.
- The federal EEOC 80% rule (4/5ths rule) is the informal disparate-impact-ratio threshold. A ratio below 0.80 indicates potentially actionable disparate impact requiring justification or remediation. The NAIC bulletin does not name this threshold, but state DOI examiners and carriers default to it in practice.
- BISG is the defensible-but-careful proxy-attribution standard. Developed by RAND, used in CFPB fair-lending compliance, probabilistic rather than deterministic, and the methodology NY DFS / Colorado / state DOIs have implicitly endorsed. Departures from BISG face heightened examiner scrutiny.
- The NY DFS proxy test is the de facto national methodology. Eight operational steps: identify candidates, BISG attribution, marginal effects, disparate-impact ratios, substitution analysis, justify or remediate, document in proxy-test memo, quarterly refresh, reference in SERFF.
- The discipline applies uniformly to triage, pricing, fraud, accelerated UW, claims, and any model producing adverse outcomes. Every adverse output gets the same workflow. The bulletin's §4.3 testing pillar applies regardless of model type, line of business, or function. The carrier's posture is one program, one methodology, one set of artifacts across all models.
Skill.re