AI for Insurance Professionals
Strategic · M5 · lesson 5 of 24 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Assess AI Readiness for Your Insurance Operation - AM Best 2026 Framework and Evident AI Index
📖
now learning

Assess AI Readiness for Your Insurance Operation - AM Best 2026 Framework and Evident AI Index

15 min

AI readiness in a 2026 insurance operation is not a sentiment. It is a measurable posture across five axes - data, talent, governance, tooling, regulatory - that the rating agency, the reinsurance treaty broker, the DOI examiner, and the board's risk committee will all interrogate inside the next twelve months. AM Best's April 2026 Best's Special Report measured that 41% of rated US carriers already use AI in one or more core functions (underwriting, claims, pricing, fraud, distribution), and roughly 60% expect a one-to-three-year transformation horizon. The same report folded an AI readiness assessment into AM Best's Performance Assessment framework - not a standalone AI capability rating (that methodology does not yet exist) but a survey-and-readiness lens that already influences the issuer-credit-rating narrative on holding companies with material AI dependency. The Evident AI Insurance Index, published in March 2026, scored fifty global insurers across four pillars - talent, innovation, leadership, and transparency - and made the gap between top-decile carriers (Ping An, Allianz, Zurich, AXA, Progressive) and the median carrier publicly visible to every reinsurer and rating analyst on the planet. This lesson is the framework for scoring your operation honestly on a 1-5 scale across the five axes, the artifacts that prove each score, and the gap-closure plan that converts a score-2 readiness profile into a score-4 inside eighteen months without setting the loss ratio on fire.

Why Readiness Now - The April 2026 Best's Special Report

AM Best published the AI Capability and Readiness Special Report on April 14, 2026. Two numbers from that report drive every L4 conversation. First, 41% of US-rated carriers report using AI in at least one core function as of Q1 2026 - up from 28% in the 2024 baseline survey. Second, approximately 60% of respondents expect a one-to-three-year horizon for what they describe as "material AI-driven transformation" of either underwriting, claims, or pricing. Those numbers became the procurement justification for every CIO budget request that quarter and the analyst-call talking point for every fourth-quarter earnings narrative. They also became the AM Best methodology team's evidentiary base for folding an AI readiness assessment into the Performance Assessment framework, which the analyst team uses to calibrate issuer-credit-rating narrative on carriers whose strategy depends materially on AI.

The structural point matters more than the headline. AM Best did not publish a standalone AI capability rating methodology in 2026. The Performance Assessment is a forward-looking management-quality lens that has always considered enterprise risk management, capital management, and operating performance; the 2026 update added AI readiness as a sub-dimension under operating performance and enterprise risk. The carrier-side narrative therefore should be calibrated to the AM Best readiness-survey categories - data readiness, model governance, talent, third-party AI risk, regulatory compliance - and not to a hypothetical AI rating product. Carriers that write board memos referencing a "future AM Best AI rating" are calibrating to something that does not exist; carriers that write to the readiness survey are calibrating to what the analyst will actually ask about during the annual rating meeting.

The Evident AI Insurance Index is the second 2026 reference. Evident scored fifty global insurers across four pillars: AI talent (sourced from LinkedIn data, patent filings, academic publications), AI innovation (product launches, technology partnerships, R&D disclosure), AI leadership (executive sponsorship, board-level governance, public commitments), and AI transparency (public disclosure of model governance, fairness testing, audit posture). The top decile - Ping An, Allianz, Zurich, AXA, Progressive, Generali, Munich Re, Swiss Re, AIA, MetLife - pulled away from the median in 2026 in a way that became visible to capital markets and reinsurers. A US specialty carrier ranked in the bottom quartile of Evident will get a different treaty-renewal conversation in 2026 than a top-decile peer; the broker now references the index in the renewal narrative.

The Five-Axis Readiness Framework

Score each axis on a 1-5 scale. 1 is "no posture at all"; 3 is "industry median"; 5 is "top-decile Evident posture with documented evidence." The five axes are data, talent, governance, tooling, and regulatory readiness. Each has a definition, a scoring rubric, and the artifacts that prove the score.

Axis 1 - Data Readiness

Data readiness is the operational ability to feed AI models clean, lineage-tracked, schema-validated, latency-appropriate data. AM Best's 2026 framework calls out three components: data quality (completeness, accuracy, timeliness), data architecture (lakehouse vs. legacy AS/400 mainframe extract, the cycle time from policy admin system to feature store), and data governance (data steward roles, lineage documentation, retention policies that satisfy GLBA and state privacy). Score 1 carriers run AI on Excel pulls from a green-screen PAS export. Score 3 carriers have a Snowflake or Databricks lakehouse, daily syncs from Guidewire / Duck Creek / Sapiens, and a documented data catalog. Score 5 carriers operate near-real-time feature stores (Tecton, Feast), have lineage instrumented across the full submission-to-bind-to-claim lifecycle, and can produce a Colorado Reg 10-1-1 explainability artifact for any production model inside thirty minutes. Artifacts that prove the score: the data catalog, the lineage diagram, the freshness SLA dashboard, the data steward roster, the retention-policy memo.

Axis 2 - Talent Readiness

Talent readiness measures the human capital actually available to design, deploy, govern, and operate AI inside the carrier. Evident's index uses LinkedIn-sourced headcount, patent filings, and academic publications; AM Best uses a survey question on accountable executives and AI-credentialed staff. Score 1: no chief data officer, no dedicated MLOps function, AI work happens in the actuarial department's spare cycles. Score 3: a CDO or chief analytics officer exists, the actuarial team has six-to-twelve credentialed analysts with predictive-modeling exam credit (CAS Predictive Analytics, CAS Online Courses 1-2), and one or two data scientists report into pricing or claims. Score 5: a chief AI officer reports to the CEO, twenty-to-fifty data scientists and ML engineers operate across pricing, claims, and distribution, the talent pipeline includes The Institutes' AIAI designation for non-technical staff, and the board has a director with explicit AI risk oversight responsibility. Artifacts: org chart, designation roster (FCAS, FSA, AIAI, CPCU, AIC), retention and turnover statistics for AI/ML headcount, board-skills matrix.

Axis 3 - Governance Readiness

Governance readiness is the framework that satisfies NAIC Model Bulletin §4 (the governance section), the Colorado Reg 10-1-1 algorithm registry requirement, and the AM Best Performance Assessment management-quality lens. Score 1: no AI policy, no algorithm inventory, no model risk management function distinct from internal audit. Score 3: an AI policy exists and references §4, an algorithm inventory is maintained in a spreadsheet, a model risk management group reviews material models annually, and a quarterly AI committee meets with mixed legal/actuarial/UW/IT membership. Score 5: the algorithm inventory lives in a registry tool (Credo AI, Holistic AI, Fiddler, or a homegrown registry), every production model has a model card with bias testing, drift monitoring, and explainability artifacts, the AI committee meets monthly with documented escalation authority, the chief risk officer signs the Exhibit B governance memo annually, and incident-response runbooks exist for the five most common AI failure modes (adverse-decision error, drift event, vendor outage, data exposure, MHPAEA NQTL violation in L&H lines).

Axis 4 - Tooling Readiness

Tooling readiness measures the AI stack actually deployed in production - not the slide-deck stack. Score 1: a single Tractable license for auto claims and nothing else. Score 3: production deployments across three-to-five vendors (Hyperscience or Indico for OCR, Cytora or Convr for triage, Akur8 or Earnix for pricing, Tractable or CCC for estimating, Shift for fraud), each with an SOC 2 Type 2 in file, monitoring instrumented, and an incident-response clause in the contract. Score 5: a coordinated platform play (Federato RiskOps as the underwriter workbench, Guidewire Predict as the embedded pricing layer, Five Sigma or Hi Marley as the claims workbench, ServiceNow GRC or Archer as the governance backbone), with a vendor management office, quarterly performance reviews against KPIs, and exit clauses tested annually. Artifacts: vendor schedule, SOC 2 file, contract addenda, performance scorecards.

Axis 5 - Regulatory Readiness

Regulatory readiness measures the operation's capacity to respond to the NAIC AI Systems Evaluation Tool, the Colorado Reg 10-1-1 compliance report (due July 1, 2026 for life insurers using external consumer data and algorithms), the New York DFS Circular Letter 7, the Connecticut DOI 2022 notice and 2024 update, the Nevada SB 234 AI-disclosure framework, the California DOI 2024 bulletin on AI in claims, and the GLBA / HIPAA / state-privacy framework. Score 1: no readiness work done, the operation is exposed to a market-conduct exam finding on any axis. Score 3: the Colorado Reg 10-1-1 compliance report draft exists, the NAIC AISET response packet (Exhibits A-D) is in draft, FCRA adverse-action language is reviewed, the GLBA Safeguards Rule artifacts are current. Score 5: every state-DOI exam request can be answered inside thirty days from existing artifacts, the AISET packet refreshes quarterly, MHPAEA NQTL testing is documented for L&H operations, the regulatory affairs function has named contacts at every domiciliary and major-write state DOI.

Scoring Honestly - The Readiness Self-Assessment

Carriers consistently overscore themselves on talent and governance and underscore themselves on data and tooling. The honest scoring discipline runs as follows. For each axis, write the definition of score 5 from the rubric above. Then ask three operational questions: (1) what artifact would you produce in an audit to prove the score, (2) when was that artifact last refreshed, and (3) who signed it. If the answer to any of the three is "we don't have one" or "I'm not sure," the score is at least one point lower than the initial self-assessment. Repeat for all five axes. A 2026 mid-market specialty carrier with $400M of premium typically lands at data 2, talent 3, governance 2, tooling 3, regulatory 2 - composite 12/25 - when scored honestly the first time. Top-decile Evident peers score 22-24/25.

The composite score matters less than the gap pattern. A carrier scoring data 2, governance 2, regulatory 2 has a data-and-control gap; the path forward is data architecture and governance build-out before any new tooling. A carrier scoring data 4, tooling 4, governance 2, regulatory 2 has a control gap on top of a strong technical base; the path forward is governance and regulatory artifacts, fast, before the next exam. Identifying the pattern is the readiness assessment's actual output - not the composite number.

Legacy Integration and Cyber - The AM Best Sub-Stack

AM Best's 2026 readiness framework explicitly calls out two substack items most carriers underweight: legacy integration and cyber. Legacy integration is the operational reality that a 2026 carrier's policy admin system is some combination of Guidewire (modern), Duck Creek (modern), Sapiens (modern-ish), and one or more legacy systems - mainframe-era AS/400, Wynsure, Point-IN-Point, custom COBOL stacks - that hold the in-force book the holding company depends on for its B+ ICR rating. AI deployments that bolt onto the modern stack but cannot read the legacy stack capture half the operation at best; the analyst will ask which book of business each AI model actually covers. Score the operation on what percentage of in-force premium is reachable by the AI stack, not on what percentage of net-new business is reachable.

Cyber is the second sub-stack. AI introduces new attack surfaces: prompt-injection vectors against LLM-based claims chat (Hi Marley, Five Sigma's Aspen agent), training-data exfiltration risk on third-party models, vendor-side breach exposure under GLBA Safeguards Rule §314.4(f) third-party oversight, and model-inversion attacks on pricing models trained on PII. The 2026 AM Best survey calls out cyber readiness as a precondition for AI readiness; a carrier with a weak cyber posture cannot honestly score above 3 on tooling regardless of the vendor count. Artifacts: SOC 2 from every AI vendor, the carrier's own SOC 2 or equivalent, penetration test reports, incident response runbook with AI-specific scenarios, cyber insurance coverage that explicitly addresses AI-related incidents.

Closing the Gap - The Eighteen-Month Readiness Roadmap

Score-2 to score-4 inside eighteen months is achievable for a $400M-$2B specialty carrier with a CEO mandate and a $4M-$12M committed budget. The path is sequenced and dependency-aware.

Months 0-3: data architecture foundation. Stand up the lakehouse (Snowflake, Databricks, or Microsoft Fabric depending on the carrier's existing footprint). Document data lineage from PAS to feature store. Hire the chief data officer if there isn't one. Score moves from data 2 to data 3.

Months 3-6: governance core. Draft the AI policy aligned to NAIC §4. Stand up the algorithm inventory in a registry tool. Convene the AI committee with monthly cadence. Draft Exhibit B governance memo. Build the FCRA adverse-action workflow if not already in place. Score moves from governance 2 to governance 3.

Months 6-12: tooling depth. Replace point solutions with platform plays where the TCO math supports it. Federato or Cytora for the underwriter workbench. Guidewire Predict or Akur8 for pricing. Five Sigma or Hi Marley for claims workbench. Negotiate the contracts with NAIC §4 third-party AI conformance, exit clauses, sub-processor disclosure, model-update notification, audit rights, BAA/DPA where applicable. Score moves from tooling 3 to tooling 4.

Months 9-15: talent depth. Hire two-to-six data scientists or ML engineers depending on scale. Roll out The Institutes' AIAI designation to all underwriters, claims adjusters, and producers in scope. Add a director-level AI risk seat on the board. Document the designation roster. Score moves from talent 3 to talent 4.

Months 12-18: regulatory and incident-response polish. Finalize the AISET response packet (Exhibits A-D). Finalize the Colorado Reg 10-1-1 compliance report. Run a tabletop exercise on the five AI incident scenarios. Refresh the SOC 2 and the cyber insurance. Score moves from regulatory 2 to regulatory 4.

Composite score at end of eighteen months: data 3, talent 4, governance 4, tooling 4, regulatory 4 - 19/25. That is the inflection point at which AM Best's Performance Assessment narrative starts characterizing the carrier as a "strong AI posture" peer rather than a "developing" peer, and the reinsurance treaty broker can credibly walk the analyst's narrative into the renewal pitch.

What the Board and the CRO Actually Ask

The readiness assessment outputs three documents the CRO and the board need: the composite scorecard (five axes, 1-5, with artifacts), the gap-closure roadmap (eighteen months, sequenced, budgeted), and the peer-benchmark exhibit (where the carrier ranks against Evident's index and against the AM Best survey median). Board-grade language anchors to outcomes, not activity. The right framing is "we are score 12 today, we will be score 19 in eighteen months, we are committing $7M over that period, the expected combined-ratio improvement is 1.8-2.4 points by month 24 driven by underwriting and claims AI deployments, and the AM Best rating analyst will see a measurable change in Performance Assessment narrative by the FY27 rating meeting." The wrong framing - "we're investing in AI" - fails at the first board-level follow-up question.

The CRO's specific questions are tighter: which models are in production today, what are the failure modes, what is the worst-case combined-ratio impact from a single model failure, what is the third-party concentration risk if Cytora or Akur8 or Tractable went offline tomorrow, what is the data-exposure risk under GLBA and state privacy, what is the unlicensed-practice or bad-faith liability exposure on AI-assisted decisions, and what does the incident-response runbook actually say. The readiness assessment must answer each of those by axis-and-artifact, not by aspiration.

Key Takeaways

  • AM Best's April 2026 Best's Special Report measured 41% of US-rated carriers using AI in core functions and ~60% expecting one-to-three-year transformation. The report folded an AI readiness assessment into the Performance Assessment framework. AM Best does not yet publish a standalone AI capability rating methodology; calibrate carrier narrative to the readiness survey, not to a non-existent rating product.
  • The Evident AI Insurance Index (March 2026) scores fifty global insurers across talent, innovation, leadership, and transparency. Top decile (Ping An, Allianz, Zurich, AXA, Progressive, Generali, Munich Re, Swiss Re, AIA, MetLife) pulled away from the median in a way visible to reinsurers and capital markets.
  • Score on five axes 1-5: data, talent, governance, tooling, regulatory. A 2026 mid-market specialty carrier typically scores 12/25 honestly the first time; top-decile peers score 22-24/25. The gap pattern matters more than the composite - data-and-control gaps require different sequencing than control-only gaps.
  • Legacy integration and cyber are the two AM Best sub-stack items most carriers underweight. Score reachable in-force premium, not net-new business reachable. Cyber posture is a precondition for tooling readiness above 3 - AI introduces prompt-injection, training-data exfiltration, vendor-side GLBA exposure, and model-inversion vectors.
  • Score-2 to score-4 in eighteen months is achievable for a $400M-$2B specialty carrier with $4M-$12M committed and a CEO mandate. Sequence: data architecture (0-3) → governance core (3-6) → tooling depth (6-12) → talent depth (9-15) → regulatory polish (12-18).
  • Honest scoring requires three operational questions per axis: what artifact proves the score, when was it last refreshed, who signed it. "We don't have one" or "I'm not sure" drops the score by one point. Carriers consistently overscore on talent and governance and underscore on data and tooling.
  • Board-grade output is three documents: composite scorecard, eighteen-month roadmap with budget, peer-benchmark exhibit against Evident and AM Best. Frame the readiness conversation in combined-ratio impact and Performance Assessment narrative change, not "we're investing in AI."
  • The readiness assessment is the prerequisite for every subsequent L4 deliverable - the roadmap (Ch1-2), the prioritization (Ch1-3), the vendor evaluation (Ch2), and the governance program (Ch5). Skip it and the roadmap is aspirational; do it and the roadmap is enforceable.