AI for Insurance Professionals
Aware · M12 · lesson 12 of 15 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
What AI Is and Isn't for an Insurance Professional
📖
now learning

What AI Is and Isn't for an Insurance Professional

15 min

It is 8:14 a.m. on a Tuesday in Dallas. A commercial property submission just hit the underwriter's inbox - ACORD 125 plus ACORD 140, a Schedule of Values with 47 buildings at $182M TIV, a five-year hard-copy loss run scanned at 200dpi with one column bled into the margin, a COPE narrative that says "mostly masonry, some frame," and a one-line broker note: "need a quote by Friday for 6/1." Forty-nine other submissions are stacked behind it. The 24-hour acknowledgment SLA started fourteen minutes ago. On the underwriter's screen, three pieces of software are humming. The first is a rating engine - deterministic, rule-bound, produces a class-plan premium from filed factors. The second is a GLM the pricing actuary built in Akur8 and filed in SERFF - predictive, statistical, returns a rate indication with a confidence band. The third is a large language model the broker used to draft the renewal narrative - generative, probabilistic, writes text. All three are called "AI" in vendor decks and DOI bulletins, and all three are doing radically different work. The April 2026 AM Best Special Report found 41% of carriers using AI in core functions and 60% expecting transformation within one to three years; the binding constraint, the report said, is governance - and governance starts with knowing which kind of AI is in which seat. This lesson is the map: deterministic vs. predictive vs. generative, on the ACORD 125, the FNOL, the Schedule of Values, and the Schedule P reserve triangle. By the end of it, you can walk onto any insurance floor and name which kind of math is making which decision.

Three Machines, One Word

The word "AI" in 2026 insurance covers three categorically different machines. They are not on a spectrum; they are different species. Underwriters, adjusters, producers, and actuaries who fail to separate them end up arguing with vendors and DOI examiners about the wrong thing.

Deterministic rules engine. The carrier's current rating engine - the one that runs a class-plan premium out of filed factors - is deterministic. Same inputs produce same outputs, every time, on every machine, forever. The logic is hand-written by a pricing analyst, encoded in IF-THEN statements (or in a rules table embedded in PolicyCenter, Duck Creek, Sapiens, or Majesco), and traceable line-by-line. When a Texas DOI examiner asks "why did this risk get this rate?", the answer is the filed manual page, the territory code, the class code, and the multiplicative factors. There is no probability and no learning. The rating engine that produced a $48,200 commercial property premium on the Dallas 47-building SOV is deterministic - every coefficient comes from the SERFF-approved rate filing and the manual page on the chief underwriter's shelf.

Predictive model. The Akur8 GLM that suggested a +12% territory factor on the Tarrant County buildings is predictive. It was fit to historical claims data - frequencies and severities by class, territory, building age, roof type, sprinkler status, COPE attributes - and it outputs a number that estimates the expected loss cost for a future risk. The output is statistical, not deterministic; the same input run on the same model produces the same number, but the model itself was learned from data rather than written by a human. Predictive models in insurance run the SERFF rate filings, the Munich Re accelerated-UW knockouts on Swiss Re Magnum and RGA AURA NEXT, the Tractable photo damage estimates, the Shift Technology fraud scores, and the Federato RiskOps triage rankings. They have AUC, KS, PSI, calibration, and SHAP - the metrics a credentialed actuary or a model-risk-management team uses to defend them. They do not write text. They produce a score, a class, or a number.

Generative model. The LLM the broker used to draft the renewal narrative - and the one the underwriter is about to use to summarize the 47-building SOV into a one-page UW memo - is generative. It produces tokens one at a time against a context window, sampled probabilistically from a distribution learned across trillions of words of training text. Same prompt, different runs, often different output. There is no AUC. There is no calibration curve in the actuarial sense. There is a system prompt, a temperature setting, a context window length, and a tendency to hallucinate ISO endorsement numbers that do not exist when pressed for citations. Generative models in 2026 insurance draft renewal stewardship narratives in Send and Outmarket, summarize FNOL transcripts in Five Sigma and Hi Marley, parse COPE narratives in Convr, write Reservation-of-Rights letter skeletons, and generate cover emails for surplus-lines submissions to Amwins and RT Specialty. They are not making the pricing decision. They are making the writing faster.

The category confusion is the source of most failures. A producer who treats a generative model like a rating engine ("the AI quoted $48,200") is in for an E&O claim. An underwriter who treats a predictive model like a deterministic engine ("the Akur8 number is the rate") is in for a SERFF objection. A DOI examiner who treats all three as equivalent is in for a productive conversation that finally separates them - which is exactly what NAIC Model Bulletin §4 was written to force.

How Each Shows Up on an ACORD 125

Trace the Dallas 8:14 a.m. submission through the three machines and the categorical differences land hard.

The ACORD 125 (Commercial Insurance Application) and the ACORD 140 (Property Section) arrive as PDFs. Hyperscience's Hypercell intelligent document processing - the IDP layer running Claude on Bedrock at the carrier's intake hub - extracts the named insured, the FEIN, the mailing address, the loss payee, the building locations, the year built, the construction code, the protection class, and the requested limits. That extraction is a predictive task: a model trained on millions of insurance documents, scoring each character cluster against a probability distribution, with a confidence number attached. Indico Data reports 99%+ extraction accuracy on the canonical ACORD set; Hyperscience publishes 99.5% accuracy with 98.0% automation. Below the confidence threshold, the document routes to a human reviewer. The IDP is not generating anything. It is reading.

The extracted record hits Cytora's triage queue. Cytora's Autopilot - the agentic capability launched post-Applied Systems acquisition - scores the submission against the carrier's appetite guide: line of business, class, territory, TIV, loss history, broker hit ratio. Federato RiskOps loads the same record into the underwriter's workbench and shows the open-portfolio exposure on a heatmap - the Tier-1 wind aggregate is 78% consumed against the annual budget, three buildings exceed the $25M single-risk treaty limit and need facultative, three are six-story wood-frame habitational that the appetite guide refuses. Each of those decisions is predictive plus rule-bound: a predictive appetite score combined with hard treaty constraints and hard appetite knockouts. There is no text being written. There is a triage ranking, a knockout list, and a portfolio impact view.

The underwriter clicks into the SOV. Convr's Risk 360 has already pulled FEMA National Risk Index data on the wind, hail, and tornado return periods for the Tarrant County zip codes, EagleView's aerial imagery has returned the roof age and material on the 29 locations where it could resolve the image, and the COPE narrative has been parsed into a structured field schema - Construction Class 1-6, Occupancy code, Protection class, Exposure flags. The Akur8 GLM ingests the structured features and produces a rate indication: a base loss cost, territory factor, COPE adjustments, roof-age factor, sprinkler credit. Each coefficient is documented for the SERFF audit trail. The model card lives in the algorithm inventory the Colorado Reg 10-1-1 compliance team maintains. This is the predictive layer. It produces a number, not a sentence.

The underwriter then drafts the quote-with-restriction memo. The carrier's enterprise LLM - Microsoft Copilot for M365 or Azure OpenAI behind the carrier's data-residency contract - takes the structured submission record, the appetite-decision reasons, the treaty-constraint flags, the Akur8 rate indication, and the named-storm-hours-clause language from the cat-XOL treaty, and drafts the memo: $50K AOP deductible, 5% Named Storm deductible, co-insurance schedule, roof-age exclusion on the 18 locations with missing roof data, three buildings declined on appetite grounds with the documented reason codes. That draft is the generative layer. The underwriter reads it, fixes the two paragraphs where the LLM invented a coverage form edition that does not exist, signs the memo, and sends it to the chief underwriter. The generative model did not price the risk. It wrote the letter that documents the pricing the predictive layer indicated and the human approved.

How Each Shows Up on an FNOL

Move to Atlanta. It is mid-morning Tuesday, and the auto-and-GL adjuster has 15 open files. A new FNOL just landed: a homeowner's water-damage loss reported via Hi Marley SMS. The text thread is 47 messages long.

The Hi Marley conversational layer is generative: the SMS responses to the insured were drafted by an LLM tuned to the carrier's claim-handling tone, sent from a human-supervised queue, and the entire transcript is preserved as a structured record. Five Sigma - the AI-native claims core - auto-drafts an FNOL summary from that transcript: date and time of loss, cause of loss, parties, witnesses, alleged damages, estimated severity. That draft is also generative. The adjuster reads it and discovers Five Sigma's coverage summary has cited the wrong policy edition - it referenced ISO HO 00 03 when the dec page shows HO 00 05 - and missed an ordinance-or-law endorsement that bears directly on the rebuild estimate. The adjuster overrides the AI's coverage position and files a note that documents the override.

Meanwhile, the predictive layer is running in the background. Shift Technology has scored the claim for fraud signals against the carrier's claims history, the insured's prior-loss footprint, the third-party-contractor network, and the geographic cluster of similar losses in the last 60 days. The score is a number with a SHAP explanation: the third-party medical clinic flagged on five plaintiff-attorney-represented soft-tissue claims in 60 days, the ISO ClaimSearch prior-loss match on the named insured, the geographic clustering. Shift's output is predictive. It is not a finding of fraud. It is a signal. The adjuster's SIU referral memo - which the adjuster will draft using the carrier's enterprise LLM as a writing aid - must enumerate the human-reviewable facts that justify the referral, separate from the Shift score. NAIC Model Bulletin §4 demands that separation. The score informs the referral; the referral cites the facts.

Tractable's photo damage estimate on the auto file is predictive: a computer-vision model trained on millions of damaged-vehicle images returns an ACV of $19,400 on the 2019 Toyota Highlander, with a confidence band and a SHAP-style attribution showing which damage zones drove the number. CCC Intelligent Solutions is in the same lane - predictive computer vision and severity scoring across 35,000+ repair facilities and 350+ insurance companies. EagleView's aerial roof report on the hail claim is predictive - image classification, roof area, age, material. When the field adjuster's photo shows architectural shingles and EagleView's classification said three-tab, the $4,200 RCV difference is a predictive-model failure, not a generative-model hallucination. The fix is a different override.

The deterministic layer on the FNOL is the carrier's claim-handling rules engine: assignment based on severity tier, diary cadence based on coverage line, ALAE thresholds that trigger supervisor escalation, statutory letter triggers (the Reservation of Rights deadline, the FCRA §615 adverse-action notice on any AI-influenced denial, the 30-day acknowledgment under the state's unfair-claims-practices act). Those rules do not learn. They run. Same input, same output.

How Each Shows Up on an SOV and a COPE Narrative

The Schedule of Values is the underwriter's primary artifact on a multi-location property submission. The 47-building SOV in Dallas is a spreadsheet - location number, address, year built, construction code, square footage, occupancy, protection class, roof material, roof age, TIV, BI exposure. Eighteen of the rows have a blank in the roof-age column. The Business-Income worksheet has a crossed-out gross-earnings line with a higher number written above it in pen and no explanation. The COPE narrative pasted into the broker's email reads: "Construction: mostly masonry, some frame. Occupancy: mixed retail/residential. Protection: Class 4 PPC, sprinklered except buildings 12 and 14. Exposure: no significant."

The deterministic step on the SOV is rating each location through the filed manual: construction class to base rate, protection class to PPC factor, territory to wind/hail factor, multiplicative coinsurance, deductible factor. Same SOV, same rating engine, same premium. Always.

The predictive step is the Akur8 GLM scoring each location's expected loss cost against the historical book - recognizing that a 1962 six-story wood-frame habitational in zip 76104 carries different statistical loss expectations than a 2015 concrete-tilt-up warehouse in zip 75201. The model card documents which features drove which coefficients. The bias-testing exhibit shows the disparate-impact ratio across geographic clusters that correlate with protected-class concentration. The Colorado Reg 10-1-1 algorithm inventory entry - required even though the policy is being issued in Texas, because the carrier's governance program is enterprise-wide - names the model, its version, its training data window, its last-tested date, its drift-monitoring status, and the human accountable for it.

The generative step is the COPE narrative parsing. Convr's NLP layer reads "mostly masonry, some frame" and asks Hyperscience's IDP to surface the underlying SOV rows that contradict that statement - three buildings coded as frame in the construction column despite the COPE saying "mostly masonry." A generative summary writes the discrepancy into the underwriter's referral note. The LLM does not decide whether the discrepancy is dispositive. The underwriter does. The LLM is the writing layer; the predictive layer is the math; the deterministic layer is the rating engine.

How Each Shows Up on a Schedule P Reserve Triangle

The reserving actuary at the same carrier is closing the quarter. The Schedule P loss development triangles - paid loss, paid ALAE, case-incurred, and reported claim count by accident year and development year - are loaded into the reserving software. The triangle for the workers compensation line shows a 24-month-to-36-month development factor of 1.18 in the most recent accident year, up from 1.11 the prior three accident years.

The deterministic layer is the Bornhuetter-Ferguson and chain-ladder methods themselves - algebraic, traceable, with every link ratio documented. The IBNR computation runs the same way every time on the same triangle.

The predictive layer is the AI-assisted commentary that surfaces unusual development patterns: a paid-loss link ratio that is two standard deviations outside the prior pattern, a case-incurred-to-paid-loss ratio that has drifted in the most recent accident year, a claim-count emergence pattern that signals severity rather than frequency. A model trained on prior years' triangles can flag the WC accident-year-2024 development as anomalous and propose alternative tail factors. The actuary reads the flag, runs the comparison against industry benchmarks, and decides whether to adjust. The model produces a signal; the actuary owns the Statement of Actuarial Opinion.

The generative layer is the opinion-supporting narrative. The Statement of Actuarial Opinion is signed by a credentialed actuary (FCAS, ACAS, MAAA), and ASOP No. 41 (Communications) attaches to every word of it. An LLM can draft the supporting narrative - describing the data, the methods, the assumptions, the reserve range, the rationale - and the actuary edits, attests, and signs. Same triangle, same models, different runs of the LLM produce slightly different prose. The math is reproducible; the prose is generated. ASOP No. 56 (Modeling) attaches to the predictive layer. ASOP No. 23 (Data Quality) attaches to the deterministic layer. ASOP No. 41 attaches to the generative layer. Knowing which standard governs which surface is what separates a defensible reserve opinion from a roundtable failure.

The Four-Persona View

The same three-machine map looks different from each desk.

The commercial property underwriter in Dallas sees deterministic in the rating engine, predictive in Cytora's triage and Akur8's pricing, and generative in the LLM that drafts the quote-with-restriction memo and the chief-underwriter referral note. The underwriter's job is to be the verifier on the predictive output, the editor on the generative output, and the human in the documented reason-code chain that NAIC §4.3 requires for every adverse decision. The 18% hit ratio target and the 11-business-day quote-to-bind cycle are not AI metrics; they are the business outcomes the three machines should be moving together.

The auto-and-GL claims adjuster in Atlanta sees deterministic in the carrier's claim-handling rules, predictive in Tractable, CCC, EagleView, and Shift, and generative in Hi Marley, Five Sigma's FNOL draft, and the LLM that drafts the Reservation-of-Rights letter and SIU referral. The adjuster's job is to override the Five Sigma coverage summary when it cites the wrong policy edition, to enumerate the human-reviewable facts that justify the SIU referral separately from the Shift score, and to file the note that documents which AI touched which artifact. The 15-file diary is not an AI metric. The 24-hour first-touch SLA is.

The retail commercial producer in Boston, working a $1.2M manufacturing renewal with a 22% incumbent rate increase, sees deterministic in the AMS data hygiene that feeds Applied Epic, predictive in Coalition's external attack-surface scan that scores the cyber risk, and generative in Send's renewal stewardship narrative, Outmarket's cover emails to seven carriers, and the LLM that drafts the BOR letter. The producer's job is to verify that the cyber attestation matches the actual control environment (because a misstatement to the cyber underwriter is an E&O claim waiting), to confirm the BOR letter complies with the venue state's e-signature requirement and the incumbent's 5-day cooling-off rule, and to write the agency-built chatbot disclaimer the CCO asked for. The four-carrier-comparison deadline is the business metric.

The pricing actuary behind both desks sees deterministic in the filed rating manual, predictive in the GLM/GBM models built in Akur8 and Earnix, and generative in the LLM that drafts the SERFF rate-filing memorandum. The actuary's job is to build the model card with SHAP, calibration, and bias-test exhibits; to pull comparable-carrier filings from Akur8 Discover and Matrisk; to document variable selection against the Colorado Reg 10-1-1 algorithm inventory; to draft the NY DFS Circular Letter 2024-7 proxy-test memo in writing; and to sign the actuarial certification under ASOP No. 41. The 30% LAE-reduction case the chief actuary is making to the board is the business case the three machines have to deliver together.

Why the Category Matters for NAIC §4

NAIC Model Bulletin on the Use of Artificial Intelligence Systems by Insurers (December 2023, adopted across 25+ jurisdictions by mid-2026) divides governance expectations across four areas: program structure, third-party AI, testing and validation, and documentation. Section 4 - the documentation and adverse-action substance - is where the category map becomes operational.

For a deterministic rating engine, the documentation is the filed manual, the SERFF audit trail, and the human-readable rate pages. NAIC §4 expectations are satisfied because the engine is, by definition, traceable.

For a predictive model, NAIC §4 demands documented data sources, documented variable selection, documented bias testing, a model card that survives a DOI examiner challenge, an algorithm inventory entry, and a drift-monitoring plan. Colorado Reg 10-1-1 made these expectations enforceable for life, private passenger auto, and health benefit plans - effective October 15, 2025, with the first compliance report due July 1, 2026. NY DFS Circular Letter 2024-7's proxy test attaches at the variable-selection layer. The model is defensible only with the supporting artifacts.

For a generative model, NAIC §4 demands that "the model said so" is never a reason code. Every adverse UW decision, every claim denial, every fraud referral, every accelerated-UW knockout, must trace to a documented, human-reviewable reason - not to LLM output. The LLM can draft the reason-code letter; the underwriter, adjuster, or actuary owns the reasons. The Colorado Reg 10-1-1 compliance report named the writer of every adverse-action letter and the variables driving the decision. The NAIC AI Systems Evaluation Tool Exhibits A-D - the 12-state pilot running early 2026 through September 2026, with re-exposure September-October 2026 and adoption expected at the NAIC Fall National Meeting in November 2026 - explicitly separates governance documentation (Exhibit B) from high-risk system detail (Exhibit C) and from the underlying AI data (Exhibit D). Which exhibit a system lands in depends on its category.

The Test You Can Run at Your Desk Tomorrow

Walk to your desk tomorrow morning, open the three or four tools you use most, and for each one, write a single line: deterministic, predictive, or generative. Then write the named regulation or standard that governs it.

The rating engine in PolicyCenter or Duck Creek is deterministic, governed by the filed manual and SERFF. The Akur8 pricing workbench, the Federato triage score, the Cytora appetite ranking, the Shift fraud score, the Tractable estimate, the EagleView roof age, the Munich Re accelerated-UW knockout, the Swiss Re Magnum routing, the RGA AURA NEXT decision, the SCOR Velogica simplified-issue path - all predictive, governed by NAIC §4, Colorado Reg 10-1-1, NY DFS Circular Letter 2024-7, and the model-risk-management standard at the carrier. The Send renewal narrative, the Outmarket cover email, the Five Sigma FNOL draft, the Hi Marley SMS, the carrier's enterprise LLM drafting the Reservation-of-Rights letter, the chatbot on the agency website - all generative, governed by the same regulations plus the agency's E&O posture and the producer's license-board obligations in Texas, Florida, New York, California, and Louisiana.

That single line, written for each tool on each desk, is the foundation. Every later lesson in this program - the cardinal reason-code rule, the hallucinated-coverage examples, the named-platform walkthroughs of Cytora and Federato and Tractable, the NAIC AI Systems Evaluation Tool exhibits, the Colorado Reg 10-1-1 algorithm inventory - sits on top of it. You cannot govern what you cannot categorize. You cannot defend a SERFF filing, a claim denial, an accelerated-UW knockout, or an SIU referral if you cannot say which machine made which call.

Key Takeaways

  • The word "AI" covers three categorically different machines in 2026 insurance: deterministic rules engines (the filed rating manual in PolicyCenter/Duck Creek), predictive models (Akur8 GLMs, Federato triage, Cytora appetite, Shift fraud, Tractable photo estimates, Munich Re and Swiss Re Magnum accelerated UW), and generative models (the LLM drafting the Reservation-of-Rights letter, Send renewal narratives, Five Sigma FNOL drafts, Hi Marley SMS). They are not on a spectrum - they are different species governed by different standards.
  • On the 8:14 a.m. Dallas ACORD 125 + 140 + 47-building SOV, all three machines run simultaneously. Hyperscience IDP at 99.5%+ accuracy extracts the fields (predictive), Cytora Autopilot triages against appetite (predictive), Akur8 prices the loss cost (predictive), the filed rating manual produces the class-plan premium (deterministic), and the enterprise LLM drafts the quote-with-restriction memo (generative).
  • On the Atlanta auto-and-GL adjuster's FNOL, the categories repeat. Tractable, CCC, EagleView, and Shift are predictive (numbers and signals); Hi Marley, Five Sigma's FNOL draft, and the LLM writing the Reservation-of-Rights letter are generative (text); the carrier's claim-handling rules and statutory deadlines are deterministic.
  • "The model said so" is never a reason code under NAIC Model Bulletin §4. Every adverse UW decision, claim denial, accelerated-UW knockout, fraud referral, and rate impact must trace to documented, human-reviewable reasons - not to LLM output and not to a fraud score alone. The Shift score informs the SIU referral; the human-reviewable facts justify it.
  • Each category maps to a different governance regime. Deterministic engines live in the filed manual and SERFF; predictive models need a model card, bias-testing exhibit, algorithm inventory entry, and drift-monitoring plan under Colorado Reg 10-1-1 (July 1, 2026 compliance report) and NY DFS Circular Letter 2024-7's proxy test; generative models need verification of every artifact they produce because hallucinated ISO endorsement numbers and fake case citations are real failure modes.
  • The four-persona view is non-negotiable. The same three machines look different from the UW desk, the claims diary, the producer's renewal cycle, and the actuary's reserve triangle. Categorizing the tools on your specific desk is the first L1 skill - every later lesson assumes you can.
  • ASOP No. 23 governs data quality (deterministic), ASOP No. 56 governs modeling (predictive), ASOP No. 41 governs communications (generative). The Statement of Actuarial Opinion, the SERFF rate-filing memorandum, and the Schedule P narrative each invoke different standards depending on which layer produced them.
  • The April 2026 AM Best Special Report numbers - 41% of carriers using AI in core functions, ~60% expecting 1-3 year transformation - are workforce numbers, not technology numbers. Data readiness, governance, cyber, and legacy integration are the binding constraints. The category map is the first move that closes the governance gap NAIC §4 was written to find.