AI for Insurance Professionals
Strategic · M10 · lesson 10 of 24 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Design a Proof-of-Concept and a Parallel-Run for Insurance AI - 90-Day POC Template
📖
now learning

Design a Proof-of-Concept and a Parallel-Run for Insurance AI - 90-Day POC Template

15 min

A 2026 insurance AI proof-of-concept that produces a defensible decision - proceed to production, kill the project, or pivot the design - runs ninety days, has five named artifacts, four named success criteria, three named kill criteria, two parallel-run patterns, and one named accountable executive. Carriers that skip the POC discipline and go directly to production deployment either accept undocumented risk that surfaces in months 7-12 as a market-conduct exam finding, a treaty audit issue, or an operational incident the CRO cannot defend - or they over-invest in vendor relationships that should have been killed at month 3. The POC discipline is the cheapest form of insurance the carrier buys in the AI portfolio: $80K-$250K invested in a 90-day evaluation prevents $1M-$5M of mis-deployment cost across a 3-year horizon. This lesson is the 90-day POC template, the champion-challenger parallel-run pattern, the worked example for a Federato deployment in a specialty UW unit at a $1.4B carrier, the explicit kill criteria language, the regulator's view (NAIC AISET Exhibit C model-level evaluation, Colorado Reg 10-1-1 algorithm inventory entry, NY DFS Circular Letter 2024-7 proxy-test memo), and the executive sign-off package that makes the POC decision defensible to AM Best, the treaty broker, and the state DOI examiner.

The 90-Day POC Template

Day 0 - Charter. The POC charter is a 2-3 page document signed by the accountable executive (CUO, CCO, Chief Actuary, or Chief Distribution Officer depending on the use case), the CRO (for risk oversight), and the vendor's account executive (for shared commitment). The charter names the scope (which LOB, which UW desks, which claims types, which producers), the success criteria (the four named metrics, see below), the kill criteria (the three named trigger conditions), the parallel-run pattern (champion-challenger or staged-shadow), the data feed (what flows from carrier systems to vendor, with privacy controls), the governance overlay (AI committee involvement, weekly cadence, escalation path), and the budget ($80K-$250K depending on use case complexity).

Days 1-14 - Integration and Baseline. Vendor integration to carrier systems on a controlled scope (single UW desk, single claims unit, single producer channel). Baseline measurement of the four success metrics on the in-scope work for the trailing 90 days using existing process - this is the comparator the POC will be measured against. Baseline measurement is the single most commonly skipped step; without it the POC produces ambiguous results because the lift cannot be computed.

Days 15-30 - Shadow Mode. The vendor's AI runs on the same submissions or claims or producer calls the human operators are working, but the AI's output is not surfaced to the human operator. The shadow output is captured for comparison to the human's decision. This is the safest POC pattern because operations are not affected; the data quality of the AI's output is established before any human consumes it.

Days 31-60 - Parallel Run. The AI's output is surfaced to the human operator alongside the human's own analysis. The human still makes the final decision; the AI provides recommendation. This is the operational test of whether AI output is useful, trustworthy, and integrable into existing workflow. Measure decision agreement rate, time-to-decision delta, and any escalation or override patterns.

Days 61-80 - Limited Production. The AI's output drives action on a controlled subset (low-stakes claims, low-premium submissions, low-volume producer channel) with human oversight on the decision. This is the operational stress test of vendor support, integration reliability, and incident-handling.

Days 81-90 - Decision Package. The decision package is assembled: baseline measurement, shadow-mode results, parallel-run results, limited-production results, vendor-performance assessment, governance and risk assessment, financial projection (3-year TCO and combined-ratio impact), and explicit go/no-go/pivot recommendation. The accountable executive presents to the executive committee; the CRO signs off; the contract proceeds to negotiation or the project terminates.

The Five Named Artifacts

The charter is the first; the others assemble through the 90 days. (1) Charter, signed Day 0. (2) Baseline measurement report, finalized Day 14 - trailing 90 days of in-scope work with the four success-criterion metrics computed against the existing process, including data-quality notes and segmentation by LOB, channel, and complexity tier. (3) Shadow-mode comparison report, finalized Day 30 - AI output captured silently versus human decisions, with disagreement analysis and reason-code classification. (4) Parallel-run operational report, finalized Day 60 - decision agreement rate, override patterns, time-to-decision delta, operator feedback survey, vendor incident log. (5) Decision package, finalized Day 88 - synthesis of the prior four, financial projection, regulatory and governance assessment, vendor scorecard, and explicit go/no-go/pivot recommendation with rationale and sign-off block. The decision package is the artifact that lives in the algorithm inventory under Colorado Reg 10-1-1, the AISET Exhibit C model-level response packet, and the AM Best meeting evidence binder. Carriers without the five artifacts cannot defend the POC decision at any external review surface.

The Four Named Success Criteria

Every POC has four explicit success criteria written before Day 0 and frozen for the 90 days. They differ by use case but follow the same structural pattern.

Criterion 1 - Primary impact metric. The one metric that anchors the use case to combined ratio (from Ch1-3's five-metric framework). For a Federato POC in specialty UW: submission throughput improvement of 15%+ on appetite-aligned submissions. For a Tractable POC in property claims estimating: first-touch estimate variance from final paid amount within 8% on 75%+ of claims. For an Akur8 POC in inland marine pricing: loss-ratio improvement projection of 2%+ on bound business based on parallel-run analysis.

Criterion 2 - Operational quality metric. The metric that proves the AI output is usable in production. Decision agreement rate above 70% on parallel-run, false positive rate below 8%, integration uptime above 99.5%, vendor response time on issues below 24 hours.

Criterion 3 - Governance and risk metric. Adverse-decision audit, fairness testing results on the carrier's actual data (NY DFS Circular Letter 2024-7 proxy-test discipline; Colorado SB 21-169 alignment for in-scope lines), NAIC §4 reason-chain documentation completeness, incident count during POC (target zero material incidents).

Criterion 4 - Adoption and change-management metric. Operator acceptance scored via end-of-POC survey, time-to-proficiency curve for trained operators, override rate trend (high overrides indicate trust issues; declining overrides indicate trust building).

Setting the Thresholds Without Rubber-Stamping the Vendor

The carrier writes the thresholds, not the vendor. A common failure mode is the vendor pre-selecting metrics they have shown to land favorably in prior deployments - Cytora's "appetite-aligned submission identification" rate, Tractable's "completion rate" without policy-form complexity stratification, Akur8's "deployment cycle time" rather than loss-ratio impact. Each of those vendor-framed metrics is defensible inside the vendor's narrative, but the carrier's success criteria must be the carrier's own combined-ratio-adjacent metrics measured against the carrier's own baseline. The discipline is enforced through the chief actuary's signoff on Criterion 1 and the chief compliance officer's signoff on Criterion 3 before the charter is finalized. Where the vendor's preferred metric and the carrier's framing diverge, the divergence is documented in the charter; the gap closes during POC execution through the carrier's data rather than the vendor's marketing collateral.

The Three Named Kill Criteria

Kill criteria are written before Day 0 and trigger termination without further executive negotiation. They prevent loyalty-to-sunk-cost dynamics from sustaining failing POCs.

Kill 1 - Primary metric failure. If at Day 60 the primary impact metric is tracking below 50% of target, the POC terminates. Example: Federato POC target 15% throughput lift; at Day 60 tracking 5%; kill.

Kill 2 - Governance or compliance failure. If at any point during the POC a material adverse-decision pattern is identified, a fairness-testing failure surfaces, or a regulatory exposure becomes apparent (FCRA, MHPAEA NQTL Tri-Agency 2024 final-rule alignment, Colorado Reg 10-1-1, NY DFS Circular Letter 2024-7 proxy-test failure, state DOI bulletin like Connecticut MC-25-8 or Nevada Bulletin 24-006), the POC terminates pending governance remediation. The remediation may extend the POC or end it depending on severity.

Kill 3 - Vendor execution failure. If the vendor produces a material incident (data exposure, model malfunction, integration outage above 48 hours), fails to meet contractual response SLAs more than twice, or makes a material organizational change (key personnel departure, ownership change, security posture degradation), the POC terminates. The vendor's incident-handling and responsiveness during the POC predicts behavior in production; weakness here is a forward indicator.

Champion-Challenger and Staged-Shadow - The Two Parallel-Run Patterns

Two parallel-run patterns serve different POC objectives. Champion-challenger compares two approaches running side-by-side on the same work; staged-shadow runs the AI silently before exposing output.

Champion-Challenger

The champion is the existing process (human UW with current tools, human adjuster with current estimating). The challenger is the AI-enabled process (UW with Federato workbench, adjuster with Tractable estimating). Both run on the same submissions or claims in parallel; the same input flows to both, both produce a decision, the human champion's decision drives the actual operational outcome, and the challenger's decision is captured for comparison. Champion-challenger is appropriate when the AI is mature enough to produce production-quality output and the carrier wants to measure the operational delta against the existing process.

Champion-challenger requires double-running infrastructure (both systems active), careful data isolation (vendor sees the submissions in shadow mode with appropriate privacy controls), and metric instrumentation that captures both decisions for analysis. The advantage is direct comparison; the disadvantage is operational cost - running two systems is more expensive than running one with a shadow capture.

Staged-Shadow

The AI runs silently in the background, observing inputs and producing outputs, but the outputs are not surfaced to human operators in early stages. After confidence is established, output is surfaced as recommendation (parallel run); then output drives action on controlled subset (limited production); then full production. Staged-shadow is appropriate when the AI is unproven on the carrier's data, the risk tolerance is low, or the use case is sensitive (claims handling, pricing decisions, adverse-action notices).

Staged-shadow has lower operational cost than champion-challenger because only one operational system is running; the shadow runs in parallel but doesn't affect operations. The advantage is risk control; the disadvantage is slower lift demonstration - operators don't experience the AI's value until parallel-run phase.

Matching Pattern to Use Case Sensitivity

The pattern selection is not stylistic. A submission-triage POC on appetite-aligned commercial submissions is low-sensitivity and benefits from champion-challenger because the carrier can see operational delta inside thirty days; a rate-engine POC on personal auto where adverse-action notices may flow is high-sensitivity and demands staged-shadow because the fairness-pipeline output must be validated before any decision reaches a consumer. A claims-estimating POC on auto physical damage with Tractable sits in between - staged-shadow for the first thirty days to validate digital-completion behavior on the carrier's policy forms, then transition to champion-challenger for the next thirty as confidence builds. The accountable executive and CRO co-decide pattern in the charter.

Worked Example - Federato POC at $1.4B Specialty Carrier

The carrier writes specialty property, professional lines, inland marine. The POC runs on the professional lines specialty UW unit - 8 UW seats, 2 UW assistants, 3 brokerages as the primary distribution. POC scope: 90 days, $185K budget, accountable executive is the VP of Professional Lines UW with the CUO as secondary accountable.

Day 0 charter. Scope: professional lines submissions only, all distribution channels, all coverage tiers. Success criteria: (1) submission throughput improvement of 15%+ on appetite-aligned submissions; (2) decision agreement rate above 75% in parallel-run; (3) zero material adverse-decision incidents, fairness testing showing no protected-class disparate impact; (4) UW operator acceptance above 70% on end-of-POC survey, override rate below 25% in last 30 days. Kill criteria: (1) at Day 60, throughput lift below 7.5%; (2) any material adverse-decision or fairness incident; (3) Federato material incident or SLA failure twice. Parallel-run pattern: staged-shadow (low risk tolerance, professional lines are sensitive).

Days 1-14 integration and baseline. Federato integrated to PolicyCenter and Cytora; baseline metrics measured: current submission throughput 14 submissions per UW per day; current decision time 87 minutes per submission; current appetite-aligned rate 62%; current declination cycle time 3.8 days.

Days 15-30 shadow mode. Federato runs silently on every professional lines submission; output captured for comparison; no UW exposure. Daily data review by carrier data science. Shadow output shows: throughput projection 17.4 submissions per UW per day if deployed (+24%); decision time projection 62 minutes (-29%); appetite-aligned rate projection 84% (+22 points).

Days 31-60 parallel run. Federato output surfaced to UWs as recommendation; UW retains final decision. Measured: actual throughput 16.8 per UW per day; actual decision time 71 minutes; actual appetite-aligned rate 81%; decision agreement rate 78%; override rate 31% (high in week 1, declining to 22% by week 8).

Days 61-80 limited production. Federato drives action on submissions below $250K limit; UW reviews and approves; full UW authority retained on $250K+ submissions. Measured: production throughput 17.1 per UW per day; integration uptime 99.7%; one minor Federato issue resolved in 6 hours.

Days 81-90 decision package. All four success criteria met: (1) throughput lift 22% on appetite-aligned (target 15%, achieved); (2) decision agreement 78% (target 75%); (3) zero adverse-decision incidents, fairness testing clean across age, geography, and broker channel; (4) UW operator acceptance 81% (target 70%), override rate 22% (target below 25%). Kill criteria not triggered. Decision: proceed to production deployment across professional lines, with subsequent extension to specialty property in Q3 contingent on professional lines stable production at month 9. Three-year combined-ratio impact projection: 1.6-2.2 points on professional lines. Contract negotiation begins Day 91.

The Regulator and the CRO Read of This POC

The decision package becomes the algorithm-inventory entry for Federato professional-lines triage under Colorado Reg 10-1-1 (the carrier files in Colorado for several professional-lines products, so the inventory is in scope). The entry records the AI capability, the deployment scope, the fairness-pipeline result, the model-card link, the named owner, and the review cadence. The same package feeds the AISET Exhibit C model-level evaluation when the carrier responds to one of the 25+ jurisdictions that adopt the NAIC Model Bulletin §4.1-§4.4 by mid-2026. The CRO's quarterly governance review cites the decision package as evidence of POC discipline; the AM Best analyst at the next rating meeting reviews it as governance posture under the survey-plus-readiness composite. A POC that lacks this evidentiary chain leaves the carrier defending a production deployment without the supporting artifacts external reviewers expect.

POC Mistakes Most Carriers Make and How to Avoid Them

Six mistakes recur. First, no baseline measurement - POC ends with results no one can compare to anything; ambiguous outcome. Always measure baseline in Days 1-14. Second, no explicit kill criteria - POC drifts into Year 1 because no one can pull the trigger; sunk-cost dynamics dominate. Write kill criteria in the charter. Third, scope sprawl - POC starts on professional lines and expands to property by Day 45 because "the vendor can do that too"; results are uninterpretable. Hold scope. Fourth, vendor-driven success criteria - vendor picks the metrics they know they'll hit; results are biased. Carrier sets criteria, not vendor. Fifth, no governance overlay - POC runs without AI-committee oversight; governance issues surface in production. Embed AI-committee review weekly. Sixth, no decision package - POC ends with informal kickoff to production; the documentation gap surfaces in the next state DOI exam or AM Best meeting. Produce the formal decision package.

Two Additional 2026 Failure Modes

Seventh, vendor concentration blindness - the POC proceeds without checking whether the vendor's success on this use case would push the carrier over the 28% vendor-concentration threshold that the CRO and AI committee have set as the operational cap. A successful POC that breaches the cap forces either a compensating-control documentation effort or a vendor substitution at production launch. Concentration is checked at the charter, not at production sign-off. Eighth, agentic-pattern under-specification - a 2026 POC on Cytora Autopilot, Shift Claims agentic, or any agentic workflow that takes actions rather than just providing recommendations needs an additional layer in the charter: explicit boundaries on the actions the agent may take, the escalation triggers, the human-in-the-loop touchpoints, and the audit-trail granularity. POCs on agentic systems that use the older recommendation-only charter template produce ambiguous results because the action boundary was never defined.

POC Budget and Team

$80K-$250K budget covers: vendor's POC pricing (often 30-50% of annual license for the 90 days, sometimes free for strategic vendors), integration engineering (carrier-side 0.4-0.8 FTE for 90 days), data science (0.3-0.6 FTE), business owner time (UW VP or claims VP carries 0.2 FTE plus accountable executive at 0.05 FTE), governance overhead (AI committee weekly), procurement and legal (contract POC terms, charter, decision package review). The accountable executive's time is the highest-value input and the most under-protected - POCs without accountable-executive attention drift and produce ambiguous outcomes.

The Treaty Broker and AM Best Evidence Loop

The treaty broker carries POC evidence into the renewal narrative. A documented Tractable POC showing 22% ALAE reduction on first-touch closeable claims becomes a treaty-broker talking point on claims-handling discipline; the reinsurer prices the ceded layer assuming the POC outcome translates to production. The AM Best analyst at the rating meeting reads the POC decision package as evidence on the third-party AI risk axis of the readiness composite (vendor evaluation framework, concentration management, incident response). Carriers maintaining the POC discipline accumulate evidence quarterly; carriers without it produce ad-hoc narrative that does not survive analyst probing or treaty-broker due diligence.

Key Takeaways

  • 90-day POC structure: Day 0 charter; Days 1-14 integration and baseline; Days 15-30 shadow; Days 31-60 parallel run; Days 61-80 limited production; Days 81-90 decision package. Budget $80K-$250K depending on use case complexity.
  • Five named artifacts: charter (Day 0), baseline measurement report (Day 14), shadow-mode comparison (Day 30), parallel-run operational report (Day 60), decision package (Day 88). The decision package feeds the algorithm inventory under Colorado Reg 10-1-1 and the AISET Exhibit C model-level response.
  • Four named success criteria written before Day 0: primary impact metric, operational quality, governance/risk, adoption/change-management. Carrier sets criteria - never the vendor; chief actuary signs off on Criterion 1 and chief compliance officer signs off on Criterion 3 before the charter is finalized.
  • Three named kill criteria written before Day 0: primary metric failure (50% of target at Day 60), governance/compliance failure (FCRA, MHPAEA NQTL, Colorado Reg 10-1-1, NY DFS 2024-7, CT MC-25-8, NV 24-006), vendor execution failure. Kill criteria prevent loyalty-to-sunk-cost dynamics.
  • Two parallel-run patterns: champion-challenger (direct comparison, higher cost, low-sensitivity use cases) and staged-shadow (lower risk, slower lift demonstration, high-sensitivity use cases). Pattern selection co-decided by accountable executive and CRO based on use case sensitivity.
  • Worked Federato example: 8-UW professional lines unit, $185K budget, 90 days, staged-shadow pattern. Result: 22% throughput lift on appetite-aligned, 78% decision agreement, 81% UW acceptance, 22% override. Proceed to production. 1.6-2.2 combined-ratio points projected on professional lines.
  • Eight recurring POC mistakes: no baseline, no kill criteria, scope sprawl, vendor-driven criteria, no governance overlay, no decision package, vendor concentration blindness, agentic-pattern under-specification. Each is preventable with charter discipline.
  • $80K-$250K POC budget is the cheapest insurance the carrier buys. Prevents $1M-$5M of mis-deployment cost across a 3-year horizon.
  • Decision package is the artifact AM Best, treaty broker, state DOI examiner, and the AI committee can all interrogate. Treaty broker carries it into renewal narrative; AM Best analyst reads it on the third-party AI risk axis; without it, production deployment lacks evidentiary base for any external review.