Intake and OCR - Hyperscience + Indico Style Document Extraction From ACORD and SOV
Lesson 1 named Step 3 as document intake and OCR with a 12-minute SLA at 99%+ accuracy on clean PDFs. This lesson is the artifact-level build of that step. The submission email lands at 7:42 a.m. with ten attachments. By 7:54 a.m., Hyperscience or Indico Data has classified each document, extracted structured fields against ACORD form templates, validated against the carrier's PAS schema, and queued the exceptions for human review. The carrier that runs this step at 99%+ accuracy moves the entire 20-step pipeline downstream cleanly; the carrier that runs it at 87% accuracy spends Steps 4-20 fighting bad data. The two most common worst-case inputs in 2026 commercial submissions are the 200-dpi loss run that bled into the margin (degraded scan with poor character contrast) and the SOV missing roof age on 18 of 47 locations (incomplete structured spreadsheet from the producer). Both fail naive OCR at 60-75% accuracy; both pass disciplined IDP (Intelligent Document Processing) at 95%+. This lesson is the workflow: queue ingestion → IDP classification → schema validation → exception triage → enrichment handoff. Field-by-field accuracy targets per ACORD form. The exception queue's SLA discipline. The carrier's tradeoff between Hyperscience's keyer-trained accuracy model and Indico's transformer-trained model on long-form unstructured documents like COPE narratives.
The Ten-Attachment Submission as Seven Document Types
The wholesale broker's 7:42 a.m. email with ten attachments doesn't contain ten documents in the IDP sense - it contains seven document types with variable representations. The IDP layer's first job is classification: which attachment is which document type, and how should each be processed?
Document type 1 - ACORD 125 Commercial Insurance Application. Standard cover page for commercial submissions. Edition matters: ACORD 125 (2016/06) vs. (2024/03) edition have different field positions. IDP classifies edition, extracts named insured, FEIN, NAICS, SIC, mailing address, contact information, prior carrier, expiring premium, expiring effective dates. Accuracy target on clean PDF: 99.5%+.
Document type 2 - ACORD 140 Property Section. Per-location property attributes including ISO construction class, occupancy, sprinkler/alarm protection, roof age, BI exposure, indemnity period, ordinance-or-law extensions. IDP extracts per-location. Accuracy target on clean PDF: 99%+.
Document type 3 - ACORD 126 Commercial General Liability Section. Per-class GL exposure including products/completed operations, premises liability, hired-and-non-owned auto if endorsed, contractual liability. Accuracy target on clean PDF: 99%+.
Document type 4 - SOV (Statement of Values). Per-building exposure spreadsheet typically in Excel or PDF-rendered Excel. Fields: building number, address, TIV, BPP (Business Personal Property), BI, construction year, square footage, occupancy class, sprinkler Y/N, roof age, alarm Y/N. Variability across producers is the IDP challenge - every producer uses slightly different column headers, row organization, and unit conventions (TIV in thousands vs. dollars, square footage in feet vs. meters). Accuracy target on clean Excel: 98%+; on PDF-rendered SOV: 95%+.
Document type 5 - Loss Run. Five-year claim history from expiring carrier. Format variability is extreme: Excel with merged cells, PDF tables, scanned faxes, carrier-specific layouts. Fields: loss date, cause of loss, paid amount, reserved amount, status (open/closed), claim number. Accuracy target on clean format: 97%+; on degraded 200-dpi scan: 90%+ with exception triage.
Document type 6 - COPE Narrative. Long-form unstructured text describing Construction, Occupancy, Protection, Exposure attributes for largest locations. Free-form prose, not structured form. Indico Data's transformer-based model outperforms Hyperscience's keyer-trained model on long-form unstructured text by 8-12 percentage points; Hyperscience outperforms Indico on form-structured fields by 2-4 points. The carrier that uses both products in stages captures the best of each.
Document type 7 - Photo packet from EagleView. Pre-pulled aerial imagery and building photos. Not a text-extraction document; routes to the EagleView integration at Step 5 enrichment rather than the IDP queue. Classified and forwarded only.
Hyperscience vs. Indico - The Architectural Choice
The 2026 IDP market for insurance has two dominant platforms. Each has a different architectural bet, different accuracy profile, different price point, and different integration story.
Hyperscience. Architectural bet: human-in-the-loop training with a keyer workforce. Hyperscience trains its extraction models on carrier-specific document templates by having human keyers correct outputs at deployment; the model learns from corrections and accuracy compounds. Strength: form-structured extraction (ACORD 125/126/140, standard SOV layouts) where field positions are reliable. Hyperscience hits 99.5%+ on ACORD forms after 2-4 weeks of keyer-supervised training. Pricing in 2026: per-page or per-document with monthly minimums; mid-size carrier ranges $200K-$800K annually. Integration: API-first, fits cleanly into Federato, ClaimCenter, PolicyCenter via REST.
Indico Data. Architectural bet: transformer-based model fine-tuned on insurance-specific corpora. Strength: long-form unstructured text - COPE narratives, broker cover letters, prior-carrier declination letters, recorded-statement transcripts. Indico's transformer architecture handles linguistic variation that Hyperscience's keyer model can't generalize over. Accuracy: 92-97% on COPE narratives where Hyperscience runs 78-85%. Pricing in 2026: subscription with per-document overage; mid-size carrier ranges $300K-$1.2M annually. Integration: API-first plus pre-built connectors for Guidewire, Duck Creek, Sapiens.
The hybrid deployment. Carriers that deploy both products use Hyperscience for ACORD forms and structured SOVs, Indico for COPE narratives, loss-run prose comments, and broker cover letters. Hybrid deployment adds $100K-$300K annually in vendor cost but recovers 4-8 LR-equivalent points through better extraction quality. Mid-2026 industry split: roughly 45% Hyperscience-only, 30% Indico-only, 25% hybrid. The hybrid share is growing as carriers recognize the architectural complementarity.
Field-by-Field Accuracy Targets Per ACORD Form
The IDP layer's accuracy is not a single number; it's a per-field accuracy distribution. The carrier sets field-level accuracy targets based on downstream consumption.
ACORD 125 fields. Named insured legal name: 99.9%+ (downstream PAS write requires exact match). FEIN: 99.9%+ (clearance match requires exact FEIN). NAICS: 99.5%+ (appetite scoring at Step 6 depends on NAICS). SIC: 99%+ (legacy systems may use SIC). Mailing address with USPS DPV: 99%+ (delivery validation). Effective date: 99.9%+ (binder timeline depends on exact date). Prior carrier: 98%+ (loss-history reconstruction). Expiring premium: 98%+ (rate comparison input).
ACORD 140 fields. Per-location address: 99%+ (geocoding for cat-model and FEMA NRI). ISO construction class: 99%+ (drives appetite and pricing). Occupancy class: 99%+ (drives appetite). Sprinkler Y/N: 99%+ (drives appetite and pricing). Alarm Y/N: 98%+. Roof age: 95%+ (often missing - see worked example below). TIV per location: 99.9%+ (cannot fabricate TIV totals; quoted from source). BI exposure per location: 99%+. Indemnity period: 99%+ (default 12 months; producer may override).
SOV fields. Building number: 99.9%+. Address (street + city + state + ZIP): 99%+. TIV: 99.9%+. BPP: 99%+. BI: 99%+. Construction year: 95%+ (often partial - only some buildings). Square footage: 98%+. Occupancy class: 98%+. Sprinkler Y/N: 99%+. Roof age: 90%+ (most missing fields). Alarm Y/N: 98%+. Distance-to-hydrant: 92%+.
Loss run fields. Loss date: 99%+. Cause of loss: 95%+ (free-text descriptions vary). Paid amount: 99%+. Reserved amount: 99%+. Status: 99%+. Claim number: 99%+.
COPE narrative fields. Construction description: 88-92%+ (long-form). Occupancy description: 92%+ (often more specific than ACORD 140 class). Protection description: 90%+. Exposure description: 88%+.
Field-level targets drive the exception triage logic. A field at 99.9% accuracy needs little human review; a field at 90% accuracy needs explicit verification at Step 4 field validation. The exception queue routes by field-level confidence.
The 200-dpi Loss Run That Bled Into the Margin
The worst-case loss-run extraction scenario: a five-year loss run scanned at 200 dpi (rather than the standard 300 dpi) with the right margin bled into the page edge such that the last 3-4 characters of paid-amount fields are truncated or smeared. Naive OCR returns "$142," instead of "$142,300." The downstream pricing model at Step 12 receives understated severity history; the rate indication comes out too low; the carrier binds at inadequate rate.
The IDP-disciplined workflow. Hyperscience's loss-run model recognizes the bleed pattern at extraction time and flags the field with low confidence (35-50% rather than 99%+). Indico's transformer can sometimes reconstruct truncated amounts from contextual cues (cause of loss + status + claim duration) and produces a higher-confidence reconstruction (75-85%) with an alternative-extraction note. The exception queue receives the flagged fields. A submissions-operations reviewer with 30 seconds per field confirms by viewing the source image; for unrecoverable fields, the producer is asked to resend a 300-dpi scan.
The 12-minute SLA absorbs the exception. 200-dpi loss runs occur in roughly 8-15% of commercial submissions in 2026. The exception-triage workflow handles them within the 12-minute Step 3 SLA: IDP extraction completes in 6-8 minutes including the exception queue write; reviewer triage completes in 3-4 minutes for the flagged fields; if producer follow-up needed, the submission flags "pending producer response" and Step 4 holds. Without IDP discipline, the same 200-dpi loss run produces 60-75% accuracy fields that flow through Steps 4-20 undetected, surfacing as adverse loss development 12-24 months post-bind.
The carrier's loss-development variance attributable to loss-run extraction quality. Industry research in 2026 from Moody's and AM Best estimates that carriers with 95%+ loss-run extraction accuracy see 2-4 LR-points better loss development on commercial property than carriers with 80-85% accuracy. The math: missed losses in the historical run lead to understated severity, lead to inadequate pricing, lead to adverse development. The 2-4 LR point variance compounds over 3-5 years before exam findings or rate-filing corrections catch it.
The SOV With Missing Roof Age on 18 of 47 Locations
The second worst-case input: an SOV where the producer left roof age blank on 18 of 47 locations. The blank fields aren't an extraction failure - IDP correctly identifies the fields as blank - but they're a data-quality failure that propagates downstream if the workflow doesn't handle it.
The workflow. Step 3 IDP completes extraction; 18 roof-age fields land in the structured output as null. Step 4 field validation flags the nulls. Step 5 enrichment queries EagleView for the 18 missing roof ages: EagleView's aerial-imagery roof-age estimation returns ages for 15 of 18 within ±2 years confidence (EagleView's stated accuracy on roof age from aerial imagery is approximately 85% within ±2 years). The remaining 3 locations have no EagleView coverage (rural locations or non-Florida flyover frequency). For those 3, the producer is queried with the Step 9 missing-info request. By Step 8 workbench load, the underwriter sees 47 locations with roof-age fields populated: 29 from the original SOV, 15 from EagleView, 3 pending producer response with flagged Status.
The EagleView roof-age estimation methodology. EagleView aerial imagery captures roof condition indicators: granule loss, color uniformity, sag patterns, structural variation. EagleView's roof-age model trained on millions of property records produces age estimates with documented confidence intervals. The carrier's appetite-scoring engine at Step 6 accepts EagleView roof age with the confidence interval as a feature; risk pricing at Step 12 incorporates the uncertainty into the rate band.
Why the producer's blank fields aren't a producer failure. Producers can't always confirm roof age - older commercial buildings change hands multiple times and the current owner doesn't have construction records. The SOV-blank-fields scenario is the normal case for older property books, not an exception. The disciplined workflow uses EagleView to fill the gap rather than blocking on producer response, which improves submission cycle time without sacrificing data quality.
The Exception Queue and Its SLA Discipline
The IDP layer's exception queue is the operational discipline that separates 99%+ effective accuracy from 87% effective accuracy. The queue receives extractions flagged below confidence threshold and routes to human reviewers with structured context.
Queue routing logic. High-confidence extractions (above 95%) flow directly to Step 4 field validation without queue stop. Low-confidence extractions (50-95%) queue for review with field-level highlighting. Very-low-confidence extractions (under 50%) route to a senior reviewer or producer follow-up. Each field in the queue carries: source document page reference, extracted value, confidence score, alternative extraction (if Indico produced one), and reviewer action options (confirm, correct, request producer follow-up).
Reviewer productivity. Trained submissions-operations reviewers handle 80-120 exception fields per hour at quality, which translates to 12-15 minutes per submission (assuming 15-25 flagged fields per submission). The 12-minute Step 3 SLA includes the exception queue review time, which means parallelism is essential - multiple reviewers handle the queue in parallel while IDP continues extracting subsequent fields.
Queue depth alerts. When queue depth exceeds 4-6 hours of review work, operational alerts fire and additional reviewer capacity routes from other teams. Queue overflow is a 12-minute-SLA failure surface; managing it is the operations manager's Monday morning metric.
The Handoff to Enrichment and the Step 5 Stage
Step 3 completes when all fields are extracted, exception queue cleared (or pending producer response flagged), and the structured output writes to the workflow engine (Federato or in-platform). The handoff to Step 5 enrichment is the §4 reason chain checkpoint for OCR completeness.
The handoff payload. Step 3 produces a structured JSON document representing the submission: account metadata (named insured, FEIN, NAICS, etc.), 47 location records (each with TIV, BI, COPE attributes, roof age status), 5-year loss-run summary (frequency, severity, cause distribution), COPE narrative summary (large-location qualitative attributes), and metadata (extraction timestamps, confidence scores per field, exception queue resolution log). The JSON is the structured input for Step 5 enrichment APIs to FEMA NRI, EagleView, D&B, Moody's Orbis, Verisk AIR/RMS/KCC.
The §4 reason chain at the handoff. File-note artifact (per Lesson 3 template) captures: AI invocation (Hyperscience or Indico or hybrid), per-field confidence distribution, exception queue review log with reviewer signatures, producer follow-up status if any, downstream Step 5 enrichment queue write timestamp. The reason chain is the structural anchor for §4.4 documentation across the rest of the pipeline. A §4 examiner six months later can trace any submission's Step 3 IDP quality through this reason chain.
Key Takeaways
- The ten-attachment commercial submission contains seven document types. ACORD 125 (cover), ACORD 140 (property section), ACORD 126 (GL section), SOV, loss run, COPE narrative, and EagleView photo packet. Each type has different extraction characteristics and accuracy targets.
- Hyperscience and Indico bet on different architectures. Hyperscience's keyer-trained model wins on form-structured extraction (ACORD forms, standard SOVs) at 99.5%+ accuracy. Indico's transformer model wins on long-form unstructured text (COPE narratives, broker cover letters, loss-run prose) at 92-97% accuracy where Hyperscience runs 78-85%. Hybrid deployment captures the best of each at $100K-$300K additional annual cost.
- Field-level accuracy targets drive the exception triage logic. Named insured legal name and FEIN at 99.9% (downstream PAS write); NAICS and effective date at 99.5%; roof age at 90% (often missing in source); COPE narrative fields at 88-92%. Per-field targets inform queue routing.
- The 200-dpi loss run is recoverable with IDP discipline. Naive OCR returns 60-75% accuracy on bled scans; disciplined Hyperscience + Indico hybrid extraction reaches 90%+ with exception triage on remaining 10%. Carriers with 95%+ loss-run extraction accuracy see 2-4 LR-points better loss development than carriers with 80-85% accuracy.
- The SOV with 18 missing roof ages on 47 locations is normal, not exceptional. Workflow: IDP extracts; Step 5 EagleView fills 15 of 18 from aerial imagery at 85% within ±2 years confidence; 3 rural locations queue for producer follow-up. Underwriter at Step 8 sees a populated SOV with mixed-source roof ages and confidence intervals.
- The exception queue's SLA discipline separates 99%+ effective accuracy from 87%. High-confidence fields (above 95%) flow directly; 50-95% queue for review at 80-120 fields per hour; under 50% route to senior reviewer or producer. Queue depth above 4-6 hours triggers operational alerts.
- Step 3 completion produces the structured JSON payload that drives all downstream steps. Account metadata, per-location records, loss-run summary, COPE narrative, metadata with confidence scores and exception log. The §4 reason chain captures AI invocation, per-field confidence distribution, exception review log, producer follow-up status, and enrichment queue handoff timestamp.
- The 2026 IDP market split - 45% Hyperscience-only, 30% Indico-only, 25% hybrid - is shifting toward hybrid as carriers recognize architectural complementarity. The 4-8 LR-equivalent point gain from better extraction quality justifies the $100K-$300K additional vendor cost at any carrier above $200M GWP.
Skill.re