Catching the Hallucinated Number
The tax return looked clean at first glance. Marcus, a loan processor at a regional bank in the Southeast, was clearing a stack of files on a Friday afternoon when the self-employed borrower's 1040 came through the AI extraction pipeline. The output showed adjusted gross income of $94,200, schedule C net profit of $91,500, and a year-over-year increase in gross receipts that made the file look well-qualified for the $280,000 refinance. Marcus accepted the extraction, keyed the income figures into the loan origination system (LOS, the software platform that routes every field from application to close), and moved the file forward to underwriting. Three days later, the underwriter pulled the actual PDF, cross-referenced the AI output against the source document, and found a field the extraction model had silently transposed: Schedule C net profit on the document was $19,500, not $91,500. One digit, transposed in position. A $72,000 difference in qualifying income on a $280,000 loan. The file did not qualify. The borrower had already been verbally told the loan was on track. And the institution had spent three days of processing, underwriting review time, and borrower expectation management on a file that should never have advanced past the income verification stage. That single uncaught transposition cost the institution hours of rework, cost the borrower a week of false confidence, and in a volume origination environment where dozens of files move simultaneously, it is precisely the type of error that does not stay isolated. This lesson is about the discipline that catches it before it moves.
Why Extracted Numbers Need a Dedicated Verification Discipline
The AI extraction model that produced $91,500 from a document that reads $19,500 was not making something up. It read the characters on the page and assembled a number. The error was a transposition in the assembly step, a known failure mode of optical character recognition (OCR, the technology that converts scanned images into machine-readable text) on certain digit combinations in small print at standard scan resolution. The model was not hallucinating in the generative sense: it was not fabricating a plausible income figure from its training data when no source data was available. It was misreading a real digit on a real document.
This distinction matters operationally, but it does not reduce the severity of the risk. Whether a number is invented outright by a generative model or misread from the source document by an OCR pipeline, the effect on the lending decision is identical: the underwriter, the underwriting system, and the credit file now contain an income figure that does not match what the borrower's tax return actually says. Under the Equal Credit Opportunity Act (ECOA, the federal statute prohibiting credit discrimination) and its implementing rule, Regulation B (Reg B, 12 CFR Part 1002), the adverse-action reason for a denial must reflect the actual factors in the applicant's file. If the income figure that drove the decision is wrong because the extraction model misread a digit, any adverse-action reason based on that figure is also wrong, a Regulation B violation independent of whether the decision itself was correct.
The verification discipline described in this lesson applies to both misread numbers and fabricated numbers, because the cross-check method is the same: confirm that the value in the LOS or in the draft document matches the corresponding source line on the actual document. The difference in error origin affects the investigation protocol after a mismatch is found (was the digit on the page ambiguous? was the field label wrong? did the model fill a gap with a plausible estimate?) but not the front-line check itself.
By 2024, roughly 38 percent of mortgage lenders were using AI or machine learning in some part of their underwriting process, up from 15 percent in 2023. That adoption rate means the volume of AI-extracted numbers flowing through LOS systems and into credit decisions has grown substantially faster than the verification disciplines designed to catch extraction errors. The institutions that build rigorous cross-check workflows before they scale AI-assisted extraction are the ones that do not discover their extraction error rate through adverse-action complaints, reprocessed files, or an OCC (Office of the Comptroller of the Currency) model-risk examination finding that their extraction tool has never been validated on production documents.
The Anatomy of a Hallucinated or Misread Number
Income and asset figures in lending documents fail in specific patterns, and knowing those patterns is what makes a verification workflow efficient rather than exhaustive. You do not need to re-read every word on every page. You need to know exactly where the high-risk extraction points are and verify those specifically.
The transposition error. The $19,500 to $91,500 case is a transposition: the model read the digits in an order that produces a plausible but wrong number. Transpositions are most common when digit pairs occupy adjacent positions on the page in small print, when the scan has uneven contrast across the number field, or when the digit pair involves visually similar characters (1 and 7, 0 and 6, 5 and 6). The check for transposition is not a plausibility check, it is a character-by-character comparison between the extracted value and the source line. Plausibility does not catch $91,500 from a file where the borrower has been in business for three years and the prior-year return showed $88,000 in net profit. That level of cross-document plausibility checking is useful as a secondary flag but cannot substitute for the source-line comparison.
The column confusion error. Paystubs and financial statements often present the same data field in two columns: current period and year-to-date, or monthly and annual. An extraction model that assigns a value to the wrong column produces an error that is structurally correct (the field label and value are both present in the document) but semantically wrong (the year-to-date gross pay of $72,150 ends up in the current-period field, and the current-period gross pay of $3,607.50 ends up nowhere). This type of error is caught by the field assignment check in the cross-check workflow: is the current-period figure smaller than the year-to-date figure by an amount consistent with the pay frequency?
The schedule miss error. For 1040 income tax returns, the most consequential extraction error is the schedule miss: the model extracts the 1040 cover page accurately but does not process the attached schedules, leaving Schedule C net profit, Schedule E rental income, and Schedule F farm income absent from the extraction output. The file appears income-complete because all extracted fields have values, but the qualifying income calculation is missing the schedule-based income adjustments that determine whether the borrower qualifies. This error is invisible unless the verifier checks the document page count against the extraction scope and confirms that every schedule present in the document was extracted.
The fabricated estimate. When an extraction model cannot clearly read a field value, some models generate a plausible estimate based on surrounding context rather than returning a "not visible" flag. A model that reads all surrounding fields correctly but finds the gross pay field obscured by a scan artifact may estimate the gross pay by inferring it from the year-to-date and the number of pay periods visible from the pay dates. The estimated value is internally consistent and passes a reasonableness check. It is also fabricated: it does not appear anywhere in the document as a discrete value. This is the overlap between extraction error and generative hallucination, and it is the error type that a source-label check catches and a plausibility check misses entirely, because the estimate is by construction plausible.
The stale document error. When a processor submits the wrong year's tax return, an older bank statement, or a paystub from a prior employer, the extraction model processes whatever document it receives. The figures it extracts are accurately read from the document, but the document is not the one the underwriter needs. This error is not an AI failure. It is a document management failure that the AI workflow does not detect, because the model reads the document it receives and has no way to know that a 2023 tax return was submitted when the guideline requires 2024. The verification step must include a date check: confirm the document year matches the guideline requirement before accepting any extracted values from it.
The Cross-Check Protocol: A Worked Example
The most effective way to internalize the verification discipline is to walk through a specific, detailed example of how the cross-check applies to a real extraction output. This example uses a Schedule C self-employed borrower because that income type presents the highest verification complexity and the highest potential for consequential extraction errors.
The borrower is a freelance software consultant. The loan file includes a 2024 1040 with a Schedule C attached, two years of supporting documentation, and two months of recent bank statements. The AI extraction returns the following output:
From 2024 Form 1040: Line 1a (Wages): $0 [from 1040, Line 1a] Line 11 (AGI): $87,340 [from 1040, Line 11] Line 15 (Taxable income): $74,890 [from 1040, Line 15] Filing status: Single [from 1040 header] Tax year: 2024 [from 1040 header] From 2024 Schedule C: Line 7 (Gross receipts): $127,500 [from Sch C, Line 7] Line 28 (Total expenses): $40,160 [from Sch C, Line 28] Line 31 (Net profit/loss): $87,340 [from Sch C, Line 31] Business name: Vance Technology Consulting [from Sch C, Part I] Business activity: Software consulting [from Sch C, Line B]
The cross-check proceeds in four steps.
Step one: confirm document scope. The 1040 is for tax year 2024. The loan file requires 2024 and 2023 returns. Is the 2023 return also present? Confirm. The extraction output covers only 2024. Is the 2023 return present in the document set? If not, flag for collection before proceeding.
Step two: source-line comparison for each extracted value. Open the 2024 1040 PDF. Go to Line 1a. Does it read $0? Confirm. Go to Line 11. Does it read $87,340? Confirm. Go to Line 15. Does it read $74,890? Confirm. Open the Schedule C. Go to Line 7. Does it read $127,500? Confirm. Go to Line 28. Does it read $40,160? Confirm. Go to Line 31. Does it read $87,340? Confirm.
In this case, all values check out. But the process reveals an internal relationship worth verifying explicitly: Line 31 (net profit) equals Line 11 (AGI). That is only consistent if Schedule C is the borrower's only income source and there are no adjustments to income on Schedule 1. Confirm that Schedule 1 is either absent or shows no other adjustments. If Schedule 1 is present and shows adjustments, the AGI does not equal Schedule C net profit, and either the extraction is missing a line or there is an additional income source not captured in the extraction output.
Step three: internal math check. Gross receipts minus total expenses should equal net profit: $127,500 minus $40,160 equals $87,340. It does. The math is internally consistent. Note that this check does not confirm the document values are correct; it confirms the extraction is internally consistent. If the underlying document had a math error, that math error would be consistently reproduced in the extraction.
Step four: qualifying income flag. The extraction output correctly captured the schedule values, but qualifying income for a self-employed borrower is not the same as Schedule C net profit. Under standard guidelines, the lender adds back certain non-cash deductions (depreciation from line 13, depletion from line 12 if applicable) and may subtract business use of home from line 30 if the borrower takes that deduction. The extraction output does not show Lines 12, 13, or 30 because the prompt did not request them. If the underwriter will apply add-back calculations, the extraction scope should be expanded to include those lines, or the processor should flag the file for a supplemental extraction run covering the add-back lines before the income calculation is performed.
This is the moment where the verification workflow produces value beyond error-catching: it surfaces the information gap that would otherwise not be visible until underwriting, two steps later in the workflow, when the cost of going back to the borrower for additional documentation is higher.
The Bank Statement Verification: The Deposit That Does Not Belong
Bank statement verification for asset purposes has a different verification structure than income document verification, because the primary concern is not a single misread field but a pattern within the transaction history. The two highest-stakes verification points on a bank statement are: the ending balance (which must match the LOS reserves figure exactly), and large or non-recurring deposits (which must be sourced according to loan guidelines before they can be counted toward qualifying assets or down payment).
Here is the verification workflow for an AI extraction that identified large deposits on a 60-day bank statement:
AI extraction output for large deposits (threshold: $1,000): 03/04/2025: $2,500 ACH DEPOSIT VANCE TECH CONSULTING LLC [Pg 1, line 14] 03/18/2025: $4,200 ACH DEPOSIT VANCE TECH CONSULTING LLC [Pg 1, line 22] 04/02/2025: $2,500 ACH DEPOSIT VANCE TECH CONSULTING LLC [Pg 2, line 6] 04/07/2025: $5,000 WIRE TRANSFER INBOUND FAMILY GIFT VARGAS [Pg 2, line 11] 04/19/2025: $2,500 ACH DEPOSIT VANCE TECH CONSULTING LLC [Pg 2, line 19]
The cross-check for this output has two stages. First, confirm each listed deposit against the page and line reference on the statement. Open the statement PDF. Go to page 1, line 14. Does it read: 03/04/2025, $2,500, ACH DEPOSIT VANCE TECH CONSULTING LLC? Confirm the amount and the description exactly. Repeat for each deposit.
Second, check for completeness: are there any deposits of $1,000 or more in the statement that do not appear in the extraction output? Scan the deposit column of the statement visually for any entry that looks like it meets the threshold. This completeness check is important because the extraction model may have missed a deposit due to page layout issues, small print, or a deposit that appears in a "pending" or "memo" section rather than the standard transaction rows.
Once the deposit list is confirmed as complete and accurate, the underwriter reviews each large deposit for sourcing. The three recurring $2,500 ACH deposits from Vance Tech Consulting align with the borrower's documented business income and are consistent with the self-employment documentation already in the file. The $4,200 and $5,000 deposits require investigation. The $5,000 wire identified as "FAMILY GIFT VARGAS" requires a gift letter from the sender and documentation that the gift is not a loan. The $4,200 ACH from the business requires confirmation that it does not represent double-counting of business income already included in the qualifying income calculation.
Notice what the verification workflow surfaces: the business deposit that might constitute double-counting is not an AI error. The AI extraction correctly identified the deposit. The human review step is what identifies the underwriting implication of a correctly extracted figure.
Building a Verification Log: The Documentation an Examiner Reads
A verification that is not logged did not happen, for purposes of any examination or dispute. The OCC (Office of the Comptroller of the Currency) model-risk framework under Bulletin 2026-13 expects institutions to maintain documentation of how AI model outputs are reviewed before they influence credit decisions. For extraction tools, that documentation is the verification log: a record that links each AI-extracted value to its source document location, confirms the match or records the discrepancy and how it was resolved, and identifies the person who performed the verification and the date it was completed.
A minimal but defensible verification log for an income document looks like this:
INCOME VERIFICATION LOG Borrower: Marcus Johnson Loan Number: [LOS reference] Document: 2024 Form 1040 with Schedule C Document Date/Tax Year: 2024 Verifier: [Name, Title] Verification Date: [Date] Field: Schedule C Net Profit (Line 31) AI Extracted Value: $87,340 Source Document Value: $87,340 Source Location: Schedule C, Line 31, Page 3 of PDF Match: Yes Field: Schedule C Gross Receipts (Line 7) AI Extracted Value: $127,500 Source Document Value: $127,500 Source Location: Schedule C, Line 7, Page 3 of PDF Match: Yes Field: Schedule C Total Expenses (Line 28) AI Extracted Value: $40,160 Source Document Value: $40,160 Source Location: Schedule C, Line 28, Page 3 of PDF Match: Yes Internal Math Check: $127,500 - $40,160 = $87,340. Consistent with Line 31. Supplemental Note: Lines 12, 13, and 30 not extracted. Underwriter to confirm add-back applicability before qualifying income calculation.
This log format records what was extracted, what the source document says, where in the document the value appears, whether they match, and any supplemental flags. If an examiner reviews this file two years later in a fair-lending examination, this log is the documentation that the income figure in the LOS was not accepted from AI output without verification. It is also the documentation that the verification identified a gap in the extraction scope, which was flagged for the underwriter rather than silently omitted.
Institutions that use AI extraction at scale should build this log format into the LOS workflow, not as a separate spreadsheet that processors may or may not maintain, but as a required step in the file progression that cannot be bypassed. A workflow that requires the processor to enter the source-document value for each extracted field before the file can advance to underwriting creates the verification record automatically, without relying on individual discipline under time pressure.
When to Reject an Extraction and Go Back to the Document
Not every extraction output is worth verifying. Some documents are sufficiently degraded, or some extraction outputs sufficiently inconsistent, that the faster path is to return the document for a better copy or to perform manual extraction rather than spending verification time on an output that has a high probability of containing errors.
The decision framework for when to reject an extraction and start over has four triggers.
Trigger one: image quality flag. If the extraction tool returns an image quality warning, or if the verifier observes that the source PDF is clearly a second-generation scan (washed-out contrast, visible copy artifacts, skewed alignment), flag the document for a replacement request before investing verification time. A degraded scan that passes extraction with warnings is a file-level risk that does not go away through verification; it only gets more expensive as it progresses.
Trigger two: more than three discrepancies on a single document. A verification workflow that consistently finds more than three mismatches between extracted and source values on the same document is encountering an extraction quality problem that is systematic rather than incidental. Three individual mismatches on a ten-field document suggest the extraction model is not performing reliably on this document type or quality level. The correct response is to escalate to manual extraction and to log the document as a validation data point for the extraction tool's model-risk file.
Trigger three: internal inconsistency that cannot be resolved by source-line verification. If the extracted values are individually confirmed against the source document but still produce an internally inconsistent result (gross receipts minus expenses does not equal net profit on the schedule, or the sum of monthly deposits does not approach the ending balance minus beginning balance on the statement), the source document itself may have a calculation or printing error. This is not an AI failure, but it is a document integrity issue that requires escalation to the underwriter before the file progresses.
Trigger four: missing pages confirmed after scope check. If the page count in the extraction output is less than the page count in the source PDF, the extraction model did not process all pages. The missing pages may contain material information. Rather than supplementing a partial extraction, return the document to the extraction pipeline with an explicit page range instruction covering all pages, and verify the rerun output from the beginning.
The Unfair, Deceptive, or Abusive Acts or Practices (UDAAP, the broad consumer protection standard under the Dodd-Frank Act) implications of a systematic rejection decision are worth noting. If the institution applies the rejection triggers inconsistently, applying them more stringently to files submitted by borrowers in certain geographic areas or with certain document formats that correlate with protected class, the differential treatment of those files is a UDAAP risk even if each individual decision appears facially neutral. The rejection protocol should be consistently documented and applied across all document types and all borrowers.
Key Takeaways
- The hallucinated or misread number in a lending extraction can be either a generative model fabricating a plausible estimate or an OCR model misreading a digit from the source page. Both produce the same downstream risk: a credit decision based on an income or asset figure that does not match what the borrower's document actually says. The verification discipline is identical for both error types.
- The five primary extraction failure modes are transposition, column confusion, schedule miss, fabricated estimate, and stale document error. Each has a specific check in the cross-check workflow. Transposition is caught by character-level source-line comparison. Column confusion is caught by field assignment verification. Schedule miss is caught by page count and scope review. Fabricated estimate is caught by source-label verification (if the label is absent, the value was inferred). Stale document is caught by date verification before any other step.
- The cross-check protocol has four steps: confirm document scope and page count; compare each extracted value to its source line by location; verify field assignments for current-period versus year-to-date figures; and check internal math consistency. All four steps complete before any extracted value enters the LOS or a draft document.
- A verification log that records the AI-extracted value, the source document value, the source location, the match result, and any discrepancy resolution is the documentation that makes the verification defensible. A verification that is performed but not logged is indistinguishable from a verification that was not performed, for purposes of an OCC 2026-13 model-risk examination.
- Bank statement large-deposit verification has two stages: confirm each extracted deposit against its source location; and check for completeness by scanning the deposit column for above-threshold entries the model may have missed. The human review step surfaces the underwriting implications of correctly extracted deposits, a function no extraction model performs.
- Reject an extraction and return to the source document when: the image quality is flagged or visibly degraded; more than three discrepancies appear on a single document; internal inconsistency cannot be resolved by source-line verification; or the extraction scope is confirmed to be missing pages. Applying rejection triggers consistently across all borrowers and document types is a UDAAP compliance requirement.
- OCC Bulletin 2026-13 requires institutions to document how AI model outputs are reviewed before they influence credit decisions. For extraction tools, this means tracking field-level accuracy by document type in the model-risk file, validated on a sample of production files, and escalating when accuracy on specific document types degrades below the institution's defined threshold.
Skill.re