โ†
AI for Banking & Lending
Proficient ยท M4 ยท lesson 4 of 19 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Catching Hallucinations in Financial Output
๐Ÿ“–
now learning

Catching Hallucinations in Financial Output

15 min

The following scenario is a composite illustration drawn from patterns common in AI-assisted lending workflows; it does not depict a specific institution or transaction. The portfolio review meeting was scheduled for 9 a.m. on a Thursday. The relationship manager had used an AI drafting tool to produce a one-page deal summary for a $3.8 million commercial real estate refinance: the borrower's company had operated for eleven years, the collateral was a stabilized mixed-use building in the bank's core market, and the debt service coverage ratio (DSCR, the ratio of the property's net operating income to its annual debt service, the core underwriting metric for income-producing commercial real estate) was stated in the AI summary as 1.42. The loan committee chair looked at 1.42, noted it cleared the bank's 1.25 minimum with comfortable margin, and asked no further questions. The committee approved the deal in eight minutes. The underwriter who picked up the file the next morning pulled the Schedule E from the most recent tax return, calculated the DSCR from the actual figures, and got 1.08. The difference between 1.42 and 1.08 is not a rounding issue. It is the difference between a clean approval and a loan that does not meet the bank's credit policy and requires either additional collateral, a reduced loan amount, a committee exception, or a decline. The AI had not lied. It had done what generative AI does: it produced a plausible number that fit the pattern of DSCR figures for mixed-use commercial deals in this size range. It had no access to the actual Schedule E. It generated what a DSCR looks like for a deal like this, and the result was confidently, consequentially wrong.

Why Financial Output Hallucinations Are Different

The word "hallucination" in the context of artificial intelligence (AI, the broad category of systems that use statistical learning to perform tasks traditionally requiring human intelligence, including large language models and generative AI systems) refers to outputs the model produces that are factually incorrect but stated with the same confidence as correct outputs. In a general content context, a hallucination might be a wrong date for a historical event or an incorrect attribution of a quote. Those are correctable and generally low-stakes. In financial output, the hallucination failure mode is materially different in two ways: the errors are numerically precise (not vague claims but specific dollar figures, ratios, and percentages), and the consequences attach to regulated decisions where an error creates credit risk, compliance risk, or fair-lending risk simultaneously.

A DSCR hallucination does not look like a hallucination. It looks like a number. It has decimal precision. It is embedded in a coherent analytical sentence. A reviewer reading for narrative quality will not catch it unless they verify it against the source document. This is what makes hallucinations in financial output more dangerous than hallucinations in general content: they are invisible to a quality read and visible only to a source check.

The categories of financial output hallucination that appear most frequently in banking AI workflows fall into four groups.

Ratio hallucinations. The model generates a plausible ratio (DSCR, loan-to-value (LTV, the loan amount divided by the appraised or purchase value of the collateral), debt-to-income (DTI, the ratio of total monthly debt obligations to gross monthly income), global cash flow coverage) based on its learned pattern of what a ratio looks like for a borrower of this type. The ratio may be directionally correct but numerically wrong. A 1.42 for a deal that should be 1.08, or a 68% LTV for collateral that should produce 74%, are both hallucination outputs that look like analytical precision.

Income figure hallucinations. The model generates income figures that fit the borrower's described profile rather than tracing to the actual tax return. Self-employment income, K-1 pass-through income, Schedule E rental income, and any income type calculated through multi-year averaging or add-back adjustments are especially vulnerable because the calculation methodology is complex and the model's pattern-matching will produce plausible but unverified results. The add-back hallucination is particularly dangerous: AI tools will frequently "add back" depreciation, depletion, or other non-cash charges at an assumed rate rather than at the actual figure from the return.

Covenant and condition hallucinations. In commercial credit memos, AI tools will generate boilerplate covenants and conditions that fit a deal of this type without these terms having been negotiated or specified. A covenant requiring a minimum DSCR of 1.25, a reporting requirement for quarterly financial statements, or a condition requiring assignment of leases may appear in an AI-generated draft even though the institution did not propose or approve those specific terms. When the credit committee approves the memo with those covenants, the approved terms include conditions the bank did not intend, creating a documentation problem at closing and a compliance problem if the covenants are never monitored.

Adverse-action reason hallucinations. For declined applications, an AI tool generating adverse-action reasons under the Equal Credit Opportunity Act (ECOA, the federal statute prohibiting credit discrimination on the basis of race, color, religion, national origin, sex, marital status, age, or receipt of public assistance) and Regulation B (Reg B, the federal regulation implementing ECOA, codified at 12 CFR Part 1002, requiring that adverse-action notices state specific, accurate reasons for denial) may generate reasons that fit the pattern of a denial for this borrower type but do not accurately reflect the actual basis for the decision. An AI that generates "insufficient income" as the first adverse-action reason when the actual primary reason was "insufficient collateral value" has produced an inaccurate adverse-action notice. Under Reg B, that inaccuracy is a compliance violation regardless of whether the denial itself was correct.

A financial hallucination looks like analysis. It has numbers, ratios, and precision. It is caught only by source verification, not by reading for quality. The verification step is not a review of the writing; it is an audit of the facts.

The Source-Document Verification Framework

The only reliable defense against financial hallucinations is a systematic source-document verification framework applied to every AI-produced financial output before it is used in a decision, submitted as an institutional deliverable, or transmitted to a borrower or regulator. "Systematic" means structured, not intuitive. A general read of the memo for coherence is not verification. Verification means tracing each figure to its source and confirming the source says what the memo claims it says.

The framework has four layers, each targeting a different category of hallucination risk.

Layer one: the quantitative trace. For every numerical figure in the AI output (ratios, income totals, balances, rates, terms, percentages), identify the source document that should contain the underlying data and verify the figure against that document. The trace is documented: a verification log that records the figure as stated in the AI output, the source document and specific location (tax return line number, financial statement line item, credit report tradeline), and the actual value found in the source. Any discrepancy is a hallucination flag that requires correction before the output is used. The trace is not optional for any figure. If the DSCR is stated in the memo, the verifier traces it to the net operating income figure from the Schedule E and the debt service calculation from the proposed loan terms. Both components are checked, not just the ratio itself.

Layer two: the policy check. Financial AI output often incorporates implicit policy assumptions that may not match the institution's current credit policy. A credit memo that states "the loan meets all policy requirements" has made a compliance assertion that must be verified against the policy document. The layer-two check confirms that any policy-referenced thresholds in the AI output (minimum DSCR, maximum LTV, maximum loan-to-cost, minimum credit score, maximum loan size) match the institution's current, version-controlled policy. A tool trained on general lending data will sometimes apply industry norms rather than the specific institution's parameters, and industry norms and institutional policy may differ in material ways.

Layer three: the consistency check. AI tools sometimes produce internally inconsistent outputs where different sections of the document calculate the same figure differently. A credit memo that states DSCR as 1.42 in the executive summary, calculates the underlying net operating income and debt service in the financial analysis section, and produces a DSCR of 1.38 from those components has an internal inconsistency. The layer-three check confirms that calculated figures in one section are consistent with the underlying components stated elsewhere. Inconsistency is itself a hallucination signal: the model is generating each section somewhat independently, and the numbers do not always agree across sections.

Layer four: the regulatory language check. For outputs with direct regulatory consequences (adverse-action notices, BSA/AML (Bank Secrecy Act/Anti-Money Laundering, the regulatory regime requiring financial institutions to maintain programs to detect, report, and prevent money laundering, codified in 31 U.S.C. 5311 et seq.) narratives, ECOA-required disclosures), verify that the regulatory language in the AI output is accurate, complete, and consistent with current regulatory requirements. Adverse-action language must cite the correct adverse-action code categories. Reg B notices must include the legally required elements. BSA/AML Suspicious Activity Report (SAR, the regulatory filing required when a financial institution suspects money laundering or other financial crimes) narratives must be complete and accurately represent the activity being reported. Regulatory language hallucinations are not just inaccurate; they are non-compliant.

Cross-Referencing the File and the Core

In a banking context, "the file" refers to the loan file: the collection of source documents associated with a specific application or credit relationship. "The core" refers to the bank's core banking system, the transaction processing platform that maintains account balances, payment history, deposit records, and the authoritative financial position of existing customers. For AI outputs that touch financial data from either source, cross-referencing means verifying the AI's claims against both the document-level file data and the system-of-record data in the core.

This cross-reference matters because AI tools may have access to one source but not the other, or may have been given document excerpts rather than complete documents. A credit memo that correctly extracts the borrower's reported income from submitted tax returns but does not check the borrower's actual deposit and payment history in the core is working from incomplete information. For an existing customer, the core's deposit data may confirm the tax-return income (a pattern-consistent inflow that matches the reported self-employment income), or it may contradict it (deposit activity inconsistent with the stated income level), and the discrepancy matters for both credit quality and BSA/AML compliance.

The specific cross-references that catch the most consequential hallucinations in banking AI output are:

Income verification: file against core. The income figures from the tax returns and pay stubs in the loan file should be compared against the actual deposit activity in the borrower's core banking accounts (when the borrower is an existing customer). If the borrower reports $180,000 in annual self-employment income on the Schedule C but the core shows deposit activity of $90,000 over the prior 12 months, the discrepancy requires resolution before the income figure is used in the underwriting. AI tools do not make this cross-reference automatically; the human verifier must do it.

Debt obligation verification: file against core and bureau. Debts listed in the AI output should be verified against both the credit bureau report in the file and the core's existing loan and deposit records for the customer. A customer who has an existing line of credit or installment loan at the originating institution that appears on the bank's core but was not prominently featured in the application may be omitted from the AI's debt inventory. The DTI calculation, and therefore the underwriting conclusion, changes if a significant obligation is omitted.

Payment history verification: file against core. AI tools analyzing an existing customer's creditworthiness will sometimes describe a payment history based on what was stated in the credit bureau report. The core's actual payment records are more granular and more recent. A commercial borrower whose existing term loan shows a 60-day delinquency in the core banking system that has not yet been reflected in the credit bureau report has a materially different risk profile than the AI's bureau-based description would suggest.

Collateral value verification: file against appraisal. The collateral value stated in the AI output must be verified against the actual appraisal report in the file, not against the AI's pattern-based estimate of what a property of this type and location should be worth. AI tools sometimes generate plausible collateral value figures when the actual appraisal is not provided in the prompt or was summarized in the input rather than attached in full. The appraisal report is the authoritative source; any AI-stated value that differs from it is a hallucination regardless of how plausible the AI's figure looks.

Building the Verification Checklist

The verification checklist is the operational tool that makes the source-document framework executable under time pressure. Without a structured checklist, verification tends toward intuitive and incomplete. The relationship manager in the opening story "reviewed" the AI memo; she did not verify it against a checklist. A checklist makes the absence of source-confirmation visible: every unchecked item is a gap, not an assumed correctness.

A verification checklist for a commercial real estate credit memo should contain the following categories, with specific items under each:

Income and NOI verification. Line items: gross rental income (verified against Schedule E line 3 or borrower operating statement), vacancy and credit loss (verified against actual historical vacancy history and Schedule E), operating expenses by category (taxes, insurance, utilities, management, maintenance, each from Schedule E or operating statement), net operating income as calculated (the sum should equal the stated NOI figure), depreciation add-back (from the applicable tax schedule, not an assumed amount).

Debt service verification. Line items: proposed loan amount (from the loan application), interest rate (from the institution's current rate sheet or commitment), amortization period and term (from the proposed terms), calculated monthly payment (verified against an amortization schedule, not the AI's stated figure), annual debt service (12 times the monthly payment), DSCR as calculated (NOI divided by annual debt service, with both components verified above).

Collateral verification. Line items: appraised or purchase value (from the appraisal report or purchase and sale agreement), loan amount (from the application), LTV as calculated (loan amount divided by collateral value, with both components verified), lien position and prior encumbrances (from the title search or existing mortgage records).

Borrower financial verification. Line items: personal income from all sources (W-2, Schedule C, K-1, Schedule E, Social Security, other, each from the applicable document), personal monthly obligations (from the credit bureau report, complete list), personal DTI as calculated, global cash flow (business income plus personal income minus all debt service obligations), credit score (from the bureau report, verified as the correct bureau and report type).

Policy compliance verification. Line items: DSCR meets minimum (stated minimum from current policy), LTV within maximum (stated maximum from current policy), loan size within approval authority or committee requirement (from current policy), credit score above floor (from current policy), any applicable exceptions documented (with the required exception approval process initiated).

Adverse-action reason verification (for declined applications). Line items: primary reason code (verified as the actual primary basis for the decline, not a pattern-based generic reason), secondary reason codes if applicable (each verified as an actual contributing factor in the file), Reg B required elements (notice includes the required elements: name and address of creditor, statement of action taken, notification of the right to the specific reasons, ECOA statement).

The checklist is completed before the AI-produced document is submitted, signed, or transmitted. Each item is marked as verified with the source document and location noted, or flagged for correction. Any item that cannot be verified against a source document is a hallucination risk that must be resolved before the output is used.

Hallucinations in Adverse Action and Fair-Lending Consequences

The fair-lending consequences of adverse-action hallucinations deserve specific treatment because the regulatory framework for adverse action is precise and the consequences of inaccuracy are independent of intent. Under ECOA and Reg B, a financial institution that denies a credit application must provide the applicant with a notice that states the specific, accurate reasons for the denial. "The model said no" is not a legally sufficient reason. Neither is a reason code generated by an AI that does not accurately reflect the actual basis for the decision.

The most common adverse-action hallucination in AI-generated output is reason code selection that fits the pattern of a denial for a borrower of this type without tracing to the actual decision factors in the file. An AI tool that has seen many credit denials will learn that certain reason codes appear together, and it will generate a plausible set of reasons for a denial based on the borrower's profile. But the reasons it generates are the typical reasons for a denial of this type, not necessarily the reasons for this specific denial. If the actual decision was driven by insufficient appraisal value but the AI generated "insufficient income" as the primary reason (because insufficient income is a more common primary reason in the model's training distribution), the adverse-action notice is factually wrong.

The fair-lending dimension of adverse-action accuracy is not abstract. If an institution's AI tool systematically generates inaccurate adverse-action reasons for a specific demographic group (for example, more frequently citing income-related reasons for minority applicants when the actual reason was collateral-related), the pattern of inaccurate notices is itself evidence of disparate impact (the legal doctrine, central to fair lending under ECOA and the Fair Housing Act, that a facially neutral policy or practice that produces statistically significant disparate outcomes for a protected class is unlawful unless justified by business necessity with no less discriminatory alternative). Inaccurate AI-generated adverse-action reasons are not just a compliance documentation problem; they can become fair-lending evidence.

The verification requirement for adverse-action notices is therefore more exacting than for other financial output. Each reason code must be verified as an actual contributing factor in the file, supported by specific evidence, and accurately ranked by its relative contribution to the decision. The human officer who signs the adverse-action notice is affirming the accuracy of the reasons. That affirmation must be based on a review of the file, not on an assumption that the AI's generated reasons are correct.

Model risk management (MRM) under OCC Bulletin 2026-13 (the April 2026 interagency update issued by the OCC (Office of the Comptroller of the Currency, the primary federal regulator for nationally chartered banks), the Federal Reserve, and the Federal Deposit Insurance Corporation (FDIC), superseding OCC 2011-12 and explicitly pulling AI and generative AI under model-risk, fair-lending, third-party, and board-governance expectations) requires that institutions document their hallucination risk assessment and controls for any AI tool used in regulated functions. For an AI tool generating adverse-action notices, the model-risk documentation must describe how the institution detects and prevents the issuance of inaccurate reasons. A verification checklist is the primary control. The model-risk file must describe the checklist, confirm it is required in the workflow, and provide evidence it is being used.

The Grounded Prompt as Hallucination Prevention

Prevention is materially more efficient than detection. A well-constructed prompt that grounds the AI on actual file data substantially reduces the frequency and severity of hallucinations, reducing the verification burden and the risk that a hallucination passes through an imperfect verification process. The grounded prompt is the upstream control; the verification checklist is the downstream control. Both are required, but reducing hallucination frequency upstream makes the downstream verification step more manageable and more reliable.

Grounding in the AI context means providing the model with specific, accurate input data from the actual source documents rather than asking it to infer or calculate from a general description. Retrieval-augmented generation (RAG, an AI architecture in which the model is constrained to generate outputs grounded in specific retrieved documents or data rather than in its general training knowledge) applied to a loan file means the AI is generating from the actual documents rather than from pattern-matching. Even without a full RAG architecture, a prompt that includes specific extracted figures from the source documents is more grounded than a prompt that describes the borrower in general terms.

The grounding principles for financial output are:

Enter key figures explicitly. Do not ask the AI to calculate DSCR from a verbal description of the property's income. Enter the actual net operating income figure from the Schedule E and the actual debt service from the proposed terms. State the source for each figure in the prompt (for example: "NOI per Schedule E, 2024 return, Line 28: $187,400; annual debt service at proposed terms: $142,200"). When the AI uses figures you have entered rather than generating its own, hallucination risk drops dramatically for those figures.

Restrict to file-grounded claims. Instruct the AI explicitly: "Base all figures and claims in this memo on the data I have provided. Do not generate estimated figures or fill in data I have not provided. If data is missing, indicate what is missing rather than estimating." This instruction does not eliminate hallucination entirely, but it substantially reduces pattern-based filling of gaps. The AI that is instructed to indicate missing data is more useful than one that generates plausible-looking estimates for the gaps.

Require source citations. Instruct the AI to cite the source document and location for every numerical claim: "For every figure you include, note the source document and line in parentheses after the figure." This creates a verification map and forces the AI to indicate when it is working from provided data versus generating from context. An AI output that includes "(Schedule E, 2024, Line 28)" after the NOI figure and "(provided loan terms)" after the debt service figure is grounded and verifiable. An AI output that states "net operating income of $187,400" without a citation is potentially hallucinated and requires verification against the source document before it can be relied upon.

Use a structured output format. Requesting the output in a structured format (for example, a JSON object with labeled fields for each verified figure, followed by the narrative memo) separates the quantitative assertions from the narrative and makes the verification layer-one check more systematic. A structured output with a DSCR field, an NOI field, a debt service field, and an LTV field, each labeled, is easier to verify than the same figures embedded in flowing prose where they may appear in multiple places with slight variations.

Hallucination Risk in BSA/AML Output

BSA/AML workflows present a distinct hallucination risk profile. The AI is not generating financial ratios from file data; it is summarizing transaction patterns, generating SAR narratives, and characterizing suspicious activity for regulatory filings. The hallucination risk here is not a wrong number but a wrong characterization of the activity, a misattribution of transactions, or a narrative that sounds complete and plausible but omits material elements that the SAR regulations require.

The BSA/AML hallucination failure modes most commonly seen in AI-assisted alert workflows are:

Transaction count and amount inaccuracies. AI summaries of transaction activity will sometimes generate transaction counts and total amounts that are approximations rather than verified figures. A SAR narrative that states "the subject made 47 cash deposits totaling $148,000 over the review period" when the actual figures were 39 deposits totaling $134,500 is factually wrong. Because SAR narratives are regulatory filings, factual inaccuracies in them carry regulatory and reputational consequences. Every transaction count and amount in an AI-generated SAR narrative requires verification against the actual transaction data.

Pattern characterization errors. AI tools trained on BSA/AML data know what suspicious activity patterns look like and will generate pattern characterizations that fit the alert type. A structuring alert summary that characterizes deposits as "appearing to be structured to avoid the $10,000 Currency Transaction Report (CTR) reporting threshold" is correct if the deposits are actually just below $10,000 in a pattern consistent with intentional structuring, but not correct if the deposits vary widely and the alert was triggered by volume rather than size pattern. The analyst must verify the pattern characterization against the actual transaction data before including it in the SAR narrative.

Subject attribute errors. AI tools sometimes generate subject attributes (occupation, business type, transaction counterparties) based on patterns rather than verified file data. A SAR narrative that describes the subject as "a cash-intensive retail business" based on pattern inference rather than verified account documentation may be inaccurate, and the inaccuracy could affect law enforcement's investigation. All subject attributes in AI-generated SAR narratives require verification against the account opening documents and transaction records.

The false-positive rate context is relevant here. Industry data indicates that BSA/AML alert workflows carry false-positive rates in the range of 90 to 95%, meaning the vast majority of alerts do not result in SAR filings. An AI tool that helps analysts triage this volume efficiently is genuinely valuable. But the efficiency gain comes from the triage function, not from the narrative generation function. The AI can help an analyst sort through 200 alerts to identify the 10 to 20 that warrant closer review. The SAR narratives for those 10 to 20 require the same source-verification standard as any other regulated output. Speed in triage does not mean speed in narrative verification.

Key Takeaways

  • Financial output hallucinations are numerically precise and invisible to a quality read. A DSCR of 1.42 that should be 1.08 looks like analysis; it is caught only by verifying the underlying NOI and debt service figures against their source documents. The verification step is an audit of facts, not a review of writing quality.
  • The four categories of financial hallucination are ratio hallucinations, income figure hallucinations, covenant and condition hallucinations, and adverse-action reason hallucinations. Each requires a specific verification check against the applicable source document or institutional policy.
  • The source-document verification framework has four layers: the quantitative trace (every number to its source), the policy check (thresholds against current credit policy), the consistency check (internal agreement across document sections), and the regulatory language check (adverse-action and BSA/AML outputs against applicable regulatory requirements).
  • Cross-referencing the file against the core catches income and payment history discrepancies that document-only review misses. For existing customers, the core's actual deposit and payment history is a verification source that AI tools cannot access and that human reviewers must apply.
  • A grounded prompt, which enters key figures explicitly from source documents and requires source citations, reduces hallucination frequency upstream and makes the verification checklist more manageable. Grounding is a prevention control; the checklist is a detection control; both are required.
  • Adverse-action reason hallucinations are fair-lending risks, not just documentation errors. An AI that systematically generates inaccurate reasons for a protected class creates disparate-impact evidence. Each reason code must be verified as an actual contributing factor to the specific decision, not a pattern-based generic reason for a denial of this type.
  • OCC Bulletin 2026-13 requires institutions to document their hallucination risk assessment and controls for AI tools used in regulated functions. A verification checklist is the primary control; the model-risk file must describe the checklist, confirm it is required, and provide evidence it is being used consistently.
  • In BSA/AML workflows, hallucination risk is in characterization (pattern descriptions, subject attributes, transaction summaries) rather than financial ratios. AI supports triage efficiently; every SAR narrative requires the same source-verification standard as any other regulated filing, regardless of the efficiency gain in the triage step.