โ†
AI for Banking & Lending
Aware ยท M1 ยท lesson 1 of 19 ยท in progress
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
AI Hallucinations in a Credit Context
๐Ÿ“–
now learning

AI Hallucinations in a Credit Context

15 min

The loan officer thought the hardest part was already behind her. The AI assistant had processed a 280-page commercial loan file overnight: pulled the key financial figures, calculated the debt service coverage ratio (DSCR, the ratio of a borrower's net operating income to its annual debt payments), flagged the covenant thresholds, and drafted a two-page credit summary. She read through it on Tuesday morning and initialed it as reviewed. The file went to senior credit committee on Thursday. Midway through the presentation, the chief credit officer asked the borrower's 2022 revenue figure and cross-checked it against the actual tax return. The number in the AI summary was $4.2 million. The number on the return was $3.1 million. The AI had hallucinated a revenue figure that made the DSCR look 35% stronger than it actually was. The deal that looked like an approval turned into a committee table and a three-week delay while the file was rebuilt from scratch. And the loan officer, who had initialed the AI summary as reviewed, spent the next six weeks explaining why her sign-off had not caught the error. The model did not know it was wrong. It was not lying. It was doing exactly what it was designed to do, and the design does not include a built-in alarm that fires when the output is fabricated.

What Hallucination Actually Means and Why It Happens

The word "hallucination" is widely used in AI discussions, often in ways that make it sound mysterious or exotic. In a lending context, it is neither. It is the predictable, documented behavior of a generative AI system that produces output that is internally coherent and confident in tone but factually wrong, unsupported, or fabricated. Understanding why it happens is the first step toward managing it.

Generative AI language models, including the kind embedded in modern loan origination systems (LOS) and document analysis tools, work by predicting the most probable next token given everything they have seen so far: their training data, the system prompt, and the document or conversation in front of them. They do not retrieve facts from a database and verify them before stating them. They generate language that is statistically consistent with patterns learned during training. That generation process is extraordinarily powerful for producing coherent, fluent text. It is not a truth-checking mechanism.

When an AI model is given a 280-page commercial loan file and asked to extract the borrower's 2022 revenue, it is doing something like this: it looks at the document, identifies text that looks like financial tables and income statements, finds numbers that appear to be annual revenue figures, and synthesizes a response. If the relevant table is on page 147, formatted differently from what the model was trained on, partially obscured by a scan artifact, or ambiguous between two possible revenue definitions (gross revenue versus net revenue, for example), the model may select the wrong number, blend numbers from different years, or extrapolate from adjacent figures. The output will be stated with the same confident tone regardless of whether the underlying extraction was accurate or not. The model has no uncertainty flag that triggers when it is guessing.

This is the core failure mode: AI does not know the difference between what it accurately extracted from the document and what it statistically inferred, approximated, or confabulated. From the model's perspective, both outputs look the same. The institution that treats AI output as accurate without verification is relying on a system that cannot distinguish its own correct extractions from its own errors.

The Three Hallucination Types That End Careers in Lending

Not all hallucinations carry equal risk. In a lending context, there are three failure modes that rise to a level of severity that creates legal, regulatory, and career consequences. Understanding each type concretely, with the specific harm it causes, is the foundation of an effective verification discipline.

Invented or distorted income figures. This is the commercial loan scenario from the opening of this lesson, replicated across the residential mortgage context as well. An AI document extraction tool reading a tax return may transpose digits, blend figures from adjacent lines, confuse gross income with adjusted gross income, or produce a number that represents something close to but not exactly the qualifying income figure under agency guidelines. In residential mortgage, the stakes are immediate: Fannie Mae, Freddie Mac, FHA, and VA all require specific income calculation methodologies, and a qualifying income figure that does not follow the correct methodology is an origination defect that can result in repurchase demands when the loan is sold into the secondary market. The institution that originated 500 mortgages with AI-extracted income figures and did not verify them faces the prospect of discovering the defects when the investor demands repurchase, years after the loans were originated and the responsible staff have moved on.

The failure mode extends to commercial lending. A DSCR calculated on a hallucinated revenue figure is a DSCR that does not reflect the borrower's actual ability to service the debt. A commercial credit committee that approves a deal on the basis of a hallucinated DSCR has made a credit risk decision grounded in fictional data. When the borrower struggles to service the debt and the credit deteriorates, the file review will reveal that the original underwriting was based on an inaccurate figure. The loan officer who did not verify the AI's extraction is the one who signed the credit memo.

Fabricated covenants and terms. Generative AI tools used to draft or summarize loan agreements, term sheets, and credit agreements are capable of producing covenant language that looks legitimate and reads as if it came from the document but does not actually appear in the source. This is a particularly insidious failure mode because covenant terms are structural elements of the credit relationship. A covenant that does not exist cannot be enforced. A covenant that is summarized inaccurately creates misaligned expectations between the lender and the borrower. And a covenant that is stated incorrectly in a credit memo used for an approval decision creates a record that does not match the legal agreement the borrower actually signed.

In commercial real estate lending, where loan agreements routinely run to hundreds of pages and contain dozens of financial maintenance covenants, debt-yield requirements, reserves triggers, and operating covenants, AI summarization tools are genuinely useful for getting a quick orientation to the deal structure. They are also capable of producing a clean, well-organized summary that attributes covenants to the wrong parties, cites threshold levels that are close to but not exactly the negotiated figures, or lists covenants that were discussed in term-sheet negotiations but did not make it into the final agreement. The practice of relying on AI summaries without tracing every material term to the actual agreement language is how an institution ends up with a credit file that does not match its own loan documents.

Wrong adverse-action reasons. This is the failure mode that most directly implicates the program's central thesis, and it is the one that can do the most regulatory damage in the shortest time. When AI is used to draft adverse-action notices or reason codes, the model may produce reason codes that are plausible and well-formatted but do not accurately reflect the factors that drove the denial.

Consider the specific scenario: a loan officer uses an AI assistant to help draft the adverse action notice for a denial. The AI reviews the file and produces four reason codes: insufficient income, excessive debt-to-income ratio, insufficient collateral, and recent delinquency. All four look like reasonable adverse-action reasons. But in this particular file, the actual reason the credit policy was not met was the property's environmental report, which revealed contamination that made the collateral unacceptable as security. The AI did not identify the environmental issue because it was buried in a third-party report attached as an exhibit; it generated reasons based on the pattern of financial figures it could read and the pattern of adverse-action notices it was trained on. The notice was sent, and the four reasons cited were all technically true about the file but were not the specific, accurate, primary reasons for the denial. Under ECOA and Regulation B (Reg B), the adverse-action reasons must be the specific, accurate reasons that actually drove the decision. Plausible reasons that are not the actual reasons are a Reg B violation.

The legal exposure from incorrect adverse-action reasons is not theoretical. ECOA allows borrowers to bring private lawsuits for actual damages, punitive damages up to $10,000 per individual action or up to $500,000 (or 1% of the creditor's net worth, whichever is less) in class actions, and attorneys' fees. A pattern of AI-generated adverse-action reasons that do not accurately reflect the actual denial factors can support a pattern-or-practice claim, which is both a private-litigation risk and a regulatory-enforcement risk. The CFPB's enforcement history includes settlements in the hundreds of millions of dollars for fair-lending and adverse-action violations. The institution that is using AI to generate adverse-action reason codes without a verification step is operating a potential ECOA violation generator at scale.

Why the Confident Tone Makes Everything Worse

The most dangerous feature of AI hallucinations in a lending context is not that they happen. It is that when they happen, the AI's output looks exactly the same as when it is correct. There is no asterisk, no uncertainty flag, no hedging language, no change in font. The model produces a revenue figure of $4.2 million in the same tone and format whether the actual figure is $4.2 million or $3.1 million. The model generates four adverse-action reason codes with the same professional language whether those codes reflect the actual denial factors or are plausible-sounding substitutes. The model drafts a covenant summary with the same confident structure whether the covenants are accurately transcribed or partially fabricated.

This is fundamentally different from the failure mode of a spreadsheet that shows a #REF! error, or a database query that returns an empty result set, or a junior analyst who says "I'm not sure about this number, let me check." Those outputs signal their own uncertainty. An AI language model that has produced incorrect output does not signal anything. It presents its output with the same confident professional tone it uses for accurate output, and it may even generate supporting context that makes the incorrect output seem more plausible.

The banking industry's historical experience with this kind of failure mode is instructive. Before AI, the closest analogy was the confident junior analyst who, under time pressure, filled in a number from memory rather than going back to the source and trusted that it was close enough. Experienced supervisors and credit committees developed instincts for spotting those situations: certain file patterns, certain number combinations, certain too-clean analyses that warrant a second look. Those instincts need to be applied, consciously and systematically, to AI output. The experienced credit professional's skepticism about numbers that look too round, analyses that hang together too neatly, and explanations that are more coherent than the underlying facts would support is exactly the right orientation for reviewing AI output.

Specific Verification Protocols for the Three Failure Modes

Knowing that hallucinations happen is necessary but not sufficient. The practical question is: what specific verification steps catch each failure type before the output becomes a decision?

For income and financial figures: every number in an AI-extracted or AI-drafted credit analysis must be traced to a specific page and line in the source document. This is not a suggestion; it is the operational definition of "verified." If the AI summary says the borrower's 2023 gross revenue was $3.4 million, the reviewer must be able to point to page 12 of the 2023 corporate tax return, line [specific line], and confirm that the figure there is $3.4 million. A number that cannot be traced to a source is an unverified number. An unverified number used in a credit decision is a credit risk and a potential ECOA problem.

In practice, this means building a verification checklist with source-document pointers into the credit review workflow. For residential mortgage, the checklist aligns with the income calculation worksheet used under agency guidelines: every figure on the worksheet gets a source citation. For commercial lending, the checklist covers revenue, EBITDA (earnings before interest, taxes, depreciation, and amortization), capital expenditures, debt service, and any covenant thresholds used in the credit analysis. The checklist takes time, but it is faster than the fourteen-month review process in the opening scenario.

For loan terms and covenants: every material term in an AI-drafted or AI-summarized loan document summary must be traced to the specific section and page of the executed agreement. Covenant thresholds, definitions, calculation methodologies, and parties must be verified against the document language rather than the summary language. AI summaries are useful for orientation; they are not a substitute for reading the relevant sections. The practice of creating a term-sheet or summary-of-terms document that cites each term's location in the executed agreement provides both a verification record and a useful working reference for ongoing covenant monitoring.

For adverse-action reasons: the specific, documented verification step that ECOA and Reg B demand is this: before the adverse-action notice is issued, a qualified human must confirm that each reason stated in the notice (1) is actually present in the file, meaning there is specific data supporting it, (2) actually contributed to the denial decision, not just appeared in the file, and (3) is the most accurate and specific statement of that contributing factor available given the file data. AI-generated reason codes should be treated as a first draft, not as a final determination. The loan officer or underwriter who signs off on the adverse-action notice is representing to the applicant and to any regulator who later reviews the file that those reasons are specific and accurate. That representation requires human verification, not just AI generation.

The Regulatory and Career Consequences When Verification Fails

The scenarios described in this lesson are not hypothetical edge cases. They are documented failure modes that have produced real regulatory consequences and real career disruptions. Understanding those consequences concretely is the motivation for maintaining the verification discipline when time pressure makes shortcuts appealing.

From a regulatory perspective, the consequences of AI hallucinations in credit files flow through three channels. Under ECOA and Reg B, incorrect adverse-action reasons expose the institution to individual claims, class actions, and enforcement actions by the CFPB, the OCC, the FDIC, the Federal Reserve, and state regulators. The damages in individual actions are capped at $10,000 plus actual damages plus attorneys' fees; in class actions, the cap is $500,000 (or 1% of the creditor's net worth, whichever is less) plus actual damages plus attorneys' fees. But the regulatory consequence of a pattern of incorrect reasons is not primarily the per-case damages; it is the consent order that requires remediation of every affected decision, a third-party audit, enhanced monitoring, and in some cases a lending freeze pending the audit. The operational cost of a consent order runs into the millions.

Under model-risk governance expectations codified in OCC Bulletin 2026-13, an institution that relies on AI to generate adverse-action reasons without a verification process has a model-risk deficiency. The bulletin is explicit that AI outputs in credit and compliance functions require human oversight and that the institution is responsible for the accuracy of those outputs. A model-risk examination that reveals an institution is sending AI-generated adverse-action notices without verification is a material finding, not a comment.

From a career perspective, the individual who initialed the AI summary as reviewed and did not verify the hallucinated revenue figure is the individual whose name is on the credit memo. The institution's investigation, when the credit deteriorates and the file is reviewed, will start with the credit memo and the initials on it. "I relied on the AI" is not a defense against the question "why did you sign a credit analysis based on an income figure you did not verify?" The obligation to verify is not transferred to the AI tool by virtue of using it. It stays with the professional who reviewed and approved the output.

Key Takeaways

  • AI hallucination is the predictable behavior of a language model that generates output that is internally coherent and confident in tone but factually wrong, unsupported, or fabricated. It is not a bug in the software; it is a characteristic of how generative AI works.
  • In a lending context, the three hallucination types with the most severe consequences are invented or distorted income figures, fabricated loan covenants and terms, and incorrect adverse-action reason codes. Each failure mode carries distinct legal and regulatory exposure.
  • Incorrect adverse-action reasons that do not accurately reflect the specific factors that drove a denial are a violation of ECOA and Reg B, regardless of how those reasons were generated. Using AI to produce reason codes does not reduce the legal obligation to provide specific, accurate reasons; it adds a verification step that must be completed before the notice is issued.
  • The most dangerous feature of AI hallucinations is that they are tonally indistinguishable from accurate output. The model does not flag its errors. The loan officer, underwriter, or credit analyst is the error-detection mechanism, which means verification discipline is not optional, it is the institution's primary defense.
  • Verification means tracing every material fact in AI output to a specific source: a page and line in a tax return, a section and page in a loan agreement, a data field in the credit file. A number that cannot be traced to a source is an unverified number and should not be used in a credit decision.
  • Under OCC Bulletin 2026-13, institutions are responsible for the accuracy of AI outputs used in credit and compliance functions. The governance expectation is human oversight of AI-generated content, not AI generation followed by automatic acceptance.
  • The professional who signs the credit memo, initials the income analysis, or issues the adverse action notice is accountable for the accuracy of that document regardless of what tool generated the first draft. Accountability does not transfer to the AI.