Recognizing Bad AI Output in Credit Work
The credit memo looked good. The debt-service coverage ratio of 1.42 was cited twice, once in the executive summary and once in the body analysis, consistent and confident. The senior underwriter initialed the file and sent it to committee. The committee approved a $1.4 million commercial real estate loan. Six weeks later, the bank's internal audit pulled the borrower's financials for a routine post-closing review. The actual DSCR was 1.09. The AI tool had extracted the wrong operating income line from the income statement, compounded it across two years, and produced a coverage ratio that cleared the bank's 1.20 minimum threshold. The correct ratio did not. The loan was within policy under a different reasonable reading of the financials, so the approval stood, but the audit finding was uncomfortable: a senior underwriter had approved a credit memo containing a figure that could not be verified against the source documents. The story of how to avoid that outcome is this lesson. Recognizing bad AI output in credit work requires knowing what the failure modes look like, where each one hides, and what a fast, repeatable skeptic's checklist catches before the output reaches a decision. (This scenario is a composite illustration drawn from common field patterns; identifying details are changed.)
What Bad AI Output Looks Like in Credit Work
Bad AI output in lending is not usually garbled text or obvious nonsense. The failure modes that cause problems in credit work are the ones that look good: well-formatted, internally consistent, professionally phrased, and wrong. Understanding the specific forms this takes is the foundation of a useful skeptic's framework.
The confident fabrication. The model produces a specific figure, a ratio, or a reason that does not come from the source document but is consistent with what a typical document of that type would contain. A DSCR of 1.42 when the actual document supports 1.09 is a confident fabrication: it sounds plausible for a commercial real estate loan, it clears the typical threshold, and it fits the narrative the memo is building. The model did not produce random noise. It produced a number that made sense in context, derived from the wrong input.
The plausible-sounding but wrong reason code. An AI-drafted adverse-action notice includes reason codes that are among the CFPB's model form reason codes (or at least sound like they are) but do not accurately reflect the specific factors that drove this particular denial. The reasons sound legitimate. They may even be true of the borrower in a general sense. They are not the specific, accurate reasons Regulation B (Reg B, 12 CFR Part 1002, the implementing regulation for the Equal Credit Opportunity Act) requires. A reason code that is technically accurate about the borrower but not causally accurate about the denial is a non-compliant reason under Reg B.
The missing negative fact. The AI reviews a credit file and drafts a narrative. It correctly identifies the borrower's strengths: stable income, low utilization, long credit history. It omits a Chapter 7 bankruptcy from five years ago that is still within the lookback period and is material to the credit decision. The omission is not a fabrication. It is a gap that produces a falsely positive characterization of a file that has a significant negative fact. Missing negatives are harder to spot than invented positives because you are looking for something that is not there.
The outdated or inapplicable rule. When asked about a regulatory requirement, the AI cites a rule, a threshold, or a guidance document that has been superseded or modified. The Qualified Mortgage ability-to-repay rules, BSA/AML threshold requirements, and model-risk guidance all change, and AI models trained on historical data may confidently cite a prior version of the rule. An AI tool that cites OCC 2011-12 model-risk guidance as though it is current is wrong: OCC Bulletin 2026-13 (the April 2026 interagency update that superseded 2011-12 and pulled AI and generative AI under model-risk, fair-lending, third-party, and board-governance expectations) is the governing framework.
The hallucinated citation. The AI cites a real-sounding source (a CFPB circular, a specific OCC bulletin number, a case name) that either does not exist or does not say what the AI claims. Regulatory references in AI-drafted compliance memos or adverse-action notices are a specific risk because the person reviewing the output typically does not verify citations they did not request.
Bad AI output in credit work does not announce itself. It arrives formatted, confident, and internally consistent. The skeptic's job is to ask: where did this number, reason, or rule come from, and can I verify it in the source?
The Income and Ratio Red Flags
Income figures and financial ratios are the load-bearing numbers in a credit analysis. An error in a ratio that a committee is evaluating against a policy threshold can be the difference between an approval and a denial, or, as in the opening story, between a finding and no finding. The red flags that indicate potential income or ratio errors are specific and fast to check.
Red Flag: The Uncited Figure
Any income figure, ratio, or financial metric in an AI output that does not include a citation to the specific document and location where it was found is a red flag. In a well-structured AI output, every figure has a citation: "Schedule C Line 31, Year 1: $42,300." A figure without a citation is either an inference, a calculation from unstated inputs, or a fabrication from training data. All three require the same response: locate the source before using the figure.
The practical test: for each financial figure in the AI output, ask "what document and what line is this from?" If the citation is missing, ask the AI to provide it. If the AI cannot provide a citation, treat the figure as unverified. If the AI provides a citation and the figure does not appear at the cited location, treat the figure as potentially fabricated and check adjacent lines for the number the model may have misread.
Red Flag: The Threshold-Clearing Figure
A ratio or metric that falls just above a policy threshold, particularly when the margin is narrow (a DSCR of 1.22 against a 1.20 minimum, for example), warrants additional verification precisely because it clears the threshold. Narrow threshold clearances are not necessarily wrong, but they are the outcomes most affected by a small extraction error. A correct DSCR of 1.19 that the AI reports as 1.22 because of a small income overstatement is the case where verification matters most, because it is also the case where the verification is most likely to be skipped on the assumption that the figure is reasonable.
Red Flag: The Directional Inconsistency
A set of financial metrics that do not tell a consistent story across the file is a red flag. A credit memo where the income figure implies strong debt-service coverage but the bank-statement analysis shows deposits inconsistent with that income level is internally inconsistent. The inconsistency may have an explanation (the borrower moves income between accounts, seasonal business patterns, noncash income items), but it requires investigation rather than acceptance. An AI tool that analyzes each document in isolation may produce internally consistent outputs for each document without noticing the cross-document inconsistency.
Red Flag: The Unexplained Positive Trend
When an AI analysis presents a favorable trend in income or ratios over the analysis period without citing the specific driver, the trend requires verification. Self-employed income that grew 40 percent year over year may be correct. It may also be an artifact of a gross-receipts extraction in Year 2 and a net-income extraction in Year 1, producing an apparent trend from an apples-to-oranges comparison. Ask what drove the trend; if the AI cannot explain it from the documents, investigate before relying on it.
The Adverse-Action Red Flags
Adverse-action reason codes are the place where an undetected AI error carries the most direct legal consequence. A wrong reason code in an adverse-action notice is not just a quality problem; it is a potential violation of ECOA and Reg B, which require specific, accurate reasons that the applicant can understand and challenge.
Red Flag: The Generic Reason Code
A reason code that could apply to many borrowers but does not trace to a specific data point in this particular file is a red flag. Reason codes like "insufficient credit history" or "excessive obligations in relation to income" are valid Reg B reason codes when they are the actual factors that drove the denial. They are non-compliant when they are generated as generic characterizations of a borrower profile without verification that they were the primary denial factors for this specific file. The test: for each reason code, identify the specific credit file data point (the credit report trade line, the debt-to-income calculation, the income figure) that supports it. If the data point does not exist in the file or does not support the reason as stated, the reason code is a red flag.
Red Flag: The Reason That Does Not Match the Credit-Policy Denial Basis
The credit decision record should show the reason(s) the institution's credit policy required the denial. The adverse-action notice should state those same reasons. When the AI-drafted reasons differ from the credit-policy denial basis, the reasons are wrong even if they are accurate descriptions of the borrower's file. Reg B requires that the stated reasons be the real reasons. A borrower who was denied because the property type requires committee approval and committee approval was not sought needs to receive that reason, not a substitute reason about income that is also technically true but was not the proximate denial factor.
Red Flag: The Reason the File Does Not Support
An AI tool asked to draft adverse-action reasons based on a file summary or a general description of the borrower may produce reasons that are plausible for a borrower with those general characteristics but are not specifically supported by the actual file data. A reason citing "length of employment" when the borrower has been at their current job for eight years is a specific file-data contradiction that a fast review should catch. The test: for each reason code, read the relevant section of the credit file to confirm the reason is supported. Any reason that contradicts file data or is not supported by file data is a red flag that requires correction before the notice is sent.
The Regulatory Citation Red Flags
AI-generated compliance memos, policy summaries, and regulatory analyses are a specific risk category because the person reviewing them often does not verify citations they did not request. An AI that cites the wrong version of a regulation, the wrong bulletin number, or a rule that has been modified since its training cutoff is providing legally inaccurate guidance that can shape institutional decisions.
The practical checklist for regulatory citations has three items:
Is the cited document real and current? For any bulletin, guidance document, or regulation cited by AI output, verify that it exists and has not been superseded. A reference to OCC 2011-12 model-risk guidance should prompt a check: that guidance was superseded by OCC Bulletin 2026-13. A reference to a CFPB circular that does not appear in the CFPB's published circular database should be treated as a potential fabrication until confirmed.
Does the citation support the stated proposition? AI tools sometimes cite real documents for propositions those documents do not actually support. An AI claiming that OCC Bulletin 2026-13 permits AI-generated adverse-action notices without human review is citing a real document to make a claim the document does not support (it requires the opposite). The verification step is not just confirming the document exists but confirming that the document says what the AI says it says.
Is the rule being applied to the right context? A rule that applies to covered institutions may not apply to credit unions or non-bank mortgage companies. A threshold that applies to Bank Secrecy Act (BSA, the statute governing financial crime reporting obligations) reporting may not apply to internal monitoring. An AI that applies a rule to the wrong institution type or the wrong regulatory context produces technically accurate regulatory language that leads to incorrect compliance conclusions. The context check is fast: read the scope section of the cited rule and confirm it covers your institution and transaction type.
The Skeptic's Checklist: A Fast Pass That Catches Most Errors
The red flags described above translate into a practical checklist that a credit analyst or underwriter can run in eight to twelve minutes on most credit work AI output. The checklist is not a comprehensive review. It is the pass that catches the failure modes that most frequently produce compliance exposures, and it is designed to be fast enough to maintain the productivity advantage that AI provides.
The checklist has seven items:
Item 1: Figure source check. For each income figure, ratio, or financial metric in the output, confirm it has a citation to a specific document and line. Flag any uncited figure for verification before use.
Item 2: Threshold-clearing check. For each ratio that falls above or within 5 percent of a policy threshold, open the source document and verify the calculation. This applies to debt-to-income, debt-service coverage, loan-to-value, and any other ratio with a pass/fail threshold in the credit policy.
Item 3: Adverse-action reason match. For each reason code in the adverse-action output, identify the specific credit file data point that supports it. If a data point cannot be identified, flag the reason for correction. Compare the reasons to the credit policy denial basis and confirm they match.
Item 4: Missing negative fact scan. Read the narrative or summary for material negative facts about the borrower that are relevant to the credit decision under the institution's policy: recent derogatory marks, bankruptcy within the lookback period, prior NSF patterns, material business losses in a recent period, or property condition issues for collateral. If the file contains a negative fact that would be material to the committee's review and the AI output does not mention it, the output is incomplete regardless of how accurate the reported positives are.
Item 5: Cross-document consistency check. If the file includes multiple financial documents (a tax return and bank statements, a credit report and a credit memo, paystubs and a 1040), confirm that the key figures are consistent across documents. A payroll figure on a paystub that is materially inconsistent with the W-2 wages on the 1040, without explanation, is a flag for both potential income misrepresentation and AI extraction error.
Item 6: Regulatory citation spot-check. For each regulatory citation in the output, confirm that the cited document exists and has not been superseded. Apply the three-item citation checklist (real and current, supports the stated proposition, right context) to any citation that appears in a compliance memo, adverse-action notice, or regulatory analysis.
Item 7: Completeness test for adverse action. Review all material factors that contributed to the denial (from the credit file and the credit policy) against the set of reasons in the adverse-action output. Confirm that no material denial factor is absent from the reasons. The Community Reinvestment Act (CRA, the statute requiring banks to meet the credit needs of their communities, including low-and-moderate income neighborhoods), geographic policy restrictions, product-type limits, and committee-approval requirements are the categories most frequently omitted by AI adverse-action tools that focus on financial ratios.
Running the seven-item checklist takes eight to twelve minutes on a well-structured credit work output. The time is front-loaded: Items 1 and 2 (figure source and threshold-clearing) take the most time for complex commercial files. Items 3, 4, and 7 (adverse-action reasons, missing negatives, completeness) take most of the time for consumer credit denials. The checklist is faster for files with specific AI citations, which is another argument for building the citation requirement into every prompt.
The Accountability and Governance Dimension
Running the skeptic's checklist is not an optional quality improvement. It is the mechanism through which the human accountability requirement under ECOA, Reg B, and OCC Bulletin 2026-13 is operationalized in daily credit work. The checklist is what the underwriter, credit analyst, or loan officer does to move AI output from "draft that a computer produced" to "credit analysis that a named human has reviewed and accepted responsibility for."
The accountability transfer happens at the moment the named reviewer signs off on the reviewed output. But the sign-off only means something if the reviewer can, in any subsequent examination, litigation, or audit, describe what they reviewed and what they checked. The seven-item checklist, when documented in the loan origination system (LOS, the software platform managing loan applications from intake through funding) as a completed step, provides that description. It is the evidence that the institution produces when asked: what oversight existed over this AI-assisted credit analysis?
Under OCC Bulletin 2026-13, the institution must demonstrate that AI-touched credit work was reviewed and validated before it was used in a credit decision. Industry experience suggests that institutions building validation evidence into their workflow as a routine step have fared better in model-risk examinations than those constructing that evidence only after a finding. The skeptic's checklist is the daily workflow version of that validation evidence.
The 38 percent of mortgage lenders who were using AI by 2024 (up from 15 percent in 2023) include both institutions that built these disciplines early and institutions that are still learning, sometimes through examination findings, that confident AI output requires verification regardless of how sophisticated the tool. The skill of recognizing bad AI output before it reaches a decision is not a skepticism about AI. It is the competence that lets AI's speed benefits be captured without the compliance liabilities that unreviewed AI output creates.
The Bank Secrecy Act and Anti-Money Laundering (BSA/AML) context adds a specific dimension to the missing-negative-fact failure mode: an AI tool summarizing a Suspicious Activity Report (SAR, the federal filing required when a bank identifies potential money laundering or financial crime) that omits a transaction pattern that a human analyst would have flagged produces a false-negative summary that may lead to a missed true positive. The 90 to 95 percent false-positive rate in BSA/AML alerts is the industry pain point; the true-positive rate is the regulatory failure mode. An AI that makes triage faster but introduces systematic omissions in the summaries it produces is trading a false-positive problem for a true-positive miss problem. The missing-negative-fact scan in the checklist is equally important for SAR narrative review as for credit memo review.
Key Takeaways
- Bad AI output in credit work does not look like random noise. It looks like a professionally formatted, internally consistent analysis that contains a fabricated figure, an inaccurate reason code, a missing negative fact, or a superseded regulatory citation. The checklist catches failure modes that do not announce themselves.
- The five categories of bad AI output in credit work are: confident fabrications (figures that look plausible but do not come from the source document), plausible-but-wrong reason codes (Reg B non-compliant because they do not trace to the actual denial basis), missing negative facts (omitted information that is material to the decision), outdated regulatory rules (prior versions of guidance cited as current), and hallucinated citations (source references that do not exist or do not say what the AI claims).
- The seven-item skeptic's checklist (figure source check, threshold-clearing check, adverse-action reason match, missing negative fact scan, cross-document consistency check, regulatory citation spot-check, and completeness test for adverse action) runs in eight to twelve minutes and catches most of the failure modes that produce compliance exposures.
- Threshold-clearing figures require extra attention, not less. A ratio that falls just above a policy minimum warrants verification precisely because it is the case most affected by a small extraction error and most likely to be accepted without checking.
- ECOA and Reg B adverse-action compliance requires that stated reasons match the credit-policy denial basis, not just that they are accurate descriptions of the borrower's file. An AI that generates plausible-sounding reasons that do not match the actual denial logic creates a Reg B exposure regardless of how well-formatted the notice looks.
- Regulatory citations in AI output require a three-part check: confirm the document is real and current, confirm it supports the proposition stated, and confirm it applies to the correct institution type and regulatory context. A citation to a superseded guidance document is as non-compliant as no citation at all.
- The missing-negative-fact failure mode affects both credit memos and BSA/AML summaries. An AI that omits a material derogatory fact from a credit narrative or a significant transaction pattern from a SAR summary produces output that is actively misleading, not just incomplete.
- Running the skeptic's checklist is the mechanism through which the human accountability requirement under ECOA, Reg B, and OCC Bulletin 2026-13 is operationalized. The documented checklist is the evidence the institution produces when asked what oversight existed over AI-assisted credit work.
Skill.re