Verifying the Reason Matches the Decision
The following scenario is a composite illustration drawn from patterns documented in actual CFPB examinations and consent orders; it is not a description of any specific institution or proceeding. In the spring of 2025, a fair-lending examiner from the CFPB (Consumer Financial Protection Bureau) sat across from the compliance director of a mid-size community bank and asked a question the director had not prepared for. The bank used an AI underwriting platform that generated adverse-action reason codes automatically and routed them to the loan officer for review before release. The examiner had sampled fifteen denied mortgage applications from the prior year and pulled both the adverse-action notices and the underlying model documentation for each. In eleven of the fifteen, the stated reason codes did not align with the decision factors documented in the model's own output for that application. The reasons were specific. They were formatted correctly. They looked exactly like compliant adverse-action language. They were not, in any verifiable sense, the reason the applications were denied. The compliance director had no documentation to demonstrate that any human had compared the stated reasons to the model's decision log before the notices were released. The bank's argument was that the loan officers were expected to review the notices before release, so review had taken place. The examiner's response was two questions: "What did they review the codes against?" and "Where is the documentation that the review happened?" Neither question had a good answer. The examination findings that followed that afternoon became a consent order eight months later. The consent order did not allege intentional discrimination. It alleged a governance failure: an institution that could not demonstrate that its adverse-action reasons were grounded in its actual credit decisions. That failure, across hundreds of applications, was the finding. The specific check the bank lacked a name for had a name. It was verification, and it was the non-negotiable step between AI-drafted reason codes and a legally defensible denial.
Why Verification Is a Distinct, Required Step
The term "verification" in the context of AI-assisted adverse action means something precise: it is the process of confirming that each stated adverse-action reason is the actual reason the application was denied, traced to a specific data point in the file and the specific provision of the institution's credit policy that the data point implicates. Verification is not a reread. It is not a formatting check. It is not approval of the AI's output because the output looks professional. It is a substantive, document-level comparison between the stated reason and the actual decision basis.
Verification is a distinct required step because the Equal Credit Opportunity Act (ECOA, 15 U.S.C. 1691 et seq.) and Regulation B (Reg B, 12 CFR Part 1002) require that adverse-action reasons be accurate, not just specific. The accuracy requirement has a meaning: the stated reason must be the actual reason. It must not be a reason that is true about the file but was not a determinative factor. It must not be a reason the AI generated based on pattern-matching against training data. It must be the reason the institution, under its credit policy, denied this application. The only way to confirm that the stated reason is the actual reason is to compare the two, one to one, before the notice goes out. That comparison is verification.
The reason verification needs to be a named, distinct step in an adverse-action workflow is that the alternative, which is what the bank in the opening scenario had, is a review process that confirms the notice looks compliant without confirming it is accurate. A loan officer who reads an AI-drafted adverse-action notice and approves it because the language is familiar and professionally formatted has done a formatting review, not a verification. Those are not the same thing, and under ECOA, only the verification satisfies the accuracy requirement.
Verification is not "does this reason code look right?" It is "is this reason code the actual reason this application was denied, and can I trace it to the file?"
The distinction also matters under OCC Bulletin 2026-13, the April 2026 interagency model-risk guidance that superseded OCC 2011-12 and extended model-risk management requirements to AI and generative AI. The 2026 bulletin requires that institutions be able to validate, at the transaction level, that their AI model's outputs are accurate representations of the factors driving individual credit decisions. For adverse-action reason codes, this means the institution must have, for each denied application, a documented record demonstrating that someone compared the stated reasons to the model's documented decision logic and confirmed they match. A general policy saying "loan officers review adverse-action notices" is not transaction-level validation. Transaction-level validation means there is a record for each application, created at the time of review, showing what the reviewer compared and what they concluded.
The Four Dimensions of a Verification Check
A complete verification of an AI-generated adverse-action reason set has four dimensions. Each addresses a different way the stated reason can fail to match the actual decision. Skipping any one of them leaves a gap that an examiner, an applicant's attorney, or a CFPB investigation can find.
Dimension one: factual grounding. Is there a specific, documented fact in the file that supports the stated reason? This dimension catches hallucinated reasons and reasons that are technically plausible but not present in the file. For "excessive obligations in relation to income," the factual-grounding check is: what is the calculated debt-to-income ratio, what are the specific obligations and income figures that produced it, and what documents support those figures? If the ratio is 43.1 percent and the institution's threshold is 43 percent, the stated reason is factually grounded. If the ratio is 38 percent and the threshold is 43 percent, the stated reason is not factually grounded, regardless of how well-formatted it looks in the notice.
Factual grounding requires the reviewer to have the file in front of them, specifically the documents that support the financial figures the reason code references. A reviewer who is checking factual grounding by memory, or by relying on the AI's summary of the file, is not performing a grounding check; they are assuming the AI's description of the file is accurate, which is the assumption the check is designed to test. The check requires the document: the tax return, the credit report, the paystub, the appraisal report, the bank statement. Each reason must trace to a page and a number.
Dimension two: policy causation. Is the stated reason a factor that, under the institution's credit policy, contributed to the denial? This dimension catches reasons that are true about the file but were not determinative under the policy. A file might have a limited credit history (6 years), a 44 percent debt-to-income ratio, and one unpaid collection account. If the institution's credit policy denies the file solely because the DTI exceeds the 43 percent threshold for the credit score tier, then "limited credit history" is true about the file but is not a policy-based denial reason. Including it in the notice is inaccurate under the ECOA accuracy requirement, because it was not a reason the application was denied, just a characteristic the AI noticed and formatted as a reason code.
Policy causation requires the reviewer to consult the institution's credit policy document and confirm that the stated reason implicates a specific policy provision that required, or contributed to requiring, the denial. For complex files with multiple contributing factors, the policy-causation check also ensures that the principal reasons under Reg B are all represented: if three policy provisions contributed to the denial, all three must appear in the notice, not just the one the AI happened to select.
Dimension three: reason-code precision. Is the stated reason as specific as the file supports? This dimension catches reasons that are factually grounded and policy-based but stated at a level of generality that fails the Reg B specificity requirement. "Credit history" fails specificity; "serious delinquency in the past 24 months" passes it. "Insufficient income" fails specificity in some contexts; "income insufficient to support the proposed monthly obligations based on documented annual income of $X" passes it. The reviewer needs to confirm that the AI's phrasing is the most specific accurate statement available, not a generic category that the AI selected because it pattern-matched to the file type rather than the specific file characteristic.
This dimension also catches the reverse problem: a reason stated with so much specificity that it is technically inaccurate. If the notice says "34-day late payment on installment account in March 2024" but the credit report shows a 30-day late, the over-specific phrasing is wrong. The general code "delinquent account" would be more accurate. The reviewer's job is to match the precision of the stated reason to the precision of the supporting documentation, erring toward accuracy over specificity when the two are in tension.
Dimension four: completeness. Does the reason set as a whole give the applicant a fair picture of the basis for the denial? This dimension catches omissions: material denial factors that the AI did not include in the reason set, either because they were policy-level reasons the AI did not have visibility into, or because the AI's pattern-matching produced a subset of the relevant factors rather than the complete set. Under Reg B's "principal reasons" standard, the notice must cover the dominant factors in the decision. A notice with one accurate, specific reason when there were three material denial factors is non-compliant for incompleteness even if the one reason it includes is perfect.
Completeness requires the reviewer to think affirmatively about what is missing, not just about whether what is present is accurate. The question is: "Considering all the reasons this application was denied under our credit policy, is there a material reason not represented in this notice?" That question requires the reviewer to work from the decision backward to the notice, not just from the notice forward to the file. It is a constructive analysis, not just a validation analysis.
The Comparison Task: Matching the Stated Reason to the Decision Log
The central act of verification is a comparison: the stated reason in the AI-drafted notice on one side, the model's documented decision output for this application on the other. Understanding what the decision log needs to contain, and how to perform the comparison, is the practical core of this lesson.
For a well-governed AI underwriting system, the model's decision documentation for each application should include: the input values used for decision-relevant features (DTI, credit score, LTV, employment status, income, specific derogatory items from the credit report), the model's output (pre-score or decision recommendation, with the feature weights or SHAP values that explain the output), and the mapping from those feature values and weights to the reason-code categories the institution uses in adverse-action notices. This documentation is the model's "decision log" for the application, and it is what the reviewer compares against the AI-drafted reasons.
The comparison itself is a line-by-line exercise. For each stated reason in the AI draft, the reviewer asks: does this reason appear in the model's decision log as a significant contributing factor? Is the feature value documented in the log consistent with what the reviewer sees in the file? Does the mapping from feature to reason code follow the institution's documented reason-code translation framework?
Where the comparison reveals a mismatch, the reviewer must determine whether the mismatch is a drafting error by the AI (wrong reason code applied to a factor that is in the decision log) or a more fundamental problem (the AI drafted a reason that has no basis in the decision log at all). A drafting error can be corrected by substituting the accurate code. A fundamental problem means the AI produced a hallucinated reason that must be struck, and the reviewer must determine whether the correct reason is the one in the decision log (which the AI missed) or whether there is a problem with the decision log itself.
For institutions that do not yet have a model-decision documentation system that produces per-application reason-code inputs, the comparison task defaults to a manual analysis: the reviewer must reconstruct the decision basis from the credit policy, the file data, and any documentation the underwriting system does produce, and compare the reconstructed basis to the AI-drafted reasons. This is more labor-intensive but not optional. If the institution cannot produce a model decision log that links the stated reasons to the model's documented output for the individual application, the institution has a model-risk governance gap that is itself a finding under OCC 2026-13, separate from any specific adverse-action violation.
The Fair Credit Reporting Act (FCRA, 15 U.S.C. 1681 et seq.) adds a parallel verification task for files where a consumer credit report was used in the decision. The adverse-action notice for those files must also identify the consumer reporting agency that provided the report, state that the agency did not make the decision and cannot give specific reasons for it, and inform the applicant of their right to obtain a free copy of the credit report within sixty days and to dispute its accuracy. The FCRA verification check is separate from the Reg B accuracy check but must be performed at the same time: the reviewer confirms that all required FCRA disclosures are present and accurate, including the correct agency name, before the notice is released.
Building a Verification Workflow That Holds Under Volume
The challenge of adverse-action verification is not understanding what needs to be done. Any underwriter who has been through a fair-lending examination understands the principle. The challenge is building a workflow that performs genuine verification on every denial, consistently, under the time pressure of a busy origination queue. Understanding the common shortcuts that degrade verification quality is how you design a workflow that survives volume.
The most common degradation: visual review without document comparison. The reviewer reads the AI-drafted reason codes and, because they look like proper adverse-action language, approves them without pulling the relevant documents to confirm the facts. This is the failure mode from the opening scenario. The fix is structural: the verification step must require the reviewer to identify the specific document page and value supporting each reason code, not just to read and approve the code. A required annotation field in the LOS, or a checklist that lists each reason code with a required "supporting document and value" field, makes the comparison step visible and mandatory rather than a matter of individual initiative.
The second common degradation: verifying against the AI's file summary rather than the file itself. Many AI underwriting platforms produce both a pre-score recommendation and a summary of the file's relevant characteristics. When the reviewer checks the AI-drafted reason codes against the AI's own file summary rather than the underlying documents, they are trusting the AI to accurately describe its own inputs, which is the assumption the verification step is designed to challenge. The AI that hallucinated a reason code can also produce an inaccurate file summary. The verification must go to the source document: the credit report for delinquency information, the tax return for income, the appraisal for property value. The AI's summary is a starting point, not the verification standard.
The third common degradation: verifying the lead reason and assuming the rest. A reviewer who confirms that the AI's primary reason code is accurate (the DTI calculation is correct and exceeds the policy threshold) may treat the remaining codes as implicitly verified. This assumption is wrong for two independent reasons. First, the AI may have produced accurate primary reasons and hallucinated secondary ones. Second, the completeness check is separate from the accuracy check: even if all stated reasons are accurate, the reviewer must still consider whether there are material denial factors missing from the stated set. Partial verification is not verification.
The fourth common degradation: deferring the verification to a second reviewer who has less file context. Some institutions route the verification to a compliance reviewer who was not the primary underwriter on the file. If that reviewer does not have access to the full file and decision documentation, the verification is a checklist exercise rather than a substantive comparison. The verification must be performed by someone who has the full file, the credit policy, and the model's decision documentation in front of them. If that is the compliance reviewer, they need the full file, not just the notice.
The workflow design that addresses all four degradation modes has four operational features: a mandatory document-reference field for each reason code, a policy-provision reference confirming each reason is policy-grounded, a completeness attestation signed by the reviewer, and a named-reviewer annotation with date in the LOS. These features make the verification substantive, structured, and documented. They add time to the process: roughly 15 to 25 minutes per denial file in a disciplined implementation. Against the cost of a consent order, a remediation program, and the civil money penalty structure for pattern-or-practice ECOA violations, 15 to 25 minutes per file is not overhead. It is risk management.
When the Stated Reason and the Actual Reason Diverge: The Redlining Structure
The regulatory stakes of the verification step become clearest when you look at what happens, at a portfolio level, when stated reasons and actual reasons diverge consistently. This is the scenario that transforms a documentation problem into a fair-lending finding, and it is the scenario that the verification step, consistently applied, prevents.
Consider an institution whose AI adverse-action drafting tool has learned to associate certain demographic-adjacent file characteristics with denial and to generate reason codes that reference those characteristics even when they are not the actual policy-based denial drivers. The institution's credit policy denies applications primarily on DTI and credit score grounds. The AI's reason-code drafting tool, trained on the institution's prior denials, has learned that files from certain zip codes tend to receive certain reason codes. When a new file from a similar zip code is denied (correctly, on DTI grounds), the AI generates reason codes that reference characteristics associated with those zip codes in its training data, rather than the specific DTI calculation that is the actual policy basis for the denial.
The stated reasons are wrong. The actual reasons are right. The institution is denying the application for a legitimate credit reason. But the notice describes the denial in terms that reflect demographic-adjacent patterns in the training data rather than the individual file's actual credit characteristics. In a fair-lending examination, the examiner pulls the denied files, compares the stated reasons to the model's decision documentation, and finds a systematic pattern: applicants from certain zip codes consistently receive reason codes that do not match the model's documented decision factors. That pattern is not just a documentation finding. It is the structure of a disparate-treatment case, because it suggests the institution is stating different (and inaccurate) reasons for denials in certain geographies than it states for economically similar applicants in other geographies.
This is not a hypothetical risk. It is the mechanism by which AI-assisted adverse-action drafting, without verification, can produce a fair-lending finding at an institution that is not intentionally discriminating. The AI's pattern-matching against training data, combined with the absence of verification, produces a systematic misrepresentation of denial reasons along geographic lines. Geographic lines in lending track protected-class concentrations closely. The CFPB's examinations look at both the accuracy of stated reasons and the geographic distribution of reason-code types. An institution where the verification step is missing is an institution that cannot detect this failure mode before the examiner does.
The Community Reinvestment Act (CRA, the federal statute requiring banks to serve the credit needs of the communities, including low-and-moderate-income neighborhoods, in which they operate) adds another dimension. An institution whose AI-generated denial reasons are systematically inaccurate in ways that track neighborhood composition may face CRA performance concerns on top of ECOA findings, because the CRA examination evaluates both the distribution of credit and the institution's responsiveness to community credit needs. Inaccurate denial reasons that suggest different credit-quality concerns than the actual denial basis complicate the institution's ability to demonstrate a CRA-compliant lending pattern.
The prevention structure is, again, the verification step. An institution that systematically compares stated reasons to model decision documentation, at the individual application level, will detect geographic or demographic patterns in reason-code inaccuracy before they accumulate into a fair-lending finding. The verification is both a per-application compliance control and a portfolio-level early-warning system for fair-lending drift.
The Accountability Record: What to Document and Why
The final element of the verification step is documentation: the creation of a record that demonstrates, for each denied application, that verification was performed, by whom, against what evidence, and with what conclusion. The documentation is not procedural overhead. It is the legal substance of the institution's governance claim.
Under OCC Bulletin 2026-13, the institution's model-risk file must demonstrate ongoing validation of the AI model's outputs, including, for adverse-action applications, evidence that the model's reason-code outputs are accurate at the transaction level. The accountability record is that evidence. Without it, the institution cannot produce the documentation the 2026 guidance requires, and the absence of documentation is not a neutral fact; it is evidence that the validation did not happen.
In an ECOA litigation or regulatory investigation, the accountability record is the institution's primary defense against an allegation that the stated reasons were inaccurate or discriminatory. An institution that can produce, for each challenged denial, a document showing: (a) the specific adverse-action reasons that were stated, (b) the specific file data and policy provision that each reason traces to, (c) the name of the reviewer who performed the comparison, and (d) the date the comparison was completed, has satisfied the core governance requirement. An institution that cannot produce this documentation has, from the examiner's perspective, no basis for claiming that its stated reasons were accurate, regardless of what the loan officers believed they were doing when they approved the notices.
The accountability record format that works in practice is a structured annotation attached to each denied application in the LOS or document management system. For each reason code in the notice, the annotation records: the reason code text as stated in the notice, the specific document page and value supporting the factual grounding (example: "DTI 51.6% per calculation worksheet attached, using $4,400 gross monthly income from 2024 W-2 and $2,270 proposed total monthly obligations"), the policy provision the reason implicates (example: "Credit Policy Section 3.4.2: maximum back-end DTI for 620-640 FICO tier is 45%"), and a completeness attestation ("I have reviewed the denied application under the credit policy and confirm the stated reasons represent the principal factors in the denial. No material denial factor has been omitted."). The reviewer signs the attestation with their employee identifier and the date of the review.
This documentation level requires 15 to 25 additional minutes per denial file in a properly staffed workflow. At institutions processing 50 to 100 denials per month, the annual time investment is 12 to 50 hours. That investment buys: a defensible adverse-action process under ECOA, a transaction-level validation record under OCC 2026-13, a fair-lending early-warning system, and the professional protection of the underwriters who sign the attestations, because their documentation demonstrates they did the job the law requires. The professional who can produce that record for every denial they reviewed is not just a compliance-aware employee; they are a documented asset to the institution's regulatory defense.
Key Takeaways
- Verification is a distinct, required step between AI-drafted reason codes and a released adverse-action notice. It is not a formatting review or an approval of output that looks professional. It is a substantive, document-level comparison between the stated reason and the actual credit-policy basis for the denial, traced to specific file data.
- The four dimensions of a complete verification are: factual grounding (is there a specific documented file fact supporting the reason?), policy causation (is the reason a factor that, under credit policy, contributed to the denial?), reason-code precision (is the reason as specific as the file supports without overstating the precision of the underlying data?), and completeness (does the reason set cover the principal denial factors, with no material omission?).
- The comparison task requires matching each stated reason to the model's documented decision output for the specific application. AI file summaries are not substitutes for source documents. The verification must go to the credit report, the tax return, the income documentation, and the credit policy, not to the AI's description of those documents.
- The four common verification degradations are: visual review without document comparison, verifying against the AI's own file summary, confirming only the lead reason and assuming the rest, and delegating to a reviewer who does not have the full file. Workflow design must structurally prevent all four.
- When stated reasons and actual reasons diverge systematically across a portfolio, the mismatch can become a fair-lending finding: geographic patterns in reason-code inaccuracy track protected-class concentrations and create the structure of a disparate-treatment or discriminatory-pretext case. Systematic verification prevents accumulation of this risk before an examination finds it.
- OCC Bulletin 2026-13 requires transaction-level validation that AI model outputs accurately represent the actual decision factors for each individual application. The accountability record, attached to each denied file, is the transaction-level documentation that satisfies this requirement.
- A compliant accountability record documents, for each reason code: the factual grounding with page and value, the policy provision implicated, the name and date of the reviewer, and a completeness attestation. This record is the institution's primary legal defense against allegations of inaccurate or discriminatory adverse-action reasons.
- The institution that performs genuine verification on every denial, documents the verification, and maintains the accountability records is the institution that can walk an examiner through any denial in its portfolio and demonstrate, reason by reason, why each stated code is the actual reason. That institution turns the adverse-action trail into a governance asset rather than a liability. That is the point of the verification step.
Skill.re