Standing Up a Fair-Lending AI Testing Program
The following scenario is a composite illustration of how examiners probe for program discipline, not a report of any specific examination. A fair-lending examination at a regional bank began with a question that the bank's new compliance director had not fully prepared for: "Walk me through your fair-lending AI testing program, not the test results, the program." The bank had a disparate-impact test. It had run one before the prior exam and another six months before this one. The results were defensible: no statistically significant adverse-action rate disparities for any ECOA-protected class. The examiner acknowledged the test results within fifteen minutes. But she spent the next two days looking for the program, because test results without a repeatable testing program are not evidence of a governance posture. They are two data points. A program is the cadence, the methodology, the roles, the trigger conditions, the documentation standards, and the escalation protocol that produce test results consistently and institutionally, not just when an exam is near. What the examiner was reconstructing was whether the bank treated fair-lending AI testing as an institution-level discipline or as an event. This lesson builds the institution-level discipline: the operating model that transforms disparate-impact testing from a pre-exam scramble into a running, documented, defensible program that any examiner can reconstruct from the artifacts it generates.
Why a Program Is Different from a Test
Disparate impact means that a facially neutral lending practice produces a statistically significant adverse outcome for a protected class under ECOA (Equal Credit Opportunity Act, the federal statute that prohibits discrimination in any aspect of a credit transaction on the basis of race, color, religion, national origin, sex, marital status, age, or receipt of public assistance income). Disparate impact does not require discriminatory intent; it requires a discriminatory effect. A model can be built entirely without reference to protected class characteristics and still produce outcomes where one racial group is denied credit at a materially higher rate than a similarly situated non-minority group. This is the central compliance risk of AI in lending, and it is the risk that a fair-lending AI testing program exists to detect, document, and address.
A test is a point-in-time measurement: you run it, you get results, you respond to what you find. A program is a system: it defines what gets tested, when, by whom, using which methodology, against which thresholds, with which escalation triggers, and producing which artifacts for governance and examination. The distinction is not academic. Under OCC Bulletin 2026-13 (the April 2026 interagency model-risk guidance, issued by the OCC, or Office of the Comptroller of the Currency, the Federal Reserve, and the FDIC, which superseded OCC Bulletin 2011-12 and made fair-lending risk a formal dimension of model risk management, or MRM), fair-lending testing for AI credit models must be embedded in the model-risk governance framework, not maintained as a standalone compliance activity. That integration requirement is what separates a program from a test.
The examiner in Phoenix was looking for several specific things a test does not provide and a program does:
Consistency: The same methodology applied on the same cadence to every AI credit model in the institution, so that the results are comparable across models and across time periods. An institution that tests one model annually and another model only when there is a regulatory inquiry has a testing gap, not a testing program.
Independence: Testing conducted by a party independent of the model development and deployment function, so that the results are not contaminated by the interests of those who built or deployed the model. The 2026-13 framework requires that fair-lending testing for AI models be part of the independent validation process, which means it must be performed by or reviewed by someone organizationally separate from the model owner's business line.
Documentation: A complete, contemporaneous record of what was tested, how, when, with what results, and what action was taken in response. An institution that ran a good disparate-impact test but did not document the methodology, the data sample, the statistical thresholds, or the follow-up actions has a test it cannot reproduce and a defense it cannot reconstruct.
Escalation: Defined trigger conditions that automatically escalate a finding to the model owner, the fair-lending officer, and the governance committee, rather than leaving the response to individual discretion. The program must specify: if the adverse-action rate ratio for any protected class exceeds a defined threshold, what happens next and who is responsible for making it happen.
A single disparate-impact test tells an examiner what you found once. A documented testing program tells an examiner that you built the discipline to find it every time, act on it consistently, and defend it under scrutiny.
The Program's Scope: Which Models, Which Classes, Which Decisions
The first design decision in building the program is scope: which AI models must be tested, which ECOA-protected classes must be analyzed, and which lending decisions fall within the testing perimeter.
Which AI models: The program should cover every AI model in the institution's model inventory that influences a credit decision, including pre-scoring models (models that produce a score or risk band used by an underwriter), decision-assist models (models that recommend an approve or decline that an underwriter reviews), automated-decision models (models whose output is the final decision with no human review), document extraction models (models that extract and populate income, asset, or liability fields used in underwriting), and GenAI models that produce adverse-action notices or reason codes (because an inaccurate or systematically biased reason code is itself a fair-lending event under ECOA and Reg B, or Regulation B, 12 CFR Part 1002, the CFPB's implementing regulation for ECOA). Models with a High risk rating in the model inventory require the most intensive testing, but all in-scope models require at minimum an annual baseline disparate-impact test.
Which protected classes: ECOA protects applicants from discrimination based on race, color, religion, national origin, sex, marital status, age (for applicants old enough to contract), and receipt of income from public assistance programs. The Fair Housing Act adds disability and familial status for residential mortgage transactions. A comprehensive fair-lending AI testing program must analyze all of these protected classes, not just the ones with historically elevated disparity rates. The institution should maintain a protected-class testing matrix that documents, for each model, which protected classes were tested, which statistical methodology was used, and what the results were.
Which decisions: The program should cover adverse action decisions (denials and counteroffers that are adverse to the applicant under ECOA), rate and terms decisions where AI influences the pricing offered (because a disparate pattern in pricing, even on approvals, is a fair-lending concern under ECOA's scope), and any decision point in the loan origination system (LOS, the software platform that manages the loan application process) where AI output is a direct input to the final disposition. The scope should be documented and reviewed annually to capture new AI deployments and changes in how existing models are used.
Testing Methodology: The Three Required Analyses
A defensible fair-lending AI testing methodology uses three analyses in combination. No single analysis is sufficient, and an institution that relies on only one is missing discrimination that the others are designed to surface.
Analysis 1: Adverse action rate ratio (univariate disparate-impact test). The adverse-action rate ratio compares the rate at which a protected class receives an adverse action to the rate at which a control group (typically white non-Hispanic applicants for race-based analyses, male applicants for sex-based analyses) receives the same adverse action. The most common threshold used by examiners is the 80% rule, also called the four-fifths rule: if the approval rate for the protected class is less than 80% of the approval rate for the control group (equivalently, the adverse-action rate for the protected class divided by the adverse-action rate for the control group exceeds approximately 1.25), there is a prima facie indication of disparate impact warranting further analysis. The four-fifths rule is an analytical screen, not a binding legal safe harbor: a ratio above 0.80 does not immunize the institution from a disparate-impact finding, and a ratio below 0.80 does not by itself establish a violation. Some institutions also track the absolute disparity (the percentage point difference in adverse-action rates) alongside the ratio, because an 80% ratio on a small absolute difference may be less significant than an 80% ratio on a large absolute difference.
The univariate test is the most straightforward to compute and the easiest to explain to a board or an examiner, but it has a known limitation: it does not control for differences in creditworthiness between the protected and control groups. A protected class that has higher adverse-action rates partly because its members have lower credit scores, higher debt-to-income ratios, or less documentation of income is not necessarily experiencing AI-generated discrimination; it may be experiencing legitimate credit-risk differences. The univariate test alone cannot distinguish between these two explanations, which is why the second analysis is required.
Analysis 2: Matched-pair or regression-controlled analysis (multivariate disparate-impact test). The multivariate analysis controls for legitimate credit-risk factors (credit score, debt-to-income ratio, loan-to-value ratio, income documentation type, and other underwriting variables) and asks: after accounting for these factors, is there still a statistically significant difference in adverse-action rates between protected and control groups? A statistically significant residual disparity after controlling for legitimate credit factors is evidence that the model is producing outcomes that cannot be explained by credit risk alone, which is the core of a disparate-impact finding.
The regression model for the multivariate analysis should be designed by the fair-lending testing team rather than borrowed directly from the credit-scoring model, because using the credit model's own features to control for risk in the fair-lending analysis is circular: if one of those features is a proxy variable for a protected class characteristic, including it in the control set masks the very discrimination being tested. Proxy variables are inputs that are not themselves protected class characteristics but that are highly correlated with them: ZIP code is the canonical example, as it is correlated with race due to residential segregation patterns. A fair-lending analysis that controls for ZIP code may inadvertently control away a portion of the racial disparity it is supposed to detect.
Analysis 3: Feature-level attribution and proxy variable screening. The third required analysis goes inside the model's output to examine which features contribute most to the adverse-action pattern for each protected class, and whether any of those features are functioning as proxy variables. SHAP values (SHapley Additive exPlanations, a technique adapted from cooperative game theory that attributes the model's output to each input feature) are the standard tool for this analysis in machine learning models. The SHAP analysis identifies the top contributors to adverse-action decisions for each demographic subgroup and flags features whose SHAP contribution is substantially different across protected and control groups, which is the signature of a potential proxy variable.
The proxy variable screen does not definitively establish discrimination; it identifies features that warrant closer examination. A feature that has materially higher SHAP values for denied minority applicants than for denied white applicants should be examined for correlation with protected class characteristics. If the feature is highly correlated with a protected class characteristic and can be replaced by a less-correlated feature with similar predictive power, the LDA (less-discriminatory alternative) analysis process (the documented search for a model configuration that achieves equivalent credit-risk prediction with less disparate impact) applies.
The Testing Cadence and Trigger Conditions
The program's cadence is the schedule that makes it a program rather than a series of one-off tests. Two types of testing activity must be scheduled: routine testing on a regular cycle, and event-triggered testing that responds to changes in the model or the population.
Routine testing cadence: High-risk AI credit models require quarterly monitoring of adverse-action rate ratios (the univariate test) and annual comprehensive testing that includes the multivariate regression and the feature attribution analysis. Medium-risk models require semi-annual monitoring and annual comprehensive testing. The quarterly monitoring for High-risk models does not require the full analytical apparatus of the annual comprehensive test; it requires running the univariate adverse-action rate ratio against the current quarter's loan applications and comparing the result to the established baseline. This is a computationally simple calculation that can be standardized into a quarterly report template and run by a fair-lending analyst, not a data scientist.
The annual comprehensive test requires more resources: a qualified data analyst or external fair-lending testing firm, a complete data set from the prior twelve months, the multivariate regression analysis, and the SHAP attribution analysis. The annual test generates the primary fair-lending documentation that becomes part of the model-risk file and is reviewed by the independent model validator as part of the annual model validation.
Event-triggered testing: Four types of events should trigger an out-of-cycle fair-lending test, in addition to the routine schedule:
Model changes: Any change to the model's feature set, threshold, training data, or decision logic triggers a new disparate-impact analysis before the changed model goes into production. This is a requirement that belongs in the model change management process, not just in the testing program. The change approval checklist should include: "Has fair-lending testing been completed on the proposed change?" A change that has not been tested for disparate impact should not be approved for production deployment.
Population composition changes: If the institution's applicant population changes materially (for example, the institution expands into a new geographic market, launches a new loan product, or changes its marketing reach), the existing fair-lending baseline may no longer be representative. A population-triggered test establishes a new baseline for the changed population context.
Monitoring alerts: If the quarterly monitoring of adverse-action rate ratios produces a result that exceeds the alert threshold established in the monitoring program, a comprehensive test must be conducted within 60 days of the alert. The alert threshold is the point at which a quarterly ratio deviation is significant enough to require deeper analysis rather than just documentation and a watchful eye on the next quarter's results.
Examination or complaint: An examiner's inquiry about AI fair-lending, a CFPB (Consumer Financial Protection Bureau) complaint that references demographic disparities, or a CRA (Community Reinvestment Act) examination finding related to the institution's lending patterns in protected-class neighborhoods should trigger a comprehensive review of the relevant models, even if the routine testing schedule has not yet reached the next cycle. The institution should not wait for the exam to compile its fair-lending evidence.
The LDA Integration: Fair-Lending Testing Meets Model Risk
The single most important integration in the fair-lending AI testing program is the connection between the disparate-impact testing results and the LDA (less-discriminatory alternative) analysis. Under OCC Bulletin 2026-13, the LDA documentation is a required component of the model-risk file for any AI model used in credit decisioning. Under the three-step disparate-impact legal framework, the LDA analysis is the institution's defense against the Step 3 argument that a fairer alternative existed and was not adopted. These two requirements converge on the same document: the LDA documentation in the model-risk file is both the compliance record and the legal defense.
The LDA analysis asks: given the disparate-impact results for this model, are there alternative model configurations that would achieve substantially equivalent credit-risk prediction with materially less adverse impact on protected classes? The analysis tests specific alternative configurations: feature removal (removing the highest-disparity-contributing features and measuring the accuracy and disparity profile of the remaining model); feature substitution (replacing a high-disparity feature with a correlated but less-discriminatory alternative); threshold adjustment (testing the disparity profile at alternative decision thresholds, using the Pareto frontier of accuracy versus disparity to identify whether a lower-disparity threshold is available at an acceptable accuracy cost); and fairness-constrained model alternatives (training a version of the model using an algorithm that explicitly optimizes for a combination of predictive accuracy and outcome equality).
The LDA analysis must reach a documented conclusion: either no less-discriminatory alternative was found that meets the institution's credit-policy minimum accuracy requirements (with all alternatives tested, their disparity and accuracy profiles documented, and the selection conclusion signed by a named responsible officer); or a less-discriminatory alternative was identified, with a documented plan for its adoption. An LDA analysis that documents testing without reaching a conclusion is not a defensible document, because it does not demonstrate that the institution made a good-faith determination about the availability of a fairer alternative.
The testing program's cadence connects to the LDA in a specific way: every annual comprehensive fair-lending test should include a review of whether the prior LDA conclusion remains valid. Since the LDA analysis is based on the state of credit-risk modeling at the time it was conducted, an annual review should ask: have new modeling techniques or data sources become available since the last LDA analysis that might produce a less-discriminatory alternative that did not exist before? If yes, those alternatives must be evaluated in the current cycle's LDA update.
Documentation Standards: The File the Examiner Reconstructs
The testing program's documentation standards determine whether the program is examinable. A program that runs rigorous tests but documents them incompletely is, from the examiner's perspective, a program that may not exist: if the examiner cannot reconstruct what was done, when, and with what result, the evidence of a defensible program is not present.
Each fair-lending test, whether routine or event-triggered, should produce a standard documentation package:
Test scope memo: A one-to-two page document that identifies the model being tested, the time period covered, the ECOA-protected classes analyzed, the data source and sample size, the statistical methodology applied, and the name of the analyst or firm that conducted the test. The scope memo is the cover document for the full test package and is the first thing an examiner reads to orient to the test.
Statistical results table: A structured table reporting the adverse-action rate for each protected class and the control group, the adverse-action rate ratio, the statistical significance of any disparity (p-value or confidence interval), and whether the ratio triggered an alert threshold. For the annual comprehensive test, the table should also include the multivariate regression residual disparity for each protected class and the top-five features by SHAP contribution to adverse-action decisions for each protected class with elevated disparity.
Conclusion and response memo: A brief document that states the test's finding (no material disparity, disparity below alert threshold, disparity above alert threshold) and the institutional response (no action required, enhanced monitoring, event-triggered LDA review, model change request). The conclusion memo must be signed by the fair-lending officer and, for any above-threshold finding, by the model owner and the head of the governance committee. The signature trail is not bureaucratic formality; it is the evidence that a responsible human reviewed the findings and made an accountable decision about the institutional response.
LDA documentation (for any above-threshold finding): The full LDA analysis as described in the prior section, documenting the alternatives tested, the results, and the conclusion, signed by the responsible officer.
Board and governance committee routing: A record showing that the test results and any LDA documentation were presented to the governance committee and included in the board reporting package for the relevant period. This routing record is what the examiner uses to verify that fair-lending testing results moved through the governance structure as required by the 2026-13 framework, rather than remaining in a compliance filing that no one with authority reviewed.
The documentation package for each test should be maintained in the model's risk file, linked to the model inventory entry, and retained for a minimum of five years (or the institution's regulatory examination cycle, whichever is longer). An institution that produces well-structured, consistently formatted documentation packages for every test is demonstrating that the testing program is institutional infrastructure, not individual effort.
Staffing the Program and Managing the Vendor Option
Running a fair-lending AI testing program at the level of rigor required by OCC Bulletin 2026-13 requires specific capabilities: statistical analysis of lending outcomes, knowledge of ECOA and the disparate-impact legal framework, familiarity with machine learning explainability techniques (SHAP analysis and related approaches), and understanding of the model-risk governance context. Community and regional banks may not have all of these capabilities in-house, which makes the vendor option important to understand.
In-house staffing model: An institution with a dedicated fair-lending AI testing capacity typically staffs a fair-lending analyst with quantitative skills (capable of running the univariate and multivariate analyses), a model-risk or data science resource with explainability expertise (capable of running SHAP analyses on the institution's AI models), and a fair-lending officer with legal and regulatory knowledge (capable of interpreting results and making the LDA determination). At larger institutions, these may be full-time dedicated roles. At smaller institutions, they may be part-time assignments within larger risk and compliance teams, supplemented by vendor support for the more technically intensive components.
Vendor option: External fair-lending testing vendors can provide the full testing package, from data extraction through statistical analysis through LDA documentation, as an outsourced service. Using a vendor does not transfer the institution's compliance accountability: the institution must still review and approve the vendor's methodology, integrate the results into the model-risk file, route the findings through the governance structure, and sign the conclusion memo. The vendor provides analytical capacity; the institution provides oversight and accountability. Institutions using a vendor for fair-lending AI testing should document the vendor's methodology in the test scope memo, retain the vendor's analysis as part of the documentation package, and conduct an annual review of the vendor's methodology to confirm it continues to meet the program's standards.
Independent validation integration: Regardless of whether the testing is performed in-house or by a vendor, the annual fair-lending test results must be reviewed by the institution's independent model validator as part of the annual model validation. The validator is not re-running the tests; the validator is reviewing the testing methodology, assessing whether the tests were appropriately scoped and correctly executed, and confirming that the findings and conclusions are supported by the evidence. This review is documented in the annual validation report and is the mechanism by which the fair-lending testing program becomes part of the institution's formal model-risk governance record.
Key Takeaways
- A fair-lending AI testing program is distinct from a fair-lending test: the program defines the cadence, methodology, roles, documentation standards, and escalation protocols that produce test results consistently and institutionally, not just before examinations.
- Under OCC Bulletin 2026-13, fair-lending testing for AI credit models is a required element of the model-risk governance framework, not a standalone compliance activity; results must be routed through the model-risk structure and included in board reporting.
- Program scope must cover all AI models that influence credit decisions (pre-scoring, decision-assist, automated-decision, and GenAI adverse-action drafting), all ECOA-protected classes, and all adverse-action and material-pricing decision points in the LOS.
- The three required analytical methods are the univariate adverse-action rate ratio (the four-fifths rule), the multivariate regression-controlled analysis (residual disparity after controlling for legitimate credit factors), and the SHAP-based feature attribution and proxy variable screen.
- High-risk AI credit models require quarterly monitoring of adverse-action rate ratios and annual comprehensive testing; event triggers (model changes, population changes, monitoring alerts, and examination inquiries) require out-of-cycle testing regardless of the scheduled cadence.
- Every annual comprehensive test must include a review of whether the prior LDA (less-discriminatory alternative) conclusion remains valid given advances in credit-risk modeling since the last search; above-threshold findings require a full LDA analysis with a documented conclusion signed by a named responsible officer.
- Each test must produce a standard documentation package (scope memo, statistical results table, conclusion and response memo with signatures, LDA documentation if triggered) maintained in the model's risk file and linked to the model inventory entry.
- Community banks may use vendors for analytical capacity, but the institution retains compliance accountability: methodology review, results integration into the model-risk file, governance routing, and the required annual independent validator review of the testing program all remain institutional responsibilities.
Skill.re