โ†
AI for Banking & Lending
Proficient ยท M3 ยท lesson 3 of 19 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Building a Disparate-Impact Testing Workflow
๐Ÿ“–
now learning

Building a Disparate-Impact Testing Workflow

15 min

The fair-lending exam began on a Tuesday morning in March, and by Thursday the examiners from the OCC (Office of the Comptroller of the Currency) had a single, precise question for the compliance team at Meridian Community Bank: "Produce your most recent disparate-impact test for the AI pre-scoring model you deployed in Q3 of last year, along with documentation of any remediation and your less-discriminatory-alternative search." The compliance director, Angela Reyes, had spent the previous eighteen months building out an AI-assisted underwriting workflow that cut cycle time from eleven days to four. She had vendor certifications, model cards, and performance metrics in binders. What she did not have was a structured, repeatable disparate-impact testing workflow that documented outcomes across protected classes as a regular operational process rather than a one-time pre-deployment check. The exam that followed cost the bank eight months of remediation work, a substantial consent-order payment, and a full rebuild of the fair-lending testing program from scratch. (Meridian Community Bank is a composite illustrative scenario; the figures are representative rather than drawn from a single actual enforcement action.) What Meridian was missing is the subject of this lesson: not whether to test for disparate impact, but how to build the testing as a workflow that produces defensible documentation before the examiner asks for it.

Why Disparate Impact Testing Must Be a Workflow, Not an Event

The word "workflow" is doing load-bearing work in this chapter's title. A fair-lending compliance team that tests for disparate impact once, at model deployment, and then moves on has built a one-time event. An institution that builds disparate impact testing as a repeatable, scheduled, documented process has built a workflow. The legal and regulatory difference between those two things, in 2026, is significant.

Disparate impact (the legal doctrine under which a facially neutral lending policy or model is discriminatory if it produces a statistically significant adverse outcome for a protected class without sufficient business justification) is not a static condition. A model that passes its pre-deployment disparate-impact test may develop disparate-impact exposure over time as the applicant population shifts, as economic conditions change the model's behavior at the margin, or as the lender modifies the model's features or thresholds. The OCC Bulletin 2026-13, the April 2026 interagency model-risk guidance that superseded OCC 2011-12 and pulled AI and generative AI under model-risk, fair-lending, third-party, and board-governance expectations, explicitly requires ongoing monitoring of AI model outcomes for fair-lending risk. "Ongoing" is the regulatory standard. The one-time deployment test meets none of that standard.

The other reason testing must be a workflow rather than an event is documentation. A well-documented workflow generates the specific type of record that a fair-lending examination requires: a testing history showing when tests were run, what populations were analyzed, what disparities were found or not found, what remediation was taken in response to any findings, and what less-discriminatory-alternative (LDA) search was conducted. An institution without that documentation cannot demonstrate compliance. It can only assert it. Assertions do not satisfy examiners; records do.

There is also a practical institutional reason. Fair-lending testing that exists as a workflow, with assigned owners, a regular cadence, defined methods, and required documentation, is testing that actually happens. Fair-lending testing that is conceptually everyone's responsibility but no one's assigned task tends to happen when someone remembers to do it, which in a busy lending shop is another way of saying "not reliably." Building the testing into the workflow structure of the compliance and model-risk functions is the governance move that converts a theoretical obligation into an operational practice.

A disparate-impact test you cannot document is a test that, in a regulatory examination, did not happen. The workflow exists to produce the record, not just the result.

Before building the workflow, it is necessary to understand what the workflow is designed to produce evidence of, because the legal standard defines the testing requirement. Disparate impact in credit lending is governed by ECOA (Equal Credit Opportunity Act, 15 U.S.C. 1691 et seq., the federal statute enacted in 1974 prohibiting credit discrimination on the basis of race, color, religion, national origin, sex, marital status, age, and receipt of public assistance income) and Regulation B (Reg B, 12 CFR Part 1002, the CFPB's implementing regulation for ECOA). The Fair Housing Act adds a parallel prohibition for real-estate-related credit.

Under the three-step legal framework for disparate-impact analysis:

Step one is the establishment of a prima facie case by showing that a facially neutral policy, practice, or model produces a statistically significant adverse effect on a protected class. "Facially neutral" is the key phrase. A model that does not use race, sex, national origin, or any other protected characteristic as an explicit input can still produce disparate impact if its neutral-seeming inputs produce outcomes that disproportionately harm a protected group.

Step two shifts the burden to the lender to demonstrate that the challenged practice is justified by a legitimate business necessity, meaning it achieves a legitimate, substantial objective that cannot be served equally well by a less-discriminatory alternative. The business-necessity standard requires specific, quantified documentation of why the challenged feature or model configuration is necessary for accurate credit risk prediction. An unquantified assertion that the feature improves the model does not satisfy the standard.

Step three allows the plaintiff or examiner to establish that a less-discriminatory alternative existed and was available. If such an alternative existed, the business-necessity defense fails even if the necessity was genuine, because the same objective could have been achieved with less discriminatory impact. This is why the LDA search is the linchpin of the defense: it is the documented record that the lender considered and evaluated available alternatives and selected the option that best balanced accuracy and fairness.

The protected classes under ECOA include race, color, religion, national origin, sex (including gender identity and sexual orientation as of 2021 CFPB guidance), marital status, age (provided the applicant has legal capacity to contract), and receipt of public assistance income. The Fair Housing Act adds familial status and disability for real-estate-related transactions. A complete disparate-impact workflow tests for each of these classes, not just race, and uses a testing method appropriate to each class's statistical distribution in the lender's market.

The 2026 regulatory environment has made one additional specification clear: the lender, not the vendor, bears the burden of demonstrating disparate-impact compliance for any AI model it deploys in credit decisioning. A vendor's certification that its model is "fair-lending compliant" or has been tested for bias does not satisfy the institution's regulatory obligation. The lender must conduct or commission its own disparate-impact testing using its own applicant and decision data, because the model's disparate-impact profile in the lender's market, with the lender's applicant population, is the relevant measure, not the model's performance in an abstract test environment.

Building the Testing Cadence and Governance Structure

The governance structure for a disparate-impact testing workflow answers four questions: who owns the testing, how often does it run, what triggers an out-of-cycle test, and who reviews and acts on the results. Each of these questions has a defensible answer that flows from the regulatory standard, and each answer needs to be documented in a policy or procedure that an examiner can read.

Ownership. Disparate-impact testing sits at the intersection of fair-lending compliance and model risk. Both functions need ownership stakes in the process. A common governance structure assigns primary ownership to the fair-lending compliance function (which owns the legal standard, the analysis methods, and the response to findings) with a required model-risk review of the testing methodology and any model changes triggered by findings. The responsible human must be named: not "the compliance team" but the Chief Compliance Officer or a named Fair-Lending Officer with specific accountability for the testing program. Under OCC 2026-13's governance expectations, board and senior management oversight of AI model risk includes fair-lending risk, meaning test results and remediation decisions need to reach a management-level governance committee on a defined schedule.

Cadence. Annual testing at a minimum is the standard floor for a model in stable operation. For high-volume models (processing more than a few hundred applications per month), quarterly testing is a defensible baseline because the statistical power to detect disparities improves with larger sample sizes and the lender accumulates enough data in a quarter to run meaningful analysis. Lenders whose volumes are lower may need to aggregate multiple months of data to achieve sufficient statistical power, but the analysis still needs to happen on a schedule, not ad hoc.

Triggers for out-of-cycle testing. Six conditions should trigger a disparate-impact test outside the regular cadence: (1) any modification to model features, thresholds, or weights; (2) any retraining of the model on new data; (3) a material change in the applicant population served (a new market entry, a new product launch, a significant shift in application volume from a particular geographic area); (4) a significant macroeconomic change that could shift model behavior at the credit margin; (5) any fair-lending complaint or referral related to the AI model; and (6) any supervisory finding or guidance that identifies the model's type or methodology as a fair-lending concern. Each of these conditions can change a model's disparate-impact profile independently of the model's documented configuration, so each requires a fresh test.

Review and escalation. Test results need a documented review process: who reviews them, what threshold triggers escalation, and what actions are required at each escalation level. A common framework uses a traffic-light approach: no material disparity (green, document and continue), disparity below a defined threshold (yellow, investigate cause, document, monitor next cycle), disparity at or above the threshold (red, initiate LDA search, prepare remediation plan, escalate to senior management and governance committee, report to board). The specific thresholds should be defined in writing in the testing policy, not left to case-by-case judgment, because documented thresholds are defensible and case-by-case judgments are not.

The Testing Methodology: Measuring Outcomes Across Protected Classes

With governance structure in place, the next layer of the workflow is the testing methodology itself. This is where the statistical work happens, and where institutions that rely solely on vendor reports or summary statistics most often leave gaps that examiners will find.

A complete disparate-impact test for an AI credit model measures outcomes across each protected class using at least two analytical approaches: a univariate disparity analysis and a multivariate regression analysis. Each serves a different purpose, and neither alone is sufficient for a fully defensible test.

Univariate disparity analysis computes the raw denial rate (or adverse-outcome rate, for models that produce rate pricing or credit limit decisions rather than binary approve/deny) for each protected class and compares it to the denial rate for the benchmark group. The standard statistical measures include:

The adverse action rate ratio (the denial rate for the protected group divided by the denial rate for the benchmark group). CFPB examination guidance has used an adverse action rate ratio of 1.5 or higher as a threshold for heightened review. An adverse action rate ratio of 2.0 means the model denies protected-class applicants at twice the rate of the benchmark group. This is a red-flag threshold for disparate-impact examination.

The disparity index (calculated in some testing frameworks as the ratio of approval rates: the approval rate for the benchmark group divided by the approval rate for the protected group). An index of 1.25 or higher, meaning the benchmark group is approved at a rate 25 percent or more above the protected group, is a common heightened-review threshold.

The statistical significance test using a chi-square test or Fisher's exact test for small samples, confirming that the observed disparity is unlikely to be the result of random variation in the sample. A p-value below 0.05 is the conventional significance threshold, though some testing frameworks use 0.01 for high-stakes models. Note that statistical significance and practical significance are different things: a very large sample can produce a statistically significant disparity that is practically trivial, and a small sample may fail to detect a meaningful disparity because the sample size lacks statistical power.

Multivariate regression analysis controls for legitimate credit risk factors (credit score, debt-to-income ratio, loan-to-value ratio, employment history, income verification status) and measures the residual adverse-outcome disparity attributable to protected class membership after those factors are held constant. This is the more legally significant analysis, because it addresses the lender's most common defense ("our model's disparate outcomes reflect legitimate creditworthiness differences, not discrimination") directly. A multivariate regression that shows a statistically significant disparity after controlling for legitimate credit factors is evidence that the model's outcomes cannot be fully explained by legitimate risk factors.

For a 2026 AI-assisted underwriting model, the multivariate regression analysis faces a specific challenge: the model's own risk score is highly correlated with the legitimate credit factors being controlled for, and including the model score in the regression as a control variable can suppress the measured disparity. The testing methodology must address this problem explicitly. Two common approaches are: (1) use only the underlying credit factors (credit score, DTI, LTV, employment) as controls rather than the model's composite output, so the analysis tests whether the model's outcomes can be explained by those factors without assuming the model's output is the best measure of them; or (2) include the model score as an additional variable alongside the underlying factors and test whether protected class membership has explanatory power beyond the model's own score, which tests for "residual discrimination" in the model's output after its stated risk logic is accounted for.

The testing must also address data completeness and proxy imputation. HMDA (Home Mortgage Disclosure Act) data provides race, national origin, and sex for most mortgage applications. For non-mortgage consumer credit applications, race and national origin are not collected, and the analysis must use proxy methods to estimate protected class membership. The CFPB's BISG (Bayesian Improved Surname Geocoding) methodology is the regulatory standard for race and ethnicity proxy in non-HMDA contexts. The methodology, and its specific implementation in the institution's testing process, must be documented.

The output of the testing methodology is a structured test report for each cycle. The report should document: the analysis period and sample size for each protected class tested; the univariate disparities with statistical significance tests; the multivariate regression results; any proxy methodology used and its documented accuracy in the institution's market; findings of material disparities and their threshold classification; and the planned next step (continue monitoring, initiate LDA search, or remediate). This report is the core document the examiner will review.

Proxy Variables and Indirect Discrimination in the Testing Scope

A complete disparate-impact workflow tests not just the model's ultimate outcomes but also the features that drive those outcomes. This is the proxy-variable analysis layer of the workflow, and it is where many institutions' testing programs stop short.

A proxy variable (a variable that is not itself a protected characteristic but that correlates with a protected characteristic strongly enough to serve as a statistical substitute for it) can drive disparate impact in an AI model even if the model never receives race, national origin, sex, or other protected-class data as an explicit input. Common proxy variables in AI credit models include geographic features (zip code, census tract, neighborhood designation), employer characteristics (employer industry code, employer location, employer size), transactional behavioral features (merchant category patterns, subscription service usage, payment timing patterns), and asset account features (which types of financial institutions hold the applicant's accounts). Each of these can correlate with race or national origin through the structural legacy of residential segregation, employment concentration, and wealth-accumulation patterns that persist from periods of explicit discrimination.

The proxy-variable analysis layer of the testing workflow uses feature attribution methods to identify which features contribute most to the model's disparity, so that the LDA search can target those features specifically. The primary tools are:

SHAP values (SHapley Additive exPlanations, a method from cooperative game theory adapted to machine learning by Lundberg and Lee) decompose each model prediction into additive contributions from each feature. By computing SHAP values across protected-class and benchmark applicants and comparing the distributions, the fair-lending analyst can identify which features contribute disproportionately to the adverse outcomes of protected-class applicants. A geographic feature that has high average SHAP magnitude for denied applicants who are predominantly from majority-minority census tracts is a candidate proxy variable.

Permutation importance with demographic stratification measures how much the model's disparity changes when a specific feature is randomly scrambled (removing its predictive power) within protected-class and benchmark subgroups. Features whose permutation increases the disparity are features that are currently helping protected-class applicants; features whose permutation reduces or eliminates the disparity are proxy variables driving adverse outcomes for that class.

Conditional independence testing directly tests whether a feature contains information about protected class membership after legitimate credit factors are controlled for. A feature that is conditionally dependent on protected class membership (meaning it provides information about whether an applicant is a member of a protected class, beyond what the legitimate credit factors already tell you) is a proxy variable by definition, regardless of whether the feature was included with discriminatory intent.

The proxy-variable analysis layer is not just a fairness concern. It is a direct legal issue. The CFPB's enforcement position, and the position of courts that have addressed the question, is that a lender cannot escape disparate-impact liability by pointing to a neutral feature name and saying "we used zip code, not race." If zip code functions as a proxy for race in the model's outcomes, the use of zip code produces disparate impact in the same way that using race directly would. This is the "we didn't use race as an input" argument that fails, and the proxy analysis is the methodology that demonstrates why it fails for any specific model.

Incorporating the proxy-variable analysis into the testing workflow serves a practical purpose beyond compliance: it identifies the specific features that need to be evaluated in the LDA search, focusing that search on the features that actually drive the disparity rather than requiring the institution to evaluate every feature in the model.

The LDA Search as Part of the Workflow

The LDA (less-discriminatory alternative) search is the documented evaluation of whether an alternative model configuration would achieve substantially equivalent credit risk prediction with materially less disparate impact on protected classes. It is the legal defense to a disparate-impact finding, and it must be documented before the examiner asks, not assembled in response to an examination finding.

Within the testing workflow, the LDA search is triggered by any test cycle that identifies a material disparity (a red-flag finding in the traffic-light framework described above). It should also be conducted proactively at model deployment and at any model modification, as part of the governance documentation, even if no disparity is found, because a proactive LDA search is stronger evidence than a reactive one: it demonstrates that the institution evaluated alternatives and found none that would reduce disparate impact without unacceptable accuracy loss, rather than simply asserting that no such alternatives existed.

The LDA search process has five documented steps:

Step one: identify candidate features for removal or modification. Using the proxy-variable analysis from the testing workflow, identify the features that contribute most to the measured disparity. These are the targets for the LDA analysis. Document the specific features and their estimated contribution to the disparity, in quantitative terms (percentage reduction in disparity from removing or modifying the feature, based on the SHAP or permutation analysis).

Step two: build and test alternative model configurations. For each candidate feature or set of features, build a version of the model that removes or modifies the feature and train it on the historical data. Measure both the disparate-impact reduction and the accuracy reduction for each alternative. Document the accuracy metric (AUROC, Kolmogorov-Smirnov statistic, Gini coefficient, or institution-specific credit metric) and the disparate impact measurement for each alternative. Use the same testing methodology as the baseline disparate-impact test so that results are directly comparable.

Step three: conduct the business-necessity analysis for each alternative. For each alternative model configuration, document the specific accuracy loss relative to the baseline model and evaluate whether that accuracy loss is acceptable under the institution's credit policy. An accuracy loss that would materially increase default rates, impair the institution's ability to originate credit within its risk appetite, or violate regulatory capital requirements is an unacceptable accuracy loss that supports retaining the challenged feature. An accuracy loss that is statistically measurable but practically trivial (a reduction in AUROC from 0.83 to 0.82, for example) does not support retaining the feature if it drives material disparate impact. The business-necessity analysis must be specific, quantified, and documented: "Removing zip code from the model reduces AUROC by X, which we estimate would increase annual expected credit losses by $Y based on our historical performance data" is business-necessity documentation. "Zip code helps the model" is not.

Step four: select the model configuration and document the rationale. Choose the model configuration that best balances accuracy and fairness, document the selection rationale, and if the baseline model is retained over a less-discriminatory alternative, document precisely why the accuracy cost of the alternative was unacceptable. The rationale must be defensible: it must show that the institution genuinely considered the alternative and rejected it for specific, quantified reasons, not that it preferred the baseline model for reasons unrelated to accuracy.

Step five: schedule the next LDA review. The LDA search is not permanent. A model that retains a disparate-impact-producing feature based on a 2026 business-necessity analysis must revisit that analysis periodically, because credit modeling techniques improve and a feature that was necessary for accuracy in 2026 may be replaceable with a less-discriminatory feature by 2027. The testing workflow should include a scheduled LDA review at each testing cycle for any model that retains a feature on business-necessity grounds.

The documentation produced by the LDA search is the core of what an examiner is looking for when they ask, as the examiners at Meridian Bank did, for "documentation of any remediation and your less-discriminatory-alternative search." The workflow produces this documentation as a natural output of each testing cycle rather than requiring the institution to reconstruct it in response to an examination request.

AI Tools in the Testing Workflow

An AI-assisted fair-lending testing workflow is not a contradiction. AI tools can meaningfully accelerate several components of the testing process: data aggregation and cleaning, proxy imputation using BISG or similar methods, SHAP value computation for large models, multivariate regression setup, and report generation. But AI tools create specific risks in a fair-lending testing context that must be actively managed.

The automation risk. An AI-assisted workflow that runs automatically on a schedule and generates automated test results without human review is a workflow with no human accountability for the testing conclusions. OCC Bulletin 2026-13 requires that model risk governance include human judgment in the interpretation of model outputs. A fair-lending test is a model output about another model: the testing methodology is itself a model, and its conclusions require human expert interpretation. The workflow must include a human review step where a qualified fair-lending or model-risk analyst reviews the testing results, evaluates the proxy-variable findings, makes the traffic-light threshold determination, and signs off on whether an LDA search is required. The AI accelerates the computation; the human makes the compliance determination.

The hallucination risk in narrative generation. If AI tools are used to generate the narrative portions of the test report (for example, summarizing test results or drafting the LDA search documentation), those narratives must be verified against the actual test data before the report is finalized. An AI-generated statement that "no material disparities were found" when the actual data shows an adverse action rate ratio of 1.48 for a protected class is a false compliance record, regardless of how it was generated. Every factual statement in the test report must be verified against the underlying data.

The model-on-model risk. When the AI tools used in the testing workflow were developed by the same vendor as the model being tested, there is a structural conflict of interest: the vendor's testing tools may not be designed to detect disparities in the vendor's model that the vendor would prefer not to disclose. The institution's testing program must use independent methods, or at minimum supplement vendor-provided testing with independently conducted analysis, to ensure the testing is genuinely designed to detect disparities rather than to document their absence.

A practical AI-assisted testing workflow might look like this: an automated pipeline runs the BISG proxy imputation, computes the univariate disparity ratios, runs the SHAP attribution analysis, and generates a structured data output with all the numerical results. A fair-lending analyst reviews the structured output, applies the traffic-light threshold framework, runs or commissions the multivariate regression, evaluates the proxy-variable findings, and writes the test report's conclusions and recommendations. The AI handles the computation-intensive steps; the analyst handles the judgment-intensive steps. The workflow log documents both contributions, so the institution can demonstrate that a human made the compliance determinations even in a heavily automated testing environment.

Documenting for the Exam

The final layer of the workflow is documentation design. The testing workflow generates value only if its documentation meets the standard of what a fair-lending examiner expects to see. Building the documentation structure as part of the workflow design, rather than after the fact, ensures that the right records are created in the right format at the right time.

A complete fair-lending testing file for a single AI model includes the following elements:

Testing policy and procedures. The written document establishing the testing cadence, the responsible parties, the trigger conditions for out-of-cycle testing, the statistical methods employed, the threshold framework for escalation, and the documentation requirements. This document establishes that the institution has a testing program, not just ad hoc tests. It should be reviewed and approved annually by senior management or the governance committee.

Model inventory entry. Each AI model used in credit decisioning should have an entry in the institution's model inventory that includes the model's purpose, deployment date, data inputs, model type, the date of its most recent disparate-impact test, the test result, any findings and their resolution, and the date of the next scheduled test. The model inventory is often the first document an examiner reviews, and a model inventory entry that shows no disparate-impact test date is a red flag that prompts immediate examination of the testing program.

Test reports. For each testing cycle, a structured report documenting the analysis period, sample sizes, protected classes tested, univariate disparity results with statistical significance, multivariate regression results, proxy-variable findings, traffic-light determination, and planned next steps. Reports should be numbered sequentially and retained indefinitely, because a multi-year history of test results is evidence of ongoing compliance monitoring.

LDA search documentation. For any cycle that triggers an LDA search (or for the proactive LDA search at deployment), the structured documentation described in the previous section: candidate features identified, alternative configurations tested, accuracy and disparate-impact metrics for each alternative, business-necessity analysis, selection rationale, and next LDA review date.

Remediation records. Where a test cycle identifies a material disparity that requires model modification, the remediation record documents the disparity found, the specific model change made in response, the remediation timeline, the re-test results confirming that the modification achieved the intended disparity reduction, and the sign-off by the responsible compliance and model-risk officers.

Governance committee minutes or management reports. Evidence that test results were presented to the appropriate oversight level (management committee, board-level risk committee) in accordance with the governance policy. A test report that was prepared but never presented to oversight has a documentation gap that examiners will note.

When Angela Reyes at Meridian Bank came through the remediation period after the March examination, the program she built was centered on exactly these documentation elements. The cost of building that program after an examination is many times the cost of building it before one. The workflow approach inverts that cost structure: the testing documentation is a byproduct of an operational process rather than a crisis response.

The metric by which a disparate-impact testing workflow succeeds is not whether it finds disparities. It is whether, when an examiner asks for the documentation of the testing program, the institution can produce a comprehensive, consistent, well-organized file that demonstrates ongoing, methodologically sound, human-reviewed testing across all protected classes with a documented LDA search and a clear escalation and remediation history. That file is what the workflow is built to produce.

Key Takeaways

  • Disparate impact testing must be a repeatable workflow with assigned ownership, a documented cadence, defined escalation thresholds, and structured reporting, not a one-time event at model deployment. OCC Bulletin 2026-13 requires ongoing monitoring of AI model outcomes for fair-lending risk, and a one-time test does not satisfy that standard.
  • The legal framework requires testing across each protected class under ECOA (Equal Credit Opportunity Act) and the Fair Housing Act, using both univariate disparity analysis and multivariate regression that controls for legitimate credit factors. Neither method alone is sufficient for a fully defensible disparate-impact test.
  • Six conditions should trigger out-of-cycle testing beyond the regular schedule: model feature or threshold changes, model retraining, material shifts in applicant population, significant macroeconomic changes, fair-lending complaints, and supervisory guidance identifying the model type as a concern.
  • Proxy-variable analysis using SHAP values, permutation importance, or conditional independence testing identifies which features drive disparate outcomes for protected classes, enabling the LDA (less-discriminatory alternative) search to focus on the features that actually matter rather than evaluating the entire feature set.
  • The LDA search must be documented before the examiner asks. A proactive LDA search at deployment and at each model modification, with quantified accuracy and disparity metrics for evaluated alternatives, is a stronger defense than a reactive search assembled after a finding.
  • AI tools can accelerate computation-intensive testing steps (proxy imputation, SHAP attribution, regression setup) but must not replace human judgment in interpreting test results, applying threshold frameworks, and making compliance determinations. The workflow log must document that a qualified human made those determinations.
  • The institution, not the vendor, bears the burden of demonstrating disparate-impact compliance for any AI model it deploys. A vendor's certification does not satisfy the lender's regulatory obligation to test its own applicant data in its own market.
  • Documentation designed for examination readiness, including the testing policy, model inventory entries, structured test reports, LDA search records, remediation files, and governance committee evidence, is the output the workflow is built to produce. The cost of building this documentation reactively after an examination finding is many times the cost of building it proactively as part of the workflow.