โ†
AI for Banking & Lending
Visionary ยท M10 ยท lesson 10 of 16 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Running Lending AI Pilots to Evidence
๐Ÿ“–
now learning

Running Lending AI Pilots to Evidence

15 min

In the spring of 2025, a $19 billion regional bank declared its AI-assisted mortgage underwriting pilot a success. The model had processed 4,200 applications over eight weeks, reduced average decision time from eleven days to three, and flagged 94 percent of subsequent defaults within its top-risk tier. The chief risk officer reviewed the summary report and asked a single question: "Where is the fair-lending analysis?" The pilot team had produced eleven pages of accuracy metrics, throughput statistics, and cost reduction projections. The fair-lending analysis occupied three bullet points at the end. The CRO returned the report unsigned. The model sat in limbo for four months while the team went back and built the evidence the CRO needed. The cost of that four-month delay, in engineer time, business-line frustration, and competitive displacement, exceeded $800,000. (The scenario above is a composite illustration; the institutions, figures, and timelines are representative rather than drawn from a single documented event.) The lesson was not that the model was wrong. The lesson was that a pilot generates two kinds of evidence: the kind that impresses an innovation steering committee and the kind that satisfies a chief risk officer. A well-designed pilot produces both simultaneously, because the discipline required to produce the second kind is the same discipline that makes the first kind credible.

What Evidence a Chief Risk Officer Actually Needs

The chief risk officer's signature on a scale decision is not an endorsement of accuracy metrics. It is a certification that the institution can deploy this model in production, own the decisions it influences, explain those decisions to any applicant who receives an adverse action, demonstrate that the model does not produce disparate impact on protected classes, and defend the deployment under examination. These are five distinct evidentiary requirements, and a pilot that addresses only accuracy metrics satisfies one of the five. The other four are governance requirements that flow directly from the institution's legal obligations under ECOA (Equal Credit Opportunity Act), Regulation B (Reg B, 12 CFR Part 1002, the CFPB's implementing rule for ECOA), and OCC Bulletin 2026-13 (the April 2026 interagency model-risk update jointly issued by the OCC, the Federal Reserve, and the FDIC, which extended model risk management (MRM) requirements explicitly to AI and machine-learning credit systems).

The accuracy evidence answers the question: does the model predict credit outcomes better than the current process? This is necessary but not sufficient. A model that predicts defaults with 85 percent accuracy while generating adverse action rate ratios above 1.4 for protected classes (the ratio of the adverse action rate for a protected class to the rate for the non-protected comparison group) is not deployable under ECOA regardless of its accuracy. The accuracy evidence must be accompanied by four additional bodies of evidence for the CRO's signature to be defensible.

The adverse-action explanation evidence answers the question: can every applicant who receives an adverse action from this model receive a specific, accurate explanation of the principal reasons for the decision, as required by Regulation B? The adverse-action notice requirement is not satisfied by a generic statement that the application did not meet credit standards. Reg B requires the specific factors that contributed most to the adverse decision, and those factors must be accurate representations of the model's actual decision reasoning. A pilot that does not verify adverse-action explanation accuracy in a sample of adverse decisions has not answered the question an examiner will ask first.

The fair-lending evidence answers the question: does the model produce disparate impact on any protected class, and if so, has the institution completed a less-discriminatory alternative (LDA) search? The LDA search is the documented inquiry into whether any model configuration achieves comparable credit-risk performance with less adverse impact on protected classes. A pilot that produces clean accuracy statistics but has not completed the LDA search has not satisfied the ECOA business-necessity framework, because the institution cannot demonstrate it chose the most defensible available configuration.

The model governance evidence answers the question: is the model documented, validated, and governed to OCC 2026-13 standards, with a designated model owner accountable for its ongoing performance? A pilot that lacks documentation of the model's development decisions, training data, validation results, and monitoring plan is not deployable under OCC 2026-13 regardless of its performance. The model governance documentation is not a bureaucratic exercise; it is the institutional record that makes accountability traceable and defensible when a specific decision is challenged.

The operational integration evidence answers the question: does the model integrate with the institution's loan origination system (LOS, the software platform through which loan applications are processed and decisioned), comply with the institution's existing credit policy, and fit within the controls that apply to the relevant credit product? A model that performs well in a sandboxed test environment but requires credit-policy exceptions or creates gaps in the existing control framework is not ready for production deployment, regardless of its pilot results.

A pilot designed to impress produces accuracy statistics. A pilot designed to satisfy a CRO produces evidence on all five dimensions simultaneously, because the discipline required for the second is what makes the first credible.

Designing the Pilot for Evidence, Not Performance

Pilot design determines what evidence the pilot can produce. A pilot designed primarily to demonstrate model performance will produce performance evidence. A pilot designed to produce all five bodies of evidence requires different design choices at the start, not appended at the end. These design choices include scope definition, sample construction, data infrastructure, control architecture, and exit criteria.

Scope definition: what the pilot will and will not answer. Every pilot has a defined scope: the product type, the applicant population, the time window, and the AI-assisted components of the decision process that are being tested. The scope definition should explicitly state which of the five evidentiary requirements the pilot is designed to address and which require supplementation from other sources. A pilot on a single mortgage product in a single metropolitan market can produce evidence on accuracy, fair-lending outcomes, and adverse-action explanation accuracy for that product and market. It cannot automatically answer the fair-lending question for a different product or a different demographic market. Scope limitations must be documented and must be reflected in the scale decision: a pilot that produced clean fair-lending results in one market does not justify enterprise-wide deployment without additional market-specific testing.

Sample construction: who is in the pilot and why. The applicant population in a lending AI pilot must be representative of the population the model will serve in production, including its protected-class composition. A pilot that processes only the top half of the credit-quality distribution will produce clean fair-lending statistics because the highest-quality applicants are the least likely to produce disparate impact, and those results will not generalize to the full production population where disparate impact is most likely to appear. The pilot sample should mirror the expected production population's credit-tier distribution, protected-class composition (as estimated by BISG, the Bayesian Improved Surname Geocoding method used for proxy estimation), geographic distribution, and product-type distribution. Deviation from the production population's characteristics must be documented and must be reflected in the limitation section of the pilot report.

Data infrastructure: what data the pilot needs and where it comes from. A pilot that processes applications with incomplete or inconsistently formatted data will produce results that do not replicate in production, where data quality problems are more variable. The pilot data infrastructure should mirror the production data environment: the same data sources, the same preprocessing pipeline, the same integration with the LOS. A pilot that uses pre-cleaned, manually curated data to avoid the complications of the production data pipeline will produce performance numbers that overstate production performance. The data infrastructure decision is not a technical detail; it is a design choice that determines whether the pilot evidence is transferable to the production environment.

Control architecture: how the AI model interacts with human decision-makers. OCC Bulletin 2026-13 requires that the institution maintain human accountability for credit decisions. In practical terms, this means the pilot must test not just the model's output accuracy but the human-in-the-loop structure that will govern production deployment. If the production design requires an underwriter to review the model's recommendation before a final adverse action is issued, the pilot must test that review process, measure how often underwriters override the model and why, and assess whether the override pattern is consistent across demographic subgroups. A pilot that tests the model in isolation from the human decision-making structure it will operate within does not produce evidence about the production system; it produces evidence about the model in a test environment that will not persist in production.

Exit criteria: what the pilot must prove to recommend scale. Exit criteria are the pre-specified evidence thresholds that the pilot must meet for the institution to proceed to a scale decision. Defining exit criteria before the pilot begins prevents the organizational momentum that builds around a performing pilot from overriding the governance requirements. Exit criteria for a lending AI pilot should include: minimum accuracy thresholds (AUROC, Gini coefficient, or the institution's primary performance metric); maximum adverse action rate ratio by protected class (the commonly used threshold is 1.25, corresponding to the 80 percent rule); adverse-action explanation accuracy standard (for a sample of adverse decisions, the percentage of explanations that accurately reflect the model's decision reasoning); LDA search completion (a documented LDA search with at least two alternative configurations tested and a conclusion documented); and model documentation completeness (all documentation required by OCC 2026-13 completed and available for independent validation review).

Defining exit criteria in writing before the pilot begins creates a governance artifact that documents the institution's pre-specified risk tolerance. When an examiner asks how the institution decided the pilot had produced sufficient evidence to proceed to scale, the answer is the exit criteria document, not a retrospective judgment. This pre-specification is a direct parallel to the pre-specified fair-lending gate framework developed in the L4 lessons, applied at the pilot stage rather than the proof-of-concept stage.

The Fair-Lending Measurement Protocol for Pilots

The fair-lending measurement component of a lending AI pilot requires more structure than a post-hoc disparity review. It requires a pre-specified measurement protocol that defines the statistical tests to be used, the demographic data sources to be used, the disparity thresholds that trigger deeper investigation, and the LDA search process that will be conducted regardless of whether the disparity test produces a material finding.

The demographic data foundation begins with available self-reported data for credit products that require HMDA (Home Mortgage Disclosure Act) reporting (primarily mortgage applications), supplemented by BISG proxy estimates for the protected-class attributes not captured in the application data. For non-HMDA products (consumer installment loans, personal lines of credit, small business loans under the HMDA reporting threshold), BISG is the primary demographic data source. The pilot protocol must specify which protected classes will be tested, which demographic data source will be used for each class, and how the probability weights from BISG will be incorporated into the statistical analysis (typically through weighted regression rather than a binary classification of applicants into protected and non-protected groups).

The disparity measurement component uses two complementary statistical approaches. The first is the univariate adverse action rate ratio: the ratio of the adverse action rate for a protected class to the rate for the control group (non-Hispanic white applicants for race and national origin analysis). An adverse action rate ratio above 1.25 (the 80 percent rule threshold) triggers a deeper analysis. The second approach is the multivariate regression analysis, which controls for legitimate credit-risk characteristics and measures the residual disparity attributable to protected-class membership after controlling for creditworthiness. The multivariate analysis distinguishes between disparity driven by credit-risk differences across groups (which may be legally defensible through business necessity) and disparity that persists after controlling for creditworthiness (which is less defensible and requires a documented LDA search).

The pilot's fair-lending measurement protocol must specify, in advance, what sample size is required to achieve statistical power adequate to detect a material disparity. A pilot with 200 adverse actions against 15 Hispanic applicants cannot detect a statistically significant disparity at a reasonable significance threshold. The minimum sample size for fair-lending analysis depends on the institution's demographic market composition and the product's expected adverse action rate, but a pilot that processes fewer than 500 adverse-action outcomes total is unlikely to produce statistically reliable fair-lending results for any protected class. This sample-size requirement has direct implications for pilot scope: a pilot too small to produce statistically reliable fair-lending evidence is not sufficient to support a scale decision, regardless of its accuracy statistics.

When the disparity analysis produces a material finding, the pilot protocol must specify how the LDA search will proceed. The LDA search during a pilot is structured differently from an LDA search conducted at production deployment. In a pilot, the institution has the opportunity to modify the model's feature set or decision logic and re-test on the pilot data before committing to a production configuration. This is the most cost-effective point in the development lifecycle to conduct the LDA search, because modifying the model before production deployment is less expensive than modifying it after deployment, and less risky than continuing to run a model with a documented disparity finding. Institutions that complete the LDA search during the pilot stage and document the results produce a governance record that demonstrates good-faith compliance before any production deployment begins.

The Pilot Report Structure That Satisfies Governance

The pilot report is the governance artifact that converts pilot activity into production authorization. A report that addresses only accuracy and throughput leaves four of the five evidentiary requirements unanswered and cannot support a CRO signature on a scale decision. A report structured to address all five requirements produces the authorization the institution needs and the documentation the examiner will expect.

The pilot report structure that satisfies governance has seven sections. The first section is the executive summary: a two-page summary of the pilot scope, exit criteria, results, and recommendation. This section is written for the CRO and board committee members who will approve the scale decision and must convey the key findings on all five evidentiary dimensions without requiring the reader to work through the technical appendices. The recommendation in this section should state explicitly whether the pilot met all exit criteria or whether conditional approval is being requested, with specific conditions attached.

The second section is the pilot design documentation: scope, sample construction, data infrastructure, and control architecture, as defined before the pilot began. This section demonstrates to an examiner that the pilot was designed systematically rather than opportunistically, and that the exit criteria were specified before results were known. Retrospective pilot design documentation (documentation assembled after the results are known) is visible to experienced examiners and damages institutional credibility.

The third section is the accuracy and operational evidence: performance metrics against exit criteria thresholds, throughput statistics, decision-time reduction, and operational integration assessment. This is the section that will receive the most attention from the business-line sponsor and the innovation team, and it should be written to accurately convey both the strengths and the limitations of the accuracy evidence (including any sample composition differences from the expected production population).

The fourth section is the fair-lending evidence: the complete disparity analysis, by protected class, using the pre-specified measurement protocol. This section must include the adverse action rate ratios and statistical significance assessments, the multivariate regression residual disparity analysis, a description of any material findings, and the LDA search results for any model configurations tested in response to material findings. A clean fair-lending result (no material disparities identified) requires a section that is complete and unambiguous: "no material disparity was identified for any protected class at the pilot's sample size" is a clear finding. "The fair-lending analysis did not identify concerns" is ambiguous and will invite examiner questions about the methodology.

The fifth section is the adverse-action explanation audit: an assessment of explanation accuracy for a random sample of adverse decisions generated during the pilot. The audit protocol should specify the sample size (commonly 100 to 200 adverse actions), the testing methodology (having a human reviewer assess whether the stated reasons accurately reflect the model's decision factors for each sampled decision), and the accuracy rate achieved. An explanation accuracy rate below 95 percent should trigger remediation before production deployment, because inaccurate adverse-action explanations are a direct Reg B violation at production volume.

The sixth section is the model governance documentation checklist: a structured review of the documentation required by OCC 2026-13, confirming that each required document is complete, that a model owner has been designated, and that the monitoring plan for production is specified. Missing or incomplete documentation items are identified as conditions for production authorization rather than deferrals.

The seventh section is the scale recommendation: a specific recommendation on whether to proceed to full production deployment, proceed to an expanded pilot, proceed with conditions, or suspend the program pending remediation. The recommendation must address the disposition of any exit criteria that were not fully met, any limitations on the scope of production authorization (for example, limiting deployment to the markets and product types represented in the pilot until additional market-specific fair-lending testing is completed), and the monitoring milestones that will trigger a post-deployment governance review.

Escalation and Decision Authority: Who Signs What

Pilot governance requires a clear escalation and decision-authority structure that specifies who reviews which sections of the pilot report, who has authority to approve a scale decision, what conditions require board-level awareness, and how dissenting views within the review committee are documented. Without this structure, the scale decision becomes an informal consensus rather than a governed authorization, and the institution cannot demonstrate to an examiner that the appropriate oversight was applied.

The model owner is the individual accountable for the model's performance and governance throughout its lifecycle. Under OCC 2026-13, the model owner is not the data science team member who built the model but a designated senior individual in the business line or risk function who owns the model on behalf of the institution. The model owner reviews all seven sections of the pilot report, signs off on the accuracy and operational evidence, and formally requests scale authorization. The model owner is on record as having reviewed and endorsed the pilot's findings, and that accountability does not transfer when the model transitions from pilot to production.

The fair-lending compliance officer reviews the fourth and fifth sections of the report (the fair-lending evidence and the adverse-action explanation audit) and provides a written opinion on whether the fair-lending evidence is sufficient to support production deployment. If the fair-lending compliance officer identifies material concerns, those concerns must be addressed before the CRO signs the scale authorization, not deferred to post-deployment monitoring. The compliance officer's review is independent of the model owner's endorsement, and both reviews are required for the authorization package to be complete.

The chief risk officer reviews the complete pilot report and the written opinions of the model owner and fair-lending compliance officer. The CRO's signature on the scale authorization certifies that the institution has reviewed the governance documentation and that production deployment is authorized within the scope and conditions specified in the pilot report. For models that are deployed in high-volume products or that target underserved markets (where the fair-lending risk is highest and the reputational stakes are greatest), the authorization should require board audit committee awareness, with a summary of the pilot findings included in the next scheduled committee report.

Dissenting views within the review committee must be documented. An institution where the fair-lending compliance officer raises concerns that are overridden by the business-line sponsor without a documented resolution has created a governance record that is worse than having conducted no formal review: the record shows that concerns were identified and dismissed without resolution. Dissenting views should be either resolved (the concern is addressed and the resolution is documented) or escalated (the concern cannot be resolved at the review committee level and is referred to the CRO or board for a decision). An unresolved dissenting view in the authorization package is a flag that will receive examiner attention.

From Pilot to Monitored Production: The Transition That Fails Most Often

The transition from pilot to production is the point in the lending AI lifecycle where governance most often breaks down. The pilot was a controlled environment with heightened monitoring, dedicated review resources, and an active governance process. Production is an ongoing operation where the model processes applications continuously, the monitoring resources are limited, and the governance attention that characterized the pilot phase gradually dissipates under the pressure of other institutional priorities. The governance design challenge is to build a production monitoring program that maintains the surveillance intensity of the pilot phase without requiring the resource investment of a dedicated pilot team.

The production monitoring plan specified in the pilot report should include four components. The first is an accuracy monitoring schedule: monthly reporting of the model's primary performance metrics against the pilot benchmarks, with pre-specified thresholds that trigger a governance review if performance degrades. Performance degradation in a credit model often precedes fair-lending deterioration, because the same population shifts that reduce model accuracy also change the model's behavior across demographic subgroups.

The second component is the fair-lending monitoring schedule: quarterly disparity reporting using the same measurement methodology as the pilot's fair-lending analysis. The quarterly cadence is the minimum defensible frequency for a production credit model under OCC 2026-13; monthly fair-lending monitoring is preferable for high-volume products where disparity can accumulate quickly. The monitoring report should use the same statistical tests, the same demographic data sources, and the same disparity thresholds as the pilot, so that production results are directly comparable to pilot benchmarks.

The third component is the adverse-action explanation quality control: a monthly sample audit of adverse-action notices generated in production, with the same accuracy methodology as the pilot's explanation audit. Explanation accuracy can degrade over time if the model is retrained or if the downstream explanation generation process is modified. Ongoing sampling maintains the evidentiary basis for the institution's Reg B compliance.

The fourth component is the change governance trigger: a pre-specified list of the model changes that require a governance review before implementation. This list should include any change to the model's feature set, any change to the model's thresholds or decision logic, any retraining of the model on updated data, and any change to the explanation generation process. The change governance trigger ensures that the model version that was validated in the pilot remains the model version in production, or that changes to the model are separately validated before deployment.

When production monitoring produces a material finding (an adverse action rate ratio above the exit criteria threshold, an explanation accuracy rate below the standard, or a performance metric below the benchmark), the institution's response protocol must be pre-specified. The response protocol should define who is notified, what investigation is initiated, what timeline applies to the investigation, and what authorization is required to continue operating the model while the investigation proceeds. An institution that does not have a pre-specified response protocol for a material monitoring finding will spend time in an emergency governance process rather than in a structured investigation, and the gap between the finding and the response will be visible to an examiner who reviews the monitoring records.

Key Takeaways

  • A CRO signature on a scale decision certifies that the institution has addressed five evidentiary requirements: accuracy, adverse-action explanation accuracy, fair-lending (disparate impact and LDA), model governance under OCC 2026-13, and operational integration; a pilot that addresses only accuracy satisfies one of the five.
  • Pilot design determines what evidence the pilot can produce: scope, sample construction, data infrastructure, control architecture, and exit criteria must be designed before the pilot begins to ensure all five evidentiary requirements can be addressed within the pilot window.
  • The fair-lending measurement protocol must be pre-specified before the pilot begins, including the demographic data sources (BISG proxy for non-HMDA products), statistical tests, disparity thresholds that trigger deeper analysis, and the LDA search process that will proceed regardless of whether the disparity test produces a material finding.
  • The minimum sample size for statistically reliable fair-lending results typically requires at least 500 adverse-action outcomes; a pilot too small to produce statistically reliable fair-lending evidence cannot support a scale decision regardless of its accuracy statistics.
  • The pilot report must address all seven sections (executive summary, design documentation, accuracy evidence, fair-lending evidence, explanation audit, governance documentation checklist, and scale recommendation) to support a CRO authorization; missing sections cannot be deferred to post-deployment documentation.
  • Dissenting views within the pilot review committee must be documented and either resolved or escalated; an unresolved dissenting view in the authorization package creates a governance record that is worse than no formal review and will receive examiner attention.
  • The transition from pilot to production requires a monitoring plan that replicates the pilot's surveillance intensity through quarterly fair-lending reporting, monthly explanation accuracy auditing, and pre-specified change governance triggers, because the governance attention of the pilot phase will not sustain itself without a structured ongoing program.