Identifying Novel, Defensible Use Cases
The chief innovation officer at a $12 billion regional bank walked into the quarterly AI steering committee with a one-page brief: a generative AI system that would analyze applicants' social media histories, public court records, and gig-economy platform ratings to construct a "behavioral creditworthiness score" to supplement traditional FICO. The brief cited a fintech competitor that had publicized similar capabilities and claimed 18 percent better default prediction on thin-file applicants. The room was quiet for a moment before the chief risk officer set down his coffee and asked two questions: "Which features correlate with race or national origin in this scoring system?" and "Have we run a less-discriminatory alternative search on it?" The innovation officer had no answers. The proposal never left the room. The exchange is instructive not because the idea was reckless but because the innovation officer had built a business case on upside without building a defense against the downside. In lending AI, that sequencing error does not produce a failed project -- it produces a consent order. This lesson teaches the discipline of identifying novel AI use cases in lending that are genuinely worth pursuing: cases where the upside is real, the compliance posture is defensible, and the institution can hold both positions in the same room.
The Innovation Asymmetry in Lending AI
Lending AI operates under a structural asymmetry that most technology sectors do not face. In e-commerce, an AI recommendation engine that performs poorly simply produces less revenue. In lending, an AI model that performs poorly on protected-class applicants produces a fair-lending violation, a regulatory examination, and potentially a consent order, even if the model is improving aggregate default prediction. The Equal Credit Opportunity Act (ECOA), enacted in 1974 and codified at 15 U.S.C. 1691 et seq., prohibits discrimination in credit transactions on the basis of race, color, religion, national origin, sex, marital status, age, and receipt of public assistance. Regulation B (Reg B), implemented by the CFPB at 12 CFR Part 1002, is the implementing regulation that translates ECOA into specific lender obligations, including the adverse-action notice requirement and the prohibition on using facially neutral practices that produce disparate impact on protected classes without business necessity.
Disparate impact under ECOA and Reg B means that a facially neutral lending practice can be unlawful if it produces statistically significant adverse outcomes for a protected class and no less-discriminatory alternative exists that serves the same credit-risk objective. The institution does not need to intend discrimination. The model does not need to contain a variable labeled "race." It needs only to produce worse outcomes for a protected class in a way the institution cannot justify through documented business necessity combined with a good-faith search for alternatives. This legal structure makes the innovation calculus in lending AI fundamentally different from the innovation calculus in other AI applications: the downside of a novel use case that imports proxy-variable risk is not a product failure but a regulatory enforcement action.
OCC Bulletin 2026-13, issued in April 2026 jointly by the Office of the Comptroller of the Currency (OCC), the Federal Reserve, and the Federal Deposit Insurance Corporation (FDIC), reinforces this calculus at the governance level. The bulletin explicitly extends model risk management (MRM) requirements to artificial intelligence (AI) and machine-learning systems used in credit decisions. MRM is the discipline of validating, monitoring, and documenting models to ensure they perform as intended within acceptable risk limits. Under OCC 2026-13, an institution that deploys a novel AI model in credit decisioning without completing an MRM-compliant validation, including a fair-lending analysis, has violated its own governance obligations regardless of whether the model produces fair-lending violations. The governance failure is a separate risk from the fair-lending risk, and an examiner who finds it will expand the scope of the examination accordingly.
The innovation asymmetry, then, has two dimensions. The upside of a novel lending AI use case is measured in default prediction improvement, cost-to-originate reduction, and reach into underserved markets. The downside is measured in consent orders, remediation costs that in documented large-institution enforcement cases have reached or exceeded $50 million, and reputational damage that compounds with each enforcement headline. The discipline of identifying novel, defensible use cases is the discipline of finding opportunities where the upside is real and the downside is contained before the institution commits resources to development or deployment.
The question is not whether a novel AI application is innovative. The question is whether it is defensible -- and defensibility is constructed before deployment, not assembled under examination pressure.
What Makes a Use Case Novel and What Makes It Defensible
Not every new lending AI application is genuinely novel in the sense that matters for this lesson. Using a large language model (LLM) to generate a first-draft adverse-action notice from a structured data file is novel in the technology sense but not novel in the credit-risk sense: it is an automation of a compliance workflow, not a new input to credit decisioning. The distinction matters because the fair-lending risk profile of a novel use case depends primarily on whether the use case introduces new data inputs or new decision logic into the credit process, not on whether the technology is new.
A use case is novel in the relevant sense when it does one or more of the following: introduces a new data source (alternative data, behavioral data, transaction data, social media signals) that was not previously used in the institution's credit decisioning; applies a new model architecture (LLM-based scoring, neural network ensembles, reinforcement learning for pricing) to a credit decision that was previously made using traditional statistical models; extends AI-assisted decisioning to a new stage of the credit lifecycle (servicing, early intervention, collections) where AI has not previously been used; or uses AI to segment applicants in a way that changes who receives which credit offer or term. Each of these forms of novelty introduces a new proxy-variable risk surface: new data inputs may correlate with protected characteristics; new model architectures may encode that correlation in ways that are harder to detect than traditional scorecard models; new decisioning stages may import disparate-impact risk into credit functions where it was not previously present.
A use case is defensible when the institution can satisfy four conditions simultaneously. First, the institution can identify every data input and explain why it is a legitimate, job-related credit-risk factor under ECOA's business-necessity framework. Second, the institution can demonstrate that each input has been tested for correlation with protected-class characteristics and that the test results support deployment rather than requiring further investigation. Third, the institution can produce a completed less-discriminatory alternative (LDA) search that evaluated candidate alternative configurations and documented why the chosen configuration represents the best available balance of credit-risk performance and disparity outcomes. Fourth, the institution can show that the use case has been through a completed MRM validation under OCC 2026-13 that includes a fair-lending component, signed off by a model owner who is accountable for the model's ongoing performance.
The four-condition test is not an abstract standard. It is the four-part answer a bank needs to give when an examiner opens the conversation about a novel AI credit application. Institutions that build use cases from the top of this framework outward -- starting with what the examiner will ask rather than starting with the model's performance metrics -- consistently produce more defensible applications than institutions that build use cases from the technology inward and add compliance review at the end. The architectural difference between these two approaches shows up in exam results, consent order tallies, and the institution's ability to continue deploying AI after the first examination cycle.
The Use Case Taxonomy: Where the Genuine Upside Lives
Mapping the lending AI use case landscape along two axes (credit-risk improvement potential and fair-lending risk exposure) produces four quadrants. The quadrant with high credit-risk improvement potential and manageable fair-lending risk is where the genuine innovation upside lives. Understanding what belongs in each quadrant helps leadership teams allocate AI development resources toward opportunities that produce sustainable advantage rather than temporary improvements followed by regulatory correction.
High upside, manageable risk: the priority tier. The use cases with the best combination of genuine innovation potential and defensible compliance posture typically share a common structure: they use AI to improve the accuracy, speed, or cost of decisions that the institution was already making using existing data, without introducing new proxy-variable risk through novel data sources. Examples include:
Income verification acceleration using AI to extract, validate, and reconcile income data from tax documents, pay stubs, and bank statements. The data sources are established, the variables are defined by credit policy, and the AI is improving the processing accuracy and speed of information the institution was already using. The fair-lending risk is low because the model is not introducing new proxies; it is improving the extraction accuracy of existing approved inputs. In illustrative deployments of this type, institutions have reported cost-to-originate reductions on mortgage applications of roughly 20 percent or more, with meaningful reductions in income-verification defect rates; these figures are composites and individual results vary by institution size, data quality, and implementation scope.
Early delinquency prediction models using transaction data already held by the bank. An institution that holds a borrower's deposit account and loan account has behavioral payment data that is a legitimate, job-related predictor of repayment risk: account balance volatility, missed minimum payment patterns, and overdraft frequency are directly related to credit performance. When the AI model is trained on this data within the bank's existing customer base and tested for fair-lending outcomes before deployment in servicing or early-intervention workflows, the use case is both genuinely predictive and defensible because the data sources are credit-behavior data, not demographic proxies.
Document intelligence for commercial loan underwriting. AI-assisted extraction and analysis of financial statements, covenant compliance data, and borrower-provided projections in commercial and industrial (C and I) lending does not involve consumer protected-class characteristics in the same way consumer mortgage decisions do. The fair-lending risk surface is narrower, the credit-risk improvement is substantial (commercial underwriters report 40 to 60 percent reductions in time-to-decision on complex credits), and the governance framework is straightforward to construct.
High upside, high risk: the caution tier. Some use cases promise genuine credit-risk improvement but import a significant fair-lending risk surface. These are not necessarily prohibited, but they require more rigorous pre-deployment governance work than the priority tier. Examples include thin-file applicant scoring using alternative data sources such as rent payment history, utility payment data, or telecom payment history. The CFPB and OCC have acknowledged that these data sources can expand credit access for underserved applicants, but the same sources correlate with neighborhood characteristics that can function as proxies for race and national origin. A use case using rental payment data requires a completed proxy-variable analysis on the specific data source in the institution's specific market before deployment, not a general-purpose assertion that the data type is credit-relevant.
Cash flow underwriting for small business lending falls in the same tier. Transaction data from a small business's bank account is a genuinely superior predictor of repayment capacity compared to traditional financial statement analysis, particularly for early-stage and informal businesses. The innovation upside is real. But cash flow patterns correlate with industry concentration in specific geographic areas, and industry concentration correlates with the demographic composition of small business ownership in those areas. The institution deploying cash flow underwriting must test for these correlations in its own market data and document the results before deployment.
Low upside, low risk: the efficiency tier. Many AI applications in lending offer modest innovation upside but minimal fair-lending risk because they do not touch credit decisioning. Document routing, compliance checklist automation, regulatory change monitoring, and adverse-action notice generation from structured decisioning data all belong in this tier. These applications deserve investment on cost-reduction grounds, but they do not belong in the innovation pipeline that leadership teams should be using to differentiate the institution's credit capabilities.
Low upside, high risk: the avoid tier. The social media scoring proposal that opened this lesson belongs here. Any use case that uses behavioral data from social platforms, public records of non-credit-related behavior, or consumer data aggregated without direct credit-behavior relevance falls into this quadrant. The credit-risk improvement claims for these data sources are empirically weak (the published studies supporting them typically use small samples and non-representative populations), and the proxy-variable risk is high because social media behavior, neighborhood activity patterns, and consumer purchase data correlate strongly with race, national origin, and other protected characteristics. The combination of weak credit-risk evidence and strong proxy risk makes these use cases indefensible under the ECOA business-necessity framework even if they pass an initial fair-lending screen, because the institution cannot sustain the business-necessity argument when the credit-risk evidence is thin.
The Proxy Variable Screening Process
Every novel data input considered for a lending AI use case requires a proxy variable screen before the development investment is made. Proxy variable screening is the process of testing whether a candidate feature correlates with protected-class characteristics at a level that would produce disparate impact when the feature is used in credit decisioning. The screen is not a one-time approval; it is a dataset-specific and market-specific analysis that must be repeated when the data source changes, when the model is retrained, or when the institution expands the use case to a new geographic market or product type.
The proxy screening process for a novel data input has four components. The first component is correlation testing: computing the statistical association between the candidate feature and BISG-estimated protected-class probabilities in the institution's own applicant population. BISG (Bayesian Improved Surname Geocoding) is the most widely accepted method for estimating applicant race and ethnicity when self-reported demographic data is not available, using applicant surname and census-tract demographics to produce probability estimates. A correlation above a threshold defined in the institution's fair-lending policy (commonly an absolute Pearson correlation of 0.25 or higher, but the threshold should be set by the institution's model-risk committee rather than adopted from a generic standard) flags the feature for deeper analysis.
The second component is disparate impact simulation: using the candidate feature in a test version of the model trained on historical data and measuring the adverse action rate ratio (the ratio of adverse action rates for a protected class to the adverse action rate for the control group) that results. An adverse action rate ratio above 1.25 (the 80 percent rule, where the protected class receives adverse action at a rate more than 25 percent higher than the control group) is the standard regulatory threshold that triggers a deeper examination. The simulation produces a pre-deployment estimate of the disparate impact the feature would contribute if deployed, allowing the institution to evaluate whether the credit-risk improvement justifies the disparity level and whether the LDA search is likely to find an alternative with comparable performance.
The third component is the marginal contribution analysis: using SHAP values (SHapley Additive exPlanations, a method for computing each feature's contribution to a model's output, derived from cooperative game theory) or permutation importance to measure how much the candidate feature contributes to the model's overall disparate impact, controlling for the other features in the model. A feature that has a high standalone correlation with protected-class characteristics may contribute relatively little to the model's overall disparate impact if the model already captures most of the credit-risk signal from safer features. Conversely, a feature with a modest standalone correlation may produce disproportionate disparate impact if it is the most predictive feature in the model and the model would lose substantial accuracy without it. The marginal contribution analysis identifies which features are the highest-priority targets for LDA substitution and which features can be retained with lower concern.
The fourth component is the documentation of conclusions: a written assessment, reviewed and approved by the model owner and the fair-lending compliance officer, that states whether the candidate feature meets the institution's proxy-variable standard for deployment, requires an LDA search before deployment, or is disqualified. This document becomes part of the model's MRM file and is available for examiner review. Institutions that conduct proxy screening informally, without contemporaneous documentation, cannot demonstrate to an examiner that the screening occurred or that it was conducted in good faith.
Building the Use Case Defense from the Ground Up
The most durable way to construct a defensible novel use case is to build the governance documentation into the development process rather than appending it at the validation stage. This approach, which mirrors the "fair lending gate" framework developed in the L4 lessons, treats the compliance posture as an architectural decision rather than a review step. When the compliance posture is an architectural decision, the institution makes it explicitly at the beginning of the development process, when changing direction is cheap. When it is a review step, the institution discovers compliance problems after the development investment has been made, when changing direction is expensive and the organizational momentum for deployment is already established.
The ground-up approach has five stages that parallel the development lifecycle. In stage one, concept definition, the innovation or product team defines the use case in terms of the credit-risk problem it solves, the data inputs it requires, and the credit decision it will influence. At this stage, the fair-lending officer reviews the concept brief for proxy-variable risks before any development resources are committed. A one-page concept brief that takes half a day to write and review prevents six months of development work from being directed at an indefensible target.
In stage two, data assessment, the data science team conducts the proxy variable screen on the specific dataset proposed for the use case. The screen produces a written assessment that categorizes each candidate feature, identifies any features that require LDA substitution, and estimates the expected disparate-impact level of the proposed model configuration. This assessment is the foundation of the use case's fair-lending posture and should be completed before model development begins.
In stage three, model development, the data science team builds the model to the specification approved after stage two, including any feature substitutions required by the proxy-variable screen. The development process includes a designated model owner (an individual accountable for the model's performance and compliance throughout its lifecycle) and maintains version control documentation that records each feature selection decision and its rationale. This documentation will be required during the MRM validation and in any subsequent examination.
In stage four, the LDA search, the team tests at least two to three alternative model configurations on the development dataset, measuring both credit-risk performance (AUROC, Gini coefficient, or the institution's primary accuracy metric) and fair-lending outcomes (adverse action rate ratio by protected class). The LDA search produces a structured comparison table and a documented conclusion: either the primary configuration represents the best available balance of performance and fairness, or a superior alternative configuration should replace it. If the LDA search identifies a less-discriminatory alternative with comparable credit-risk performance, the institution deploys the alternative rather than the primary configuration. This conclusion is not a failure; it is the LDA framework working as designed.
In stage five, MRM validation, the institution's independent validation team reviews the model, the proxy-variable screen, the LDA search, and the model documentation against OCC 2026-13 requirements. The validation is completed before any production deployment, and the validation findings are incorporated into the model's ongoing monitoring plan. A model that completes stage five with a clean or conditionally clean validation has a defensible governance trail from concept to production.
Common Innovation Errors and How to Avoid Them
Certain patterns recur in lending AI innovation programs that produce regulatory problems. Understanding these patterns allows leadership teams to recognize them before the development investment is made rather than after the examination begins.
The performance benchmark error. This is the mistake of adopting a vendor's published performance claims as the justification for deploying a novel use case without verifying those claims on the institution's own data. Vendor performance benchmarks are typically derived from the vendor's training data, which may differ significantly from the institution's applicant population in geographic composition, product mix, credit tier, and demographic profile. A performance claim derived from a national dataset may not hold in a regional bank's predominantly rural market, or may hold for the overall population while masking underperformance on specific demographic subgroups. Every vendor performance claim must be verified on the institution's own data before it is cited in a use case approval decision.
The speed-to-market compression error. Innovation timelines in lending AI are subject to competitive pressure, and teams under competitive pressure sometimes compress the governance steps to meet a deployment deadline. The governance timeline for a novel credit AI use case (concept review, proxy-variable screen, model development, LDA search, MRM validation) typically requires four to eight months for a well-resourced team. Compressing this timeline does not eliminate the governance obligations; it defers them to the examination. Deferred governance costs more to correct under examination pressure than it does to complete before deployment, and the institutional credibility loss from a governance gap discovered during an examination compounds the direct remediation cost.
The vendor-as-compliance error. A related pattern is assuming that a vendor's fair-lending documentation satisfies the institution's own MRM and fair-lending obligations. OCC Bulletin 2026-13 is explicit on this point: the institution's obligations under model risk management and fair-lending are not transferable to a vendor. The institution must conduct its own independent validation of a vendor model, its own proxy-variable screen on the vendor's feature set, and its own LDA search on the vendor's model in the institution's market. A vendor who characterizes its own fair-lending documentation as sufficient for the institution's compliance purposes is either unaware of OCC 2026-13 or is misrepresenting what the bulletin requires.
The classification scope error. Some novel use cases start as workflow automation (the low-risk tier) and migrate into credit decision influence without a governance reassessment. An AI tool deployed initially to summarize loan files may gradually begin highlighting specific risk factors that influence underwriter decisions, effectively becoming a component of the credit decisioning process without having been validated as one. The governance framework for a document summary tool is different from the governance framework for a model that influences credit decisions, and the migration between these categories requires a fresh governance assessment. Institutions that do not have a mechanism for detecting and flagging this migration will discover it during an examination when an examiner traces an adverse action back through the AI-assisted workflow and asks what validation was completed on the AI component that influenced the decision.
The inherited bias assumption. A subtler error is assuming that a novel use case inherits the clean fair-lending record of the traditional process it replaces. A bank that has never had a fair-lending finding in its manual underwriting process may assume that an AI system replicating the underwriter's decision logic will also produce clean fair-lending results. This assumption ignores the possibility that the manual underwriting process contained unmeasured disparate impact that was below the detection threshold of the institution's traditional fair-lending testing, but which the AI system amplifies by applying the same decision logic at higher volume and with greater consistency. Deploying AI into an underwriting process without pre-deployment fair-lending testing on the AI system's outputs is not a safe replication of a clean process; it is a multiplication of an untested process.
Key Takeaways
- Lending AI innovation operates under a structural asymmetry: the downside of a novel use case that imports proxy-variable risk is a regulatory enforcement action, not a product failure, which means the compliance posture of a novel use case must be assessed before the development investment is made.
- A use case is defensible when the institution can simultaneously identify every input as a legitimate credit-risk factor, demonstrate proxy-variable testing results that support deployment, produce a completed LDA (less-discriminatory alternative) search, and show a completed MRM (model risk management) validation under OCC Bulletin 2026-13.
- The genuine innovation upside in lending AI lives primarily in use cases that improve the accuracy, speed, or cost of decisions using existing approved data sources, not in use cases that introduce novel data inputs with unverified proxy-variable risk profiles.
- Proxy variable screening is a dataset-specific, market-specific analysis that must be completed for every novel data input before development resources are committed; a correlation test, a disparate-impact simulation, and a marginal-contribution analysis together constitute a defensible screen.
- The most common innovation errors include adopting vendor performance benchmarks without market-specific verification, compressing governance timelines to meet competitive deadlines, assuming vendor compliance documentation satisfies institutional MRM and fair-lending obligations, and allowing workflow automation tools to migrate into credit decision influence without governance reassessment.
- The governance gate for novel use cases must be chaired by senior risk leadership and calibrated to risk tier: priority-tier use cases require standard documentation review, while caution-tier use cases with borderline disparity findings require board audit committee awareness.
- Accountability must stay human throughout the novel use case governance process: the individuals who authorize a novel AI credit application for production are on record as having reviewed the governance documentation, and that accountability cannot be delegated to the model, the vendor, or the validation team.
Skill.re