Proxy-Variable Detection
The model did not use race. The compliance team knew this, had documented it, and had been told it by the vendor in a certification letter. The institution was a mid-sized regional lender processing about 3,400 mortgage applications per year in markets that included both predominantly white suburban counties and majority-minority urban zip codes. In the eighteen months after the AI underwriting model went live, the approval rate in majority-minority census tracts declined by 11 percentage points compared to the prior system, while approval rates in predominantly white tracts improved by 4 points. When the CFPB (Consumer Financial Protection Bureau) examiner asked about the model's inputs, the compliance team's answer was accurate and completely beside the point: "We don't use race, national origin, or any protected characteristic." The examiner's response was brief: "Tell me about your zip code feature, your commute-pattern data, and your grocery purchase frequency variable." Three months later, the institution was in a consent order. (This scenario is a composite illustrative example; the specific figures are representative rather than drawn from a single named enforcement action.) The lesson here is not that "we didn't use race" is a lie. It is that "we didn't use race" is not a defense under the disparate-impact theory of fair lending, and that detecting the proxy variables that produce indirect discrimination requires a specific set of analytical tools applied systematically before deployment, not in response to an examiner's questions.
What Proxy Variables Are and Why They Matter
A proxy variable is a variable that is not itself a protected characteristic under ECOA (Equal Credit Opportunity Act, the federal statute prohibiting credit discrimination) or the Fair Housing Act but that correlates with a protected characteristic strongly enough to function as a statistical substitute. Using a proxy variable in a credit model can produce the same discriminatory effect as using the protected characteristic directly, through a mechanism that does not require intent to discriminate and that may not be visible without specific analytical tools designed to detect it.
The word "proxy" is borrowed from the language of measurement: in research settings, a proxy measure is an observable variable used when the quantity of actual interest cannot be measured directly. Applied to credit discrimination, the protected characteristic (race, national origin, sex, or another ECOA class) is what the proxy correlates with. When a credit model uses zip code as a feature, and zip code is highly correlated with race in the markets the model operates in, the model is effectively using race indirectly, through the correlation. It is getting information about race from zip code and using that information in its predictions, even though the model's code contains no reference to race.
This matters under the disparate-impact framework because the doctrine tests outcomes, not inputs. The question is not what variables were fed into the model but what effects those variables produce on protected-class applicants. A model that produces statistically significant adverse outcomes for Black applicants because it uses zip code as a feature, and zip code correlates strongly with race in its market, has produced disparate impact. The lender's defense that "we didn't use race" is irrelevant to the disparate-impact finding; it may, however, be relevant to a disparate treatment (intentional discrimination) claim, which is a different legal theory with a different evidentiary structure.
Why do proxy variables exist in real datasets? Largely because of the structural legacy of decades of discriminatory policy. Residential segregation by race in U.S. cities is not a naturally occurring phenomenon. It was produced by federal redlining policy (the federal government literally drew red lines on maps of cities identifying minority neighborhoods as high-risk for government-backed mortgage insurance), restrictive racial covenants in property deeds, racially discriminatory real estate practices, and explicitly discriminatory public housing policies. These policies persisted into the 1970s and their structural effects persist today: residential segregation patterns from the 1950s and 1960s are still measurable in census data. Any geographic feature, whether zip code, census tract, neighborhood designation, or commute-time calculation, in a U.S. credit model operates against this historical backdrop. It will correlate with race and national origin through the residential segregation patterns that history produced.
"We didn't use race" is a statement about inputs. Disparate impact is a doctrine about outputs. The two sentences answer different legal questions, and answering the first does not answer the second.
The Taxonomy of Proxy Variables in AI Credit Models
Understanding the full range of proxy variables that appear in AI credit models is necessary for building a proxy-detection workflow that actually catches the indirect discrimination that produces disparate impact. Proxy variables in 2026 lending AI models fall into five major categories, ranging from those that have been known and studied for decades to those that have only become possible with the expansion of alternative data in credit underwriting.
Geographic proxies. Zip code is the most thoroughly documented proxy variable in U.S. credit modeling. The correlation between zip code and race or national origin in most U.S. metropolitan markets is strong enough that models using zip code effectively receive information about protected class membership. The same applies to census tract, neighborhood-level designations, property appraisal zone assignments, school district identifiers, and any other geographic variable that captures the residential segregation patterns produced by historical discrimination. Geographic proxies are not new: they were the basis for the original redlining cases of the 1970s and 1980s. What is new in 2026 is the degree to which AI models, with their capacity to use hundreds of features simultaneously, can encode geographic proxy information through combinations of seemingly neutral variables that each carry some geographic signal.
Employer and employment proxies. Industry code, employer name, employer location, employer size category, and job title can each correlate with race and national origin through the historical concentration of certain demographic groups in particular industries and employers. This concentration reflects decades of discriminatory hiring and occupational segregation that are well-documented in labor economics. A model that assigns different credit risk scores to applicants from different employer industries, independent of income and employment stability, may be producing disparate impact through employer-industry proxies. Job title features carry a similar risk: job titles that disproportionately appear in one demographic group, for reasons related to historical occupational segregation rather than inherent credit risk, are proxy variables.
Financial behavior proxies. Transaction data, when used as a credit model feature, can produce proxy-variable disparate impact through multiple channels. The specific merchants an applicant patronizes, the types of financial services they use (payday lending, check-cashing services, community development financial institutions, credit unions with historically minority membership bases), the timing and regularity of deposits, and the categories of spending patterns encoded in bank transaction data all carry demographic information that reflects both individual financial behavior and the legacy of discriminatory exclusion from mainstream financial services. Applicants from communities with historically limited access to mainstream banking will have different transaction patterns than applicants from communities with generations of mainstream banking access, for reasons that are structural rather than predictive of credit risk.
Name and language proxies. Where any AI model in the origination process, whether the credit model itself or a document extraction or communication tool, has access to applicant names, it has access to a feature that predicts national origin and ethnicity with meaningful accuracy through naming patterns. Models trained with name-based features can learn to predict national origin from first names, surnames, and naming conventions. Language features, such as whether an application was completed in English or another language, which branch or channel was used to apply, or the word patterns in any applicant-generated text fields, can similarly serve as proxies for national origin. In the context of AI document extraction and analysis tools used alongside the credit model, this category of proxy becomes particularly important: even if the credit model itself does not use names, an upstream AI tool that processes application documents and systematically performs differently on documents from applicants with names associated with particular national origins can introduce proxy-variable bias into the overall decision process.
Credit history behavior proxies. The specific patterns of credit history that appear in credit bureau data can encode proxy information. Which types of credit products an applicant has historically used (subprime products, bank products, credit union products), which geographic markets their credit history was established in, and the timing patterns of their credit-building reflect the history of differential access to credit that decades of discrimination produced. Models that treat "credit history from subprime products" as categorically different from "credit history from prime products," independent of the actual delinquency patterns in each, may be encoding a proxy for the historical exclusion of certain demographic groups from prime credit markets.
Methods for Detecting Proxy Variables
Proxy-variable detection is not a single test but a suite of analytical approaches, each designed to reveal a different aspect of the relationship between a model's features and the protected characteristics those features may be proxying. A complete proxy detection workflow uses multiple methods because no single method detects all forms of proxy relationship, and some forms of proxy encoding are only visible when multiple methods are applied in combination.
SHAP value analysis for disparate outcome attribution. SHAP values (SHapley Additive exPlanations) decompose each model prediction into additive contributions from each feature. When SHAP analysis is applied across protected-class and benchmark applicant populations, it reveals which features contribute most to the adverse predictions for protected-class applicants. A feature that has high negative SHAP magnitude (contributing toward denial) in the protected-class population at a rate materially higher than in the benchmark population is a candidate proxy variable: it is having a disproportionate adverse effect on protected-class applicants.
The SHAP analysis must be examined at the population level, not just for individual predictions. The relevant question is not whether a specific applicant's denial was influenced by a particular feature but whether the distribution of feature contributions differs systematically between protected-class and benchmark populations. If zip code has a mean SHAP value of minus 0.12 for Black applicants in the model's population and plus 0.03 for white applicants, zip code is contributing adversely to Black applicants' scores at a rate materially above what is affecting white applicants. This is a proxy-variable signal.
Conditional independence testing. A proxy variable, by definition, provides information about protected class membership that goes beyond what legitimate credit factors already convey. Conditional independence testing directly measures whether a feature contains information about protected class membership after controlling for legitimate credit factors (credit score, DTI, LTV, employment history). If a feature is conditionally independent of protected class after controlling for those factors, it is not a proxy variable under the statistical definition. If a feature remains conditionally dependent on protected class after those controls, it is a proxy variable, regardless of how it is named or what credit rationale was offered for including it.
Formally, this test typically takes the form of a regression in which the protected class variable is the outcome, the legitimate credit factors are the controls, and the candidate feature is the variable of interest. If the candidate feature has a statistically significant coefficient in this regression, it provides information about protected class membership beyond what the legitimate credit factors already capture: it is a proxy variable. The test can be run for each candidate feature and also for combinations of features, which can reveal proxy encoding that is distributed across multiple features and invisible in single-feature tests.
Permutation importance with demographic stratification. Permutation importance is a model agnostic method that measures how much a model's prediction accuracy (or, in this context, disparity) changes when a specific feature's values are randomly scrambled within the prediction population. In the standard application, permuting a feature that is important for model accuracy causes a large accuracy drop. In the proxy-detection application, the key test is permuting a feature within the protected-class applicant population and measuring how much the model's disparity changes: does the model's adverse action rate ratio for the protected class decrease when the feature is randomized? If it does, the feature is contributing to the disparity and is a candidate proxy variable.
The demographic-stratification element is critical: the permutation must be applied within demographic subgroups, not across the full population, to isolate the effect of the feature on protected-class applicants specifically rather than measuring its general predictive contribution. A feature can be important for model accuracy in the overall population while also being a proxy variable that contributes disproportionately to adverse outcomes for a specific protected class. The overall permutation importance test misses this nuance; the demographically stratified version captures it.
Marginal effects analysis across geographic units. For geographic proxy variables specifically, a useful supplementary test is examining the model's marginal effect on application decisions as a function of the geographic unit's demographic composition. If the model systematically produces higher denial rates as the minority share of a census tract increases, controlling for the applicant-level credit factors, the model is producing a geographic disparity that may reflect proxy-variable encoding of race through geographic features. This test requires geocoded application data (matching applications to their census tract of the property or applicant residence) and census demographic data for each tract. The regression tests whether census-tract minority share predicts denial rates independently of applicant credit factors; if it does, the geographic feature set is functioning as a race proxy at the market level.
Intersectional proxy testing. AI models can produce proxy-variable disparate impact through feature interactions that are invisible to single-feature proxy tests. An intersection of employer industry and geographic location, neither of which is a significant proxy individually, may together strongly predict race or national origin in a particular market. Intersectional proxy testing examines pairs and triplets of features for joint conditional dependence on protected class membership. This testing is computationally more intensive than single-feature tests but is increasingly important as AI credit models grow in feature complexity: a model with 200 features has nearly 20,000 pairwise interactions to examine, and some of the most significant proxy effects may live in those interactions.
Building the Proxy Detection Workflow
Proxy detection cannot be a one-time exercise performed before deployment and then abandoned. Like disparate-impact testing, it must be a repeatable workflow with assigned ownership, a defined cadence, structured documentation, and a clear path from finding to remediation. OCC Bulletin 2026-13, the April 2026 interagency model-risk guidance, requires ongoing monitoring of AI model outcomes for fair-lending risk; proxy-variable analysis is the primary tool for understanding why those outcomes exist, which means proxy detection belongs in the ongoing model monitoring schedule alongside the outcome testing.
The proxy detection workflow has four operational phases:
Phase 1: Feature inventory and pre-classification. Before any statistical testing, the institution's model-risk or fair-lending team should conduct a manual review of the model's feature set and pre-classify each feature by its proxy-variable risk. Features in the high-risk category include any geographic feature, any employer-related feature, any financial behavior feature that reflects access to specific financial services, and any feature that captures information from text fields that could contain name or language information. Features in the medium-risk category include any feature derived from alternative data sources that have not previously been tested for demographic correlation. Features in the low-risk category are those based on direct financial performance measures (actual delinquency history, verified income, credit utilization rate) that have been used in traditional credit scoring for decades and for which extensive disparate-impact testing exists.
The pre-classification is not a substitute for statistical testing; it is a triage tool that focuses the intensive testing on the features most likely to be proxy variables. An institution with a 200-feature model that applies every proxy-detection method to every feature will exhaust its compliance resources quickly. The pre-classification focuses resources on the features most likely to require remediation.
Phase 2: Statistical proxy testing for high-risk features. Applying the suite of proxy-detection methods described above to each high-risk feature (and combinations of high-risk features), working from the pre-classification. The output is a structured results table for each tested feature showing: the SHAP disparity analysis (mean SHAP contribution for protected-class versus benchmark applicants), the conditional independence test result (coefficient estimate and statistical significance for the feature in the protected-class prediction regression), and the permutation importance test result (change in adverse action rate ratio when the feature is randomized within the protected-class population). Features that show significant proxy signals in two or more of these tests are confirmed proxy variables requiring LDA (less-discriminatory alternative) evaluation.
Phase 3: LDA referral and documentation. Every feature confirmed as a proxy variable through Phase 2 testing is referred to the LDA search process. The proxy-detection findings are the input to the LDA search: they specify which features to evaluate as candidates for removal, substitution, or modification, and they provide the quantitative evidence of proxy correlation that the LDA documentation needs to establish why the feature requires evaluation. Without the proxy-detection findings, the LDA search has no principled basis for selecting which features to evaluate. With them, the LDA search can focus on the features that statistical analysis has identified as the primary sources of disparate-impact exposure.
Phase 4: Ongoing monitoring and trigger-based re-testing. Proxy variables can emerge or intensify as the model evolves, as new features are added, as the model is retrained on new data, or as the demographic composition of the applicant population shifts. The ongoing proxy detection workflow should be scheduled on the same cadence as the disparate-impact testing (annual at minimum, quarterly for high-volume models) and should be triggered by the same events: model modification, model retraining, material population shifts, and any supervisory guidance identifying the model type or feature type as a fair-lending concern. The proxy detection results from each cycle should be documented in the same file as the disparate-impact test results, so the institution's testing history shows both the outcome-level analysis (disparity ratios) and the feature-level analysis (proxy variable findings) for each cycle.
The "We Did Not Use Race" Defense: Why It Fails
No aspect of AI fair-lending compliance is more frequently misunderstood in practice than the "we didn't use race" argument, so it is worth examining why it fails specifically, at the level of the legal doctrine, the evidentiary standard, and the practical examination context.
Under the disparate-impact theory, the legal question has two parts. First: does the challenged practice produce a statistically significant adverse effect on a protected class? Second: if so, is that effect justified by business necessity, or does a less-discriminatory alternative exist? The question of whether race was an explicit input to the model is relevant only to whether the case is a disparate-impact case or a disparate-treatment (intentional discrimination) case. Disparate impact does not require intent; it does not require that the protected characteristic appear in the model. It requires only that the outcomes the model produces are disproportionately adverse for protected classes, regardless of the mechanism through which those outcomes were produced.
The argument "we didn't use race" is actually a partial defense only to disparate treatment, not to disparate impact. A lender accused of intentional discrimination might successfully defend by showing that the protected characteristic was not used in the decision. A lender producing a 1.8 adverse action rate ratio for Black applicants through a model that uses zip code as a proxy for race cannot defend the ratio by pointing to the absence of race in the input list. The ratio exists. The model produced it. The mechanism (zip code as a proxy for race) is the subject of the LDA analysis, not a defense to the outcome.
In practice, the "we didn't use race" argument often reflects a confusion between the credit model's input variables and the model's effective behavior. A model that uses a feature highly correlated with race is effectively using race in a statistical sense: the model's predictions contain information derived from race, mediated through the proxy feature. The legal standard for disparate impact does not care how the information about protected class membership entered the model's predictions, only that it affected those predictions in a way that produced disparate outcomes.
The more sophisticated version of the "we didn't use race" argument is the business-necessity defense: the features that correlate with protected class were included because they are predictive of credit risk, not because they correlate with race. This argument is legally available and can succeed if: (1) the accuracy contribution of the feature is specific and documented; and (2) no less-discriminatory alternative exists that achieves comparable accuracy. This is the business necessity and LDA analysis described in the previous lesson. It is a strong argument when it is supported by specific, quantified evidence. It is a weak argument when it relies on general assertions that "all our features are credit-relevant," because the proxy-detection workflow is designed to show that credit-relevant features can also be race-correlated, and the legal question is whether the correlated feature is necessary for the credit-risk objective or whether a less race-correlated alternative would serve the same purpose.
Practical Examination Preparation: What Examiners Look For
Preparing for the examiner's questions about proxy variables requires building the right documentation as part of the ongoing workflow, not assembling it in response to an examination notice. Based on the examination approach that OCC Bulletin 2026-13's integrated model-risk and fair-lending framework produces, examiners reviewing an institution's proxy-variable governance are typically looking for the following elements:
A feature inventory with proxy-risk classification. The institution should be able to produce, from its model documentation, a list of all features used in the AI credit model with each feature's proxy-risk classification and the basis for that classification. This document shows that the institution has thought systematically about its feature set from a proxy-variable perspective, not just from a credit-risk accuracy perspective. A model with 200 features and no proxy-risk classification is a model whose compliance governance has not applied the proxy-variable lens to the feature selection decision.
Statistical proxy-testing results for high-risk features. For each feature classified as high-risk, the institution should have structured testing results showing the SHAP disparity analysis, the conditional independence test result, and the permutation importance result. These results, presented in a structured table, allow the examiner to verify that the institution identified potential proxy variables through rigorous statistical testing and addressed them appropriately in the LDA process.
LDA documentation for confirmed proxy variables. Every feature confirmed as a proxy variable through statistical testing should have a corresponding LDA entry documenting the alternatives evaluated and the business-necessity analysis for any retained proxy. The absence of an LDA entry for a confirmed proxy variable is a specific finding: the institution identified a proxy variable and failed to evaluate whether a less-discriminatory alternative existed.
Ongoing monitoring records. Proxy-detection testing results from each monitoring cycle, showing that the institution has maintained ongoing surveillance of its model's proxy-variable exposure rather than conducting a one-time pre-deployment analysis. The examiner will look for testing across multiple cycles, evidence that the results were reviewed and acted upon, and documentation of any proxies that were remediated in response to testing findings.
An explanation of how non-credit-model AI components were evaluated. In 2026, AI is not limited to the credit scoring model. Institutions use AI tools for document extraction, income verification, communication drafting, pre-screening, and servicing. Each of these AI components can introduce proxy-variable bias into the overall decision process, independent of the credit model itself. The institution's proxy-variable governance should address the full AI pipeline, not just the primary credit model. An examiner who identifies that a document extraction AI performs differently on applications from applicants with names associated with particular national origins, but finds no evidence that the institution evaluated this component for proxy-variable bias, will note a governance gap that may be as significant as the credit model finding.
The Pipeline Problem: Proxy Variables Beyond the Credit Model
The most commonly overlooked source of proxy-variable bias in AI-assisted lending is the AI pipeline that surrounds the credit model. By 2026, institutions typically deploy multiple AI tools in the origination process, and each can introduce proxy bias that affects credit outcomes without appearing in the credit model's input list.
AI document extraction tools that process income verification documents (tax returns, W-2s, pay stubs, bank statements) are a specific concern. If an extraction tool systematically performs better on documents from certain applicants (cleaner formatting, recognizable employer names, familiar income structures) and worse on documents from others (self-employment income, multiple part-time employers, income from community-based sources), it may produce systematic differences in income verification quality that correlate with protected class through the structural patterns described above. An applicant whose income is systematically extracted less accurately may have a higher rate of income verification failures, which feeds into the credit decision as a risk flag unrelated to the applicant's actual creditworthiness.
AI communication tools that are used to draft application-stage communications, request additional documentation, or deliver adverse-action notices can also introduce proxy-variable effects if their output differs systematically by applicant demographic characteristics. A communication tool that generates different levels of engagement or different clarity of explanation for different applicant groups, for whatever reason, can affect which applicants successfully complete the application process, producing a demographic composition effect on the approved application pool that was not produced by the credit model itself.
AI pre-screening or lead-scoring tools used to prioritize follow-up with prospective applicants are another pipeline concern. A pre-screening tool that systematically identifies lower engagement potential for prospective applicants from majority-minority neighborhoods may reduce loan officer follow-up with those applicants, producing a geographic disparity in which applications that could have succeeded never reach the credit model.
Addressing the pipeline problem requires extending the proxy-variable detection framework to each AI component in the origination process: building a complete map of AI-assisted steps from initial contact through final decision, identifying where each step could produce demographic differences in outcomes, and applying the proxy-detection methodology at each step as well as at the overall decision level. The overall disparity testing catches the cumulative effect; the pipeline analysis identifies which components are contributing to it.
Key Takeaways
- A proxy variable is a variable not itself a protected characteristic under ECOA but that correlates with one strongly enough to function as a statistical substitute, producing the same discriminatory effect as using the protected characteristic directly. "We didn't use race as an input" addresses disparate treatment (intentional discrimination); it does not address disparate impact, which tests outputs rather than inputs.
- Geographic features (zip code, census tract, neighborhood designation) are the most thoroughly documented proxy variables in U.S. credit models, because residential segregation produced by decades of discriminatory policy makes geographic location a strong correlate of race and national origin in most U.S. markets. Any geographic feature in a credit model carries proxy-variable risk that must be tested and addressed through the LDA search.
- Employer characteristics, financial transaction behavior, name and language features, and credit history behavior patterns are also established proxy-variable categories that AI models can encode, either individually or through interaction effects that are invisible to single-feature testing.
- The core proxy-detection methods are SHAP disparity analysis (identifying features with disproportionate adverse SHAP contributions for protected classes), conditional independence testing (directly testing whether a feature predicts protected class membership after controlling for legitimate credit factors), and permutation importance with demographic stratification (measuring disparity reduction when a feature is randomized within protected-class populations).
- Proxy detection must be a repeatable workflow, not a one-time pre-deployment test. The workflow includes feature inventory with proxy-risk pre-classification, statistical proxy testing for high-risk features, LDA referral for confirmed proxies, and ongoing monitoring on the same cadence as disparate-impact outcome testing.
- The business-necessity defense against proxy-variable disparate impact is available but must be specific and quantified: the accuracy contribution of the retained proxy feature must be documented, and a less-discriminatory alternative must be evaluated and found inadequate. An unquantified assertion that "all our features are credit-relevant" does not satisfy the standard.
- Proxy variables can be introduced at any step in the AI pipeline, not just in the credit model itself. Document extraction tools, communication tools, and pre-screening tools are all potential sources of proxy-variable bias that require evaluation as part of the institution's comprehensive proxy-detection governance.
- OCC Bulletin 2026-13 requires ongoing monitoring of AI model outcomes for fair-lending risk. In 2026, examiners reviewing AI credit governance expect to see a feature inventory with proxy-risk classification, statistical testing results for high-risk features, LDA documentation for confirmed proxies, monitoring records across multiple cycles, and evidence that the proxy-detection analysis covered the full AI pipeline, not just the primary credit model.
Skill.re