โ†
AI for Banking & Lending
Proficient ยท M13 ยท lesson 13 of 19 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Pre-Scoring with a Defensible Human Decision
๐Ÿ“–
now learning

Pre-Scoring with a Defensible Human Decision

15 min

On a Wednesday in March, a residential mortgage team at a regional bank completed 47 underwriting reviews before noon. Three months earlier, the same team averaged 12 to 15 complete reviews in a full day. The difference was not a larger headcount or longer hours. The difference was a pre-scoring workflow that sorted the queue every morning before the first underwriter arrived. Clean files, with all policy thresholds met and no exception triggers, were grouped and tagged. Exception files, with specific policy flags attached, were routed to senior underwriters with the exception details pre-populated. The underwriters were reviewing, deciding, and documenting, not hunting for the debt-to-income ratio or trying to figure out why the pre-score kicked a file to the exception pile. The bank's chief lending officer reported the throughput gain to the board the following month. The compliance officer presented alongside her, showing the adverse-action documentation for every denied file in the same period: specific, accurate, file-grounded reasons, reviewed and confirmed by a named human underwriter. Volume had more than tripled. The audit trail was intact. That combination, throughput and defensibility, is what a well-designed pre-scoring workflow produces. Building one requires understanding exactly what the pre-score is, what it is not, and where the human decision boundary must sit to keep accountability where the law requires it. (The scenario above is a composite illustration drawn from operational patterns reported across multiple institutions; it does not depict a specific bank or transaction.)

What Pre-Scoring Is, and What It Cannot Be

A pre-scoring model is a computational tool that evaluates a loan application's characteristics against a set of criteria and produces a summary assessment of the application's fit with the institution's credit policy. The assessment might take the form of a numeric score, a categorical rating (clean, exception, decline-probable), a set of policy-flag indicators, or some combination of these. The assessment is generated before the human underwriter reviews the file, and its purpose is to organize the underwriter's work queue so that human attention is concentrated where it adds the most value.

What the pre-scoring model is: a triage and routing tool that processes structured data faster than a human, applies consistent policy criteria across a large volume of files, and identifies the specific policy dimensions on which each file needs human attention. A pre-scoring model built around an institution's credit policy will evaluate every file against the same DTI thresholds, the same credit score tiers, the same derogatory-item rules, and the same property guidelines, without fatigue or variation. For a community bank processing 40 applications in a day, that consistency is a quality control improvement as well as an efficiency gain.

What the pre-scoring model is not: a credit decision-maker. This distinction is not semantic. Under the Equal Credit Opportunity Act (ECOA, the federal statute prohibiting credit discrimination on the basis of race, color, religion, national origin, sex, marital status, age, or receipt of public assistance) and Regulation B (Reg B, 12 CFR Part 1002, the CFPB's implementing regulation for ECOA), a credit decision triggers an obligation to provide specific, accurate adverse-action reasons if the decision is a denial. That obligation rests on the institution, and it requires that the stated reasons accurately reflect the actual factors driving the denial. An AI model's output, whether a score, a rating, or a probability, is not a statement of specific reasons. It is an output derived from weights applied to input variables, and those weights are not always interpretable as specific, applicant-meaningful reasons without human translation and verification.

OCC Bulletin 2026-13, the April 2026 interagency model-risk guidance that superseded OCC 2011-12 and pulled AI and generative AI explicitly under model-risk, fair-lending, third-party, and board-governance expectations, reinforces this distinction. The bulletin requires that any model used in a credit decision be validated to confirm its outputs accurately represent the factors driving the decision, and that the institution be able to trace a denial to specific, file-grounded reasons that satisfy ECOA's adverse-action requirements. A pre-score that routes a file to a decline queue without a human decision and documented reasons is not a compliant adverse-action process. The pre-score is a step in the workflow. The human decision is the event that triggers the legal obligation.

A pre-score is a routing signal that organizes the underwriter's queue; the credit decision, with its documented reasons, belongs to the human underwriter at the end of the review.

Designing the Pre-Score Decision Boundary

The decision boundary is the structural element that makes a pre-scoring workflow defensible: the explicit, documented design choice about which outcomes the AI is authorized to produce and at what point human authority must be exercised. Getting the decision boundary right is the primary design task for any institution building or deploying a pre-scoring system.

A decision boundary that is too early in the workflow wastes the efficiency gain. If the pre-score does nothing more than confirm that a completed application arrived, the model adds no value. A decision boundary that is too late in the workflow erodes accountability. If the pre-score produces a pass-fail output and the human's role is to enter the AI's result into the LOS (the loan origination system used to manage the mortgage application from intake through closing), the human is a data-entry function, not a decision-maker. Under OCC 2026-13, a workflow in which the human exercises no independent judgment over the credit decision is not a compliant human-in-the-loop design.

A well-positioned decision boundary gives the AI authority over routing and pre-qualification signaling, and gives the human authority over the credit decision, the adverse-action reasons, and any exception analysis. In practice, this means:

The AI is authorized to: Classify and extract documents. Order and parse credit reports. Calculate DTI, LTV (loan-to-value ratio, the loan amount divided by the appraised or purchase price of the collateral), and other policy-relevant ratios from verified inputs. Compare calculated ratios to policy thresholds. Generate a queue-routing signal (clean, exception, or decline-probable) based on the policy comparison. Flag specific policy exceptions with the exception details attached. Draft preliminary adverse-action reason language for human review.

The human is required to: Confirm that extracted data matches source documents. Review the file holistically, including risk factors the model may not have weighted. Exercise exception authority if appropriate and authorized. Make the final credit decision. Document the specific, accurate reasons for the decision. Review and confirm any AI-drafted adverse-action language before release. Sign the file, taking responsibility for the decision and its documentation.

This boundary is explicit in the workflow design: the LOS should not permit a file to proceed from the pre-score stage to the adverse-action notice stage without a human sign-off event recorded in the system. If the file is in the decline-probable bucket, the system should require a named underwriter to review the file, enter the final decision, and document the specific reasons before the adverse-action notice can be generated. The pre-score can pre-populate the adverse-action reasons for the underwriter's review and correction, but the underwriter confirms or corrects each reason, and the confirmation is logged.

The design principle behind the boundary placement is legal asymmetry: an approval needs no specific explanation, but a denial under ECOA and Reg B requires specific, accurate reasons. This asymmetry creates a different boundary design for approvals and denials. For approvals, the pre-scoring model's clean-queue signal can substantially reduce the human review burden: the underwriter confirms the key fields, checks for anything unusual, and approves. For declines, the human review burden is higher, because the underwriter must produce the specific, accurate reasons and document their verification. The queue routing should reflect this: clean-queue files get a lighter review protocol, and decline-probable files get a more intensive review protocol. Routing the most complex accountability work to senior underwriters, and the most standard work to junior staff with clear protocols, is one practical implementation of this design principle.

Building the Clean-Queue Protocol

The clean-queue protocol defines how a human underwriter reviews a pre-scored file in the clean queue: a file in which all policy thresholds are met, no exception triggers are present, and the pre-score signals a standard approval. The protocol must be specific enough to ensure genuine human review, and efficient enough to deliver the throughput gain the pre-scoring model enables.

A clean-queue protocol for a standard residential mortgage file has five steps, performed in sequence before the underwriter approves the file.

Step one: income verification confirmation. The underwriter reviews the extracted income figure alongside the source document (W-2 or pay stub for employed borrowers, federal tax returns for self-employed borrowers). The LOS presents the extracted figure and the source document image side by side. The underwriter confirms the figure matches the source or corrects it if it does not. For a clean-queue file, this step takes three to five minutes. If the correction changes the DTI materially (for example, the extracted income was higher than the actual W-2 figure, producing a DTI within threshold on the pre-score but above threshold after correction), the file should be re-routed from the clean queue to the exception queue. The LOS should automate this re-routing when a correction to a key ratio field would change the pre-score outcome.

Step two: credit report review. The underwriter reviews the credit report summary, confirming the credit score, the absence of disqualifying derogatory items, and the accuracy of the tradeline information reflected in the DTI calculation. The AI will have pre-populated the revolving account balances and minimum payments into the DTI calculation. The underwriter confirms that the AI's liability inventory matches the credit report. Common errors at this step: an AI that missed a recently opened account, or included a closed account still showing on the credit report, or used the credit limit rather than the balance for a revolving account's DTI contribution.

Step three: collateral and property review. For a purchase loan, the underwriter reviews the preliminary LTV calculation against the purchase price and loan amount. Before the appraisal arrives, this is a preliminary check. For a refinance, the underwriter reviews the LTV against the AVM (automated valuation model, the computer-generated property value estimate) or the existing appraisal. The underwriter confirms the property type and occupancy code are correct: primary residence, second home, or investment property classifications have different guideline requirements, and a misclassification here can produce an incorrect pre-score.

Step four: AUS finding review (if applicable). For agency-eligible loans, the AUS finding is reviewed and the underwriter confirms the file meets any conditions noted in the finding. An AUS "approve/eligible" finding does not mean the file automatically closes: it means the file meets baseline agency criteria, and the underwriter must apply the institution's overlay policies (stricter requirements the institution imposes above the agency minimum) and review any AUS conditions that require additional documentation or analysis.

Step five: holistic review and approval entry. The underwriter reads the complete file summary, including any borrower narrative, and confirms that nothing in the file raises a concern the model's policy-threshold analysis would not have captured: a non-arm's-length transaction, a property with unusual characteristics, an income source that is real but non-qualifying under the institution's guidelines, or a compensating-factor situation where a borderline characteristic is offset by a particularly strong one. If the file is genuinely clean, the underwriter approves, documents the approval, and submits the file to the next stage. The total review time for a standard clean-queue file using this protocol: 15 to 25 minutes, compared to 60 to 90 minutes for a manual review of the same file without AI pre-processing.

The Exception Queue and Human Judgment

The exception queue contains the files that matter most for accountability design: the ones where the pre-score identified a specific policy flag, and where the human underwriter's judgment determines whether the flag is disqualifying, whether exception authority applies, or whether the analysis of the flag changes the decision. Exception underwriting is where the value of experienced human judgment in an AI-integrated workflow is highest, and where the governance design must be most careful.

A policy exception is a case in which the institution grants credit outside its standard credit policy guidelines, based on documented compensating factors and the exercise of authorized exception authority. Common exception categories in residential mortgage underwriting include: DTI above the standard threshold but within the institution's maximum exception limit; credit score below the standard guideline but above a minimum floor, with documented compensating factors; non-traditional income documentation (such as a borrower whose primary income is from a Schedule K-1 and whose income calculation requires multiple tax return schedules); and property characteristics that fall outside standard guidelines (such as a mixed-use property or a property in a rural area where comparable sales are limited).

The exception queue protocol differs from the clean-queue protocol in several important ways. The review is more intensive: the underwriter reads the specific exception flag, the policy guidance on exception authority for that flag type, the compensating factors available in the file, and the precedent or committee direction on this exception category. The decision is more consequential: an exception granted incorrectly creates model-risk and regulatory exposure; an exception declined incorrectly may create fair-lending exposure if similarly situated borrowers in non-protected classes were granted exceptions in comparable situations.

The fair-lending implication of exception decisions is a critical governance point under OCC Bulletin 2026-13. Disparate impact analysis of an AI-integrated pipeline must include exception outcomes, not just pre-score outcomes. If the institution's AI pre-scoring model routes protected-class borrowers to the exception queue at higher rates than non-protected-class borrowers with similar credit characteristics, and those borrowers are subsequently denied through the exception process at higher rates than comparable non-protected-class borrowers, the institution has a potential disparate-impact pattern that the pre-score alone would not reveal. The audit trail must capture exception routing decisions, exception authority citations, and exception outcomes by borrower demographic, so the population-level analysis can detect this pattern.

The exception decision documentation in the LOS should record: the specific exception flag from the pre-score, the underwriter's analysis of the compensating factors, the exception authority cited (whether from the credit policy manual, a credit committee directive, or a specific officer-approval level), and the final decision with reasoning. This documentation is the institution's defense in a challenge to an exception denial: evidence that the exception was evaluated under a consistent, documented framework, not on the basis of subjective factors correlated with protected class.

The 38% adoption rate for AI in mortgage origination (up from 15% in 2023) creates a comparative fair-lending context that every exception process must account for: as AI-integrated workflows become more common, examiners will compare exception rates and outcomes across institutions, and an institution that cannot document consistent, policy-grounded exception decisions will be at a disadvantage in any fair-lending examination.

Governing the Pre-Scoring Model Under OCC 2026-13

The pre-scoring model is the AI element of the pipeline with the most direct model-risk governance obligations under OCC Bulletin 2026-13. Understanding those obligations is not optional for an institution deploying a pre-scoring model: it is the foundation of the institution's ability to defend the model in an examination.

OCC 2026-13 extends the model-risk management framework established in OCC 2011-12 to AI and generative AI tools, with specific additions for fair-lending testing, third-party vendor governance, and explainability requirements. For a pre-scoring model, the governance obligations translate into four concrete requirements.

Validation. The pre-scoring model must be validated: tested to confirm that its outputs are accurate, consistent, and grounded in the credit factors it is intended to measure. Validation for a pre-scoring model includes: confirming that the model's pre-score for a given file is consistent with the credit decision the institution would have reached through manual underwriting for a representative sample of files; testing the model's stability over time (does its scoring distribution shift as the portfolio changes?); and testing for disparate impact (do protected-class applicants receive systematically different pre-scores than comparably situated non-protected-class applicants, and if so, can the disparity be explained by credit factors?).

Documentation. The model's documentation must be sufficient for an examiner to understand how it works: the input variables, the variable weights or decision rules, the training data used (if it is a trained model), the validation results, and the limitations the validation identified. For a vendor model, the institution may not have access to the full model documentation, but it must have enough documentation to perform the validation tests described above, and it must document its due-diligence process for the vendor model as part of its third-party governance record under OCC 2026-13's vendor-oversight requirements.

Ongoing monitoring. The model must be monitored for performance drift. A pre-scoring model that was validated in 2024 against a portfolio with a certain credit-quality distribution may perform differently in 2026 if the portfolio's distribution has shifted. Monitoring means running periodic validation checks against current performance, flagging the model for re-validation if performance metrics change materially, and documenting the monitoring results in the model-risk record.

Explainability. For any file that receives an adverse pre-score that contributes to a denial, the institution must be able to explain the pre-score's contribution to the decision in terms that are meaningful to the human underwriter at the decision stage and, ultimately, to a compliance examiner reviewing the file. "The model scored the file below threshold" is not an explanation. The explanation must identify the specific input variables that drove the pre-score below threshold, so the human underwriter can verify that those variables accurately describe the file and that they are consistent with the adverse-action reasons stated in the notice.

For institutions using vendor pre-scoring models (which describes the majority of community banks and regional lenders, since building a proprietary model requires data science resources that most smaller institutions do not have), OCC 2026-13's third-party requirements add a layer to each of these obligations. The vendor agreement must give the institution access to the documentation, validation data, and explanation outputs it needs to meet the governance requirements. An institution that deploys a vendor pre-scoring model without contractual rights to this information is not meeting OCC 2026-13's third-party governance obligations, regardless of how well the model performs on average.

Key Takeaways

  • A pre-scoring model is a routing and triage tool that organizes the underwriting queue: it is not a credit decision-maker. The credit decision, with its specific, accurate, ECOA-compliant adverse-action reasons, belongs to the human underwriter at the end of the review process.
  • The decision boundary, the explicit design choice about which AI outputs are authorized and where human authority must be exercised, is the primary governance design task for any institution deploying a pre-scoring workflow. The boundary must be documented in the workflow design and enforced in the LOS so that no file can be declined without a named human decision and documented reasons.
  • The legal asymmetry of lending, approvals need no explanation while denials under ECOA and Reg B require specific, accurate reasons, shapes the boundary design. Clean-queue files (all policy thresholds met) have a lighter human review protocol; decline-probable files require a more intensive review, because the underwriter must produce the specific, accurate reasons and document their verification against the file.
  • The clean-queue protocol (income verification confirmation, credit report review, collateral review, AUS finding review, holistic review) reduces a standard file review from 60 to 90 minutes to 15 to 25 minutes, while preserving the human verification controls that make the approval defensible. The throughput gain is real only if all five steps are performed on every clean-queue file.
  • Exception underwriting is the highest-stakes component of the exception-queue workflow: exception decisions must be documented with the specific exception flag, the compensating factors analyzed, the exception authority cited, and the reasoning. Fair-lending testing of the AI pipeline must include exception routing and outcome rates by protected class, not just pre-score distributions.
  • OCC Bulletin 2026-13 requires that pre-scoring models be validated, documented, monitored, and explainable for individual file decisions. For vendor models, the institution's third-party governance program must include contractual rights to the documentation and explanation outputs needed to meet these requirements.
  • The 38% AI adoption rate in mortgage origination (up from 15% in 2023) sets the competitive baseline: lenders not using pre-scoring are clearing fewer files per underwriter than those that are. The differentiator between a productivity gain and a regulatory exposure is the decision boundary design: does the human at the end of the pipeline have genuine authority, or is the AI making the calls?
  • Accountability stays human regardless of where the AI sits in the pipeline. "The pre-scoring model declined the file" is not a legally sufficient adverse-action reason, and it is not a defense to a Reg B examination finding. The underwriter who reviewed the file and entered the decision is the institution's accountable party, and the documentation of that review is what makes the decision defensible.