Underwriting Governance and Audit Trail
The examiner arrived at the regional bank's lending operations center on a Tuesday morning in October and asked, in the way examiners do, for a sample. Specifically, she wanted the complete underwriting record for 25 mortgage applications denied in the prior six months in which the bank's AI pre-scoring system had assigned a decline-probable rating. She wanted to see the pre-score outputs, the human review records, the adverse-action notices, and the model-risk documentation for the pre-scoring model. She had a laptop, a checklist, and several days. The bank's compliance officer had a moment of quiet confidence she had not expected to feel. Three months earlier, the bank had completed a governance project that had cost four weeks of staff time and several hundred hours of IT work: redesigning its loan origination system (LOS, the software platform managing the mortgage application from intake through closing) to automatically capture the pre-score output, the data extraction log, the underwriter's verification record, the exception documentation where applicable, and the final decision with reasons, all timestamped, linked to the specific file, and queryable by reviewer, model version, and decision type. The compliance officer pulled the sample in eleven minutes. The examiner read it for two days. She had questions about two files. She had no findings. The bank had built the audit trail as the work happened, not assembled it after the exam was scheduled. That is the difference between a governance program and a governance response, and it is the distinction at the heart of what OCC Bulletin 2026-13 requires from institutions that integrate AI into their underwriting process. (The scenario above is a composite illustration drawn from examination patterns reported across multiple institutions; it does not depict a specific bank or examination.)
What OCC Bulletin 2026-13 Actually Requires of the Audit Trail
OCC Bulletin 2026-13, the April 2026 interagency model-risk guidance issued jointly by the OCC (Office of the Comptroller of the Currency), the Federal Reserve, and the FDIC, superseded OCC 2011-12 and explicitly extended the model-risk management framework to AI and generative AI tools. For the underwriting audit trail specifically, the bulletin's requirements translate into four categories of documentation that an institution must be able to produce for any AI-assisted underwriting decision.
The first category is model identity and version documentation. For each credit decision in which an AI model contributed, the record must identify which model was used, what version of the model was active at the time the decision was made, and what the model's governance status was at that time (validated, in monitoring, under re-validation, or flagged for concern). This documentation serves a specific examination function: if the institution later identifies a performance problem with a model version, the examiner can determine which decisions were affected by tracing the model-version log. An institution that cannot identify which model version processed a specific application cannot respond to this question, and the inability to respond is itself an examination finding.
The second category is input and output documentation. For each AI-assisted decision, the record must capture the inputs the model used (the structured data fields, their values, and their sources) and the outputs the model produced (the pre-score, the queue-routing signal, the exception flags, any draft adverse-action language). This input-output record has two purposes. The first is validation support: the model-risk team needs the input-output data to run periodic validation tests confirming that the model's outputs are consistent with what a manual underwriter would have produced. The second is individual-decision explainability: when an applicant or examiner asks why a specific application received a specific pre-score, the institution needs the input record to trace the model's output to the specific data values that drove it.
The third category is human decision and verification documentation. For each file, the record must capture the human underwriter's verification of AI-extracted data, the human's credit decision, the reasons for the decision (specific, accurate, and file-grounded per ECOA and Regulation B, which is Reg B, 12 CFR Part 1002, the CFPB's implementing regulation for ECOA), any exceptions invoked with their authority citation, and the underwriter's identity and timestamp. This documentation is the accountability record: it demonstrates that a human exercised genuine judgment over the credit decision and does not merely transmit the AI's output.
The fourth category is fair-lending testing documentation. For the AI pipeline as a whole, the institution must document the periodic disparate-impact testing it performs: the methodology used, the results across protected classes, any disparities identified, the investigation conducted to explain them, and the documentation of the less-discriminatory alternative search where disparities were found. This documentation is not generated file by file; it is generated at the population level, on a schedule defined by the institution's model-risk governance policy. But it draws on the file-level data captured in the first three categories, which is why the file-level documentation must be structured in a way that supports population-level analysis.
The audit trail is not documentation that gets assembled when an examiner arrives; it is the record that is created automatically as the work happens, structured to answer the questions an examiner will ask before they ask them.
Building the Model-Risk Record
The model-risk record is the institutional file for the AI model itself: not the record of individual decisions, but the record of the model's governance, validation, monitoring, and change history. Every AI model used in the underwriting pipeline needs a model-risk record, whether it was built in-house or purchased from a vendor. The record is the institution's documentation that it knows what the model is, how it works, what its limitations are, and how it is being governed.
A model-risk record for a pre-scoring model used in mortgage origination has eight elements, each of which must be kept current as the model's governance status changes.
Element one: model inventory entry. The model must be registered in the institution's model inventory: a centralized list of all models in use, with their purpose, developer or vendor, approval date, current governance status, and assigned model owner. The model inventory is the starting point for any examination of the institution's model-risk program. An AI tool used in origination that does not appear in the model inventory is an examination finding regardless of how well the tool performs.
Element two: model documentation. The documentation describes how the model works: the inputs, the methodology (whether it is a rules-based engine, a statistical model, or a machine learning model), the training data (for trained models), the output format, and the intended use. For vendor models, the documentation includes the vendor's technical documentation (to the extent the vendor provides it), the institution's assessment of what it was and was not able to verify about the model's design, and the due-diligence record from the vendor selection process. OCC 2026-13's third-party governance requirements specify that the institution is responsible for understanding the model it is deploying, not the vendor. "The vendor didn't explain it" is not a governance defense.
Element three: validation report. The validation report documents the pre-deployment testing of the model: the tests performed, the results, the comparison to manual underwriting on a representative file sample, the disparate-impact test results, and the limitations identified by the validation. The validation report must be performed by a function independent of the model's development team: either an internal model-risk validation team or a third-party validator. A model that was not independently validated before deployment does not meet OCC 2026-13's requirements, and deploying it exposes the institution to both model-risk and fair-lending examination findings.
Element four: approval record. The model must be approved for use by an appropriate authority level within the institution: the model-risk committee, the chief risk officer, or a designated approval body defined in the institution's model governance policy. The approval record documents who approved the model, when, under what conditions, and what remediation or monitoring requirements were attached to the approval. A model approved without formal governance approval is not a compliant model deployment under OCC 2026-13.
Element five: monitoring log. The monitoring log captures the ongoing performance testing that occurs after the model goes into production: periodic checks of the model's score distributions, comparison of model outputs to actual credit outcomes, re-running of key validation tests on current file samples, and any alerts or flags generated by the monitoring program. The monitoring frequency should be defined in the model's governance approval: monthly for high-volume, high-risk models; quarterly for lower-volume or lower-risk models. Any performance drift, score distribution shift, or fair-lending concern identified in monitoring must be escalated to the model-risk committee and documented in the monitoring log.
Element six: change and update log. Every change to the model (updates to input variables, recalibration of weights, new training data, vendor-pushed version updates) must be documented in the change log and, depending on the magnitude of the change, may require a re-validation. Minor parameter updates may require only enhanced monitoring after the update; major structural changes require full re-validation before redeployment. The change log is the institution's defense in the event that a model performance problem is identified: the log enables the institution to trace when the change occurred, what changed, and what governance process was followed.
Element seven: fair-lending testing record. The fair-lending testing record documents the disparate-impact analyses performed on the model's outputs, including the methodology, the time period covered, the protected classes tested, the results, and any investigative or remediation steps taken in response to identified disparities. For a mortgage origination pre-scoring model, this record should include HMDA (Home Mortgage Disclosure Act, which requires collection and reporting of mortgage lending data by race, sex, income, and other applicant characteristics) data analysis comparing pre-score distributions and final approval/denial rates across applicant demographics. The record is updated with each periodic testing cycle.
Element eight: model retirement or replacement record. When the model is replaced, retired, or significantly rebuilt, the record documents the transition: which model version was in use for which time period, the reason for the transition, the governance process for the new model, and the handling of in-process applications during the transition. This record is relevant to any retrospective examination of decisions made during the transition period.
The Individual-File Audit Trail
The model-risk record governs the model. The individual-file audit trail governs each decision. Both are required; neither substitutes for the other. An institution with an excellent model-risk record and poor individual-file audit trails is not meeting OCC 2026-13's requirements, because the individual-file trail is what enables the examiner to verify that the governance described in the model-risk record was actually applied to individual decisions.
The individual-file audit trail for an AI-assisted underwriting decision has five components, each of which is created during the underwriting process and stored in the LOS or document management system.
Component one: AI contribution log. A timestamped record of every AI-generated output that contributed to the file's processing: the document extraction results (with the extracted fields, values, and source document identifiers), the credit report parsing summary, the pre-score output (the score or routing signal, the model version, the input values used), and any AI-drafted language (adverse-action reasons, approval conditions, borrower communication drafts). This log is the factual record of what the AI did with this file.
Component two: human verification record. A timestamped record of the underwriter's verification steps: the extracted fields that were confirmed against source documents, any corrections made to extracted values (with the corrected values and the source documents they came from), and the reviewer's identity. For a clean-queue file with a straightforward income verification, this record may be brief. For a complex commercial exception file, it may be extensive. The record must exist for every file, regardless of complexity.
Component three: decision record. The credit decision, the underwriter's identity, the decision date, the reasons (specific, accurate, file-grounded reasons for denials; approval conditions and terms for approvals), and any exceptions invoked with the authority citation. This record is the core of the adverse-action compliance trail: the documentation that enables the institution to demonstrate, for any denied file, that the stated reasons are grounded in the actual file data and were confirmed by a named human underwriter before release.
Component four: exception documentation (if applicable). For files in the exception queue, the documentation of the exception analysis: the specific exception flag, the compensating factors reviewed, the exception authority cited, the reviewer's analysis of whether the compensating factors support the exception, and the exception decision. This documentation is the institution's defense in a fair-lending challenge to an exception denial: evidence that the exception was evaluated under a consistent, documented framework rather than on the basis of factors correlated with protected class.
Component five: disclosure and notice record. The adverse-action notice or approval communication issued to the applicant, with its issuance date, and the record of the human review that confirmed the notice content before issuance. For adverse-action notices, this includes the confirmation that the stated reasons match the decision record in component three. A notice that was issued without this confirmation creates a trail gap that an examiner will flag.
The combined individual-file audit trail enables the institution to answer, for any specific file, the three questions an examiner always asks: What did the AI do? What did the human do? What decision was made, and why? An institution that can answer these questions for every file in the exam sample is in a defensible position. One that cannot answer them for any file has a compliance finding.
Building the Trail in Real-Time, Not in Retrospect
The most important operational principle in underwriting governance is that the audit trail must be built as the work happens, not assembled after an examination is announced. This principle sounds obvious, but it is consistently violated in institutions that treat governance as a documentation task separate from the underwriting work, rather than as a design feature of the workflow itself.
The failure mode is familiar: an institution deploys an AI pre-scoring tool, the tool works well, volume increases, and the underwriting team moves faster. Documentation habits from the manual-underwriting era, which involved paper files and physical document review, are not automatically adapted to the AI-integrated workflow. Underwriters confirm AI outputs verbally or through informal review, but the verification step is not logged in the LOS because no one redesigned the LOS workflow to require logging it. Three years later, when an examiner asks for the verification record on a denied file, the institution discovers that the verification was performed but not documented. The file is defensible in substance but not in documentation. The examination finding is avoidable and expensive.
The design solution is to make documentation the path of least resistance, not an additional step. Every verification action the underwriter performs should be captured through a workflow tool that records the action automatically: a structured checklist in the LOS where confirming each extracted field creates a timestamped log entry, a review screen where the underwriter marks each pre-score flag as reviewed and enters the analysis, a decision entry form that requires the specific adverse-action reason codes and their file-data citations before the decision record can be saved. When the workflow tool is designed so that the underwriter cannot complete their work without creating the required documentation, the documentation gets created consistently.
The adoption rate data reinforces the urgency of getting this right: 38% of mortgage lenders are using AI or machine learning in their origination or underwriting process as of 2024, up from 15% in 2023. The examination community is catching up. Examiners who are reviewing AI-integrated workflows in 2026 are more sophisticated about what the audit trail should contain than those reviewing such workflows in 2024. An institution that built its AI workflow in 2024 without an intentional audit trail design and has not revisited that design since is behind the examination expectation.
The RAG (Retrieval-Augmented Generation, the technique of grounding an AI model's output in retrieved documents from a specific knowledge base at generation time) approach to adverse-action notice drafting, discussed in the context of the origination workflow, has a specific audit-trail benefit: when the adverse-action notice is generated by an AI tool that retrieves the specific data from the LOS file (the verified income, the DTI calculation, the policy threshold that was breached, the credit report data cited), the notice's content is inherently traceable to the LOS record. The retrieval log shows which LOS data fields were used to generate the notice. The human reviewer's confirmation that the notice accurately represents the LOS record creates a documented chain from the file data to the notice content. This chain is the individual-file fair-lending audit trail for the adverse-action notice.
Reporting Governance to the Board and to Regulators
OCC Bulletin 2026-13 does not just set requirements for the audit trail and model-risk record; it extends those requirements to the board governance level. The bulletin requires that the institution's board of directors, or a board-level committee with delegated authority, receive regular reporting on the institution's AI model-risk program: the inventory of models in production, the validation and monitoring status of each model, any fair-lending concerns identified in testing, and the institution's remediation activity for any identified deficiencies.
Board reporting on AI governance is not a quarterly checkbox for the compliance calendar. It is the mechanism by which the institution's leadership confirms that the AI program is operating within the risk appetite the board has defined and that the governance controls described in the model-risk policy are actually functioning. An institution where the board approves an AI deployment policy but receives no regular reporting on whether the policy is being followed does not have board governance; it has board approval of a paper that the staff implements without oversight.
The reporting package that satisfies OCC 2026-13's board governance requirements for an AI-integrated origination program includes: the model inventory with current governance status for each model; the results of the most recent validation and monitoring cycles, including any performance concerns; the fair-lending testing results, including disparate-impact findings and remediation status; the exception activity report (volume of exceptions by type, grant rates, and any fair-lending patterns in exception outcomes); and any model-related incidents, complaints, or regulatory inquiries received during the period.
For examiners, the board reporting record is evidence of the institution's governance posture. An examiner who asks "how does the board know whether your AI models are performing appropriately?" should receive a package that demonstrates ongoing, substantive oversight, not a description of a process the board delegated entirely to management. The 2026 guidance explicitly raises the governance expectation to the board level for AI tools used in credit decisions, and examination practice in 2026 reflects this expectation.
Chief risk officers and chief compliance officers who have built this reporting capability before their first AI-era examination report a consistent finding: the discipline of building the reporting package forces the internal governance conversations that identify gaps before the examiner does. The quarterly governance report is not just an examination artifact. It is the internal feedback loop that makes the governance program self-correcting.
Key Takeaways
- The audit trail is a design requirement, not a documentation task: it must be built into the LOS workflow so that the record of AI contributions and human decisions is created automatically as the work happens. An institution that assembles the audit trail after an examination is announced is not meeting OCC Bulletin 2026-13's governance expectations.
- OCC 2026-13 requires four categories of documentation for each AI-assisted underwriting decision: model identity and version, input and output records, human decision and verification documentation, and fair-lending testing documentation. Each category serves a specific examination function; none can substitute for the others.
- The model-risk record for each AI model in the origination pipeline requires eight elements: model inventory entry, model documentation, validation report, approval record, monitoring log, change and update log, fair-lending testing record, and model retirement record. Keeping these elements current is an ongoing governance obligation, not a one-time deployment task.
- The individual-file audit trail has five components: AI contribution log, human verification record, decision record, exception documentation where applicable, and disclosure/notice record. These components, taken together, enable the institution to answer the three examination questions for any specific file: what did the AI do, what did the human do, and what decision was made and why.
- Making documentation the path of least resistance, through LOS workflow design that creates the required log entries automatically as the underwriter works, is the operational principle that ensures consistency. Documentation that requires extra steps will be skipped under time pressure; documentation that is required by the workflow will be created every time.
- Board governance of the AI model-risk program is a specific OCC 2026-13 requirement: the board must receive regular reporting on model inventory, validation and monitoring status, fair-lending testing results, exception activity, and any model-related incidents. Board approval of an AI policy without ongoing reporting does not satisfy the bulletin's governance expectations.
- The fair-lending testing documentation required by OCC 2026-13 draws on the file-level audit trail data: the individual-file records of pre-score inputs, routing decisions, exception outcomes, and final decisions provide the data for the population-level disparate-impact analysis. A file-level audit trail that is not structured to support population-level analysis does not fully satisfy the bulletin's requirements.
- The combined model-risk record and individual-file audit trail is both the institution's examination defense and its internal governance feedback loop: the quarterly reporting exercise that forces the governance conversation is also the mechanism that identifies and corrects gaps before an examiner sees them. The 38% AI adoption rate (up from 15% in 2023) means the examination community has sufficient experience with AI-integrated origination to know what a defensible governance record looks like, and the bar will only move higher as adoption continues.
Skill.re