Documentation Standards Under OCC 2026-13
The following scenario is a composite illustration drawn from examination patterns observed under model-risk guidance; it does not depict a specific institution, examiner, or examination. The examiner set her laptop on the conference room table at 8:47 a.m. and opened the bank's model-risk inventory. She was looking for three things: the model-risk file for the commercial credit scoring model deployed fourteen months earlier, the fair-lending testing results for the past four quarters, and the override log with demographic annotations. The model-risk manager, seated across from her, pulled up a shared drive and began searching. Twenty minutes later he had found a PDF of the original validation report, a spreadsheet of approval rates that may or may not have included the four required test periods, and an email thread with subject line "re: re: re: re: model updates." There was no override log with demographic annotations. There was no consolidated model-risk file. The examiner closed her laptop and said, "Let's talk about documentation." That conversation lasted three days. It produced eleven findings. The institution had, in fact, conducted much of the required analysis. It had done the work. What it had not done was build an audit-grade record that could survive examination. Under OCC Bulletin 2026-13, issued in April 2026 by the Office of the Comptroller of the Currency (OCC), the Federal Reserve, and the Federal Deposit Insurance Corporation (FDIC), doing the work and documenting the work are the same obligation.
What OCC 2026-13 Requires in Plain Terms
OCC Bulletin 2026-13 superseded OCC Bulletin 2011-12, the prior model-risk management guidance that had governed how banks documented, validated, and governed quantitative models since the aftermath of the 2008 financial crisis. The 2026-13 update was driven by three developments that the 2011-12 framework did not adequately address: the proliferation of artificial intelligence (AI) and generative AI (GenAI) tools in credit decisioning and operational workflows; the integration of third-party and vendor model components that are partially or wholly opaque to the institution using them; and the interagency recognition that fair-lending testing, third-party risk oversight, and board governance obligations are not separate compliance topics layered on top of model risk management but are constituent elements of a unified governance framework.
The documentation requirements in 2026-13 flow from this integration. A model-risk file under 2026-13 is not a folder of validation reports. It is a living record that spans the model's entire lifecycle from initial proposal through production use through retirement, and that captures every material decision about the model, every test result, every finding, and every remediation action, organized so that an examiner who has never seen the model can reconstruct its history from the file alone. The regulation uses the phrase "audit-grade records" in describing the standard. The meaning of audit-grade is operational: a record that can be traced, verified, and reproduced. An email that says "we discussed the bias results and they look fine" is not audit-grade. A testing report that shows the methodology, the sample, the results by demographic stratum, and the escalation determination is audit-grade.
The four documentation domains that 2026-13 organizes the requirements around are: model development and validation records; production monitoring and performance records; fair-lending and compliance records; and governance and decision records. Each domain has specific required elements, each must be current (not a snapshot from deployment that has never been updated), and each must be organized within the consolidated model-risk file rather than distributed across departmental records systems.
Model Development and Validation Records
The development and validation record is the foundation of the model-risk file. Under 2026-13, it must document not only the model's current state but the rationale for the decisions made at each stage of development. This is a higher standard than 2011-12's requirement for a validation report. The rationale requirement means that a reviewer reading the file must be able to understand not only what the model does but why it was built the way it was, which features were considered and why certain features were selected or rejected, what alternatives were evaluated and why the chosen approach was preferred.
For AI and machine learning (ML) models specifically, the development record must address explainability. OCC 2026-13 defines explainability as the ability to describe, in terms that are accessible to non-technical reviewers including examiners and board members, how the model produces its outputs. For tree-based ensemble models such as gradient-boosted machines (GBM), the explainability record typically includes feature importance analysis, partial dependence plots, and at least a subset of individual prediction explanations using a method such as Shapley Additive Explanations (SHAP). For neural network models, the record must address the limitations of available explainability tools and document how the institution manages the gap between the model's internal complexity and the interpretability needed for governance.
The validation record under 2026-13 must include the following elements:
- Conceptual soundness assessment: a documented evaluation of whether the model's theoretical foundation is appropriate for the intended use case, including any assumptions embedded in the model architecture and the evidence supporting those assumptions.
- Data quality assessment: documentation of the data sources used in model development, the period of data used for training and testing, any data cleaning or preprocessing steps applied, and an assessment of whether the training data is representative of the model's intended deployment population.
- Performance benchmarking: the model's performance metrics on the validation holdout sample, including relevant discriminatory power measures (area under the receiver operating characteristic curve, or AUC), calibration measures, and comparisons to the challenger model or the prior model being replaced.
- Sensitivity analysis: documentation of how the model performs under stressed or out-of-distribution conditions, including economic stress scenarios for credit models and distributional shift scenarios for models trained on historical data.
- Model limitations: a documented enumeration of the model's known limitations, the conditions under which those limitations are most likely to affect model performance, and the compensating controls in place.
- Approval record: the formal sign-off from the model risk management function, including the risk rating assigned, any conditions or requirements attached to approval, and the schedule for the next periodic revalidation.
For vendor models and third-party AI components, where the institution does not have access to the model's source code, training data, or internal architecture, the documentation requirement does not disappear. It shifts to compensating documentation: the due diligence record for the vendor, the contractual terms securing access to validation-relevant information, the results of independent outcomes testing conducted by the institution, and the documentation of any limitations identified through that testing and the institution's compensating controls.
Production Monitoring and Performance Records
A model-risk file that reflects only the state of the model at deployment is not compliant with 2026-13. The regulation requires that the production monitoring record be maintained as a rolling log of the model's performance throughout its operational life, updated at each monitoring cycle and incorporating any out-of-cycle reviews triggered by the monitoring program's alert thresholds.
The production monitoring record must include the following elements for each monitoring period:
- Population stability index (PSI): a measure of whether the distribution of model inputs in the current production population has shifted materially from the distribution in the model's training or validation population. A PSI above 0.25 typically indicates a material shift requiring investigation. The monitoring record must show the PSI by key feature group, not only an aggregate PSI, so that the source of any shift can be identified.
- Performance metric tracking: the model's current discriminatory power and calibration metrics compared to the metrics at the last validation, with a documented assessment of whether any decline is within acceptable bounds or constitutes a performance degradation trigger.
- Outcome rate monitoring: for credit models, the approval rate, denial rate, and modification rate for the current period, shown by product type and by geographic segment at minimum, and by demographic stratum where required for fair-lending purposes.
- Exception and override tracking: the number and rate of manual exceptions to the model's recommendations, the direction of those exceptions (approving above the model recommendation or declining below it), and the performance of exception decisions versus model recommendations over time.
- Alert status and escalation record: whether any threshold triggers were activated during the monitoring period, what investigation or remediation actions were taken, and the current status of any open findings from prior monitoring periods.
The override tracking element deserves particular attention because it is one of the areas most frequently cited in 2026-13 examination findings. Many institutions have robust model monitoring programs that track model performance but do not systematically track the rate, direction, and demographic distribution of manual overrides. OCC 2026-13 requires override tracking as a model-risk governance obligation because the override pattern can reveal two types of problems simultaneously: a credit risk problem (if overrides are consistently overperforming or underperforming the model, the model's thresholds may be miscalibrated) and a fair-lending problem (if overrides are distributed unevenly across demographic groups, the override process may be introducing disparate treatment on top of any model-driven disparity). An override log that does not capture demographic distribution cannot serve the fair-lending governance function that 2026-13 requires.
Fair-Lending and Compliance Records Within the Model-Risk File
The integration of fair-lending compliance records into the model-risk file is the most significant structural change that 2026-13 introduced relative to prior practice in many institutions. Under prior practice, fair-lending testing for credit models was conducted by the fair-lending compliance function and the results were maintained in the compliance department's records. The model-risk team received a summary or a notification when a finding was escalated. The 2026-13 framework requires that the full fair-lending testing record for each AI credit model appear in that model's risk file, because the regulation treats fair-lending performance as a component of model performance, not as a separate compliance question.
The fair-lending record within the model-risk file must include, at minimum:
- The testing methodology for each fair-lending test conducted since the model's deployment, including the basis for comparator group construction, the demographic estimation methodology used (such as Bayesian Improved Surname Geocoding, or BISG), the statistical tests applied, and the sample size for each credit quality stratum tested.
- The testing results for each period, showing approval rate ratios by demographic group and credit quality stratum, the statistical significance of any identified disparities, and the disparity ratio relative to the institution's alert thresholds.
- Proxy variable analysis results, including the correlation analysis, the remove-and-retest findings, and the geographic disparity analysis.
- The less-discriminatory-alternative (LDA) search documentation for any identified proxy features, including the alternative features tested, the testing methodology, the results showing the trade-off between disparate-impact reduction and business performance, and the disposition decision with supporting rationale.
- The escalation and remediation record for any findings, including the nature of the finding, the investigation conclusions, the remediation plan, the implementation status, and the post-remediation test results confirming the effectiveness of the remediation.
An institution whose fair-lending compliance function is conducting thorough testing but maintaining the results exclusively in the compliance team's records, with the model-risk file receiving only a summary notification, is not meeting the 2026-13 documentation standard. This is a structural gap, not a minor administrative deficiency. It means that the model-risk governance chain -- including the model risk management function, the risk committee, and ultimately the board -- is not receiving the fair-lending performance information required by the regulation's governance framework. It also means that an examiner reviewing the model-risk file will not find the fair-lending record, which is itself a finding.
Governance and Decision Records
The governance record within the model-risk file documents the human decisions made about the model at each stage of its lifecycle. This is the record that demonstrates the accountability chain that OCC 2026-13 requires for AI models: a traceable path from each material model decision to the individual or body that made it, the information that was available at the time of the decision, and the rationale for the decision made.
The governance record must include:
- The initial model approval, including the model risk rating, the approving body, and any conditions attached to approval.
- The deployment authorization, including who authorized production deployment, on what date, and whether any conditions from the approval process had been satisfied or deferred.
- Periodic revalidation decisions, including the date of each revalidation, the findings and risk rating update, and the approving body.
- Material model changes, including any changes to the model's feature set, thresholds, or operating parameters since deployment, the business and risk rationale for each change, the validation status of each change, and the approval record.
- Threshold and policy decisions, including any changes to the model's use (such as expanding or restricting the loan types or geographies to which it is applied), the rationale for those changes, and the approval record.
- The model owner record, including the current named model owner, any ownership transitions since deployment, and the model owner's attestation at each monitoring cycle that they have reviewed the monitoring results and any open findings.
The model owner attestation requirement is significant. OCC 2026-13's human accountability principle requires that every AI model have a named individual owner who is personally accountable for the model's performance, and that this accountability be demonstrated through a recurring attestation process, not merely through an organizational chart assignment. An attestation that says "I have reviewed the Q2 2026 monitoring results for the commercial credit scoring model. No threshold triggers were activated. Two open findings from prior periods are being remediated on schedule. Signed: [name and title]" is the type of record 2026-13 requires. A governance record that shows the model owner's name in the inventory but no attestations for the past three quarters is a documentation gap.
The Consolidated Model-Risk File Structure
The four documentation domains described above must be organized within a consolidated model-risk file for each model. The consolidated file structure is not merely an administrative convenience. It is the mechanism by which the governance framework functions: if the validation record is in one system, the monitoring record is in another, the fair-lending record is in a third, and the governance record exists only as email threads, no single reviewer can see the model's full history, and the institution cannot demonstrate the integrated oversight that 2026-13 requires.
A compliant consolidated model-risk file structure typically includes the following components, organized by model and maintained as a living document updated at each monitoring cycle:
Model identification section: Model name and version identifier, model type (predictive/AI/GenAI), intended use, in-scope products and geographies, current risk rating, deployment date, current model owner with contact information, next scheduled revalidation date, and link or reference to the most recent validation report.
Lifecycle events log: A chronological record of all material lifecycle events since the model's creation: initial approval, deployment, each revalidation, each material change, each ownership transition, and any significant incidents or findings. Each entry includes date, event type, summary, approving body or individual, and a reference to supporting documentation.
Current monitoring dashboard: A summary of the most recent monitoring period results, including population stability, performance metrics, outcome rates, override rates, fair-lending ratios, and alert status. This section is updated at each monitoring cycle and retains a history of prior periods for trend analysis.
Open findings register: A list of all open findings associated with the model, including the finding description, the source (internal monitoring, external examination, audit, or self-identification), the severity, the remediation plan, the responsible owner, the target remediation date, and current status. Closed findings remain in the register with their closure date and post-remediation test results.
Complete documentation archive: The full text of all validation reports, monitoring reports, fair-lending testing reports, LDA search documentation, and governance approvals, organized by document type and date.
The model-risk information management system that houses this file must support version control (so that earlier versions of documents are preserved and the history of changes is traceable), access control (so that access to sensitive model information is restricted to authorized personnel while remaining accessible to examiners), and audit logging (so that a record exists of who accessed the file, when, and what changes were made). For institutions that maintain model-risk files in document management systems or enterprise risk platforms, these system requirements must be validated as part of the overall model governance framework.
Documentation for Generative AI Tools
OCC 2026-13 explicitly addresses generative AI (GenAI) tools used in banking and lending workflows, requiring documentation standards that reflect the distinctive characteristics of GenAI compared to traditional predictive models. The documentation requirements for GenAI tools are structured around three characteristics that distinguish them from predictive models: they produce natural-language outputs rather than numeric scores; their behavior can vary in ways that are not fully predictable from their architecture; and their use of retrieval-augmented generation (RAG), fine-tuning, or prompt engineering creates documentation obligations that do not exist for traditional models.
For a GenAI tool used in credit-related workflows -- such as a credit memo drafting assistant, an adverse-action reason code generator, a compliance question-answering tool, or a loan document extraction tool -- the model-risk file must include:
- Intended use definition: A precise statement of the specific task or tasks the GenAI tool is authorized to perform, the specific outputs it is authorized to produce, and the specific workflows it is authorized to participate in. The intended use definition must explicitly state what the tool is not authorized to do, since GenAI tools are capable of producing outputs beyond their intended scope if not constrained by prompt engineering, RAG configuration, or output validation.
- Human review requirements: Documentation of the review process that applies to GenAI outputs before they reach a credit decision, a regulatory filing, a customer communication, or any other consequential use. Under 2026-13, every GenAI output that affects a regulated credit function must have a documented human review step. The human review documentation must specify what the reviewer is checking (factual accuracy, regulatory compliance, fair-lending compliance, or other specific criteria), not merely that review occurs.
- Output validation testing: Documentation of pre-deployment testing of the GenAI tool's outputs for accuracy (are factual claims in the output accurate?), regulatory compliance (do the outputs comply with applicable requirements such as Regulation B's adverse-action requirements?), and fair-lending compliance (are there systematic differences in output quality or accuracy across demographic groups?). This testing must be documented with the methodology, the sample, and the results, to the same standard as predictive model validation.
- Prompt and configuration record: Documentation of the system prompts, retrieval configurations, grounding documents, and output constraints applied to the GenAI tool, with version control so that any changes to these configurations are tracked. A system prompt is a model configuration decision. A change to the system prompt is a material model change that requires the same documentation and approval process as a change to a predictive model's feature set or threshold.
- Production output monitoring: Documentation of the ongoing monitoring program for GenAI outputs, including what monitoring is conducted (sampling and review of production outputs), the frequency of monitoring, the criteria applied in monitoring review, and the findings from each monitoring cycle.
The prompt and configuration record requirement addresses a gap that has appeared in institutions that deployed GenAI tools before 2026-13 was effective. In many of those deployments, the system prompt and retrieval configuration were set up by a technology team and not documented as model configuration decisions. When the system prompt was later modified to improve output quality or address a problem, the modification was treated as a software configuration change rather than a model change, and the documentation and approval process applicable to model changes was not followed. Under 2026-13, this is a model governance failure. The system prompt is the specification of the model's behavior. Modifying the system prompt without following the model change management process is equivalent to modifying a credit scoring model's coefficient without revalidation.
Key Takeaways
- OCC Bulletin 2026-13, issued in April 2026 by the OCC, Federal Reserve, and FDIC, superseded OCC 2011-12 and established that doing the work and documenting the work are the same obligation; an institution that conducted required analysis but produced no audit-grade records has not satisfied the regulation.
- The 2026-13 documentation framework organizes requirements across four domains: model development and validation records, production monitoring and performance records, fair-lending and compliance records, and governance and decision records, all of which must be consolidated in a single model-risk file maintained as a living document throughout the model's lifecycle.
- Fair-lending testing results must appear in the model-risk file itself, not only in the compliance department's records; an institution whose fair-lending function conducts thorough testing but routes results only through the compliance reporting chain is not meeting the 2026-13 documentation standard.
- Override tracking is a required element of the production monitoring record under 2026-13, capturing the rate, direction, and demographic distribution of manual exceptions, because the override pattern can simultaneously reveal credit risk problems (miscalibrated thresholds) and fair-lending problems (disparate treatment introduced by the override process).
- The governance record must include model owner attestations at each monitoring cycle, not merely an organizational assignment; a model owner's name in the inventory with no attestations for recent quarters is a documentation gap that will be identified in examination.
- GenAI tools used in credit-related workflows require documentation of intended use definition, human review requirements, output validation testing, prompt and configuration records with version control, and production output monitoring; a change to the system prompt is a material model change requiring the same documentation and approval process as a change to a predictive model's feature set.
- The consolidated model-risk file must be maintained in a system that supports version control, access control, and audit logging, so that an examiner can reconstruct the model's complete history, including every material decision, test result, finding, and remediation action, from the file alone.
- The test of a compliant model-risk file is not whether it exists but whether an examiner who has never seen the model can, within a reasonable amount of time, understand the model's purpose, design, performance history, fair-lending record, governance decisions, and current status from the file without requiring explanation from bank staff.
Skill.re