โ†
AI for Banking & Lending
Proficient ยท M1 ยท lesson 1 of 19 ยท in progress
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
AML Model Governance
๐Ÿ“–
now learning

AML Model Governance

15 min

The following scenario is a composite illustration; the figures are drawn from documented model-governance examination findings and do not reflect a specific institution or individual. Marcus had been the chief model risk officer at a $7.4 billion regional bank for six years when his institution deployed its first AI-assisted transaction monitoring system. The deployment went smoothly. The vendor delivered a well-packaged product with solid performance metrics, and the project team celebrated a successful go-live. Fourteen months later, sitting across from an OCC examination team, Marcus realized that "smooth deployment" and "defensible governance" were not the same thing. The examiners asked for the model inventory entry, the validation report, the ongoing performance monitoring records, and the documentation of the two threshold adjustments made in months four and nine. Marcus could produce the vendor's pre-sale validation report, a single email chain approving the threshold changes, and quarterly alert-volume dashboards. What he could not produce was an institution-level validation, a documented performance-monitoring framework with defined trigger thresholds, or a governance record showing that the BSA compliance committee had evaluated the trade-offs in the threshold adjustments. The model risk team had treated the AI monitoring tool as a software product, not as a model. That distinction, in the OCC's view, was the central failure.

What Model Governance Means in the BSA/AML Context

The word "governance" carries a lot of weight in financial regulation, but in the BSA/AML context it has a specific and concrete meaning. Model governance is the set of policies, processes, and documentation that allows an institution to demonstrate: that it understands what its AML monitoring models do, that it tests them before and after deployment, that it monitors their ongoing performance, and that it can reconstruct the reasoning behind every material change made to those models. Governance is not the model itself. It is the institutional infrastructure around the model that makes the model's use defensible to an examiner.

The governance framework for AML models in the United States is anchored in OCC Bulletin 2026-13, issued in April 2026. This bulletin superseded OCC 2011-12 (the prior model risk management guidance, often called "SR 11-7" in its Federal Reserve incarnation) and updated the framework to explicitly address AI and machine learning models in all banking functions, including BSA/AML compliance. The core architecture of model risk management that 2011-12 established (model development, validation, and ongoing monitoring as three distinct functions) is preserved in 2026-13, but the newer guidance adds specific requirements around explainability, third-party vendor accountability, and fair-lending overlay for models that touch credit decisions. For BSA/AML AI tools, the relevant provisions are model inventory, validation, ongoing monitoring, and documentation -- all of which apply regardless of whether the model was built internally or purchased from a vendor.

The gap that caught Marcus -- treating the AI monitoring tool as a software product rather than a model -- is a specific governance failure that OCC 2026-13 names directly. An institution that purchases an AI transaction monitoring system from a third-party vendor does not inherit the vendor's validation work. The vendor's pre-sale performance report may demonstrate how the model performs on the vendor's test dataset. It cannot demonstrate how the model performs on the institution's specific customer mix, transaction patterns, geographic footprint, and risk profile. Institution-level validation is a separate and mandatory step, and the institution retains full accountability for the model's performance regardless of who built it.

This accountability principle runs throughout OCC 2026-13 and throughout BSA/AML compliance law. The named BSA compliance officer who certifies the institution's annual BSA program assessment is certifying the outputs of every system that feeds that program, including the AI models that prioritize the alert queue, draft SAR (Suspicious Activity Report) narratives, and score customer risk. There is no regulatory mechanism that transfers that certification obligation to a vendor. "The vendor validated it" is not a defense when an examiner finds that the institution deployed a model that was miscalibrated for its specific customer segments and then operated it for 18 months without institution-level testing.

The Model Inventory: The First Requirement

Every model that materially influences a BSA/AML outcome must appear in the institution's model inventory. This includes not just the primary transaction monitoring system (TMS) rule engine, but also any AI overlay that scores or re-ranks alerts, any machine learning model that assigns customer risk ratings, any natural language processing (NLP) tool that assists in SAR narrative drafting, and any anomaly detection system that identifies unusual customer behaviors for enhanced review. The scope is broader than many institutions initially assume.

The reason that scope matters is not bureaucratic. It is substantive. An AI alert scoring model that re-ranks 900 alerts per month and determines which 15 percent the analyst team actually investigates is making consequential decisions about which suspicious activity gets detected. If that model is not in the model inventory, it is not subject to periodic validation, it is not receiving ongoing performance monitoring, and changes to its configuration are not being governed through the model risk management process. The model's influence on BSA program outcomes is real regardless of whether the institution has formally classified it as a model.

The model inventory entry for an AML AI tool should capture several categories of information. The first is identifying information: the model name, version, vendor (if applicable), deployment date, and the BSA/AML function the model supports (alert prioritization, customer risk scoring, narrative drafting, and so on). The second is purpose and scope: what the model does, which customer segments and transaction types it covers, and what decisions or outputs it generates. The third is technical documentation: the model's architecture at a level sufficient to understand how it produces its outputs, the input features it uses, and the training data on which it was developed. The fourth is validation status: when the model was last validated, by whom, and what the validation found. The fifth is ownership: which individuals or functions are responsible for the model's ongoing performance and governance.

This documentation serves two purposes simultaneously. First, it is the evidence record that an examiner reviews to confirm that the institution understands its own tools. Second, it is the institutional memory that prevents the common failure mode where a model's original rationale and configuration are forgotten over time and the model begins operating as a black box that nobody inside the institution can explain. Marcus's failure was partly attributable to this second problem: the AI monitoring system had been deployed by a project team that subsequently dispersed, and the model risk team that inherited the tool had incomplete documentation of its design parameters and governance history.

Validation: What It Requires and Why the Vendor Cannot Do It for You

Model validation is the process of confirming that a model does what it is supposed to do, using the institution's own data and the institution's own independent evaluators. Under OCC 2026-13, validation has three core components: conceptual soundness review, outcomes analysis, and ongoing monitoring review. For an AML AI tool, each component has specific meaning.

Conceptual soundness review evaluates whether the model's approach is theoretically appropriate for its intended use. For an AI alert scoring model, this means reviewing whether the model's architecture and training methodology are suited to the task of identifying suspicious activity patterns in transaction data. It means examining the training data for representativeness: was the training dataset drawn from transaction patterns similar enough to the institution's portfolio that the model's learned associations apply? It means assessing whether the features the model uses to score alerts are logically related to suspicion (a model that assigns high risk scores based on features unrelated to actual suspicious activity is conceptually unsound, even if it has been trained on a large dataset). And it means evaluating whether the model's outputs can be explained at a level sufficient to satisfy OCC 2026-13's explainability requirements.

The explainability requirement deserves particular emphasis. OCC 2026-13 does not require that every AI model be a simple rule-based system. It does require that the institution be able to explain, in terms that a compliance professional and an examiner can understand, how the model produces its outputs. For an alert scoring model, this means the institution should be able to answer the following question at the case level: why did this account receive a low risk score for the past 14 months while exhibiting the transaction patterns that the institution is required to monitor? "The model gave it a low score" is not an answer that satisfies OCC 2026-13. The institution must be able to describe what features drove the score, whether those features were accurate for this account, and whether the model's architecture was capable of detecting the patterns present in this account's transaction history.

Outcomes analysis tests whether the model performs as intended when deployed on the institution's own data. This is the component that most directly addresses Marcus's failure. An outcomes analysis for an AML alert scoring model examines how well the model's scores predict which alerts are true positives: does the model assign higher scores to alerts that result in SAR filings than to alerts that are investigated and closed as non-suspicious? It examines performance across customer segments to identify systematic miscalibration: does the model perform similarly for commercial accounts and retail accounts, for domestic wire transfers and international remittances, for long-tenured customers and recently onboarded ones? And it examines the model's coverage of FinCEN (Financial Crimes Enforcement Network) typologies: do the model's scores reflect the financial crime patterns that FinCEN has identified as prevalent in the institution's markets?

A segment-level outcomes analysis is not optional. An AI model can produce excellent aggregate performance statistics while performing systematically poorly for specific customer segments. A model trained predominantly on commercial account data may score retail account alerts poorly. A model trained on domestic transaction data may be poorly calibrated for international wire activity. The aggregate metrics will not reveal these gaps; only segment-level analysis does. OCC 2026-13 requires the institution to confirm that the model performs appropriately across the populations it serves, and that requires looking below the aggregates.

Ongoing monitoring review is the third validation component and, in many ways, the most practically important because it is the discipline that prevents validated models from degrading between validation cycles. Ongoing monitoring means tracking defined performance metrics on a defined frequency, comparing current performance to the validated baseline, and escalating to formal revalidation when performance deviates materially from that baseline.

For AML AI tools, the performance metrics subject to ongoing monitoring should include: SAR filing rate trends (overall and by activity category), the distribution of model scores for historically confirmed true positives, the distribution of model scores for dismissed alerts, segment-level alert generation rates, and any significant changes in the underlying data (customer mix, transaction volumes, product usage patterns) that might cause model drift. The frequency of monitoring should be calibrated to the model's risk level and operational significance: a model that scores 1,200 alerts per month and determines which ones get investigated warrants monthly monitoring of its key metrics, not an annual review.

The validation cadence for AML AI models under OCC 2026-13 is at minimum annual, and more frequent if the institution's customer mix, geographic footprint, or business strategy changes materially. "Changes materially" means something specific: a significant acquisition, entry into a new product line, expansion into a new geographic market, or a change in the customer segments the institution serves. These events can cause the model's learned associations to diverge from the current portfolio in ways that are not visible from aggregate performance metrics alone. The institution must have a defined process for determining when a material change has occurred and triggering revalidation accordingly.

Third-Party Vendor Governance: Accountability Does Not Transfer

The majority of smaller and mid-sized institutions that deploy AI-assisted transaction monitoring are purchasing products from third-party vendors rather than building models internally. This is a rational choice: vendor products benefit from development resources and training datasets that most community and regional banks cannot match. But vendor deployment creates a specific governance challenge that OCC 2026-13 addresses directly: the institution retains full accountability for the model's performance and governance regardless of who built it.

This accountability principle has several practical implications. The first is institution-level validation. As described in the previous section, the vendor's pre-sale validation report and performance documentation are inputs to the institution's validation process, not substitutes for it. The institution must conduct its own outcomes analysis using its own transaction data to confirm that the vendor's model performs as intended for the institution's specific customer population. The vendor's testing on its training dataset, or on a pooled dataset of its client institutions, may not capture the institution's specific characteristics.

The second implication concerns the model inventory and documentation requirements. Even when a model is purchased from a vendor, the institution must be able to explain how it works at a level sufficient to satisfy OCC 2026-13. A vendor contract that restricts the institution's access to model documentation creates a governance problem: the institution cannot validate what it cannot examine, and it cannot explain to an examiner what it does not understand. Vendor contracts for AI tools used in BSA/AML compliance should include provisions for model documentation access, performance data sharing, and cooperation with institution-level validation activities.

The third implication involves ongoing monitoring and change management. When a vendor updates the model -- releasing a new version, retraining on updated data, changing scoring algorithms -- the institution must treat that change as a material model modification subject to its model risk management process. A vendor software update that silently changes the model's scoring behavior is not exempt from governance review just because the change originated with the vendor. The institution must know when model changes occur, assess whether the changes require revalidation or additional testing, and document its review in the model-risk file.

This third implication deserves a specific example. Suppose a vendor retrains its transaction monitoring model quarterly using aggregated case data from across its client base. Each quarterly release may shift the model's scoring weights in ways that change how it performs on the institution's specific customer segments. If the institution is not monitoring its key performance metrics at a frequency that would detect such shifts, a quarterly vendor retrain could degrade the institution's true-positive detection rate for three months before the next monitoring cycle reveals the problem. The governance structure must be designed with this risk in mind.

The fourth implication of third-party accountability is that vendor commitments in sales materials do not satisfy institutional obligations. A vendor that promises "95 percent false-positive reduction" or "best-in-class AML detection" is making representations about its model's general performance characteristics. Those representations do not translate into regulatory compliance commitments for the institution. The institution cannot represent to its examiner that its BSA program is effective because a vendor promised excellent performance. The institution can only represent that its program is effective if it has tested and documented that performance using its own data and its own validation processes.

Documentation: The Governance Record Examiners Read

Documentation in model risk management is not paperwork. It is the institutional memory that allows governance to function across personnel changes, time, and regulatory scrutiny. An institution whose AML governance is sound but whose documentation is inadequate will fail an examination almost as definitively as an institution whose governance is actually deficient, because the examiner's access to the institution's practices runs entirely through the document record.

The documentation that must be maintained for each AML AI model includes several layers. The first layer is the model inventory entry, described earlier, which identifies the model, its purpose, its scope, and its ownership. The second layer is the development and validation documentation: the conceptual soundness review, the outcomes analysis, the ongoing monitoring methodology, and the validation conclusions. For a purchased model, this includes the institution's institution-level validation documentation (not just the vendor's report) and any supplemental testing conducted to address gaps in the vendor's documentation.

The third layer is the change management record. Every material change to the model's configuration, including threshold adjustments, parameter changes, scenario additions or deletions, and model version updates, must be documented in the change management record. The record must capture what changed, why it changed, what testing was conducted before the change was deployed, what governance body approved the change, and what post-deployment monitoring was conducted to confirm the change performed as expected. This is the layer that Marcus could not produce: the two threshold adjustments in months four and nine had been made informally, documented only in an email chain, without formal retrospective simulation testing or compliance committee approval in the documented governance record.

The fourth layer is the ongoing performance monitoring record: the monthly or quarterly metrics reports, the trend analyses, the governance reviews, and any escalation or remediation actions taken when metrics deviated from baseline. This layer creates the time-series evidence that the institution was actively watching its model's performance throughout its operational life, not just at the validation checkpoints.

The fifth layer is the examination-ready summary that consolidates the governance history into a narrative an examiner can follow. This is not a compliance theater document. It is the practical answer to the question: if an examiner walks in tomorrow and asks about the AML model we deployed 22 months ago, can we tell the story coherently? The story should include: why we deployed this model, what we validated before deployment, how we monitored it after deployment, what changes we made and why, and what the model's current performance indicators show. An institution that can tell this story clearly is demonstrating that its model governance program is substantive rather than nominal.

Documentation must also address the human oversight layer. OCC 2026-13 requires that AI-assisted processes in BSA/AML programs maintain meaningful human review of model outputs. The documentation record should show not just that analysts reviewed alerts but that the review was meaningful: that analyst overrides of model scores are recorded and their rationale documented, that systematic patterns in analyst disagreement with model scores are tracked and assessed for model recalibration implications, and that the BSA compliance officer's certification of the annual program assessment is based on a substantive review of the model's performance, not a rubber stamp of the model's output.

Governance Structure: Bridging Model Risk and Compliance

The most common institutional failure in AML model governance is a structural one: model risk management and BSA compliance operate as separate functions with separate reporting lines, separate governance committees, and separate performance metrics. Model risk validates the model's technical performance. BSA compliance monitors the program's regulatory outcomes. Neither function is watching the place where the two intersect: the alignment between what the model is optimized to do and what the BSA program is required to accomplish.

This gap creates a specific failure mode. Model risk can conclude that the AI alert scoring model is technically accurate, performing within its validated parameters, with acceptable precision and calibration metrics, while the BSA compliance program quietly degrades because the model is optimized for a metric (alert volume reduction) rather than a compliance objective (detecting and reporting suspicious activity). The model risk function, reviewing model-level metrics, sees nothing wrong. The BSA compliance function, reviewing program-level metrics, sees SAR filing rates holding steady (or perhaps not tracking the right metrics to notice a problem). Neither function catches the disconnect until an examiner does.

OCC 2026-13 responds to this structural problem by requiring that model risk governance be integrated with the functions the model serves, not just the model risk function itself. For AML AI models, this means the model's governance committee should include both model risk and BSA compliance representation, that the performance monitoring framework should include both model-level metrics (precision, calibration, score distribution) and program-level metrics (SAR filing rates, typology coverage, investigative quality), and that revalidation decisions should be informed by compliance performance data, not just model performance data.

In practice, bridging this gap requires two structural changes that many institutions have not yet made. The first is a shared metrics framework: a single dashboard that shows both model performance and BSA program performance together, with defined relationships between the two (for example, an alert score distribution metric adjacent to the SAR filing rate by score decile, so that changes in the score distribution are immediately visible alongside their impact on filing rates). The second is a defined escalation path: when BSA compliance performance metrics decline, the governance framework must have a mechanism for triggering model risk review, not just a BSA program review, because the root cause may be model-level rather than process-level.

The board governance requirement in OCC 2026-13 also applies here. The board, or its designated risk committee, must receive periodic reports on material AI models used in BSA/AML compliance, including both their technical performance and their program outcomes. Board-level reporting that presents only operational metrics (alert volumes, cost per investigation, analyst headcount) without program-quality metrics (SAR filing rates, typology coverage, recall by segment) does not satisfy the board's governance oversight obligation. The board cannot provide meaningful oversight of risks it cannot see.

Building a Defensible Program

A defensible AML model governance program is not necessarily a large program. Even at a community bank with two BSA analysts and a single purchased transaction monitoring system, it is possible to build a governance program that satisfies OCC 2026-13 and produces genuinely useful management information. The core elements are the same regardless of institution size; what scales is the depth and formality of each element.

The first element is a complete model inventory that includes every tool fitting the definition of a model in the BSA/AML context: the primary TMS, any AI overlay, any customer risk scoring tool, and any AI-assisted narrative tool. For a small institution with a single purchased product, this might be a two-page inventory entry. For a larger institution with multiple systems, it might be a structured database with dozens of entries. What matters is completeness, not size.

The second element is an institution-level validation conducted before or shortly after deployment, using the institution's own transaction data. For a small institution, this may require external expertise: engaging a qualified independent reviewer who can conduct the outcomes analysis and produce the validation report. The cost of an institution-level validation is real but substantially smaller than the cost of an OCC examination finding that the institution deployed an AML AI model without validating its performance for its own customer population.

The third element is an ongoing monitoring framework with defined metrics, defined measurement frequency, and defined trigger thresholds that escalate to formal revalidation or governance review. The metrics must include both model-level indicators (score distribution, precision by segment) and program-level indicators (SAR filing rates by activity category, typology coverage). The triggers must be defined in advance, not evaluated retrospectively, to prevent normalization of declining metrics.

The fourth element is a change management process that applies to every material configuration change, regardless of whether it originates internally or from a vendor update. Every change should require: documentation of what changed and why, pre-deployment testing that quantifies the impact on true-positive detection, governance approval at the appropriate level (BSA compliance committee for threshold changes; model risk committee for architectural changes), and post-deployment confirmation that the change performed as expected.

The fifth element is the documentation record that consolidates all of the above into a format that an examiner can review. This is not about creating paper for its own sake. It is about ensuring that the institutional knowledge embedded in the governance process is captured in a durable form that survives personnel changes, system changes, and the passage of time. Marcus's core problem was not that his institution's governance was entirely absent. It was that it was not captured in the documentation record in a form that could be reconstructed by an examiner who arrived 14 months after the deployment with no knowledge of what had been discussed in project meetings and informal emails.

The final element is regular reporting to the BSA compliance committee and the board risk committee that presents both model performance and program performance together. This reporting creates the accountability loop: governance bodies see the metrics that reveal program quality, not just the metrics that reveal operational efficiency, and they are in a position to ask the questions that Marcus's compliance committee should have been asking throughout the 14-month period between deployment and examination.

Key Takeaways

  • OCC Bulletin 2026-13, issued April 2026, superseded OCC 2011-12 and explicitly requires that AI and machine learning tools used in BSA/AML programs be governed as models under the full model risk management framework, including inventory, validation, ongoing monitoring, and documentation.
  • An institution's accountability for the performance of a vendor-supplied AML AI model is complete and non-transferable: the vendor's pre-sale validation does not substitute for institution-level validation, and the vendor's model performance on its training dataset does not establish performance on the institution's specific customer population.
  • The model inventory requirement applies to every tool that materially influences a BSA/AML outcome, including AI alert scoring overlays, customer risk rating models, and AI-assisted SAR narrative tools, not just the primary transaction monitoring rule engine.
  • Ongoing performance monitoring must track both model-level metrics (score distributions, precision by segment) and program-level metrics (SAR filing rates by activity category, typology coverage), because model performance and program performance can diverge in ways that neither metric set alone would reveal.
  • Every material model change, including vendor-originated updates, requires pre-deployment testing that quantifies the impact on true-positive detection, documented governance approval, and post-deployment confirmation, all captured in the model-risk file.
  • The structural gap between model risk management and BSA compliance creates a specific failure mode: model validation concludes technical accuracy while the program degrades because the optimization target diverges from the compliance obligation; governance must integrate these two functions to close the gap.
  • Board-level reporting on AML AI models must present program-quality metrics (SAR filing rates, typology coverage, recall by segment) alongside operational metrics, because boards cannot govern risks they cannot see.
  • Documentation is the institutional memory that makes governance auditable: the model inventory entry, the validation report, the change management record, the ongoing monitoring record, and the examination-ready narrative must together allow the institution to reconstruct the complete governance history for every model it operates.