โ†
AI for Banking & Lending
Capable ยท M12 ยท lesson 12 of 21 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
Keeping the Analyst Accountable
๐Ÿ“–
now learning

Keeping the Analyst Accountable

15 min

The examination team arrived on a Tuesday morning. Three examiners from the OCC (Office of the Comptroller of the Currency) were reviewing the BSA/AML program at a $1.9 billion community bank. They requested the case files for a sample of 40 Suspicious Activity Report (SAR) filings from the prior 12 months and 30 closed-alert records for alerts that were investigated and dismissed as false positives. The bank had deployed an AI transaction monitoring triage system eight months earlier, and the lead examiner wanted to understand how it worked. Within two days, the examination team had found a pattern. In the closed-alert records, the documentation was thin: the AI score appeared, the alert type appeared, and a line reading "reviewed and closed" with a date appeared. In most cases, there was no record of what the analyst actually found when they reviewed the account activity, what specific facts led to the conclusion that the alert was a false positive, or what policy basis supported the dismissal. For the SAR filings, the narratives were complete and accurate, but the underlying investigation records were less clear about the investigative steps taken before the filing decision. The examiner's conclusion: adequate SAR filing, inadequate investigation documentation. The bank received a finding that required a remediation plan addressing documentation standards across the BSA program. The AI system was not the problem. The documentation culture around the AI system was. This lesson is about building the documentation infrastructure that keeps the analyst accountable to the examiner, and why that infrastructure is not optional in an AI-assisted BSA program.

What the Examiner Is Actually Reading

A federal banking examiner reviewing a BSA/AML program is reading for evidence of a functional compliance program, not for evidence of sophisticated technology. The FFIEC (Federal Financial Institutions Examination Council) BSA/AML Examination Manual, which provides the framework for BSA examinations, describes an adequate program in terms of its outcomes: does the institution identify suspicious activity? Are SAR filing decisions appropriate? Is the monitoring system calibrated to the institution's risk profile? Are investigations meaningful and complete?

When an examiner opens a BSA case file, they are trying to reconstruct what actually happened. They want to see: what alert or other indicator triggered the investigation, what data the analyst examined, what the analyst concluded from that data and why, whether any explanations were sought from the customer and what those explanations were, how the investigation conclusion was reached, and who made the final decision (for SAR filings, who certified the SAR; for closed alerts, who approved the dismissal). An examiner who can reconstruct that chain of reasoning from the case record has seen adequate documentation. An examiner who cannot reconstruct it has seen a documentation gap, regardless of whether the underlying investigation was actually thorough.

This is the core insight about documentation in an AI-assisted BSA program: the documentation must demonstrate the analyst's reasoning, not just record the AI's output. An AI score of 0.87, by itself, does not explain why the analyst found the activity suspicious. An AI-generated summary, pasted without annotation into the case record, does not demonstrate that the analyst evaluated the underlying facts. A SAR narrative drafted by AI and filed without recorded verification does not show that the certifying officer independently confirmed its accuracy. The examiner is looking for evidence of human judgment applied to facts. If the documentation shows only AI output and a sign-off, the examiner cannot tell whether the analyst exercised judgment or delegated it.

OCC Bulletin 2026-13, issued in April 2026, superseded OCC 2011-12 and updated the interagency model-risk management framework to explicitly include AI and generative AI tools across all bank functions, including BSA/AML. The bulletin requires model-risk documentation, validation, and ongoing governance for AI models used in compliance functions. In the context of a BSA examination, this means the examiner may ask not just about the BSA program itself but about the governance of the AI tools used in that program: is the triage model in the bank's model inventory? When was it last validated? Who approved its use? What performance metrics are monitored? Institutions that deployed AI tools without corresponding governance documentation face a two-front examination issue: inadequate program documentation and inadequate model-risk documentation.

The Documentation Standard for Closed Alerts

A closed alert is a transaction monitoring alert that was investigated and determined not to warrant a SAR filing or any other action. In a program with a 90 to 95 percent false-positive rate, the vast majority of alerts in any given period are closed. The documentation standard for a closed alert is lower than for a SAR filing, but it is not zero, and in an AI-assisted program it has specific requirements that go beyond what many institutions currently produce.

A complete closed-alert record should contain: the alert type and the specific rule that triggered it, the AI score assigned to the alert (if an AI triage system is used), the key facts from the account and transaction history that were relevant to the investigation, the analyst's assessment of those facts (why did these facts, in context, lead to a conclusion that the activity was not suspicious?), any customer-level context that explained the activity (established business pattern, documented seasonal variation, known customer relationship with a counterparty), and the approval notation: who closed the alert, on what date, under what authority.

The "analyst's assessment" element is where most documentation gaps appear in AI-assisted programs. When an AI summary accurately describes a customer's activity and the activity is clearly explained by the customer's documented business profile, an analyst under time pressure may simply note "explained by customer profile, no further action" and move on. That notation is not wrong. It is incomplete. The examiner who reads it cannot see which aspects of the customer's profile explain the activity, whether the analyst actually reviewed the underlying records (or just the AI summary), or what specific facts the analyst found convincing.

A better version of the same conclusion, written to the documentation standard an examiner will accept, looks like this: "Alert triggered by three cash deposits totaling $27,400 in a 10-day period. Customer is documented as a retail florist with established daily cash deposit pattern. AI summary confirmed: account history shows consistent weekly cash deposits averaging $8,000 to $12,000 per week for 14 months. Reviewed three months of account statements. Deposit amounts consistent with seasonal holiday volume (Valentine's Day period). No structuring pattern (deposits varied between $4,200 and $9,800, timing irregular over business days). Alert closed as false positive, explained by documented business activity."

The difference between these two versions is not elaborateness. It is traceability. The second version records the specific facts the analyst examined, the specific conclusion they drew, and the specific reasoning that connects the facts to the conclusion. An examiner can reconstruct the investigation from that record. They can evaluate whether the reasoning was sound. They can confirm that a human analyst applied judgment to the specific facts, not just reviewed an AI summary.

Documenting the AI Role in the Case Record

In an AI-assisted BSA program, the case record should explicitly acknowledge where AI contributed to the investigation and where the analyst independently verified or supplemented the AI's output. This is not a bureaucratic requirement for its own sake. It is the documentation that allows the institution to respond to the examiner's question about how AI is used in the BSA program with a specific and accurate answer drawn from the case records.

The AI contribution documentation should note: the AI score assigned to the alert and the system that generated it, the AI-generated summary or contextual enrichment used in the investigation, which elements of the AI-generated summary the analyst verified against source records, any discrepancies between the AI summary and the underlying records that were identified during verification, and any additional data the analyst pulled independently beyond the AI-generated summary. This documentation does not need to be lengthy. A single paragraph in the case notes that identifies what the AI provided, what the analyst confirmed, and what the analyst independently found is sufficient for most closed-alert records.

For SAR filings that used AI for narrative drafting, the verification documentation should be more explicit. The case record should reflect that the AI-drafted narrative was reviewed against the case file, that specific factual elements were confirmed (and ideally which elements were confirmed), that any corrections were made to the AI draft before certification, and that the named compliance officer reviewed and certified the final narrative as accurate. A SAR narrative that lacks any verification documentation record is a filing that appears to have been generated by AI and filed without human review, regardless of whether verification actually occurred. The verification must be documented to be defensible.

Who Owns the Conclusion: The Accountability Chain

The most fundamental principle in BSA/AML compliance, unchanged by AI and reinforced by every regulatory framework that has addressed AI in financial services, is this: every consequential decision in the compliance program belongs to a named, accountable human. AI can triage alerts, enrich context, and draft narratives. The SAR filing decision, the alert dismissal decision, the enhanced due diligence designation, the account restriction, the customer relationship termination: each of these belongs to a human being who can be named, whose judgment can be evaluated, and who can be asked to explain the reasoning behind the decision.

This is not a philosophical position. It is embedded in the legal framework. The BSA requires a designated compliance officer or senior management to certify SAR filings. The certification is a personal attestation with criminal penalty exposure for knowing false statements. The institution's BSA officer cannot certify a SAR filed automatically by an AI system because the certification requires a personal judgment that the filing is accurate and complete to the officer's knowledge. That judgment requires a human being to actually review the filing and form the relevant belief.

For alert dismissals, the accountability principle is equally clear even though there is no explicit certification requirement analogous to the SAR. The institution's BSA program is evaluated, in part, by whether its dismissal decisions are appropriate: are alerts being closed for valid reasons that a reasonable BSA professional would find adequate? An institution that has delegated dismissal decisions to an AI model (by auto-closing alerts below a certain score threshold without human review) has removed the human accountability from those decisions. If an examiner later finds that a significant suspicious activity pattern was missed because the alerts were auto-closed, the institution cannot point to a human investigator's reasoning as a defense. The institution made no decision. A machine made one.

The accountability chain in an AI-assisted BSA program has three links. The first is the analyst who conducts the investigation, reviews the AI-generated summary, supplements it with independent research, applies judgment to the facts, and records their conclusion in the case record. The second is the BSA officer or supervisor who reviews the analyst's work for significant cases (SAR filings, enhanced due diligence designations, account-level actions) and confirms that the investigation was adequate and the conclusion was supported. The third is the compliance officer who certifies SAR filings and is ultimately accountable to the regulator for the quality of the BSA program.

AI tools can assist at every point in that chain: they can help the analyst investigate faster, they can help the supervisor review more efficiently, they can help the compliance officer spot patterns across the program. They cannot substitute for any link in the chain. The examiner, looking at the documentation, must be able to see each link: the analyst's investigation record, the supervisor's review notation (where required), and the compliance officer's certification. If any link is missing, the documentation is incomplete.

Building the Examination-Ready Documentation Framework

An examination-ready BSA documentation framework in an AI-assisted program has four components: a case documentation standard, a model governance record, a quality assurance process, and a SAR program review. Each component serves a different examination audience but all four must be present and current when the examiners arrive.

The case documentation standard is the written policy and accompanying template or guidance document that tells analysts what the case record must contain. It should specify the minimum documentation requirements for different case types (low-complexity false-positive closures, complex pattern investigations, SAR filings), describe the AI contribution documentation requirement, define who must approve each case type, and establish the retention period and storage location for case records. A case documentation standard that lives only in the BSA officer's knowledge and has never been written down is a governance gap. The examiner will ask to see the written standard, and the written standard sets the baseline against which the actual case records are evaluated.

The case documentation standard should be reviewed when the AI tools used in the program change. If the institution deploys a new AI triage tool, or adds a generative AI drafting capability, the documentation standard should be updated to address those tools specifically. An outdated documentation standard that does not mention AI, in a program that uses AI extensively, creates an impression that the institution's governance has not kept pace with its technology deployment.

The model governance record is the model-risk documentation required by OCC Bulletin 2026-13 for each AI tool used in the BSA program. For a transaction monitoring triage model, the governance record includes the model inventory entry (purpose, inputs, outputs, approval date, owner), the validation report (assessment of the model's performance metrics, comparison to the prior system, coverage of customer segments), the performance monitoring records (ongoing metrics reviewed at defined intervals), and the change management log (documentation of any updates to the model and the revalidation that accompanied them). For a generative AI SAR drafting tool, the governance record includes the same structural elements adapted to the tool's specific function, plus documentation of the verification requirement and the policy language that governs its use.

The model governance record does not need to be a separate document for each model. Many institutions maintain a consolidated model-risk management file with sections for each model in the inventory. What matters is that the documentation exists, is current, and is organized in a way that allows the examiner to find the information they need without a scavenger hunt through multiple systems.

The quality assurance process is the periodic review of a sample of case records to confirm that documentation standards are being followed in practice. A BSA program that has strong written standards but weak actual documentation is a program where the gap will surface in examination. A quality assurance review of 50 to 100 case records per quarter, conducted by the BSA officer or a compliance quality-control function, checks for completeness (all required elements present), substance (the analyst's assessment is recorded, not just the AI output), and accuracy (the facts stated in the record are consistent with the source data). Findings from the quality assurance review should be documented, shared with the analyst team, and used to update the training and documentation guidance as needed.

In an AI-assisted program, the quality assurance review should specifically assess the AI contribution documentation: is the AI tool's role in each investigation recorded? Is there evidence that the analyst independently verified the AI-generated summary or narrative? Are corrections to AI-generated content noted? If the quality assurance review finds that the AI contribution is rarely documented or that verification is seldom recorded, the program has a systematic documentation gap that will recur in examination.

The SAR program review is the retrospective evaluation of SAR filing quality that the FFIEC BSA/AML Examination Manual expects as part of an ongoing compliance program. A SAR program review looks at a sample of SARs filed in a period and a sample of investigations that resulted in no SAR filing, and evaluates whether the filing decisions were appropriate given the evidence. In an AI-assisted program, the SAR program review should also assess whether AI-assisted narratives meet the FinCEN quality standards for completeness and specificity, whether the verification process is producing narratives that accurately reflect the investigation record, and whether there are patterns of error in AI-drafted narratives that suggest the drafting tool's calibration needs adjustment.

The Analyst's Professional Accountability

Beyond the institutional governance framework, the individual BSA analyst carries a professional accountability that is not reduced by the availability of AI tools. The analyst who certifies a case record as representing their investigation of the alert is attesting, within the institution's program, that they reviewed the relevant facts and applied their professional judgment. An analyst who delegates that judgment to an AI system, accepts the AI's output without verification, and records the conclusion as their own assessment is both producing inadequate documentation and making a false attestation.

This matters practically as well as ethically. In the event of a regulatory investigation or a law enforcement matter involving a SAR, the analyst who investigated the case may be asked to describe what they found and how they reached their conclusion. An analyst whose case record consists of AI-generated summaries with minimal personal notation will struggle to answer that question credibly. An analyst whose case record reflects their own reasoning, the specific facts they examined, the questions they asked, and the analysis they applied will be able to testify credibly to the adequacy of the investigation.

The professional development implication is that AI tools should be used to accelerate the investigation, not to replace the investigative thinking. The analyst who uses AI to quickly assemble the customer context, and then applies their own judgment to evaluate that context against what they know about the customer, the institution's community, and the applicable typologies, is developing and demonstrating BSA expertise that will be professionally valuable in any examination, any legal proceeding, and any career discussion. The analyst who uses AI to do the investigation and then records the AI's conclusions as their own is not developing expertise; they are taking on personal accountability for an investigation they did not actually conduct.

When the AI Is Wrong: Documentation of Disagreements and Overrides

AI triage models are not perfectly calibrated. A model that gives a high score to an alert that investigation reveals to be a clear false positive, or a low score to an alert that investigation reveals to be a significant suspicious pattern, is not a system failure. It is the expected behavior of a statistical model operating on imperfect information. The institution's response to these cases is where governance quality is demonstrated most clearly.

When an analyst's investigation leads to a different conclusion than the AI model's scoring predicted, that disagreement should be documented. If a high-score alert is investigated and found to be a clear false positive, the case record should note that the AI scored the alert highly, that the investigation found a specific explanation (and what the explanation was), and that the analyst concluded the activity did not warrant a SAR. This documentation serves two purposes. It allows the institution to track patterns in the AI model's performance: if the model consistently over-scores a particular customer segment or alert type, the performance monitoring process should surface that pattern and trigger a revalidation review. It also demonstrates to the examiner that the analyst exercised independent judgment rather than simply accepting the AI's recommendation.

When an analyst's investigation leads them to file a SAR on an alert that the AI model scored low, the documentation is equally important. The case record should note the low AI score, explain what additional factors the analyst identified that the model did not capture in its scoring, and document the analyst's independent assessment that the activity crossed the suspicion threshold. This kind of override documentation is a marker of a healthy BSA program: it shows that the AI is a prioritization tool, not a gatekeeper, and that human investigators can and do find suspicious activity in the lower-scored portions of the queue.

A pattern of overrides in a particular direction (analysts consistently finding that low-score alerts contain SAR-worthy activity in a specific category, or consistently finding that high-score alerts in a specific customer segment are false positives) is also valuable feedback for model performance monitoring and revalidation. OCC Bulletin 2026-13 expects ongoing performance monitoring with a process for triggering revalidation when the model's performance falls below acceptable thresholds. Override patterns, tracked systematically in the case management system, are an input to that performance monitoring process.

Key Takeaways

  • An examiner reviewing a BSA/AML program in an AI-assisted institution is looking for evidence of human judgment applied to facts, not evidence of AI deployment. Documentation that shows only AI scores and automated workflow steps, without recorded analyst reasoning, cannot demonstrate that the compliance program is actually functioning. The documentation must show the analyst's assessment, not just the AI's output.
  • A complete closed-alert record documents the specific facts examined, the analyst's reasoning connecting those facts to the dismissal conclusion, and the approval by a named individual. "Reviewed and closed" with an AI score does not satisfy this standard. Traceability of reasoning is the documentation requirement, not word count.
  • The accountability chain in an AI-assisted BSA program has three links that must all be visible in the documentation: the analyst who investigated and recorded their reasoning, the supervisor who reviewed significant cases, and the compliance officer who certified SAR filings. AI tools assist each link but cannot replace any of them. If any link is absent from the documentation, the program has a governance gap.
  • No SAR may be auto-filed from AI output. The BSA certification requirement requires a named human compliance officer to personally attest to the accuracy of the filing. That attestation cannot be made on behalf of an automated process. The verification step before SAR filing must be documented in the case record to demonstrate that the certification is grounded in a human review of the draft's factual accuracy.
  • OCC Bulletin 2026-13 applies to all AI tools used in BSA programs and requires a model inventory entry, validation documentation, performance monitoring records, and a governance trail for each AI tool. The examination-ready BSA documentation framework must include model governance records alongside case documentation, quality assurance records, and SAR program review findings.
  • When an analyst's investigation reaches a different conclusion than the AI model's scoring predicted (an override in either direction), the case record should document the disagreement: the AI score, the specific factors the analyst identified that differed from the model's assessment, and the analyst's independent conclusion. Override patterns are a governance input to model performance monitoring and a demonstration that human judgment remains active in the program.
  • Individual BSA analysts carry professional accountability for the investigations they record as their own. Using AI output as a starting point and then applying independent judgment is the appropriate use of AI in BSA work. Accepting AI output without verification and recording it as a personal conclusion is both inadequate documentation and a professionally risky practice in any examination, legal proceeding, or performance review.