AI for Risk, Compliance & Audit
Strategic · M24 · lesson 24 of 26 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Testing and Monitoring AI-Related Controls
📖
now learning

Testing and Monitoring AI-Related Controls

15 min

Introduction

Learn how to test that AI-related controls are operating effectively and how to monitor AI performance over time. This lesson covers both one-time testing (as part of audit) and ongoing monitoring.

At the Workflow Integration level, you are designing and implementing AI-enhanced processes across your function. You need to think systematically about how AI fits into existing workflows, what controls are necessary, and how to measure the effectiveness of AI-integrated processes at scale.

This lesson is designed to be accessible to professionals at all experience levels while providing the depth needed for practical application. Whether you are encountering these concepts for the first time or building on existing knowledge, the material ahead will strengthen your ability to navigate AI governance challenges with confidence and competence.

Core Concepts

Practical Use Cases

Use Case 1: Testing AI-Enhanced Audit Controls

Internal audit team implements AI transaction flagging; testing plan includes:

Quarterly Testing Plan:

  • Input Validation Testing (1 audit day)
  • - Obtain transaction log for previous quarter
  • - Sample 50 transactions randomly
  • - Verify each transaction is in source system with matching amount, date, user
  • - Document any discrepancies
  • - Assess: Are input data controls operating?
  • Configuration Testing (1 audit day)
  • - Obtain documented AI rules/parameters (what criteria trigger a flag)
  • - Verify rules match the documented control procedures
  • - Test a sample of flagged transactions to confirm they match stated rules
  • - Assess: Is the AI configured as intended?
  • Output Review Testing (2 audit days)
  • - Obtain all transactions flagged by AI in one week (e.g., 200 transactions)
  • - Sample 20 transactions randomly
  • - For each sampled transaction, confirm:
  • - An auditor reviewed the transaction details
  • - Auditor documented their assessment (control deficiency or not)
  • - Review appears substantive (not auto-approved)
  • - If any reviews appear cursory, interview auditor
  • - Assess: Are auditors actually reviewing output, or rubber-stamping?
  • Performance Testing (2 audit days)
  • - Obtain list of all transactions tested by AI in a full audit
  • - For a sample of 30 non-flagged transactions, perform independent testing to confirm no control issue was missed
  • - For flagged transactions, review audit conclusion to assess whether classification was correct
  • - Calculate: False positive rate (items flagged but not deficiencies), false negative rate (items not flagged but were deficiencies)
  • - Assess: Is the AI achieving its intended objective of identifying control deficiencies?

Annual Bias Testing (1 audit day) - Analyze AI output by location/business unit - Compare false positive rate by location; if one location has significantly higher FP rate, investigate whether it's real or bias - Assess: Is the AI showing systematic bias?

Total audit effort: 7 days per year (can be integrated into regular audit testing; not additional work)

Result: Audit committee can confirm that AI controls are operating effectively.

Use Case 2: Monitoring AI-Enhanced Compliance Screening

AML compliance team implements monitoring dashboard for transaction screening AI.

Real-time Monitoring: - Daily: Automated job checks that screening completed successfully; alert if it failed - Daily: Automatic calculation of flag rate (% of 50,000 transactions flagged); alert if outside 2-5% range - Weekly: Operations team reviews alert log; escalates any issues

Performance Metrics Dashboard (updated weekly, reviewed by AML team): - Total transactions screened (target: 50,000 per day) - Flag rate: % flagged as high-risk (target: 2-5%) - SAR filing rate: % of flagged transactions resulting in SAR (target: 15-25%) - False positive rate estimate: If we have historical data on which flagged transactions turned out to be suspicious vs. legitimate, calculate this - Processing time: Average time to screen all transactions (target: 90%)

Monthly Review Meeting (1 hour, AML leadership): - Review metrics dashboard - Investigate any trends (e.g., if flag rate is climbing, why?) - Discuss root causes and corrective actions - Document discussion and decisions

Quarterly Bias Analysis (4 hours): - Segment output by customer geography, industry, customer size - Calculate false positive and false negative rates for each segment - If one segment has significantly different rates, investigate (is it real or bias?) - Document findings and any corrective actions

Annual Model Validation (8 hours): - Obtain test dataset of known suspicious and legitimate transactions - Run model; calculate accuracy, false positive/negative rates - Compare to prior year's performance; identify any degradation - If performance is degrading, consider retraining model - Document validation results

Total monitoring effort: ~2 hours per week (standing agenda item) + 4 hours per quarter + 8 hours per year = ~30 hours per year for a mid-sized compliance team (worth the effort for a critical control)

Result: AML leadership has visibility into AI performance; issues are detected and addressed quickly.

Anti-patterns / Misuse Risks

Anti-Pattern 1: Testing Once, Never Again Testing AI controls during initial audit; then not testing again because "we know it works."

Risk: Controls degrade over time; AI performance drifts; problems are not detected.

Prevention: Build testing into annual audit plan; test regularly.

Anti-Pattern 2: Monitoring Without Action Collecting metrics but not reviewing them or acting on findings.

Risk: Problems are visible but not addressed; metrics become pointless.

Prevention: Schedule regular review meetings; assign accountability for investigation and corrective action.

Anti-Pattern 3: Overly Granular Monitoring Tracking too many metrics; dashboards become unreadable; noise obscures signal.

Risk: Important issues are missed because there are too many metrics.

Prevention: Focus on key metrics that matter; avoid vanity metrics.

Anti-Pattern 4: No Baseline Starting to monitor without establishing a baseline of normal performance.

Risk: Can't tell if changes are significant; constant false alarms.

Prevention: Before implementing AI, establish baseline metrics; define normal range; set alert thresholds.

[Practical Tip]

As you work through these concepts, consider how each one applies to your current role. Think of a specific scenario from your recent work where this concept would have been relevant. Building these mental connections between theory and practice is the fastest way to internalize new knowledge and make it actionable in your daily responsibilities.

Human Judgment Checkpoints

Checkpoint 1: Testing Plan Sufficiency Are you testing all critical AI controls? Are tests happening regularly enough? Are tests designed to be meaningful?

Checkpoint 2: Monitoring Usefulness Are metrics being reviewed regularly? Is action taken when issues are detected? Or is monitoring just for show?

Checkpoint 3: Root Cause Investigation When monitoring reveals a problem, is there a process to investigate and address it?

Traceability / Defensibility Considerations

Audit Documentation - Maintain testing working papers (procedures, samples, results) - Document monitoring results (dashboards, metrics, trends) - Document investigation and corrective actions when issues are detected

If auditors ask "How do you know your AI controls are working?", you can show testing results and monitoring data.

[Practical Tip]

As you work through these concepts, consider how each one applies to your current role. Think of a specific scenario from your recent work where this concept would have been relevant. Building these mental connections between theory and practice is the fastest way to internalize new knowledge and make it actionable in your daily responsibilities.

Responsible AI and Control Considerations

Continuous Bias Assessment - Testing and monitoring should include bias assessment - Bias can emerge over time as data or business changes; monitoring helps detect it

Practice / Reflection Prompts

  • Testing Plan: For each AI control in your function, design a test. What would you test? How would you test it? How often?
  • Monitoring Metrics: What metrics would tell you if the AI is working as intended? What is the normal range? What would trigger investigation?
  • Monitoring Frequency: How often should metrics be reviewed? Who should review them? What actions should they trigger?
  • Audit Testing: How would you design audit tests for AI controls? What sample size? What procedures?
  • Dashboard Design: Draft a monitoring dashboard for an AI system in your function. What metrics would you include? What should the threshold be for each metric?

Detailed Examples

The following examples illustrate how the concepts from this lesson play out in real-world oversight scenarios. Each example is designed to help you recognize similar situations in your own work and respond with appropriate professional judgment.

Example 1: Comprehensive Control Testing Audit testing plan for AI-assisted controls: - Input validation testing: Verify data quality - Configuration testing: Verify model rules and parameters - Output review testing: Verify humans are actually reviewing outputs - Performance testing: Verify the AI is accurate - Bias testing: Verify no systematic bias

Testing is integrated into regular audit activities; is performed quarterly or annually; results inform audit conclusions.

Example 2: Weak Monitoring (Anti-Pattern) Compliance team has no monitoring of transaction screening AI: - No real-time alerts if system fails (someone notices 3 days later when transactions haven't been screened) - No tracking of flag rate or SAR rate (no one knows if performance is changing) - No bias analysis (no one checks whether model is biased) - Problems are discovered late, if at all

Result: Poor visibility into AI performance; issues are not detected; organization is surprised when something goes wrong.

Putting It Into Practice

Workflow integration requires systematic thinking about how these concepts fit into broader organizational processes:

  • Design with controls in mind: When integrating AI into workflows, build verification checkpoints and quality controls into the process from the start -- not as afterthoughts.
  • Measure effectiveness: Establish metrics that track both the efficiency gains from AI integration and the quality of AI-assisted outputs over time.
  • Train and support others: As you integrate AI into team workflows, ensure that all team members understand the controls, verification requirements, and escalation procedures.
  • Iterate based on evidence: Use data from your monitoring processes to continuously improve AI-integrated workflows. What works well? Where do errors occur? How can controls be strengthened?

Key Takeaways

  • Regular testing: Build AI control testing into annual audit plan; test quarterly or annually
  • Meaningful monitoring: Track metrics that matter; review regularly; act on findings
  • Baseline establishment: Know what "normal" looks like before you can detect abnormal
  • Drift detection: Monitor performance over time; detect and address degradation early
  • Bias assessment: Include bias testing and monitoring in your plan
  • Documentation: Maintain audit trail of testing and monitoring results
  • Continuous improvement: Use testing and monitoring findings to improve controls and AI systems

Chapter Summary

In this chapter, you learned:

  • Framework adaptation: How to extend COSO, COBIT, and other frameworks to address AI-specific risks
  • Control design: Specific controls for AI inputs, processing, and outputs
  • Testing approach: How to test AI controls as part of regular audit activities
  • Monitoring: How to establish ongoing monitoring of AI performance and control operation

Together, these elements create a control framework that governs AI-assisted processes and demonstrates their safety and effectiveness to auditors and regulators.


Glossary / Key Terms

Control: A process, system, or procedure designed to prevent, detect, or correct a problem

Control activity: A specific action taken to mitigate a risk (e.g., data validation, peer review)

Detective control: A control that identifies problems after they occur (e.g., audit review)

Drift: Degradation of AI model performance over time, often due to changes in data or business conditions

False negative: A failure to detect a problem that exists (e.g., missing a suspicious transaction)

False positive: Incorrectly flagging something as a problem when it is not (e.g., flagging a legitimate transaction as suspicious)

Model validation: Testing an AI model before deployment to ensure it performs well and is not biased

Preventive control: A control that stops a problem before it occurs (e.g., data validation before processing)

Root cause analysis: Investigation to determine why a control failed or problem occurred

Segregation of duties: Distributing control responsibilities so that no one person can commit and conceal an error


Links to Related Lessons

  • Chapter 1: "Designing AI-Integrated Oversight Workflows" (framework for integrating AI into processes)
  • Chapter 3: "Governance Reporting with AI Support" (reporting on control framework to oversight bodies)
  • Chapter 4: "Continuous Monitoring and AI-Enhanced Surveillance" (specific application of control framework to monitoring)
  • L3: "Assessment and Governance" (foundational governance concepts)

As you continue through this credential program, you will build on the foundation established in this lesson. Each subsequent lesson adds new dimensions to your understanding and expands your capability to work effectively with AI in oversight roles.