AI in Control Testing
Introduction
To clarify what AI can and cannot do in control testing, develop the capability to design effective control testing plans that leverage AI for efficiency and coverage while maintaining audit rigor and professional judgment, and understand the boundaries of AI in evidence gathering and conclusion-making.
At the Independent Application level, you are expected to apply AI tools and techniques without direct supervision in routine scenarios. You should be able to independently assess AI output quality, identify when outputs require additional review, and produce work products that meet professional standards with AI assistance.
This lesson is designed to be accessible to professionals at all experience levels while providing the depth needed for practical application. Whether you are encountering these concepts for the first time or building on existing knowledge, the material ahead will strengthen your ability to navigate AI governance challenges with confidence and competence.
Core Concepts
Practical Use Cases
Use Case 1: High-Volume Authorization Control Testing Scenario: An organization processes 50,000+ monthly invoices. All invoices must be approved before payment. Approval authority depends on invoice amount and vendor type. There is an incentive to test this control (if it operates effectively, substantive testing of invoices is minimized). However, testing 50,000 invoices manually is infeasible.
Traditional approach: Sample 50-100 invoices, test that each was approved by appropriate authority. Results in 99.5% confidence in control operating for sample, but low visibility into full population.
AI-supported approach: 1. Export approval data for all 50,000 invoices: vendor ID, invoice amount, approved authority level, approval date 2. Configure AI to compare each invoice against the approval matrix (required authority level based on amount and vendor type) 3. AI identifies all invoices approved by non-authorized authority (either too low or inappropriate type) 4. Test findings: - If 100% compliant: Control operates effectively. Substantive testing of invoice accuracy can be minimal. - If X% non-compliant: Investigate root causes and make materiality judgment about whether control is effective enough to rely on.
Advantages: - Full population visibility (not just sample) - Objective, verifiable testing - More efficient than sample-based testing for large populations - Identifies systematic patterns (e.g., specific approver, specific vendor type consistently out of compliance)
Limitations: - Cannot assess whether approval authority exercised appropriate judgment (only whether they approved at all) - Does not verify that the invoice was reviewed for reasonableness, not just authority-checked - Cannot detect fraud if the perpetrator has appropriate authority (e.g., authorized approver approves fictitious invoices) - Results in audit reliance only for the specific control tested (approval authority); other invoice controls still require testing
Use Case 2: Segregation of Duties Control Testing Scenario: Your organization has a segregation of duties control: no individual should be able to both create a purchase order and approve payment for that PO (conflict of interest). The system should prevent this, but you want to verify that the system restriction is actually operating.
Traditional approach: Sample transactions, verify that the same individual did not both create and approve.
AI-supported approach: 1. Export transaction data showing who created the PO and who approved payment 2. AI compares creator and approver for all transactions 3. Flags any transaction where the same individual performed both functions 4. Test results indicate whether the control is operating system-wide
Advantages: - Full population testing (100% of transactions covered) - Objective, verifiable result - Identifies any exceptions to segregation of duties across the full period
Limitations: - Cannot assess whether there is compensating control (e.g., management review of all transactions by the same individual) - May flag exceptions that are approved deviations (e.g., a specific transaction was deliberately approved with waiver) - Does not verify that the segregation is effective in preventing fraud; just that it is technically operating
Anti-patterns / Misuse Risks
Anti-pattern 1: Testing Volume Over Judgment Risk: You use AI to test a high-volume, objective control and conclude the control is effective without testing design adequacy or understanding the broader control environment.
Why it fails: - A control that operates consistently but is poorly designed is not an effective control - You may be testing that a system restriction works, but the system may have been misconfigured at the outset - Full population testing of a poorly designed control does not ensure control effectiveness
Example of misuse: "AI tested all 50,000 approval transactions and found 100% compliance with authority levels. Control is effective."
Better practice: "AI tested all 50,000 approval transactions and found 100% compliance with authority levels. Separately, we assessed control design by reviewing the approval matrix, confirming it aligns with organizational risk tolerance, and validating that system configuration enforces the matrix correctly. Overall, the authorization control is effective in design and operation."
Anti-pattern 2: Confusing Population Testing With Judgment Testing Risk: You use AI to test a high-volume, objective control and assume that you have tested all aspects of that control.
Why it fails: - Testing that approvals occurred does not test whether approvals were made with appropriate judgment - You may be testing the mechanics while missing the intent - You may have false confidence in control effectiveness
Example of misuse: "AI tested all approval transactions. Control is effective."
Better practice: "AI tested that all approval transactions were approved by appropriate authority levels (population testing = 100% compliance). We also tested a sample of approvals to assess whether approvers exercised appropriate judgment in assessing reasonableness of the transaction. Results indicate [XY]."
Anti-pattern 3: Treating AI Anomalies as Complete Control Testing Risk: AI identifies anomalies in transaction processing, you investigate a few, and you conclude control testing is complete.
Why it fails: - Anomalies are not a comprehensive test of control operation - You may miss patterns that are not statistical outliers - You have not tested the control as a whole, just specific exceptions
Example of misuse: "AI identified 47 anomalies in the transaction population. We investigated these. Control is effective."
Better practice: "AI identified 47 anomalies and we investigated them to identify root causes. We also conducted systematic testing of [specific control objective] across [population] using [methodology]. Based on these combined results, we concluded the control is effective [with noted exception XY]."
[Practical Tip]
As you work through these concepts, consider how each one applies to your current role. Think of a specific scenario from your recent work where this concept would have been relevant. Building these mental connections between theory and practice is the fastest way to internalize new knowledge and make it actionable in your daily responsibilities.
Human Judgment Checkpoints
Checkpoint 1: Control Design Assessment Before you test a control operationally, confirm: - Is this control appropriately designed to address the risk? - Is the control objective clearly documented? - Does the control depend on judgment, or is it objective and rule-based? - Would a higher risk profile warrant a different control design?
Checkpoint 2: AI Testing Scope When designing AI-supported control testing: - What aspect of control operation can AI test objectively? (e.g., compliance with a rule) - What aspects require professional judgment or direct observation? - Have I designed testing that covers both objective compliance and judgment-based aspects? - Is my AI testing at population level sufficient, or do I also need sample-based testing?
Checkpoint 3: Interpretation of Results When reviewing AI control testing results: - Do the results indicate that the control operates as designed? - Have I investigated any exceptions to understand root causes? - Do the results support my conclusion that the control is effective? - Is there any evidence that the control was circumvented or that users do not understand the control?
Checkpoint 4: Control Reliance Decision Before you reduce substantive testing based on AI-supported control testing: - Am I comfortable relying on this control based on the testing? - Do I understand the control environment and any broader concerns about management override? - Have I tested both design and operation? - Is my conclusion defensible if questioned?
Traceability / Defensibility Considerations
Document Control Testing Scope and Limitations: Your control testing documentation should clearly state: - What control was tested and what control objective was being assessed - How AI supported the testing (e.g., "AI analyzed all [X] transactions to assess compliance with authorization levels") - What aspects of the control were tested by AI and what aspects were tested through other procedures - What limitations exist in the testing (e.g., "AI testing confirmed compliance with authority rules but did not assess judgment in approvals")
Separate Design and Operating Testing: Clearly distinguish: - Design testing: Assessed that the control is designed to address [risk] through [mechanism] - Operating testing: Tested that the control operated as designed across [population] using [methodology] - Overall conclusion: Based on design and operating testing, the control is effective/has exceptions
Maintain Population Testing Documentation: For AI-supported population testing: - Document the population tested (date range, transaction types, volume) - Explain the AI testing logic (what criteria were used to identify exceptions) - Document results (number of exceptions, percentage, patterns) - Explain root causes of exceptions - State your conclusion about control effectiveness
[Practical Tip]
As you work through these concepts, consider how each one applies to your current role. Think of a specific scenario from your recent work where this concept would have been relevant. Building these mental connections between theory and practice is the fastest way to internalize new knowledge and make it actionable in your daily responsibilities.
Responsible AI and Control Considerations
Bias in Anomaly Detection: AI-based anomaly detection may over-identify or under-identify control issues in certain transaction types, user groups, or time periods due to data patterns. Mitigate by: - Validating that AI's anomaly-flagging logic makes sense for your control objectives - Confirming that detected patterns are actually control issues, not just statistical variations - Supplementing AI analysis with professional judgment about where control failures are most likely to occur
Automation Bias: There is a risk of over-relying on AI testing results without questioning them. Mitigate by: - Always validating a sample of AI-identified exceptions to confirm they are genuine - Investigating root causes of any exceptions, not just tallying them - Maintaining healthy skepticism about AI results
False Confidence in Negative Results: If AI finds 0 exceptions in a large population, this can provide false confidence. Mitigate by: - Validating that the AI testing logic is sound and would catch control failures if they existed - Performing a small manual validation sample to ensure AI testing is working correctly - Considering whether the absence of exceptions is realistic given historical patterns
Practice / Reflection Prompts
- Current Testing: Identify a control you currently test. Is it objective and rule-based, or does it depend on judgment? Would AI-supported population testing be appropriate?
- Testing Design: If you were to use AI for control testing, what would you test (objective compliance)? What would you test separately using judgment-based testing?
- Boundaries: Describe the boundaries between what AI should test and what you, as an auditor, must assess through professional judgment.
- Documentation: Outline how you would document control testing that integrated AI with traditional procedures.
- Risk Scenarios: What control failures do you think AI-based testing might miss? How would you design testing to catch those?
[Practical Tip]
As you work through these concepts, consider how each one applies to your current role. Think of a specific scenario from your recent work where this concept would have been relevant. Building these mental connections between theory and practice is the fastest way to internalize new knowledge and make it actionable in your daily responsibilities.
Glossary / Terms
- Control testing: The process of evaluating whether a control is effectively designed and operates as intended.
- Operating effectiveness: The degree to which a control is executed consistently and achieves its control objective.
- Judgment control: A control that depends on human judgment or professional assessment (e.g., review for reasonableness) vs. rule-based or automated control.
- Population testing: Testing that covers all items in a population (100%) as opposed to sample testing.
Related Lessons
- Lesson 2: Using AI for Data Analysis in Audit and Compliance Testing (techniques for data analysis supporting control testing)
- Lesson 3: Maintaining Testing Rigor with AI Assistance (quality standards for AI-supported testing)
- Chapter 1, Lesson 2: AI-Assisted Issue Spotting (identifying control exceptions and issues)
- Chapter 3, Lesson 1: Advanced Critical Review (frameworks for reviewing control testing performed by others)
Detailed Examples
The following examples illustrate how the concepts from this lesson play out in real-world oversight scenarios. Each example is designed to help you recognize similar situations in your own work and respond with appropriate professional judgment.
Example 1: AI-Supported Control Testing -- When It Works Well
Control: All credit memos (refunds) must be approved before issuance. Approval authority depends on credit memo amount: >$50,000 requires VP approval; $10,000-$50,000 requires manager approval; <$10,000 requires supervisor approval.
Testing approach: 1. AI analyzes all 2,000 credit memos issued in the period 2. AI compares approval authority against amounts 3. AI flags any credit memo approved by lower authority than required 4. Finding: 5 credit memos approved by supervisors where manager approval was required. Root cause: delegation matrix update was not communicated to users. 5. Professional judgment: Control design is sound; execution failure was isolated and due to training/communication gap, not systemic weakness. Control is effective overall with noted improvement opportunity.
Why this works: The control is objective (authority levels are clear), operating testing is feasible at population level (system data is available), and root cause can be identified.
Example 2: AI-Supported Control Testing -- When It Has Limitations
Control: All capital projects must be reviewed for strategic alignment before approval. Finance must ensure the project aligns with strategic priorities identified in the five-year plan.
Testing approach: AI analyzes project approvals and compares against the five-year plan.
Problem: Strategic alignment is a judgment call. A project may explicitly be in the five-year plan, but the executive sponsor may have decided to deprioritize it. Or a project may not be explicitly in the plan, but is approved as a strategic pivot. AI can flag that a project was not explicitly in the plan, but cannot assess whether the lack of explicit mention represents a control violation.
Better approach: Interview project sponsors and finance leadership about how they assess strategic alignment. Test a sample of approvals to understand the decision-making process. Determine whether there is documented evidence (e.g., steering committee minutes) of strategic alignment discussions.
Lesson: AI can supplement judgment control testing but cannot replace it.
Putting It Into Practice
Independent application requires a disciplined approach to integrating these concepts into your workflow:
- Establish personal standards: Define your own quality criteria for AI-assisted work products. What level of verification satisfies you professionally? Document these standards and apply them consistently.
- Build verification routines: Create repeatable processes for checking AI outputs against source materials, professional standards, and organizational requirements.
- Exercise professional judgment: Identify situations where AI assistance is appropriate and where human judgment must prevail. This discernment is the hallmark of Level 3 competence.
- Contribute to organizational learning: Share your experiences -- both successes and challenges -- with your team. Your practical insights help improve AI governance for everyone.
Key Takeaways
- AI is well-suited for testing high-volume, objective, rule-based controls (authorization, segregation of duties, compliance with thresholds). It is less suited for testing judgment-dependent controls.
- Population-level testing with AI can provide more comprehensive coverage than sample-based testing, but does not replace professional assessment of control design and judgment.
- Control testing requires both design assessment (is the control appropriately designed?) and operating assessment (does it work?). AI can assist with operating assessment of objective controls but must be supplemented with professional review.
- Exceptions identified by AI require investigation to understand root cause and determine materiality. The number of exceptions alone does not answer the question "Is the control effective?"
- Always validate that AI testing logic makes sense for your control objectives and maintain professional skepticism about results.
As you continue through this credential program, you will build on the foundation established in this lesson. Each subsequent lesson adds new dimensions to your understanding and expands your capability to work effectively with AI in oversight roles.
Skill.re