AI for Risk, Compliance & Audit
Proficient · M23 · lesson 23 of 26 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Statistical Sampling in AI-Supported Control Testing
📖
now learning

Statistical Sampling in AI-Supported Control Testing

15 min

LECTURE TRANSCRIPT

Statistical Sampling in AI-Supported Control Testing

Level 3: Independent Application -- Chapter 2, Lesson 5

AI for Risk, Compliance, Audit & Governance Credential

Duration: ~25 minutes

Generated: March 2026


Statistical sampling is a fundamental auditing technique: test a sample of transactions and extrapolate results to the population. AI changes this practice. AI can perform control testing on much larger samples--potentially even entire populations rather than samples. But AI's involvement introduces questions: How do statistical concepts apply when AI is performing the testing? How do you adjust sample sizes when AI is involved? What confidence can you place in AI-generated test results? This lesson focuses on applying statistical sampling methodologies to AI-supported control testing and understanding how to adjust sample design and interpretation when AI is involved.


STATISTICAL SAMPLING FUNDAMENTALS

Statistical sampling enables auditors to make inferences about populations based on samples. Key concepts include:

Population: The complete set of items you want to draw conclusions about. For a control test, the population might be "all purchase transactions in a month" or "all account reconciliations performed."

Sample: A subset of the population selected for testing.

Sample Size: How many items from the population are selected. Larger samples provide more confidence but cost more to test.

Confidence Level: The probability that the true population error rate falls within the estimated range. Common confidence levels are 90%, 95%, or 99%.

Tolerable Error Rate: The maximum error rate that is acceptable for the control to be considered effective. If you can tolerate a 5% error rate, that is your threshold.

Expected Error Rate: Based on prior knowledge, what error rate do you expect? This affects sample size calculation.

Precision: The range around the sample result. If you test a sample and find a 2% error rate with precision of 1%, you conclude that the true error rate is between 1% and 3%.

These concepts apply in AI-supported testing, though the dynamics change when AI conducts testing rather than auditors.


HOW AI CHANGES STATISTICAL SAMPLING

When AI performs control testing, several traditional assumptions change.

Sample Size Implications: AI can test much larger samples at lower cost than manual testing. Instead of testing 30 transactions, AI can test 3,000. What does larger sample size mean for statistical inference?

Larger samples provide tighter confidence intervals. If you manually test 30 items and find 2 errors (6.7%), the 95% confidence interval around that result is wide--perhaps 1% to 15%. If AI tests 3,000 items and finds 200 errors (6.7%), the 95% confidence interval is narrow--perhaps 5.8% to 7.6%. Larger samples reduce uncertainty. This is good--it provides more precise estimates.

However, larger samples make it harder to meet "tolerable error rate" criteria. If you tolerate a 5% error rate and AI's precise estimate is 6.7%, you miss the tolerance by a smaller margin than if your estimate had wide confidence intervals. Precision cuts both ways.

Consistency and Reproducibility: AI should test transactions consistently. Every transaction is evaluated by the same logic. Auditors' testing can be inconsistent--one auditor might interpret a control differently than another. AI's consistency enables more reliable inference.

Bias and Fairness Concerns: If AI is biased (performs differently for different populations), statistical sampling can mask the bias. If AI performs at 5% error on most transactions but 20% error on transactions from a particular business unit, a sample-based estimate might miss the pattern. Awareness of potential AI bias affects how you interpret sample results.

Data Quality Dependence: AI's results are only as good as the data it tests. If input data is incomplete or inaccurate, AI testing results are unreliable. Traditional manual testing shares this risk but auditors' direct examination of transactions sometimes reveals data quality issues that automated testing might miss.


SAMPLE SIZE DETERMINATION FOR AI-SUPPORTED TESTING

How do you determine appropriate sample size when AI performs testing?

Traditional Approach: Traditional auditing provides formulas. Sample size = (reliability factor x population size) / (tolerable error rate x expected error rate). Using these formulas:

  • Testing 1,000 transactions in population
  • Reliability factor of 3 (95% confidence)
  • Tolerable error rate of 5%
  • Expected error rate of 1%

Sample size = (3 x 1,000) / (5% x 1%) = 60,000 / 0.05 = 600 transactions

This formula suggests testing 600 items. But this is testing more than half the population. In this scenario, testing the entire population (1,000 items) might make sense--you gain precision at minimal additional cost because AI can test all items efficiently.

Risk-Based Adjustment: When using AI, adjust sample size based on:

  • Cost of Testing: If AI can test items nearly free, test more items or the entire population.
  • Risk of Significant Error: For high-risk controls, use larger samples for greater confidence.
  • Expected Reliability: If AI is expected to be highly reliable, smaller samples may be adequate. If AI reliability is uncertain, use larger samples.
  • Data Quality: If data quality is poor, AI testing is less reliable; larger samples don't compensate for poor data quality.

Population Coverage: Consider whether to test the entire population rather than a sample. If population size is reasonable and AI can test all items at low cost, 100% testing eliminates sampling risk. You know the true population error rate, not an estimate. This shift from sample-based to population-based inference is significant.


PRECISION AND CONFIDENCE INTERVALS

When you test a sample and find results, the statistical question is: What does the sample tell us about the population?

Precision Calculation: If you test a random sample and find a 5% error rate, the true population error rate is not exactly 5%. It is probably in a range around 5%. The width of that range is precision.

Confidence intervals establish the range. "The sample error rate is 5%; we are 95% confident that the true population error rate is between 3% and 7%." The 3%-7% range is the 95% confidence interval.

Precision improves with larger sample sizes. A sample of 30 items might yield 95% confidence interval of 1%-10% (wide). A sample of 300 items might yield 95% confidence interval of 4%-6% (narrow). AI's ability to test large samples dramatically improves precision.

Decision Implications: Compare the confidence interval to your tolerable error rate. If you tolerate a 5% error rate and the confidence interval is 3%-7%, the true error rate might exceed tolerance (if it is 6%) or might be acceptable (if it is 4%). You cannot be certain based on the sample. If confidence intervals overlap the tolerance boundary, you may need larger samples or additional testing to reduce uncertainty.


ADDRESSING AI BIAS IN SAMPLING INTERPRETATION

AI bias creates special challenges for statistical sampling.

Stratified Sampling: If you suspect AI might perform differently for different populations, use stratified sampling. Divide the population into strata (e.g., by transaction type, by business unit, by size). Test each stratum separately. This enables you to detect whether AI performs differently for different strata.

AI Performance Testing: Before relying on AI for control testing, test AI's performance on a known set of items where you can verify AI results. If AI is supposed to identify control exceptions, test AI on a sample where you know which items are exceptions. Does AI correctly identify them? This validation testing establishes AI reliability before relying on AI sampling results.

Comparative Testing: For important controls, use AI to test a large sample and use manual testing to verify a subsample. Compare AI results to manual results. If AI and manual results align, you gain confidence in AI. If they diverge, investigate why.


DOCUMENTING STATISTICAL SAMPLING WITH AI

When you use statistical sampling with AI-supported testing, document your approach.

Population Definition: Document what constitutes the population. "Population: All purchase requisitions submitted in March 2026; N = 2,847."

Sample Selection Method: Document how the sample was selected. "Random sample of 500 items selected by AI system using random number generation, stratified by purchase category."

Sample Size Justification: Document why you chose this sample size. "Sample size of 500 selected to achieve 95% confidence with precision of 2% around estimated error rate, adjusted for AI reliability testing which showed 98% accuracy."

AI Testing Methodology: Document what AI testing involved. "AI system evaluated each transaction against control rule: 'All purchases over $50,000 require documented three-bid approval.' AI extracted approval documentation from procurement system and flagged transactions missing required approvals."

Results: Document results clearly. "Sample of 500: 23 exceptions identified (4.6% error rate). 95% confidence interval: 2.8% - 6.4%."

Conclusion: Document your conclusion about population control effectiveness. "Sample results indicate control exception rate within tolerable range of 5%; control is considered operating effectively."


1. AI TESTING WITHOUT SAMPLE SIZE ADJUSTMENT

Using traditional sample size formulas without adjusting for AI's ability to test large volumes. You continue testing samples of 30-50 items even though AI could efficiently test 3,000 items. Failure to leverage AI's ability to test large samples means missing the precision advantage AI provides.

2. BLIND RELIANCE ON AI RESULTS

Accepting AI testing results without validating AI reliability. You assume AI is accurate and extrapolate results to population. But AI has undetected bias or systematic errors. Sampling conclusions are wrong. Address by validating AI performance before relying on it for sampling.

3. IGNORING AI BIAS RISK

Testing a sample of 5,000 items with AI without considering whether AI might perform differently for different populations. If AI is biased, sample-based inference is unreliable. Address by using stratified sampling and testing AI performance across strata.

4. INSUFFICIENT DOCUMENTATION

Not documenting your sampling approach when AI is involved. Future reviewers cannot understand your methodology or verify your conclusions. Address by documenting population, sample selection, sample size justification, AI testing methodology, and conclusions.


PRACTICE PROMPTS

  1. You want to test a control over approval of expense reports. Population is 10,000 expense reports in a month. Tolerable error rate is 3%. Expected error rate is 1%. Using traditional sample size formulas, what sample size would you use? How would this change if AI conducts testing?
  2. Design an AI-supported sampling approach for testing a control. What would the population be? How would you select the sample? What sample size would you use? How would you validate AI results?
  3. You test a sample of 1,000 items using AI and find a 2.5% error rate. Calculate the 95% confidence interval. If your tolerable error rate is 2%, what conclusion would you reach?
  4. Design a stratified sampling approach for testing a control where you suspect AI might perform differently for different transaction types.

KEY TAKEAWAYS

  1. Statistical sampling enables inferences about populations based on samples; traditional concepts apply to AI-supported testing but dynamics change when AI conducts testing.
  2. AI's ability to test large samples at low cost enables shift from sample-based to potentially population-based testing, dramatically improving precision of estimates.
  3. Sample size should be adjusted for AI-supported testing, considering cost of testing, risk of control failure, AI reliability, and data quality.
  4. AI bias risk requires consideration of stratified sampling and validation of AI performance before relying on AI testing results.
  5. Documentation of sampling methodology--population, sample selection, sample size justification, AI testing approach, and conclusions--is essential.

GLOSSARY

Confidence Interval: A range around a sample estimate within which the true population value is likely to fall, expressed with a probability level (e.g., 95%).

Precision: The width of the confidence interval; narrower intervals indicate greater precision.

Reliable Error Rate: The maximum error rate in a control that is considered acceptable; controls exceeding this rate are considered not effective.

Sample Stratification: Dividing the population into subgroups (strata) and sampling from each stratum separately to ensure representation and enable detection of differences across strata.

Tolerable Error Rate: The maximum error rate acceptable for control conclusions; synonymous with "reliable error rate."


SYNTHESIS AND APPLICATION

The shift from sample-based to potentially population-based testing, enabled by AI, is significant. It changes the nature of audit evidence. Traditionally, auditors worked with estimates and ranges. With AI-enabled population testing, auditors can have precise knowledge of actual population error rates. This is more definitive. But it also raises the bar--if you know the actual error rate, you cannot claim your estimate is uncertain.

This shift also affects audit efficiency and audit scope. If you can test 100% of transactions at reasonable cost, what is your audit scope? Do you test only the most important transactions (risk-based testing) or do you test everything? The answer depends on your risk assessment and your audit objectives.


REFLECTION EXERCISE

  1. For a key control in your organization, how is it currently tested? Could AI-supported testing improve the audit?
  2. If you used AI to test a control on 100% of transactions, how would your conclusions change compared to traditional sample-based testing?
  3. What data quality issues exist in your organization that would affect AI-supported control testing?

CLOSING REMARKS

Statistical sampling in the AI era requires rethinking traditional approaches. AI's ability to test large volumes efficiently changes sample size decisions and enables shift toward population-based testing. Auditors who understand how to leverage this capability can provide more definitive assurance while improving audit efficiency.


End of Transcript

KEY TAKEAWAYS

  1. Statistical sampling enables inferences about populations based on samples; traditional concepts apply to AI-supported testing but dynamics change when AI conducts testing.
  2. AI's ability to test large samples at low cost enables shift from sample-based to potentially population-based testing, dramatically improving precision of estimates.
  3. Sample size should be adjusted for AI-supported testing, considering cost of testing, risk of control failure, AI reliability, and data quality.
  4. AI bias risk requires consideration of stratified sampling and validation of AI performance before relying on AI testing results.
  5. Documentation of sampling methodology--population, sample selection, sample size justification, AI testing approach, and conclusions--is essential.

GLOSSARY

Confidence Interval: A range around a sample estimate within which the true population value is likely to fall, expressed with a probability level (e.g., 95%).

Precision: The width of the confidence interval; narrower intervals indicate greater precision.

Reliable Error Rate: The maximum error rate in a control that is considered acceptable; controls exceeding this rate are considered not effective.

Sample Stratification: Dividing the population into subgroups (strata) and sampling from each stratum separately to ensure representation and enable detection of differences across strata.

Tolerable Error Rate: The maximum error rate acceptable for control conclusions; synonymous with "reliable error rate."


SYNTHESIS AND APPLICATION

The shift from sample-based to potentially population-based testing, enabled by AI, is significant. It changes the nature of audit evidence. Traditionally, auditors worked with estimates and ranges. With AI-enabled population testing, auditors can have precise knowledge of actual population error rates. This is more definitive. But it also raises the bar--if you know the actual error rate, you cannot claim your estimate is uncertain.

This shift also affects audit efficiency and audit scope. If you can test 100% of transactions at reasonable cost, what is your audit scope? Do you test only the most important transactions (risk-based testing) or do you test everything? The answer depends on your risk assessment and your audit objectives.


REFLECTION EXERCISE

  1. For a key control in your organization, how is it currently tested? Could AI-supported testing improve the audit?
  2. If you used AI to test a control on 100% of transactions, how would your conclusions change compared to traditional sample-based testing?
  3. What data quality issues exist in your organization that would affect AI-supported control testing?

CLOSING REMARKS

Statistical sampling in the AI era requires rethinking traditional approaches. AI's ability to test large volumes efficiently changes sample size decisions and enables shift toward population-based testing. Auditors who understand how to leverage this capability can provide more definitive assurance while improving audit efficiency.


End of Transcript

<?