AI for Risk, Compliance & Audit
Proficient · M4 · lesson 4 of 26 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Assessing Completeness, Accuracy, and Relevance of AI Outputs
📖
now learning

Assessing Completeness, Accuracy, and Relevance of AI Outputs

15 min

Introduction

To develop systematic frameworks for assessing the three critical dimensions of AI output quality: Is it complete? Is it accurate? Is it relevant to the audit objective? Each dimension requires different assessment approaches and professional judgment.

At the Independent Application level, you are expected to apply AI tools and techniques without direct supervision in routine scenarios. You should be able to independently assess AI output quality, identify when outputs require additional review, and produce work products that meet professional standards with AI assistance.

This lesson is designed to be accessible to professionals at all experience levels while providing the depth needed for practical application. Whether you are encountering these concepts for the first time or building on existing knowledge, the material ahead will strengthen your ability to navigate AI governance challenges with confidence and competence.

Core Concepts

Practical Use Cases

Use Case 1: Assessing Completeness of AI-Generated Risk Assessment Scenario: AI generated a regulatory risk assessment for a new product line. The assessment identified 18 regulatory requirements. You need to assess whether this is complete.

Completeness assessment approach:

  • Expected scope definition:
  • - What regulatory requirements should apply to this product line?
  • - Which regulators have authority? (Federal, state, industry-specific?)
  • - Which product attributes drive regulatory requirements? (Value, customer type, transaction type)
  • Comparison to AI output:
  • - AI identified 18 requirements across [X] regulators
  • - You expected approximately [Y] requirements
  • - If 18 is close to Y: likely complete. If 18 is significantly below Y: potential gaps.
  • Gap analysis:
  • - Identify any expected regulatory areas that AI did not address
  • - For each gap, determine: Is it genuinely missing, or did AI address it under different terminology?
  • - Interview regulatory and compliance experts: "Looking at this list of 18, are we missing anything important?"
  • Validation against benchmarks:
  • - Compare AI's list to regulatory guidance from agencies, industry associations, or peer organizations
  • - If peer organizations identified additional requirements: Investigate whether those are relevant to your product
  • Root cause of gaps (if any):
  • - Did AI have access to all relevant regulatory sources?
  • - Are there recent regulatory updates that AI sources may not have captured?
  • - Are there organization-specific factors (exemptions, product modifications) that eliminate some requirements?

Outcome: Assess completeness as adequate, adequate with noted limitations, or incomplete. Document any gaps and their implications for the audit plan.


Use Case 2: Assessing Accuracy of AI-Identified Control Exceptions Scenario: AI identified 42 control exceptions in a transaction population. You need to assess whether these are actually control exceptions or whether AI is making errors.

Accuracy assessment approach:

  • Validation sample:
  • - Select 20 exceptions (random sample from the 42)
  • - For each, verify: Is this genuinely a control exception, or did AI make an error?
  • - Confirm exception is not a data quality issue, documented override, or misinterpretation
  • Rate analysis:
  • - If 20/20 are confirmed exceptions: Confidence is very high. Likely the remaining 22 are accurate.
  • - If 18/20 are confirmed: 90% accuracy rate. Significant but not perfect. Recommend validating remaining exceptions or noting accuracy limitation.
  • - If 15/20 are confirmed: 75% accuracy rate. Meaningful error rate. Stop and investigate what is causing errors (AI logic problem, data quality issue, misunderstanding of control).
  • Error pattern analysis:
  • - Of the exceptions that proved inaccurate, what caused the error?
  • - Are errors random or patterned? (Same vendor type, same approver, same exception category?)
  • - If patterned: May indicate systematic AI error (logic issue)
  • - If random: May indicate data quality issues or AI limitations
  • Materiality of inaccuracies:
  • - If AI flagged 42 exceptions but 8 are false positives, does it matter?
  • - Depends on: How are you using the exceptions? Are you investigating all 42, or using 42 as a population rate?
  • - For a finding: False positives are problematic and should be removed before reporting
  • - For a population rate estimate: False positives inflate the error rate but can be corrected using validation results

Outcome: Assess accuracy as high, acceptable, or concerning. If concerning, investigate root cause before relying on AI results.


Anti-patterns / Misuse Risks

Anti-pattern 1: Assuming Completeness Without Validation Risk: AI generates a comprehensive-looking analysis, and you assume it is complete without actively looking for gaps.

Why it fails: - "Comprehensive-looking" is not the same as complete - AI may have gaps in data access, scope definition, or knowledge - You may be missing material risks or issues - You have not exercised professional judgment about completeness

Example of misuse: "AI provided a detailed risk assessment. Assuming it is complete."

Better practice: "AI provided a risk assessment covering [X areas]. Validated completeness by (1) comparing to prior assessments, (2) interviewing risk owners about gaps, (3) benchmarking against peer organizations. Found [Y] gaps; recommend adding [Z]."


Anti-pattern 2: Assuming Accuracy Because Output Looks Reasonable Risk: AI output appears reasonable and is well-formatted, so you assume it is accurate without validation.

Why it fails: - Formatting and reasonableness of presentation are not the same as accuracy - AI can produce plausible-sounding but inaccurate conclusions - You have not validated facts or logic - You may be basing decisions on inaccurate information

Example of misuse: "Output looks professional and well-organized. Assuming it is accurate."

Better practice: "Reviewed output for reasonableness. Validated accuracy by [sampling approach]. Confirmed [X]% of findings are accurate; investigated [Y] discrepancies."


Anti-pattern 3: Accepting Relevance Without Questioning Context Risk: AI generated analysis that is relevant to a generic audit objective, but you did not confirm that it is relevant to your specific organization and context.

Why it fails: - Generic analysis may not fit your specific risk profile, control environment, or organizational context - You may be basing conclusions on analysis that does not apply to your situation - You have not validated that the analysis accounts for organization-specific factors

Example of misuse: "AI generated a compliance assessment for financial services firms. We are a financial services firm, so the assessment is relevant."

Better practice: "AI generated an assessment for financial services firms. Validated relevance by (1) confirming our product line is covered, (2) checking whether organization-specific exemptions or factors are considered, (3) assessing whether assessment scope matches our audit scope. Found [relevance issues]; recommend [modifications]."


[Practical Tip]

As you work through these concepts, consider how each one applies to your current role. Think of a specific scenario from your recent work where this concept would have been relevant. Building these mental connections between theory and practice is the fastest way to internalize new knowledge and make it actionable in your daily responsibilities.

Human Judgment Checkpoints

Checkpoint 1: Completeness Reality-Check Before accepting AI output as complete: - Based on my knowledge of this area, what should be included? - Is anything obviously missing? - Have I asked subject matter experts whether the output is complete? - Would I expand the list if I were doing this analysis manually?

Checkpoint 2: Accuracy Validation For material AI conclusions: - Can I verify key facts against primary sources? - Have I tested logic with examples? - Is there a reasonable accuracy validation sample I could perform? - Would a peer expert agree with AI's characterizations?

Checkpoint 3: Relevance Confirmation Before relying on AI output: - Does this address my specific audit question? - Are there organization-specific factors that change the relevance of AI conclusions? - Would this analysis look the same if conducted for a different organization? - Are there context assumptions in the analysis that I should validate?

Checkpoint 4: Proportionality Consider whether your assessment effort is proportionate: - How much effort should I invest in validating AI output? - Is this a critical decision where high confidence is essential, or a supporting analysis? - Is the cost/benefit of detailed validation justified? - Can I accept some uncertainty about completeness/accuracy, or must I validate thoroughly?

Traceability / Defensibility Considerations

Document Assessment Approach: For significant AI outputs, document: - How you assessed completeness (what comparison or validation was performed) - How you assessed accuracy (what validation sample was tested) - How you assessed relevance (what context or organization-specific factors were considered) - What gaps, errors, or limitations you identified - What modifications you made based on assessment - Your overall conclusion about the usability and reliability of the AI output

Maintain Assessment Evidence: Keep: - List of expected completeness areas compared to AI output - Documentation of gap analysis - Validation sample with confirmation results - Any subject matter expert input on accuracy or completeness - Notes on relevance assessment and context assumptions

Qualify Output Appropriately: When presenting AI-assisted analysis: - State the assessment you performed - Disclose any gaps, limitations, or areas where accuracy is uncertain - Explain any modifications you made - Communicate your confidence level in the output

[Practical Tip]

As you work through these concepts, consider how each one applies to your current role. Think of a specific scenario from your recent work where this concept would have been relevant. Building these mental connections between theory and practice is the fastest way to internalize new knowledge and make it actionable in your daily responsibilities.

Responsible AI and Control Considerations

Bias in Completeness Assessment: You may unconsciously assess completeness in a biased way: - Accepting completeness because you expect AI to be comprehensive - Rejecting completeness because you distrust AI - Over-emphasizing gaps that align with your prior concerns - Missing gaps because you are satisfied with what you see

Mitigate by: - Using explicit comparison frameworks (prior assessments, benchmarks) rather than intuition - Actively looking for both strengths and gaps in AI output - Documenting your assessment logic - Seeking independent input on completeness from subject matter experts

Fairness in Accuracy Assessment: Ensure your accuracy validation is fair: - Validate a random sample, not just areas you suspect are wrong - Acknowledge AI strengths in areas where accuracy is high - Investigate root causes of errors, not just count them - Communicate results fairly to others

Practice / Reflection Prompts

  • Completeness Approach: If you received an AI-generated risk assessment, how would you assess whether it is complete? What comparison points would you use?
  • Accuracy Validation: Design a validation approach for an AI analysis in your audit area. What would constitute adequate validation?
  • Relevance Assessment: Think of a generic audit analysis you have seen. How would you assess whether it is relevant to your specific organization?
  • Documentation: Draft assessment documentation that you would maintain for an AI-supported analysis.
  • Communication: How would you communicate to a peer or supervisor the completeness, accuracy, and relevance limitations of an AI analysis?

[Practical Tip]

As you work through these concepts, consider how each one applies to your current role. Think of a specific scenario from your recent work where this concept would have been relevant. Building these mental connections between theory and practice is the fastest way to internalize new knowledge and make it actionable in your daily responsibilities.

Glossary / Terms

  • Completeness: Whether an analysis or output covers all relevant categories, risks, or items expected within its scope.
  • Accuracy: Whether the facts, interpretations, and logic in an analysis are correct and free of errors.
  • Relevance: Whether an analysis or output addresses the specific audit objective and is applicable to the organization's context.
  • Validation sample: A representative sample of output items tested to assess accuracy or reliability of the overall output.

Related Lessons

  • Lesson 1: Advanced Critical Review: Beyond Basic Verification (broader critical review framework)
  • Lesson 3: Peer Review and Quality Assurance for AI-Assisted Work Products (collaborative review)
  • Chapter 2, Lesson 3: Maintaining Testing Rigor with AI Assistance (rigor in testing with AI)
  • Chapter 4, Lesson 1: What Makes an AI-Assisted Work Product Defensible (defensibility standards)

Detailed Examples

The following examples illustrate how the concepts from this lesson play out in real-world oversight scenarios. Each example is designed to help you recognize similar situations in your own work and respond with appropriate professional judgment.

Example 1: Completeness Gap Discovery

AI Output: Risk assessment identified 12 compliance risks: regulatory reporting, data privacy, KYC, sanctions, AML, market conduct, product governance, conflicts of interest, operational resilience, information security, outsourcing, business continuity.

Completeness Assessment: - Reviewed audit history and prior risk assessments for this business line - Prior assessments identified 15 risk areas - Comparing: AI list includes most prior areas but is missing: "consumer protection," "record-keeping," "disclosure standards" - Investigation: Are these genuinely missing, or are they embedded in other categories? - Finding: "Consumer protection" is embedded in "product governance." "Record-keeping" and "disclosure standards" are separate risks not addressed in AI output. - Outcome: AI output is incomplete. Recommend adding consumer protection, record-keeping, and disclosure requirements to assessment.


Example 2: Accuracy Validation

AI Output: "Identified 156 travel expense violations: meals >$75, hotel >$250/night, taxi >$40."

Accuracy Assessment - Validation Sample: - Selected 20 flagged expenses randomly - Reviewed each with documentation and policy - Finding: 18 violations confirmed, 2 appear to be false positives - Violation 1: Meal flagged as $76, but receipt shows $75.50 (rounded to $76 in system, actually under limit) - Violation 2: Hotel flagged as $252, but is a destination with limited options and documented exception was approved - Accuracy rate: 90% (18/20 confirmed) - Data quality issue: System rounding is causing false positives on borderline expenses - Documented exception: There is an exception approval system, but AI did not flag whether exceptions were documented

Outcome: Accuracy is adequate but not perfect. Recommend: (1) clarify rounding/threshold logic with AI, (2) validate all flagged exceptions that are close to threshold, (3) confirm AI can access exception approval records to reduce false positives.


Putting It Into Practice

Independent application requires a disciplined approach to integrating these concepts into your workflow:

  • Establish personal standards: Define your own quality criteria for AI-assisted work products. What level of verification satisfies you professionally? Document these standards and apply them consistently.
  • Build verification routines: Create repeatable processes for checking AI outputs against source materials, professional standards, and organizational requirements.
  • Exercise professional judgment: Identify situations where AI assistance is appropriate and where human judgment must prevail. This discernment is the hallmark of Level 3 competence.
  • Contribute to organizational learning: Share your experiences -- both successes and challenges -- with your team. Your practical insights help improve AI governance for everyone.

Key Takeaways

  • Assess completeness by comparing AI output to expected categories and validating with subject matter experts. Gaps often indicate areas where AI data access or scope was limited.
  • Assess accuracy through validation sampling. Even small sample validation (e.g., 20 of 500 items) can indicate reliability of AI output.
  • Assess relevance by confirming scope alignment and validating organization-specific context assumptions. Generic analysis may not apply to your situation.
  • Document your assessment approach and findings. Transparent communication about limitations and modifications builds confidence in AI-assisted work.
  • Invest assessment effort proportionate to the materiality of the AI contribution and the criticality of the decision.

As you continue through this credential program, you will build on the foundation established in this lesson. Each subsequent lesson adds new dimensions to your understanding and expands your capability to work effectively with AI in oversight roles.