AI for Risk, Compliance & Audit
Proficient · M5 · lesson 5 of 26 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Case Studies

15 min

Introduction

Comparative case studies are one of the most efficient ways for an experienced practitioner to internalize defensibility. Reading principles in the abstract is one thing; seeing the same engagement reach two very different outcomes depending on rigor and documentation is another. This lesson sets up four parallel scenarios drawn from common audit and compliance contexts -- control testing, regulatory risk assessment, transactional analytics, and management reporting -- and walks through each twice: first as it is often produced when AI is treated as a shortcut, and second as it should be produced when AI is treated as a tool subject to professional standards.

At Level 3 (Independent Application) you are expected to operate without supervisor babysitting on routine engagements. That independence is exactly what makes defensibility critical. You are signing the workpaper. You are answering the regulator's email. You are explaining the conclusion to the audit committee. AI-assisted output that looks fine in a draft will not stand up to the kind of probing your peers, your firm's quality reviewers, or external regulators apply once the work is in the field.

The pattern you will see repeated in every case study is consistent. Indefensible work tends to share four traits: AI logic is not validated before the population is processed, exceptions are not investigated, professional judgment is delegated to the model, and documentation describes the result without describing the method. Defensible work systematically inverts each of these: pre-test validation, exception triage, explicit professional judgment, and traceable documentation that an external reader can follow without ever speaking to the auditor.

You will also notice that the additional effort required to convert an indefensible workpaper into a defensible one is not large. It is usually a few hours of disciplined documentation and validation -- a fraction of the time the original AI-assisted analysis saved. The risk-adjusted return on that effort is enormous. The work papers that follow are written in the format you can adopt directly in your own engagements; treat them as templates as well as instructional examples.

Core Concepts

Case Study 1: Control Testing -- Defensible vs. Indefensible

Engagement context. The internal audit function is testing approval authority over the procure-to-pay cycle. The control: every purchase order above $10,000 must be approved by a director; every PO between $5,000 and $10,000 must be approved by a manager; vendors must appear on the approved-vendor master file. Population: 50,000 invoices in fiscal year 2025 totaling $847M. The auditor uses a generative AI assistant with code-execution tooling to flag potential exceptions.

Indefensible version. The auditor uploads the population, asks the model to identify any approvals that violate the matrix, and receives back a list of 12 exceptions out of 50,000 transactions. The workpaper records: 'AI analysis identified 12 exceptions of 50,000 invoices (0.024%). Control is operating effectively.' No further work is performed. The audit committee sees a green rating.

Why it fails. Six independent defects converge: (1) The AI's logic was never validated against the actual approval matrix; the auditor cannot demonstrate the model implemented the control correctly. (2) None of the 12 exceptions were inspected to confirm they are genuine; some may be false positives caused by stale vendor master records or system-of-record lag. (3) Root cause is unaddressed; the auditor does not know whether the violations indicate execution lapses, design weaknesses, or data quality artefacts. (4) Methodology is undocumented; how the prompt was constructed, what data fields were exposed to the model, and what tool calls the model issued are all absent from the workpaper. (5) Independent quality review is not evident; no second pair of eyes has confirmed the AI's logic. (6) The conclusion overstates what the evidence supports; a 0.024% rate is presented as proof of effectiveness without applying the firm's materiality threshold.

What happens under examination. A regulator or external auditor asks: 'How do you know AI was correctly configured?' The auditor has no answer. 'Did you investigate the 12 flagged items?' No. 'What is the dollar value and is it material?' Not analyzed. The conclusion collapses. The work has to be redone with proper procedures, often under time pressure, and the firm absorbs the reputational cost of having issued a flawed assertion.

Defensible version. The same auditor, working with the same model, structures the engagement differently. First, the prompt and tool calls are versioned and stored: 'Identify invoices where the approving manager's delegation limit is below the invoice amount per the approval matrix effective [date]. Output approver email, approver delegation limit, invoice amount, vendor ID, and the matching matrix row.' Second, before population analysis, the auditor manually re-performs the test on a 25-invoice subsample (15 known compliant, 10 known exceptions) and confirms AI flagging matches manual judgment 25-for-25; this pre-test validation is signed and dated. Third, all 50,000 invoices are processed; AI flags 12 exceptions; the auditor inspects all 12, contacts approvers where needed, and confirms each as a genuine policy violation rather than a master-data artefact. Fourth, root cause is established for each: 3 attributable to delegation limits not updated after a managerial promotion, 5 to system lag between approval and invoice booking, 4 to invoices coded to a cost center whose approval requirements are stricter than the originating cost center. Fifth, materiality is computed: the 12 exceptions represent $173,400, of which $94,200 exceeds the engagement's $50,000 performance materiality. Sixth, the audit conclusion is calibrated to the evidence: the control is operating with deficiencies in execution; the design is adequate; remediation is recommended for vendor master timeliness and approval delegation maintenance. Seventh, an independent reviewer signs off on pre-test validation, exception inspection, materiality assessment, and overall conclusion.

Why it stands. Under the same regulatory examination, every question has an answer grounded in workpaper evidence. The auditor can show that AI was validated, exceptions were investigated, root causes were diagnosed, materiality was applied, and judgment was exercised. The 0.024% number does not appear in isolation; it appears alongside the analysis that gives it meaning. The same engagement, the same data, the same model, but a methodology that survives scrutiny.

Case Study 2: Regulatory Risk Assessment -- Defensible vs. Indefensible

Engagement context. A new consumer financial product is launching in five jurisdictions. Compliance must produce a regulatory risk assessment before go-live. The compliance officer uses an LLM to synthesize regulatory guidance from regulators, industry associations, and peer benchmarks.

Indefensible version. The compliance officer prompts the LLM with: 'List the regulatory requirements applicable to a consumer credit product launching in [jurisdictions].' The model returns 22 requirements. The officer copies them into the assessment template, marks each 'in remediation,' and submits the deliverable. The product goes live.

Why it fails. (1) Source provenance is missing; nothing in the deliverable shows which regulatory documents the model drew from or whether they were current. (2) Applicability has not been tested; some requirements may not apply to the specific product variant launching, while others critical to the variant may have been omitted. (3) Organization-specific context is absent; the model has no awareness of the firm's existing controls, exemptions, or registrations. (4) Completeness is unverified; no comparison to prior assessments, regulator examination guidance, or peer benchmarks. (5) Interpretation is unvalidated; for each requirement, the wording in the assessment is the model's paraphrase, which may diverge subtly from the regulation's actual text. (6) Documentation is thin; the basis for each conclusion lives only inside the model's hidden state, not in the workpaper.

What happens under examination. A regulator asks: 'Where did you identify these 22 requirements?' AI synthesized them. 'Did you confirm they all apply to the product?' Assumed yes. 'Are there requirements you missed?' Unknown. 'Show me your interpretation of [specific requirement].' The model produced it. Each answer reduces credibility. The regulator concludes that the assessment is unsupported and requests a redone analysis, often with a written remediation plan and follow-up examination.

Defensible version. The compliance officer scopes the assessment first: product features, customer types, geographies, and existing registrations are documented. Regulatory domains and authoritative sources are listed (e.g., CFPB guidance, state attorney general consent orders, applicable banking-agency interpretive letters, the firm's prior assessments for adjacent products). The LLM is then used as a synthesis aid: it ingests the source documents and produces a candidate list of 30 obligations. The compliance officer reviews each candidate against the primary source: 22 confirmed applicable, 5 determined not applicable based on documented product features, and 3 added as omissions discovered during peer-benchmark comparison. For each of the final 25 obligations, the workpaper includes the regulatory citation, a verbatim quote of the operative language, the officer's professional interpretation, the impact on the product, and the current control maturity (adequate / needs enhancement / gap). Completeness is verified by cross-checking against two prior risk assessments for similar products and against the regulator's most recent examination manual. Legal counsel reviews interpretations for the five highest-risk obligations and signs off in writing. Compliance leadership signs the deliverable.

Why it stands. The deliverable now reads as a regulator-ready document. Each obligation has a citation, an interpretation, an applicability determination, a control maturity rating, and a remediation owner where needed. The model's contribution is acknowledged transparently as synthesis; the professional judgment is the compliance officer's. Under examination, the officer can produce the source documents, walk through the applicability determinations, and explain why three obligations the model missed were nonetheless captured. The product launch can proceed with confidence.

Case Study 3: Transactional Analytics -- Defensible vs. Indefensible

Engagement context. An external audit team is performing journal-entry testing on a manufacturing client. The team uses an AI-assisted analytics tool to score every general-ledger entry posted during the year against a set of fraud-risk attributes (manual entries, round-dollar amounts, reversal patterns, weekend or after-hours postings, unusual user-account combinations).

Indefensible version. The tool returns 480 entries above the model's risk threshold. The audit senior selects 25 of these for further investigation, finds nothing material, and writes: 'AI-assisted journal-entry testing identified 480 high-risk entries. Sample of 25 was investigated; no exceptions noted. Journal-entry fraud risk is concluded to be low.' The workpaper does not document the threshold, the model's scoring features, the sampling method, or the basis for stopping at 25.

Why it fails. The conclusion ('fraud risk is low') is unsupported by the work performed. A 25-of-480 sample, not statistically constructed, cannot rule out fraud in the 455 unexamined entries. The threshold may have been set in a way that excluded the most suspicious entries entirely (high specificity, low sensitivity). The scoring features are not described, so a reviewer cannot judge whether they are appropriate for this client's posting patterns. There is no acknowledgment that 480 of 1.4 million entries means 99.97% of entries were excluded with no documented basis.

Defensible version. The audit senior performs the same analytics but treats the tool's output as a starting point for sampling, not as a conclusion. The workpaper documents the risk attributes, the threshold (and the rationale for it, calibrated against last year's known issues and an industry benchmark), and the model's coverage. From the 480 high-risk entries, the senior stratifies (large dollar, manual posters, period-end timing) and selects 60 entries using a defined sampling methodology that yields a stated confidence level. Each of the 60 entries is investigated to source documentation; one is identified as a $4.2M reclassification booked the last business day of the year by a non-routine user, prompting an expanded investigation that ultimately surfaces a misstatement requiring a management adjustment. The remaining 420 high-risk entries are addressed by analytical review (trend analysis, scan for similar attributes), not assumed clean. The 1.4 million entries below the threshold are addressed by a separate completeness check: the senior pulls a haphazard sample of 25 entries from below the threshold to confirm the model's threshold is not systematically blind. Conclusion: journal-entry testing performed with [stated coverage]; one issue identified and resolved; residual risk is low based on the stated procedures.

Why it stands. Every number in the conclusion is traceable to a procedure described in the workpaper. The reviewer can recompute the sample, re-pull the threshold, and reproduce the result. The model is positioned correctly: a powerful sampling tool that requires the auditor to design the procedures around it, not a substitute for procedures.

Case Study 4: Audit Committee Reporting -- Defensible vs. Indefensible

Engagement context. The CAE prepares a quarterly audit committee deck. AI is used to summarize the underlying audit reports and to draft the narrative slides on emerging risks.

Indefensible version. The CAE feeds the underlying reports into an LLM, asks for a 'one-page summary of the key risks,' and pastes the output verbatim into the deck. The committee receives a confident-sounding summary that includes a fabricated reference to a regulatory development that has not actually occurred and an over-precise estimate ('losses prevented: $14.2M') that has no basis in the underlying reports.

Why it fails. The CAE is signing a document that contains assertions she has not verified. Hallucinated regulatory developments and fabricated quantitative estimates are not edge cases; they are common LLM failure modes when the model is asked to be concise and confident. If the committee acts on the summary -- and they do, because they trust the CAE -- the organization may make decisions on a foundation that does not exist. If a single error is detected, the CAE's credibility with the committee is durably damaged.

Defensible version. The CAE uses the LLM as a drafting partner, not a publisher. The model produces a draft summary; the CAE checks every factual assertion against the underlying audit reports; any quantitative claim must trace to a specific reported number; any reference to external developments must be confirmed against a primary source (regulator's website, news outlet of record, internal legal memo). The slide is annotated with footnotes that point to the specific report or external citation. The deck includes a brief methodology note: 'AI was used to draft summary language; all factual assertions were verified by the CAE; AI did not introduce facts not present in the underlying reports.'

Why it stands. The committee receives a summary that combines the speed of AI drafting with the assurance of professional verification. If a member asks 'where did this $14.2M come from,' the CAE can point to the slide footnote and the underlying report. If a member asks 'is this regulatory development real,' the CAE can pull the citation. The trust relationship is preserved.

Reflection across the four cases. The pattern is invariant. Indefensible workpapers describe the result; defensible workpapers describe the method that produced the result. Indefensible conclusions delegate judgment to the model; defensible conclusions show the human professional applying judgment to the model's output. Indefensible documentation hides the AI; defensible documentation acknowledges it transparently and then demonstrates the controls that made its output trustworthy. The defensible versions in each case study took 20-40% more time to produce and produced work products that survive scrutiny -- a return on investment that any reasonable risk-adjusted view would accept.

Case Study 1: Control Testing -- Defensible vs. Indefensible

INDEFENSIBLE EXAMPLE:

The Work: Auditor uses AI to test approval authority in accounts payable process. AI analyzes 50,000 invoices and flags 12 exceptions. Auditor reports: "Control is operating effectively. AI found 12 exceptions out of 50,000 invoices (0.024%), indicating strong control operation."

Why it is indefensible: 1. No pre-test validation: Did AI logic match the control requirement? Not documented. 2. No confirmation of exceptions: Are the 12 flagged items genuine exceptions or false positives? Not investigated. 3. No root cause analysis: Why did the 12 exceptions occur? Not explained. 4. No documentation of methodology: How was AI configured? Not explained. 5. No quality review: Was testing reviewed by independent person? Not evident. 6. Over-confident conclusion: 0.024% error rate is presented as proof of control effectiveness without professional context.

If questioned by regulator: - Regulator: "How do you know AI was correctly configured?" - Auditor: "I assumed it was working correctly." - Regulator: "Did you validate any of the 12 exceptions?" - Auditor: "No, I trusted AI." - Regulator: "What would this control failure represent? Is 0.024% material?" - Auditor: "I didn't think about that."

Outcome: Work is not defensible. Regulator would likely require additional testing to validate AI results and auditor's conclusions.


DEFENSIBLE EXAMPLE:

The Work: Auditor uses AI to test approval authority in accounts payable process. Before conducting full population testing:

  • Pre-test Validation:
  • - Documented: "AI testing logic: Flag all invoices where approving manager's delegation limit is below the invoice amount, per approval matrix in [system date]."
  • - Validation: Reviewed 25 invoices (15 compliant, 10 exceptions) manually and confirmed AI flagging matched manual review. Result: AI logic is correct.
  • - Documented: "Pre-test validation confirmed that AI testing logic correctly identifies approval authority exceptions."
  • Population Testing:
  • - Analyzed 50,000 invoices across [time period]
  • - AI flagged 12 exceptions
  • - Reviewed all 12 exceptions to confirm they are genuine (not data quality issues, system lags, or approved exceptions)
  • - Result: All 12 confirmed as genuine exceptions
  • Root Cause Analysis:
  • - Investigated all 12 exceptions
  • - Root causes: [Specific factors identified, e.g., "3 exceptions: delegation limits not updated after approver promotion; 5 exceptions: system lag between approval and invoice entry; 4 exceptions: invoices recorded in wrong cost center, changing the approval requirement"]
  • - Conclusion: [Assessment of whether root causes indicate control design issue or execution lapses]
  • Materiality Assessment:
  • - Exceptions represent $[amount] in [%] of total invoices tested
  • - Applied audit materiality threshold of $[X]: [X] exceptions exceed threshold; [Y] exceptions are below threshold
  • - Conclusion: Control is [effective / has deficiencies] based on [specific assessment]
  • Documentation:
  • - All of the above documented in audit workpapers
  • - Testing summarized in audit report: "We tested authorization controls in accounts payable by analyzing 100% of [X] invoices processed in [period]. We identified [Y] exceptions where invoices were approved by authority levels below those required by policy. We investigated all exceptions and determined [conclusion about control effectiveness]."
  • Quality Review:
  • - Another auditor reviewed:
  • - Pre-test validation approach and results
  • - Exception investigation and root cause analysis
  • - Materiality assessment
  • - Overall conclusion
  • - Approval: "Testing methodology is sound. Validation is adequate. Exceptions are confirmed and appropriately investigated. Conclusion is supported by evidence."

If questioned by regulator: - Regulator: "How do you know AI was correctly configured?" - Auditor: "We performed pre-test validation by manually reviewing 25 transactions and confirming AI flagging matched our manual review. The validation confirmed AI logic was correct." - Regulator: "What about the 12 exceptions?" - Auditor: "We investigated all 12 to confirm they were genuine control exceptions vs. data quality issues. All 12 were confirmed. Root causes were [specific]. We assessed whether they indicate control design issues or execution lapses." - Regulator: "Is 12 exceptions material?" - Auditor: "The 12 exceptions represent $[amount] and [percentage] of population. Applied materiality threshold of $[X]. [X] exceptions exceed materiality; we determined control is [effective/has deficiencies] based on [assessment]."

Outcome: Work is defensible. Regulator can understand methodology, validate pre-test approach, see investigation of exceptions, and concur with conclusions based on evidence.


KEY DIFFERENCES: - Indefensible: "AI found exceptions. Control is effective." - Defensible: "We validated AI logic, investigated exceptions, assessed root causes, applied materiality judgment, and documented the process."

The difference is not the tool; it is the rigor and documentation.


Case Study 2: Risk Assessment -- Defensible vs. Indefensible

INDEFENSIBLE EXAMPLE:

The Work: Auditor used AI to conduct a regulatory risk assessment for a new product. AI generated a list of 22 regulatory requirements. Auditor reported: "Compliance has assessed the regulatory environment. 22 material regulatory requirements have been identified. Control implementation is in progress."

Why it is indefensible: 1. No input validation: What sources did AI use? Are they current? Is this comprehensive? 2. No assessment of applicability: Are all 22 requirements applicable to our product? 3. No organization-specific context: Does AI know our specific regulatory exemptions, product modifications, or customer types? 4. No completeness assessment: Did we validate that AI didn't miss any requirements? 5. No accuracy check: Did we verify that the 22 identified requirements are correctly interpreted? 6. No control assessment: What is current control maturity? What needs to be implemented? 7. Insufficient documentation: How would we explain this assessment if questioned?

If questioned by regulator: - Regulator: "Where did you identify these 22 requirements?" - Auditor: "AI synthesized regulatory guidance and identified them." - Regulator: "Did you validate that all 22 apply to your product?" - Auditor: "I assume they do." - Regulator: "Are there other requirements you might have missed?" - Auditor: "I'm not sure. AI's list seemed comprehensive." - Regulator: "Show me your analysis of this requirement [pointing to one]. How did you interpret this regulation?" - Auditor: "AI provided the interpretation."

Outcome: Assessment is not defensible. Regulator would require significant additional work to validate requirements, interpret regulations, and assess control implementation.


DEFENSIBLE EXAMPLE:

The Work: Auditor conducted a regulatory risk assessment for a new product:

  • Scope Definition:
  • - Documented: "Assessment scope: Regulatory requirements applicable to [Product], sold to [Customer Types], in [Geographies], with [specific features]."
  • - Regulatory domains: [List all applicable regulators]
  • - Sources: [Regulatory agency websites, guidance documents, industry associations, peer assessments]
  • Requirements Identification:
  • - AI synthesized [X] regulatory documents to identify material requirements
  • - Documented: "AI identified [Y] potential requirements. Compliance professional reviewed each against primary regulatory sources. [Y] confirmed as applicable; [Z] determined not applicable based on [specific product/customer factors]."
  • - Peer review: Another compliance professional reviewed [sample] of requirements for accuracy of interpretation
  • Completeness Assessment:
  • - Compared AI list to:
  • - Prior regulatory assessments for similar products
  • - Industry guidance and peer benchmarks
  • - Regulatory agency examination guidance
  • - Documented: "AI identified 22 requirements. Comparison to [benchmarks] indicates [assessment of completeness -- likely complete / potentially incomplete in [area]]."
  • - Gap analysis: If gaps identified, specific follow-up was performed
  • Accuracy Validation:
  • - For each material requirement:
  • - Regulatory citation provided
  • - Specific regulatory language quoted or summarized
  • - Professional interpretation of what the requirement means for our product
  • - Documented evidence that interpretation is accurate per regulatory source
  • Control Maturity Assessment:
  • - For each material requirement: Assessed current control design and maturity
  • - Example: "Requirement X: Current status is [Gap/Design not yet implemented]. Estimated remediation timeline: [X months]. Owner: [responsible party]."
  • Risk Rating:
  • - Each requirement rated for priority/materiality
  • - Professional judgment applied about which requirements must be implemented before product launch vs. which have more flexibility
  • Documentation:
  • - Comprehensive assessment report including all of the above
  • - Supporting workpapers with regulatory sources, interpretations, completeness comparisons, control maturity assessments
  • Quality Review:
  • - Compliance leadership reviewed assessment
  • - Legal counsel reviewed regulatory interpretations for significant requirements
  • - Sign-off: "Assessment reviewed by [compliance director] and [legal counsel]. We concur with 22 identified requirements and control implementation plan."

If questioned by regulator: - Regulator: "Where did you identify these 22 requirements?" - Auditor: "We reviewed regulatory guidance from [sources]. AI helped synthesize the guidance; we then validated each requirement through [specific process]. All 22 have documented regulatory citations and professional interpretation." - Regulator: "Did you validate that all 22 apply to your product?" - Auditor: "Yes. Each requirement was assessed for applicability based on [specific product characteristics]. We documented why each applies." - Regulator: "Are there other requirements you might have missed?" - Auditor: "We compared our list to [benchmark sources]. We identified [any gaps]. We addressed [gaps] through [specific follow-up]." - Regulator: "Show me your analysis of this requirement. How did you interpret it?" - Auditor: "Here is the regulatory citation [specific guidance]. Here is our professional interpretation of what it means for our product [explanation]. Here is our current control implementation plan [details]."

Outcome: Assessment is defensible. Regulator can validate sources, understand requirements, validate applicability, and concur with interpretation and control plan.


KEY DIFFERENCES: - Indefensible: "AI identified 22 requirements. We're implementing them." - Defensible: "We identified 22 requirements through [methodology]. We validated each requirement for applicability [process]. We assessed current control maturity [assessment]. Control implementation plan is [details]."


[Practical Tip]

As you work through these concepts, consider how each one applies to your current role. Think of a specific scenario from your recent work where this concept would have been relevant. Building these mental connections between theory and practice is the fastest way to internalize new knowledge and make it actionable in your daily responsibilities.

Reflection and Analysis

For Each Case Study:

  • Identify the Gap:
  • - What would a regulator immediately question about the indefensible work?
  • - What documentation is missing?
  • - What professional judgment is absent?
  • - What validation was not performed?
  • Understand the Fix:
  • - What specific steps converted indefensible work to defensible work?
  • - How much additional effort did defensibility require?
  • - What was the benefit of the additional effort?
  • Apply to Your Practice:
  • - Do your current work products resemble the defensible or indefensible examples?
  • - What improvements would be needed?
  • - What would you change about your current practices?

Key Takeaways from Case Studies

  • Defensibility Requires Documentation: Indefensible work is often not deficient in substance; it is deficient in documentation and transparency. Defensible work shows the process, not just the result.
  • Professional Judgment Must Be Visible: Indefensible work often delegates judgment to AI ("AI found this, so I reported it"). Defensible work shows independent professional judgment and validation.
  • Validation Is Non-Negotiable: Defensible work includes explicit validation steps. Pre-test validation, exception investigation, completeness assessment, accuracy check -- these are the hallmarks of defensible work.
  • Transparency About AI Builds Confidence: Defensible work explains AI's role clearly. This is not hiding AI; it is demonstrating that you maintained professional control while using AI as a tool.
  • Additional Effort Is Worthwhile: The defensible examples required more documentation and rigor than the indefensible ones. But the defensible work products are trustworthy and would withstand regulatory scrutiny. The effort is worth the defensibility.

[Practical Tip]

As you work through these concepts, consider how each one applies to your current role. Think of a specific scenario from your recent work where this concept would have been relevant. Building these mental connections between theory and practice is the fastest way to internalize new knowledge and make it actionable in your daily responsibilities.

Glossary / Terms

  • Defensibility: The quality of a work product being able to withstand scrutiny and challenge from peers, management, regulators, and other stakeholders.
  • Validation: Testing or review performed to confirm that a methodology, system, or conclusion is appropriate and accurate.
  • Professional judgment: The application of expertise, training, and experience to make reasoned conclusions.

Related Lessons

  • Lesson 1: What Makes an AI-Assisted Work Product Defensible (defensibility principles)
  • Lesson 2: Creating Audit Reports and Compliance Deliverables with AI Support (practical application)
  • Lesson 3: Maintaining Professional Standards in AI-Assisted Output (professional standards)
  • Chapter 1: All lessons on risk assessment and issue identification
  • Chapter 2: All lessons on control testing and monitoring
  • Chapter 3: All lessons on review of AI outputs

End of Chapter 4

Estimated time commitment: 3 hours for all four lessons

Reflection checkpoint: Review your current work products. How would you characterize them on the defensible-to-indefensible spectrum? What specific improvements would strengthen defensibility?

Putting It Into Practice

Translating the case studies into your own practice requires a small set of explicit habits that you apply on every AI-assisted engagement. These habits are not bureaucratic overhead; they are the behaviors that distinguish a defensible workpaper from an indefensible one, and they take less time than reworking a flawed deliverable under audit-committee scrutiny.

Establish a personal validation protocol. Before any AI assistant analyzes a population, validate its logic against a curated subsample of known answers. The subsample should include both positive and negative cases (i.e., items you expect the AI to flag and items you expect it to ignore) and should be sized so that a single misclassification is statistically meaningful (typically 20-40 items). Document the subsample, the AI's outputs, your reconciliation, and your conclusion that the AI's logic matches the control or risk you intend to test. Treat any reconciliation gap as a stop-the-engagement event until resolved.

Build an exception-investigation discipline. Every item the AI flags must be inspected before it appears in a finding or conclusion. False positives caused by stale data, system lag, or model misunderstanding are common; failure to triage them produces overstated findings that crumble on review. Maintain a workpaper that for each flagged item records: the AI's reason for flagging, the auditor's investigation steps, the source documents reviewed, the conclusion (genuine exception / false positive / data-quality artefact), and the basis for the conclusion. This workpaper is the single most powerful defense against an external challenge to the engagement.

Document AI's role transparently. The workpaper should state, in plain language, what the AI did, what data it received, what the prompt or configuration was, what tool calls or external lookups it issued, and what professional review was applied. Hiding the AI's role is a defensibility risk; describing it precisely is a defensibility asset. Adopt a standard 'AI methodology' section in every workpaper template so that the discipline becomes automatic rather than a per-engagement decision.

Apply professional judgment explicitly. Every AI-assisted conclusion should contain a sentence that says, in effect: 'In our professional judgment, based on the procedures described above and the evidence inspected, [conclusion].' The phrase 'in our professional judgment' is not boilerplate; it is the load-bearing element of the workpaper. It signals that a qualified human is taking responsibility for the conclusion, that the AI's output was input rather than output, and that the engagement standards governing the firm's licensure were applied. Without this phrase, the workpaper reads as if the AI is the practitioner -- which it is not, and cannot be.

Contribute to organizational learning. Each engagement is a chance to refine your firm's AI methodology. When you encounter a model failure (a hallucinated citation, a missed exception, a misclassified item), record it in a shared learnings log and bring it to the firm's AI governance forum. When you find a particularly effective prompt or validation step, share it. The defensibility of your firm's AI-assisted work is the sum of the defensibility of each engagement; raising the floor is a collective effort.

Key Takeaways

The four case studies converge on a small number of durable principles. Internalize them and you have most of what you need to apply Level 3 standards to AI-assisted work.

Defensibility is a property of method, not output. Two workpapers can have identical conclusions and identical numbers; one is defensible because the method that produced the numbers is documented and validated, the other is not. Method is what reviewers, regulators, and external auditors examine. Method is what survives a quality challenge.

Validation is non-negotiable. Pre-test validation, exception investigation, completeness checks, and accuracy verification are not optional flourishes; they are the work that makes the AI's output trustworthy. Skipping them does not save time; it transfers risk from the AI to the practitioner who signed the workpaper.

Professional judgment must be visible. The most common failure mode in indefensible AI-assisted work is delegation of judgment to the model. The corrective is not to use less AI; it is to make the human's judgment explicit at every step where judgment is required. Materiality is judgment. Applicability is judgment. Severity is judgment. The model can support these decisions but cannot make them.

Transparency builds confidence. Disclosing that AI was used, in detail, with the controls that validated its output, is more credible than presenting AI-assisted work as if it were purely manual analysis. Sophisticated reviewers (audit committees, regulators, external auditors) increasingly know that AI is in the toolkit; what they want to know is how it was governed in this specific engagement.

Defensibility cost is small; defensibility benefit is large. The conversion from indefensible to defensible cost roughly 20-40% additional effort across the four cases. The cost of an indefensible workpaper that fails on review is much larger: rework under time pressure, regulatory follow-up, reputational damage, and personal accountability for the practitioner who signed it. The risk-adjusted choice is unambiguous.

As you proceed through the remainder of this credential, you will encounter additional dimensions of L3 practice -- monitoring AI quality over time, recognizing when AI assistance is insufficient, and operating in higher-stakes contexts such as regulatory examinations and litigation support. The defensibility frame established in these case studies is the foundation that makes those advanced practices possible.