Chapter 4: Building Defensible AI-Assisted Work Products
What "Defensible" Actually Means for AI-Assisted Work
Picture this: your AI-assisted audit report is challenged -- by a regulator during an inspection, by management disputing a finding, by opposing counsel in litigation, or by your firm's quality review team. "Defensible" means you can withstand that challenge. You can demonstrate that the work was performed with due professional care, that the conclusions are supported by sufficient appropriate evidence, and that the AI tool was used responsibly within established governance frameworks.
Defensibility is not about proving the AI was right -- it is about proving you were right to rely on the AI in the way you did, and that you exercised independent professional judgment throughout. Under PCAOB AS 1015 (Due Professional Care), the auditor must exercise the skill and care of a reasonably prudent professional. Under IIA Standard 11.1 (Proficiency), internal auditors must possess the knowledge, skills, and competencies needed for their responsibilities. When AI assists your work, defensibility requires demonstrating that you understood the AI's capabilities and limitations, provided appropriate inputs, critically evaluated outputs, and took professional responsibility for the final product.
The bar is rising. The PCAOB's 2025 inspection priorities explicitly include the evaluation of how firms use technology in the audit, including AI. The DOJ's 2024 corporate compliance guidance asks whether companies can demonstrate effective controls around AI use. Building defensible AI-assisted work products is not a nice-to-have -- it is a professional survival skill.
The Defensibility Chain: From Prompt to Deliverable
Defensibility is not a single attribute -- it is a chain where every link must hold. Break any link and the entire work product becomes vulnerable to challenge. The chain has six links:
Link 1: Authorized tool use. You used an approved AI tool within your organization's governance framework. You can point to the policy that authorizes this tool for this type of work.
Link 2: Appropriate input. The data and instructions you provided to the AI were relevant, accurate, and did not violate confidentiality or data classification requirements. Your prompt log documents exactly what was submitted.
Link 3: Traceable processing. You recorded the AI tool, model version, date, and any parameters that affected the output. A reviewer could theoretically reproduce your analysis.
Link 4: Critical evaluation. You reviewed the AI output using a structured methodology (CAR framework, peer review, hallucination checks). Your workpapers document the specific verification steps performed and their results.
Link 5: Professional judgment. You modified, supplemented, or overrode the AI output based on your expertise and knowledge of the entity's specific context. The modifications are documented with rationale.
Link 6: Supervisory approval. A qualified supervisor reviewed the final work product with knowledge of the AI's involvement and confirmed it meets professional standards.
Test your work product against each link. If you cannot demonstrate any one of these six elements, your work product has a defensibility gap that could be exploited during challenge.
Creating Audit Reports and Compliance Deliverables with AI Support
AI can accelerate audit report drafting significantly, but the approach matters. Here is a workflow that produces defensible deliverables.
Phase 1: Structure first, draft second. Before involving AI, define the report structure: executive summary, scope and objectives, methodology, findings (condition, criteria, cause, effect, recommendation), and management response. Map your evidence to each section. This ensures the AI is drafting within a framework you control, not inventing its own structure.
Phase 2: Section-by-section drafting. Draft each section individually rather than asking AI to produce a complete report. For findings, provide the AI with your documented evidence and prompt: "Draft an audit finding narrative using the following evidence. Structure it as Condition (what we found), Criteria (what should be), Cause (why it happened), Effect (what impact it has), and Recommendation (what to fix). Use precise professional language appropriate for a [SOX/internal audit/compliance] report. Do not add facts not supported by the evidence provided."
Phase 3: Evidence-linking review. After drafting, verify that every claim in the report traces back to documented evidence. This is the step most teams skip, and it is the step that makes the difference between defensible and indefensible. Create a simple cross-reference: each paragraph references the workpaper or exhibit that supports it. If a paragraph cannot be traced to evidence, it either needs supporting documentation or removal.
Phase 4: Tone calibration. Review the complete draft for tone. AI tends to either overstate findings (using alarm language for minor issues) or understate them (hedging on serious deficiencies). Calibrate finding severity to match your organization's rating criteria and the actual significance of the issue.
Maintaining Professional Standards in AI-Assisted Output
Every AI-assisted work product must comply with the professional standards governing your function. These standards were written before AI, but they apply with full force to AI-assisted work. Here is how key standards translate to AI-assisted practice.
PCAOB AS 1215 (Audit Documentation): "The auditor must prepare audit documentation in sufficient detail to provide a clear understanding of its purpose, source, and the conclusions reached." For AI-assisted work, this means documenting the AI's role as a source, preserving the AI output as evidence, and distinguishing AI-generated content from auditor-authored content.
IIA Standard 2310 (Identifying Information): "Internal auditors must identify sufficient, reliable, relevant, and useful information to achieve the engagement's objectives." AI-generated information is not automatically reliable -- reliability must be established through verification, just as you would evaluate any other information source.
ISA 500 (Audit Evidence): Audit evidence must be sufficient and appropriate. Appropriateness includes reliability, which depends on the source. AI-generated analysis is less reliable than primary source documents, so AI outputs should corroborate, not replace, primary evidence.
IESBA Code of Ethics (Integrity and Professional Competence): Using AI does not diminish your personal responsibility for the quality of your work. If you lack the competence to evaluate AI output in a specific domain, you must either develop that competence or seek assistance from someone who has it. Delegating to AI is not delegation to a qualified professional -- it is using a tool that you remain fully responsible for.
Case Studies: Defensible vs. Indefensible AI-Assisted Work
Case A (Indefensible): An auditor asks ChatGPT to "write an audit finding about weak access controls" with no context about the entity, no supporting evidence, and no specific control tested. The AI produces a generic finding citing "industry best practices" and "common frameworks." The auditor adds the client's name and submits it. This work product is indefensible because: no actual testing was performed, the finding is not based on entity-specific evidence, the criteria are vague, and a reviewer cannot determine what audit procedure generated this finding.
Case B (Defensible): An auditor completes access control testing, documenting 15 segregation of duties conflicts identified through analysis of SAP role assignments. The auditor provides Claude with the testing results, the company's access control policy, and the relevant COBIT control objectives. The AI drafts the finding narrative. The auditor verifies every statement against the test results, adds context about compensating controls that partially mitigate three of the conflicts, adjusts the severity rating based on management's remediation timeline, and documents all modifications. The workpaper includes the prompt log, the V0 AI output, the tracked-changes V1, and the final V2 with supervisory sign-off.
The difference is not whether AI was used -- it is whether the work product is supported by evidence, reflects professional judgment, and maintains a complete documentation trail. Case B would withstand PCAOB inspection, management challenge, or peer review. Case A would not survive any of them.
Treating AI Output as Evidence: Classification and Weight
How much evidentiary weight should you assign to AI-generated content? This question does not have a single answer, but a framework helps you make consistent judgments.
Think of AI output as occupying one of three evidentiary tiers. Tier 1: AI as drafting aid. The AI helps you write -- structuring narratives, suggesting language, formatting data. The evidentiary value comes from the underlying data and analysis, not from the AI's contribution. This is the lowest-risk use case and requires standard documentation.
Tier 2: AI as analytical tool. The AI performs substantive analysis -- pattern recognition, anomaly detection, gap analysis, risk scoring. The AI output itself becomes part of your evidence base. This requires additional documentation: validation of the AI's analytical method, independent verification of results, and explicit assessment of the AI output's reliability as audit evidence under ISA 500 or PCAOB AS 1105.
Tier 3: AI as subject matter input. You rely on the AI's "knowledge" to inform your analysis -- asking it about regulatory requirements, industry practices, or technical accounting guidance. This is the highest-risk tier because AI knowledge may be outdated, incomplete, or fabricated. Never treat Tier 3 AI output as primary evidence. Always corroborate with authoritative sources (official regulatory texts, professional standards, published guidance). In your workpapers, the corroborating source is the evidence; the AI output is merely the research mechanism.
Classify every AI interaction by tier and document accordingly. This classification discipline ensures your evidence hierarchy remains sound.
Legal Defensibility: When AI-Assisted Work Faces Litigation
AI-assisted audit and compliance work products may eventually face legal scrutiny -- in securities litigation, regulatory enforcement proceedings, or professional liability claims. The legal landscape is evolving rapidly, but several principles are already clear.
Disclosure obligations. Courts and regulators will increasingly expect transparency about AI use. If you used AI to perform a substantive analytical procedure and do not disclose it, the non-disclosure itself may become a liability. The SEC's 2025 guidance on AI use in financial reporting emphasizes the importance of transparency about technology's role in the audit.
Standard of care. The legal standard of care for professionals using AI tools is likely to be: did the professional use the AI tool as a reasonably competent professional would, including exercising appropriate oversight and verification? This means that both over-reliance on AI (failing to verify) and under-use of AI (ignoring available tools that could have caught an error) may create liability exposure.
Privilege considerations. AI prompts and outputs may be discoverable in litigation. If you share privileged or work-product-protected information with an external AI tool, you risk waiving privilege. Enterprise AI deployments with appropriate confidentiality protections are essential for litigation-sensitive work.
Your protection is documentation. The best legal defense for AI-assisted work is a complete, contemporaneous record showing authorized tool use, appropriate inputs, critical evaluation, professional judgment, and supervisory approval -- the six-link defensibility chain from earlier in this chapter.
Building Team-Wide Standards for AI-Assisted Deliverables
Individual excellence is insufficient -- defensible AI-assisted work requires consistent team standards. Without them, one team member's sloppy AI use creates risk for the entire engagement.
Establish these team-wide standards: Approved use cases. Define which deliverable types may use AI assistance and which may not. Consider starting with a "green list" (approved for AI) and a "red list" (prohibited from AI), with everything else requiring case-by-case approval from the engagement lead.
Minimum documentation requirements. Specify the minimum documentation for each use case tier (drafting aid, analytical tool, subject matter input). Make templates available and require their use -- do not rely on individual judgment about what to document.
Review protocol by deliverable type. Define which review layers (self-review, peer review, supervisory review) are required for each deliverable type. Document this in your engagement planning memo or audit charter.
Quality gate checklist. Before any AI-assisted deliverable is finalized, require a quality gate review using a standardized checklist that covers: defensibility chain completeness, evidence cross-referencing, hallucination check confirmation, professional standards compliance, and appropriate disclosure of AI use.
Periodic calibration. Monthly, review a sample of AI-assisted workpapers across the team. Identify best practices and common deficiencies. Share findings in a team meeting. This creates a learning loop that raises quality continuously. The IIA's Quality Assurance and Improvement Program (QAIP) requirements under Standard 12.1 provide a natural framework for this ongoing calibration.
Try This Now
Take an existing AI-assisted work product (or create one for this exercise) and stress-test its defensibility:
- Apply the six-link defensibility chain. For each link (authorized tool, appropriate input, traceable processing, critical evaluation, professional judgment, supervisory approval), document whether the link is present, partially present, or missing. Be honest -- this is a diagnostic exercise.
- Classify the AI use by evidence tier. Was the AI used as a drafting aid (Tier 1), analytical tool (Tier 2), or subject matter input (Tier 3)? Is the documentation appropriate for that tier?
- Perform the evidence cross-reference test. For every factual claim in the work product, can you point to a specific piece of evidence in your workpapers? Mark any unsupported claims.
- Run the hostile reviewer test. Imagine a PCAOB inspector, a plaintiff's attorney, or a skeptical audit committee member reviewing this work product. What questions would they ask? Can you answer each one from your documentation alone, without relying on memory or verbal explanation?
- Identify your top three defensibility gaps and create an action plan to close them. Prioritize gaps that would be most damaging if exploited during a challenge.
This exercise is best done in pairs -- have a colleague stress-test your work and vice versa. The discomfort of finding gaps now is far preferable to discovering them during an inspection.
Key Takeaways
- Defensible means your AI-assisted work can withstand challenge from regulators, management, opposing counsel, or quality reviewers -- with documentation alone, not verbal explanations.
- The six-link defensibility chain (authorized tool, appropriate input, traceable processing, critical evaluation, professional judgment, supervisory approval) must be complete for every deliverable. One broken link compromises the entire product.
- Draft reports section-by-section with AI, always providing your evidence as input. Never ask AI to generate findings without entity-specific evidence -- the result is indefensible by definition.
- Professional standards (PCAOB AS 1215, IIA 2310, ISA 500, IESBA Code) apply fully to AI-assisted work. AI does not reduce your personal responsibility; it adds new obligations around tool verification and output validation.
- Classify AI use by evidence tier (drafting aid, analytical tool, subject matter input) and scale your documentation and verification requirements accordingly.
- Legal defensibility requires disclosure of AI use, compliance with the evolving standard of care, protection of privileged information, and complete documentation.
- Build team-wide standards: approved use cases, minimum documentation, review protocols, quality gate checklists, and periodic calibration against a sample of AI-assisted workpapers.
Skill.re