Internal AI Audit - Annual Plan & Quarterly Reviews
At the Q1 2026 Audit Committee meeting, the Acme Inc. chair, a former PwC audit partner now on three boards, turned to the Chief Audit Executive and asked one sentence: "Show me the AI audit plan." The CAE opened the binder, paged through SOX, cyber, financial controls, ESG, vendor risk, and stopped. There was no AI track. The cyber audit plan covered identity, endpoint, cloud, and incident response: none of it touched the seven LLM-powered systems the CAIO had stood up in 2025, the Article 9 RMS, the Article 17 QMS, the Article 27 FRIAs, the SR 11-7 model risk management work, the ISO 42001 binder, or the AI Governance Committee charter. The CAE said, on the record: "We do not yet have a dedicated Internal AI Audit track. We will have one before the next quarterly review." That commitment, captured in the minutes, circulated to the external ISO 42001 surveillance auditor, and referenced in the next SOC 2 + AI engagement, became the founding act of Acme's Internal AI Audit (IAA) function. By Q3 2026 the IAA had a documented annual plan, two FTEs, a Big-4 co-source bench, and the first three reviews of an eventual nine-review calendar in the field. This lesson is the L4 leadership playbook for the Third Line of Defense view of AI - IAA as 3L distinct from 1L model owners and 2L AI Risk Office / MRM, planned per IIA International Professional Practices Framework (IPPF) Standards adapted for AI, executed across an 8-section annual plan, 12 typical reviews per year, and a 3-week quarterly cycle from kickoff to issue.
Why Internal AI Audit in 2026 - Three Lines, the IPPF, and the 2026 Expectation
The Three Lines Model, published in revised form by the Institute of Internal Auditors (IIA) in July 2020 as the successor to the older "Three Lines of Defense" framing, distinguishes three roles in an organization's risk and control architecture. First line (1L) owns the risk and operates the controls; for AI, the 1L is the model owner, the MLOps team, the data-science function, the product team. Second line (2L) provides oversight, expertise, and challenge of the 1L's risk management; for AI, the 2L is the AI Risk Office, the Model Risk Management (MRM) function under SR 11-7 / OCC 2011-12 framing, the AI Governance Committee (AIGC), the Compliance function for AI-specific regulatory requirements. Third line (3L) provides independent assurance to the governing body (Audit Committee, Board) on the design and operating effectiveness of governance, risk management, and control, including 1L control operation and 2L oversight effectiveness. For AI, the 3L is Internal AI Audit, an independent function reporting administratively to the CEO and functionally to the Audit Committee chair.
The 3L role is shaped by the IIA International Professional Practices Framework (IPPF), the global standard for internal audit. The 2024 update (effective January 9, 2025) reorganized the Standards into Domains covering Purpose, Ethics, Governance, Managing the Internal Audit Function, and Performing Internal Audit Services. Core IPPF Standards that apply to AI work, practitioners commonly cite by the 2017-era numbering still in active reference, include: 1100 Independence and Objectivity; 1200 Proficiency and Due Professional Care (for AI, this means AI literacy on the audit team); 2010 Planning (the CAE must establish a risk-based plan); 2200 Engagement Planning; 2400 Communicating Results. The 2024 reorganization preserved these obligations within Domains III-V of the new architecture.
By Q1 2026 the expectation of a formal IAA function with documented annual plan and quarterly reviews comes from three converging directions:
- Audit Committee expectation. Audit committees expect a dedicated AI audit track within the annual internal-audit plan. The IIA's 2026 AI-Audit benchmarking survey (released March 2026) found 71% of Fortune 1000 audit committees expect a documented IAA plan by FY2027 close; 38% expect it by FY2026 close. "The cyber audit plan does not cover AI" is the line that has surfaced in dozens of 2026 audit-committee minutes globally.
- EU AI Act competent authority expectation under Articles 26, 27, 71, 72. National competent authorities and the EU AI Office, when conducting Article 89 information requests or post-market monitoring reviews, request evidence that the deployer or provider has an internal-audit function reviewing AI control operation. Article 17 QMS (provider) and Article 26 deployer obligations include internal-control verification; internal audit is the natural Third Line of that verification.
- ISO 42001 surveillance auditor expectation under Clause 9.2 + A.9.2. ISO/IEC 42001:2023 Clause 9.2 requires the organization to conduct internal audits at planned intervals; A.9.2 reinforces "internal audit" as a management-system requirement. Surveillance auditors (Schellman, A-LIGN, BSI, KPMG, BDO) request the IAA annual plan, completed reviews, findings, and closure evidence; absence of a documented IAA programme is itself a major finding.
IAA is complementary to but distinct from external assurance (ISO 42001 certification, SOC 2 + AI attestation, notified-body conformity assessment). External assurance opines for the customer / regulator audience; IAA reports for the Audit Committee. External assurance is performed by independent third-party firms at defined cadences; IAA operates continuously throughout the year against a risk-based plan and reports quarterly. The two are mutually reinforcing, IAA workpapers feed the external assurance evidence base; external assurance findings feed the IAA risk assessment. By Q1 2026 the operational data is consistent: organizations running coordinated IAA + external assurance reduce duplicative testing by 40-60% versus organizations running them in isolation.
The 8-Section IAA Annual Plan
The IAA annual plan is the load-bearing artifact of the function. It is prepared by the Chief Audit Executive (CAE) with input from the AI Risk Officer, the CAIO, the General Counsel, and external assurance partners; presented to and approved by the Audit Committee; shared with the ISO 42001 surveillance auditor, the SOC 2 + AI auditor, and competent authorities on request. The 8-section structure that has stabilized across Big-4 internal-audit practices and ISACA AI Audit working-group benchmarking by Q1 2026:
- Section 1 - Risk assessment. The plan opens with the IAA's risk assessment, informed by but independent of 2L risk products. Inputs include the AI risk heat map (lesson 085) covering inherent and residual risk per system; the FAIR-based Annualized Loss Expectancy quantification (lesson 086) translating risk into dollar exposure; the board AI KPI dashboard (lesson 087); AIGC minutes for emerging concerns; external-assurance findings; regulator-engagement records; incident logs from Article 73 reporting; and the model inventory with risk tiering. Outputs a prioritized list of audit topics with rationale.
- Section 2 - Audit universe. Lists every entity the IAA could audit: every AI system in the inventory with criticality tier (high / medium / low); every 2L process (AIGC, MRM, AI Risk Office, Compliance for AI); the AIGC governance machinery itself; the AI literacy programme (Article 4); the AI incident-response programme; the third-party AI risk programme; the AIMS and its sub-processes. Rebuilt annually. A typical mid-market AI-vendor universe by Q1 2026: 8-15 in-scope systems, 6-10 2L processes, 4-6 cross-cutting governance programmes, total 18-31 auditable entities.
- Section 3 - Audit topics for the year. From the risk-assessment + universe, the CAE selects the topics. Typical coverage by Q1 2026: 8-12 substantive reviews plus 4 quarterly surveillance reviews (continuous-auditing assessments of high-frequency controls). Topics rotate on a 2-3 year cycle so every universe entity is audited at least once in the cycle. High-criticality systems and high-risk 2L processes are audited annually; medium-criticality every 18-24 months; low-criticality every 36 months.
- Section 4 - Resource plan. FTE plan, budget, and co-source arrangement. Typical mid-market enterprise IAA capacity for AI in Q1 2026: 1-3 FTE (often growing from 1 in 2025 to 2-3 by 2026); plus a Big-4 or specialist firm co-source bench (Deloitte, EY, KPMG, PwC, Schellman, BDO, RSM, Crowe, Protiviti) priced at $250-$450 / hour. Regulated-industry IAA runs larger, 4-8 FTE plus dedicated co-source. Budget envelope mid-market: $400K-$900K annual.
- Section 5 - Calendar with quarterly milestones. The audit calendar lays substantive reviews across the year: typically 2-3 reviews per quarter, each 4-8 weeks engagement time, with kickoff, fieldwork, draft-report, and final-report dates. Coordinates with the external assurance calendar to avoid evidence-collection collisions. Quarterly milestones: end-of-Q1 progress report; end-of-Q2 mid-year review; end-of-Q3 progress report; end-of-Q4 annual closing summary plus next-year plan draft.
- Section 6 - Skills and training plan. The 3L capability development plan. Auditors must collectively have AI literacy to test AI controls (IPPF 1200; EU AI Act Article 4 applies to internal-audit staff as deployer personnel). The plan identifies skill gaps and the training that closes them: IAPP AIGP certification, IIA Certified Internal Auditor (CIA) for non-CIA auditors, MLOps vendor training, MITRE ATLAS / OWASP LLM Top 10 familiarity, fairness-metrics and prompt-engineering literacy. Typical 2026 training investment: $5K-$15K per FTE.
- Section 7 - Coordination with external assurance. Documents how IAA workpapers are shared with ISO 42001 surveillance auditors (AIMS portal read access), SOC 2 + AI auditors (cross-walked evidence), notified bodies (Article 31 designated-body relationship), and competent authorities on request. Names overlap areas, establishes shared evidence-collection cadence, and reduces duplicative testing 40-60%.
- Section 8 - Audit Committee approval and reporting cadence. Presented to the Audit Committee at fiscal-year start for approval; quarterly progress reports follow; annual closing summary at year-end. Reports cover (a) reviews in flight, (b) findings raised this quarter by rating, (c) findings closed, (d) overdue findings with escalation, (e) emerging concerns, (f) external-assurance coordination notes, (g) capability development progress.
The annual plan is a living document. Material change to the AI risk profile (new high-risk system deployment, regulator inquiry, major incident, new framework obligation) triggers a plan amendment with Audit Committee notification. By Q1 2026 it is standard practice to formally re-baseline the plan at mid-year if material change has accumulated; ad-hoc reviews outside the plan are documented as plan additions, not as informal work product.
The 12 Typical IAA Reviews - From AIGC Effectiveness Through ISO 42001 Evidence Quality
The 12 substantive reviews that recur across mid-market and enterprise IAA programmes in 2026: each 4-8 weeks of engagement time, each producing a written workpaper file and a finding report graded by severity, each cross-walked to specific IPPF Standards and external-assurance frameworks:
- Review 1 - AIGC effectiveness (lesson 072). Tests the AI Governance Committee against its charter: composition; cadence; decision quality (sampled decisions reviewed for documentation, alternatives, dissent, escalation traceability); follow-through. Common findings: charter scope ambiguous on tier-2 approvals; quarterly cadence missed; decision-tracking incomplete. Cross-walks IPPF 2200, ISO 42001 Clause 5.3, EU AI Act Article 26.
- Review 2 - AIRA breach response (lesson 074). Tests the AI Risk Appetite framework against operational breaches: appetite-breach events identified, escalated, and remediated per policy; AIGC notified within documented timing; board KPI dashboard updated. Common findings: breach events identified late; escalation timing missed. Cross-walks IPPF 2200, NIST AI RMF Govern 1.
- Review 3 - RACI execution (lesson 075). Tests the AI RACI matrix against accountability traceability for sampled high-risk decisions. Common findings: Accountable role ambiguous; Consult roles not engaged; documentation gaps. Cross-walks IPPF 2200, ISO 42001 Clause 5.3, EU AI Act Article 26.
- Review 4 - CAF function effectiveness (lesson 077). Tests the Chief AI Officer function against the CAF charter: scope-of-authority adherence; reporting-line independence; decision documentation; coordination with CRO, CISO, CDO, GC. Common findings: scope creep into 1L territory; decision documentation thin. Cross-walks IPPF 1100, ISO 42001 Clause 5.3.
- Review 5 - Article 17 QMS effectiveness (lesson 079). Tests the Quality Management System for high-risk AI providers per Article 17: 13 documented elements operating in evidence; nonconformity tracking; corrective action; management review. Common findings: nonconformity records sparse; corrective-action timing slipped. Cross-walks IPPF 2200, EU AI Act Article 17.
- Review 6 - Article 72 PMM effectiveness (lesson 080). Tests Post-Market Monitoring plan execution: monitoring data collected and analyzed per plan; drift breaches triaged; fairness refresh on cadence; PMM annual report drafted. Common findings: monitoring-data analysis lags; PMM annual report missing. Cross-walks IPPF 2200, EU AI Act Article 72, ISO 42001 A.6.1.5.
- Review 7 - Red team operating model (lessons 081-084). Tests AI red team independence and coverage: independence from 1L model development; coverage of OWASP LLM Top 10, OWASP Agentic Top 10, MITRE ATLAS techniques; intake → scoping → testing → reporting workflow operating; findings traceable to remediation. Common findings: independence partial; coverage uneven (OWASP covered; Agentic Top 10 sparse). Cross-walks IPPF 1100, EU AI Act Article 15, NIST AI RMF Measure 2.7.
- Review 8 - SR 11-7 MRM application (lesson 067). For financial services: tests the MRM framework's application to AI/LLM models: model inventory, materiality assessment, IMV, monitoring, lifecycle. Common findings: AI models not on MRM inventory; materiality assessment thin; IMV partial. Cross-walks IPPF 2200, SR 11-7 governance pillar, ISO 42001 Clauses 6.1.2/6.1.3.
- Review 9 - IMV independence (lesson 068). Tests the independent model validation function's independence from model development: reporting line; scope-of-work segregation; methodology documentation; finding sign-off authority. Common findings: IMV reports to same VP as model development; methodology variation across models. Cross-walks IPPF 1100, SR 11-7 effective challenge.
- Review 10 - Vendor DDQ and evidence review (lessons 070-071). Tests the third-party AI risk programme: supplier inventory; per-supplier DDQ on file; AI-specific addendum completed; foundation-model provider attestation documented evaluation; substantial-modification change-control evidence. Common findings: supplier inventory incomplete; DDQ refresh overdue; foundation-model attestation reliance not documented with gap analysis. Cross-walks IPPF 2200, EU AI Act Article 25 + Annex XII, ISO 42001 A.10.
- Review 11 - Article 73 incident response (lessons 088-089). Tests the AI-specific incident-response runbook: AI incident classification against Article 73 timing; runbook tabletop completed annually; regulator-contact log maintained; closure evidence with root-cause analysis. Common findings: classification ambiguous on edge cases; tabletop with thin scenario coverage; root-cause analysis lacks AI-specific failure-mode framing. Cross-walks IPPF 2200, EU AI Act Article 73, ISO 42001 A.8.4, NIST AI RMF Manage 4.
- Review 12 - ISO 42001 evidence-pack quality (lesson 091). Tests the AIMS evidence pack for ISO 42001 surveillance readiness: tab-by-tab evidence completeness; cross-walk register currency; document version-control; control-owner attestation timing. Common findings: tabs lag operational reality by 30-60 days; cross-walk register out of date; control-owner attestations missing. Cross-walks IPPF 2200, ISO 42001 Clause 9.2 + A.9.2.
The 12 reviews are not all executed every year. The CAE selects 8-12 based on the risk assessment and the rotation cycle; the remaining slots are filled by quarterly surveillance reviews of high-frequency controls (e.g., monthly review of model-release gate operation; quarterly review of fairness refresh completion) and reserve capacity for ad-hoc inquiries from the Audit Committee or in response to incidents.
The Quarterly Review Pattern, Finding Rating, and Closure SLAs
Each substantive IAA review follows a standard 3-week-cycle pattern from kickoff to issue, with a 4-8 week fieldwork window between kickoff and draft. The pattern that has stabilized in 2026 mid-market practice:
- Week minus 2 - Pre-engagement planning. CAE and engagement lead define scope, objective, criteria, key risks, and evidence-request list. Engagement letter drafted naming 1L and 2L counterparts. IPPF 2200 documentation: objective, scope, resources, work programme.
- Week 0 - Kickoff. Engagement letter issued. Kickoff with the 1L process owner and the 2L oversight counterpart. Evidence-request list issued with a 5-business-day initial response window. Management self-attestation requested, the 1L's written assertion that the control is operating as designed.
- Weeks 1-4 - Fieldwork. Auditor reviews the self-attestation; samples evidence per the work programme; conducts walkthrough interviews; reperforms control activities; tests independence assertions; reviews cross-walk references to external-assurance evidence. AI-specific testing: re-running red team probes against the sampled population, reviewing model cards against Annex IV expectations (lesson 016), sampling drift-monitoring threshold reviews, sampling FRIA executions.
- Week 5 - Preliminary findings. Auditor drafts the preliminary finding list with severity rating, root-cause analysis, recommended remediation, target closure date. Findings reviewed with 1L for factual accuracy.
- Week 6 - Draft report. Auditor issues the draft: executive summary, scope, methodology, findings with severity ratings and management responses, conclusions, recommendations. Circulated to 1L, 2L, CAE.
- Week 7 - Management response and final report. 1L drafts the formal management response per finding (acceptance, remediation plan, owner, target date) or contests with documented rationale. CAE finalizes; final report issued to 1L, 2L, AIGC, and Audit Committee at the next quarterly meeting.
- Post-issue - Closure tracking. Findings tracked to closure per the SLA tied to severity. Closed findings re-validated by IAA; partial closure flagged. Quarterly Audit Committee report includes finding-aging metrics.
Finding rating follows a four-level scheme that has consolidated in 2026 practice (IIA terminology with AI-specific examples):
- Critical, control failure that creates immediate material risk to the organization (regulatory non-compliance with imminent penalty exposure; control breakdown enabling significant financial loss; safety-of-life risk). AI examples: an Article 73-reportable incident not reported within timing standards; an undocumented Article 27 FRIA on a deployed high-risk system; an AIGC that has not met for 12+ months. Closure SLA: 30 days, with weekly Audit Committee escalation until closed.
- Significant, control deficiency that materially weakens the control environment but does not create immediate material risk. AI examples: IMV records sparse for material models; fairness refresh more than 90 days stale; red team coverage uneven across OWASP/ATLAS; vendor DDQ refresh overdue. Closure SLA: 60 days, with monthly Audit Committee progress reporting.
- Moderate, control weakness that should be remediated but does not materially affect the overall control environment. AI examples: AIGC charter scope ambiguous on tier-2 approvals; quarterly cadence missed once or twice; drift threshold reviewed but signed by wrong owner; evidence-pack tab lagging operational reality by 30-60 days. Closure SLA: 90 days, with quarterly Audit Committee progress reporting.
- Observation, process improvement opportunity that is not a control failure. AI examples: cross-walk register would benefit from additional columns; AIMS portal evidence-pack tab order could be optimized; training material could be updated with 2026 framework updates. Closure SLA: 180 days, tracked in IAA log without formal Audit Committee reporting.
The closure SLAs are policy-level; individual findings can negotiate longer closure with documented rationale (e.g., a critical finding requiring system replumbing rather than configuration fix may take 90 days with weekly progress reporting; the rationale and Audit Committee acknowledgment are documented). Overdue findings escalate per the policy, overdue critical to monthly board reporting; overdue significant to quarterly board reporting; overdue moderate to next AIGC meeting.
The quarterly review pattern aggregates across the function: at the end of each quarter, the IAA reports to the Audit Committee with (a) reviews completed this quarter with summary findings, (b) reviews in flight with expected completion date, (c) findings raised this quarter by rating, (d) findings closed this quarter, (e) overdue findings with escalation status, (f) emerging concerns from preliminary fieldwork, (g) coordination notes with external assurance, (h) plan-amendment requests where material change has accumulated. The Audit Committee chair signs the quarterly summary into the minutes.
Independence, Capability Buildout, Tooling, and External Assurance Coordination
Independence preservation is the load-bearing constraint of the IAA function. IPPF Standard 1100, preserved within Domain II of the 2024 Global Internal Audit Standards, requires that internal audit be free from conditions that threaten objectivity. For AI work this means four operational rules that have stabilized in 2026 practice:
- IAA cannot be involved in 2L control design. The IAA does not draft the AIRA policy, AIGC charter, IMV methodology, red-team operating model, or AI incident-response runbook. Where IAA observes gaps in 2L design, the finding is communicated through the audit report; the 2L function performs the design work.
- IAA reviews 2L process effectiveness. Audits of the AIGC, AI Risk Office, MRM, and red team are reviews of 2L process operating effectiveness, not co-development. The IAA can observe that the AIGC charter is ambiguous on tier-2 approvals (finding); the IAA cannot draft the revised charter language (2L work).
- IAA cannot test its own work. Where the IAA has been substantively involved in an area, that area is audited by external co-source until the involvement-clock has run for at least two years.
- IAA reports to the Audit Committee chair, not the CRO or CAIO. Administrative reporting to the CEO is acceptable; functional reporting must be to the Audit Committee. IAA reporting to CRO breaches independence (CRO is a 2L function the IAA may audit); IAA reporting to CAIO breaches independence (CAIO is the senior 1L owner of AI systems). The 2026 audit-committee survey data: 100% of large-cap public-company audit committees expect IAA functional reporting to the Audit Committee chair.
IAA capability buildout in 2026 follows two hiring archetypes. The audit-background-first archetype hires Certified Internal Auditors (CIA) or Big-4-trained audit professionals and trains them in AI literacy (IAPP AIGP; vendor MLOps training; OWASP / MITRE ATLAS familiarity; fairness-metrics training); methodology strong day one, AI literacy lags 12-18 months. The AI-background-first archetype hires data scientists, ML engineers, or AI risk specialists and trains them in IPPF Standards (CIA exam preparation; Big-4 methodology training; IIA workshops); AI strong day one, methodology lags 12-18 months. Mid-market enterprise IAA functions for AI in 2026 typically run 1-3 FTE, often a CAE-lead with audit background plus one or two AI-literate audit-trained staff, supplemented by Big-4 or specialist co-source. Co-source firms: Big-4 (Deloitte, EY, KPMG, PwC) for large-cap; mid-tier (RSM, BDO, Crowe, Grant Thornton, Protiviti) for mid-market; specialist (Schellman, A-LIGN, Coalfire) for niches; boutique advisory for specialty IMV and red-team. Co-source rates $250-$450/hour; engagement budgets $40K-$120K per substantive review.
2026 audit tooling falls into two categories. Workpaper management platforms, AuditBoard, Workiva, TeamMate+ (Wolters Kluwer), Diligent HighBond, MetricStream, ServiceNow IRM, host the audit universe, the annual plan, engagement workpapers, findings tracking, and Audit Committee reporting. AuditBoard and Workiva dominate the U.S. mid-to-large market by Q1 2026. The 2026 AI-augmentation (AuditBoard Copilot, Workiva AI Assist) provides natural-language evidence search, draft-finding language, and cross-framework citation lookup. AI-specific testing tools are the second category. The IAA auditing red-team operating-model effectiveness re-runs a sample of OWASP LLM Top 10 probes (Promptfoo, Garak, PyRIT, Inspect AI) and MITRE ATLAS techniques against the deployed system; results compared with the 2L red-team's documented coverage. The IAA auditing model documentation reads sampled model cards against Annex IV expectations (lesson 016) using AI-assisted cross-reference. The 2026 expectation by Audit Committee chairs is that IAA testing is at minimum partially automated.
Coordination with external assurance turns three parallel audit programs into one coordinated programme with three output deliverables. The ISO 42001 surveillance auditor receives the IAA annual plan, completed reviews with findings, closure evidence, and cross-walk references; where IAA review of AIGC effectiveness covers the same evidence as the surveillance auditor's review of ISO 42001 Clause 5.3 roles, the surveillance auditor can rely on IAA workpapers as cross-reference rather than re-test from scratch, reducing surveillance fieldwork 15-25%. The SOC 2 + AI auditor and IAA coordinate on overlapping controls; AICPA SSAE 18 guidance permits reliance on internal-audit work where the auditor concludes the IAA function has appropriate competence, objectivity, and quality (reducing SOC 2 fieldwork 15-25%). The notified body reviews internal-audit evidence as part of Annex VII Module H AIMS audit. Competent authorities exercising Article 89 information requests request IAA workpapers; the IAA function maintains retention policy (7 years; 10 years for high-risk system audits aligned with Article 11 Annex IV retention). The 40-60% duplication reduction comes from one-evidence-many-frameworks: the same red-team probe re-runs serves IAA review 7, ISO 42001 A.6.1.5 surveillance, SOC 2 + AI fairness/bias and security assertions, and notified-body conformity assessment.
Six Common 2026 IAA Failures and the Acme Worked Example
Six common 2026 IAA failures identified across Big-4 IAA practice data and ISACA AI Audit working-group benchmarking by Q1 2026:
- Failure 1 - Treating IAA as a compliance check rather than strategic risk reduction. The IAA audits the AIGC and produces a finding that the charter is missing one signature; the underlying decision quality remains untested. Remediation: orient IAA reviews to operating effectiveness of risk-reduction outcomes, not document completeness.
- Failure 2 - IAA reports to the CRO or CAIO, not the Audit Committee chair. Either reporting line breaches IPPF 1100 independence and renders the IAA work product invalid for external-assurance cross-reference. Remediation: re-line reporting to administrative-to-CEO + functional-to-Audit-Committee-chair within 90 days.
- Failure 3 - No AI literacy on the audit team. The IAA cannot test AI controls without AI literacy (IPPF 1200 proficiency; EU AI Act Article 4). An audit team without IAPP AIGP, MLOps familiarity, OWASP / MITRE ATLAS familiarity misses material control weaknesses. Remediation: structured training plan within 90 days; co-source until in-house literacy reaches threshold.
- Failure 4 - No annual plan; ad-hoc reviews. The function operates reactively. Coverage is uneven; surveillance auditors and SOC 2 + AI auditors cannot rely on the work because it lacks documented planning per IPPF 2010. Remediation: stand up the 8-section annual plan; obtain Audit Committee approval; publish to external assurance partners.
- Failure 5 - Skipping coordination with external assurance. The IAA, ISO 42001 surveillance auditor, and SOC 2 + AI auditor operate in isolation. Duplication wastes 40-60% of testing effort; control owners experience audit fatigue. Remediation: shared evidence portal; coordination plan in Section 7 of the annual plan; cross-walk all IAA work programmes to external-assurance frameworks.
- Failure 6 - Testing only documentation, not operating effectiveness. The IAA reviews the AIRA policy document and concludes the policy exists; the IAA does not test whether the policy operates against breach events. Remediation: every IAA work programme includes operating-effectiveness testing (sample of breach events, decisions, incidents, monitoring outputs).
Following the Q1 2026 Audit Committee commitment, Acme Inc. stood up the Internal AI Audit function over the next six months. The CAE, a former PwC audit partner now in-house, supported by the Audit Committee chair, drafted the 8-section annual plan, recruited two FTEs (one CIA with banking-audit background trained into AI literacy via IAPP AIGP; one former ML engineer trained into IPPF Standards via CIA exam preparation), and contracted KPMG as the Big-4 co-source bench. The plan was approved by the Audit Committee in May 2026 and rolled out as follows.
2026 plan summary: 9 substantive reviews + 4 quarterly surveillance reviews. FTE budget 2.0; co-source budget $280K; training and tooling budget $35K; total plan envelope $720K. Cross-walked to Schellman (ISO 42001 surveillance + SOC 2 + AI) and BSI (Article 31 notified body for the Annex III high-risk system) via shared evidence portal.
- Q1 2026 (kickoff and planning). Plan development, Audit Committee approval, FTE recruiting, co-source contracting, tooling setup (AuditBoard implementation). No substantive reviews this quarter; planning report to Audit Committee.
- Q2 2026 (Reviews 1 + 7). Review 1, AIGC effectiveness (June): two moderate findings, charter scope ambiguous on tier-2 approval thresholds (4 sampled tier-2 decisions ambiguous on accountability); quarterly cadence missed twice in 2025 (June 2025 cancelled due to off-site; October 2025 informally held without minutes), with documented commitment to never miss two consecutive quarters going forward. Both findings closed within 60 days. Review 7 - Red team operating model (July): no significant findings; independence preserved with the red-team function reporting to CISO independent of CAIO; coverage of OWASP LLM Top 10 strong; one moderate finding on OWASP Agentic Top 10 coverage uneven (ASI06 Memory Poisoning and ASI07 Inter-Agent Spoofing not yet tested for the multi-agent customer-support pilot). Finding closed within 60 days.
- Q3 2026 (Reviews 11 + 9 + 12). Review 11, Article 73 incident response (August): one significant finding, runbook tabletop completed in 2025 but scenario coverage thin (only one scenario tested; missing the multi-system cascading-failure scenario and the third-party model incident scenario); commitment to expand tabletop scope in Q1 2027. Review 9, IMV independence (September): one moderate finding, IMV function reports to Risk Officer but two of five FTE were previously model developers within the last 18 months; commitment to backfill assignments away from models the FTE developed until the 24-month involvement-clock runs. Review 12, ISO 42001 evidence-pack quality (September, coordinated with Schellman surveillance): zero findings, evidence pack synchronized with operations within the 30-day lag tolerance; cross-walk register current; control-owner attestations on cadence. Result fed directly into Schellman's surveillance audit, reducing surveillance fieldwork by ~18%.
- Q4 2026 (Reviews 10 + 8 + 4 + 2 + year-close). Review 10, Vendor DDQ (October): two significant findings, foundation-model provider attestation reliance documented but gap analysis incomplete on two of three providers; vendor DDQ refresh overdue on four of nine secondary vendors. Review 8, SR 11-7 MRM application (November): one significant finding, three of seven AI/LLM models on the MRM inventory but materiality assessment not refreshed in 14 months. Review 4, CAF function effectiveness (November): one observation, CAIO decision-tracking platform could integrate with AIMS portal. Review 2, AIRA breach response (December): two moderate findings, two breach events identified late in Q3. Year-close annual summary to Audit Committee in December: 9 substantive reviews complete; 12 findings raised across the year (0 critical, 4 significant, 6 moderate, 2 observations); 9 of 12 closed by year-end; 3 in progress within SLA.
2026 outcomes: IAA function stood up, operating, and producing measurable value. Audit Committee chair confirmed in December 2026 minutes: "The IAA function has matured rapidly. The coordination with external assurance reduced duplicative testing materially. Findings have driven concrete improvement in AI governance operations. The 2027 plan is approved with expanded scope including all 12 typical reviews." Schellman surveillance audit confirmed the IAA workpapers as cross-reference; SOC 2 + AI Type II issuance in November 2027 referenced IAA testing for the operating-effectiveness assertions on 6 of 12 AI-TSC control areas. KPMG co-source utilization 1,200 hours across the year, $340K against the $280K budgeted (overrun absorbed via 2027 reallocation).
Cross-walk register for Acme's IAA programme: each substantive review mapped explicitly to IIA IPPF Standards (1100 Independence, 1200 Proficiency, 2010 Planning, 2200 Engagement Planning, 2400 Communicating Results); EU AI Act Articles 17, 26, 27, 71, 72, 73, 89; NIST AI RMF Govern 6.2 + Manage 4.1; ISO 42001 A.9.2 internal audit + Clause 9.2; SR 11-7 governance pillar; PRA SS1/23 principle 4. The cross-walk register is the operational manifest enabling IAA workpapers to serve as cross-reference for ISO 42001 surveillance, SOC 2 + AI Type II, and notified-body conformity assessment without duplicative testing, the central L4 operational efficiency of the 2026 Three-Lines AI architecture.
Key Takeaways
- Internal AI Audit is the Third Line of Defense for AI: independent of 1L model owners and 2L AI Risk Office / MRM / AIGC; reports administratively to CEO and functionally to the Audit Committee chair; planned per IIA International Professional Practices Framework (IPPF) Standards adapted for AI; complementary to but distinct from external assurance (ISO 42001 surveillance, SOC 2 + AI attestation, notified-body conformity assessment).
- 8-section annual plan: risk assessment; audit universe with criticality tiers; audit topics for the year (8-12 substantive + 4 quarterly surveillance); resource plan (FTE + co-source + budget); calendar with quarterly milestones coordinated with external assurance; skills and training plan; coordination with external assurance; Audit Committee approval and reporting cadence. Plan re-baselined at mid-year if material change accumulates.
- 12 typical IAA reviews: AIGC effectiveness; AIRA breach response; RACI execution; CAF function effectiveness; Article 17 QMS; Article 72 PMM; red team operating model; SR 11-7 MRM application; IMV independence; vendor DDQ and evidence review; Article 73 incident response; ISO 42001 evidence-pack quality. Each 4-8 weeks, cross-walked to specific IPPF Standards and external-assurance frameworks. CAE selects 8-12 per year based on risk + rotation cycle.
- Quarterly review pattern, 3-week cycle from kickoff to issue with 4-8 week fieldwork. Finding rating four-level: critical (30-day SLA, weekly board escalation), significant (60-day, monthly), moderate (90-day, quarterly), observation (180-day, IAA log).
- Independence preservation, four operational rules: IAA cannot be involved in 2L control design; IAA reviews 2L process effectiveness; IAA cannot test its own work (24-month involvement-clock); IAA reports functionally to Audit Committee chair, not CRO or CAIO. Breaching any rule invalidates IAA work product for external-assurance cross-reference.
- Capability buildout, two hiring archetypes (audit-background-first + AI literacy training vs. AI-background-first + IPPF training); 1-3 FTE for mid-market + Big-4 or specialist co-source bench ($250-$450/hour); regulated industries 4-8 FTE; budget envelope $400K-$900K annual. Tooling: AuditBoard / Workiva / TeamMate+ for workpaper management; Promptfoo / Garak / PyRIT / Inspect AI plus AI-assisted document-comparison for AI-specific testing.
- External assurance coordination reduces duplicative testing 40-60%: ISO 42001 surveillance auditor receives IAA workpapers as cross-reference (15-25% surveillance reduction); SOC 2 + AI auditor relies on IAA per SSAE 18 internal-audit-reliance guidance (15-25% SOC 2 reduction); notified body uses IAA workpapers as Annex VII Module H inputs; competent authorities request IAA workpapers under Article 89.
- Six common 2026 IAA failures: treating IAA as compliance check; IAA reporting to CRO or CAIO; no AI literacy on audit team; no annual plan / ad-hoc reviews; skipping external-assurance coordination; testing only documentation not operating effectiveness. Each has documented remediation with 90-day cure window.
- Cross-walk register: IIA IPPF Standards 1100/1200/2010/2200/2400; EU AI Act Articles 17, 26, 27, 71, 72, 73, 89; NIST AI RMF Govern 6.2 + Manage 4.1; ISO 42001 A.9.2 + Clause 9.2; SR 11-7 governance pillar; PRA SS1/23 principle 4. Penalty exposure indirect, IAA finding closure delays surface in ISO 42001 Stage 2 surveillance; Article 89 requests; class-action discovery scope includes IAA records. Workpaper retention 7 years (10 years for high-risk system audits).
- Acme 2026 worked example: 9-review calendar, 2 FTE + KPMG co-source ($280K budgeted / $340K actual), AuditBoard tooling. Q1 founded; Q2 ran AIGC (2 moderate findings, charter scope + missed cadence) + red team (1 moderate finding, ASI Agentic coverage uneven); Q3 ran incident response (1 significant, thin tabletop) + IMV (1 moderate, 18-month involvement clock) + ISO 42001 evidence (0 findings, reducing Schellman surveillance ~18%); Q4 ran vendor + SR 11-7 + CAF + AIRA reviews; year-close 12 findings (0 critical, 4 significant, 6 moderate, 2 observations) with 9 of 12 closed by year-end. 2027 plan expanded to all 12 typical reviews.
Skill.re