Hosting an ISO 42001 Stage 2 Audit - Evidence Pack & Operating-Effectiveness Walks
T-minus-seven days. The Schellman engagement letter sat open on the Acme Inc CAIO's screen: three auditors on-site Monday for four days, ISO/IEC 17021-1 + ISO 19011 methodology, operating-effectiveness testing across the 38 Annex A controls and Clauses 4-10. The AIMS lead had spent the morning walking the 19-tab evidence pack with the document controller and surfaced three problems that should not exist seven days out: Tab 5 model-inventory query results were 96 days stale; Tab 11 was missing system-cards for two of the four Tier-1 systems the auditors had pre-named in scope confirmation; Tab 16 Article 73 incident log referenced a Q1 tabletop AAR never signed by the AIGC chair. The CAIO opened a calendar invite titled "Stage 2 war-room, daily 0730 standup" and forwarded it to the eight names that would run the audit. The internal pre-audit (lesson 058) had been done; the binder was real; the role-owners had been rehearsed. What remained was the hosting discipline. This lesson is the L4 host-the-audit playbook complementing lesson 058 practitioner-tier readiness: what the CAIO, AIMS lead, document controller, and 1L system owners do from T-30 through the closing meeting and into the 30-60-90 day finding-closure cycle.
Stage 2 Audit Context and Pre-Audit Logistics - T-30 to T-0
The Stage 2 audit is governed by ISO/IEC 17021-1 and ISO 19011. The certification body, Schellman, A-LIGN, BSI, KPMG most commonly for mid-2026 enterprise volume, with DEKRA, TÜV SÜD, LRQA, DNV active for industrial/notified-body-coordinated work, fields two to three auditors (one lead with ISO 42001 competence, one to two AI-specialist sector auditors) over three to five days on-site or hybrid. Stage 2 tests operating effectiveness: sampling evidence, walking workflows with role-owners, looking for the gap between documented procedure and operating reality.
The hosting discipline starts at T-30. Logistics confirmation: the certification body's pre-audit plan (2-4 weeks prior) names the lead auditor, sector auditors, scope confirmation, on-site/remote/hybrid decision, daily agenda by clause/control area, role-owner interview slots, expected evidence requests, and closing meeting time. The AIMS lead acknowledges in writing and confirms each named role-owner is available. On-site/remote decision: U.S. tech firms run 60-70% of Stage 2 audits remote or hybrid (Schellman, A-LIGN); BSI tends hybrid with on-site Day-1 opening and Day-N closing; KPMG tends fully on-site.
Secure-room provisioning: for on-site, a dedicated audit room with adjacent war-room space, named host who escorts auditors, signed NDAs, and a secure evidence-staging area (SharePoint or Box scoped to audit team plus document controller). For remote, a dedicated Zoom/Teams bridge with consented recording, screen-share discipline, and a chat channel for evidence requests with documented response-time targets (<60 minutes routine, <15 minutes in-walk).
Participant-availability matrix: every named interviewee: Responsible AI Officer (CAIO), AI Compliance Officer, AIGC chair, AI Risk Officer, AI Internal Audit Lead (3LoD), 1L system owners for sampled systems, data lead, legal lead, third-party-risk lead, security partner, HR partner (A.4 competence sampling): has a confirmed slot, a backup name, and a 30-minute pre-walk briefing 24-48 hours prior. Mock walkthroughs run T-7 to T-3: AI Compliance Officer plays auditor; 1L owner walks the workflow end-to-end; document controller traces evidence. Owners failing the first mock get a second pass at T-2; second-pass failures get the system de-scoped where possible or an external advisor sitting silently as backstop.
Evidence-pack final preparation: T-7 through T-0 is the integrity check. The document controller walks the 19 tabs, verifies version currency, confirms cross-references resolve, tests retrievability by someone other than the original author, and produces a sign-off memo to the CAIO + AIMS lead at T-3. Gaps at T-3 trigger remediation through T-0; gaps at T-1 trigger lead-auditor pre-disclosure (better to disclose than have the auditor find it).
The 19-Tab Evidence Pack - Master Structure with Named Documents
The 19-tab structure consolidates the lesson 058 Stage 1 binder with Stage 2 operating-effectiveness extensions (Tabs 14-19) and serves as the single source of truth for the audit:
- Tab 1 - AIMS policy + scope statement (Clause 4). Signed AIMS policy with version history and distribution acknowledgments; scope statement naming in-scope systems by system-card identifier, EU AI Act actor classification per system (Article 3(3) provider / 3(4) deployer / both), risk tier, deployment geographies, organizational units, explicit exclusions; signed by Responsible AI Officer and acknowledged by General Counsel.
- Tab 2 - Leadership commitment evidence (Clause 5; A.3). Board AI subcommittee minutes (last four quarters); AIGC charter (lesson 014); AI Risk Appetite Statement (AIRA, lesson 030) with board approval; Responsible AI Officer authority statement; Article 17(1)(k) designated QMS person; resource-allocation evidence demonstrating leadership commitment as documented action, not aspiration.
- Tab 3 - Planning + risk register (Clause 6; A.5). AI risk-assessment methodology (ISO/IEC 23894:2023-aligned); risk register; risk heat-map (lesson 085) with 12-month movement arrows; FAIR quantification (lesson 086) producing aggregate ALE-95 EUR; AIRA breach log naming KRI crossings and board-disclosure triggers; AI impact-assessment procedure (Clause 6.1.4) + Annex of completed assessments.
- Tab 4 - Resources + competence (Clause 7; A.4). RACI; role descriptions; AI-literacy training records keyed to roles (Article 4 deployer literacy; lesson 033); hiring records for AI-specialist roles; AI literacy program (lesson 032) curriculum + completion records + assessment evidence (lesson 034); tooling inventory; infrastructure inventory; data-resource catalogue.
- Tab 5 - Operations (Clause 8). Model inventory (lesson 015) live-query response timestamped within prior 30 days; FRIA registry (Article 27; lesson 042) listing all completed FRIAs by system + version + decision date + reviewer signature; Independent Model Validation (IMV) reports (lesson 044) for Tier-1 systems; substantial-modification change-control register (Article 43(4)); intended-purpose statements per system; deployment approval records.
- Tab 6 - Performance evaluation + monitoring (Clause 9; A.9). Post-market monitoring (PMM) dashboards (Article 72; lesson 080); drift methodology (KS test for input drift, PSI for prediction drift, performance-decay for concept drift) with documented thresholds; red-team metrics (lesson 083), coverage, severity, MTTR, reuse, with quarterly trend; incident logs; internal-audit reports (Clause 9.2); management-review minutes (Clause 9.3); CISO-CAIO joint quarterly review records.
- Tab 7 - Improvement + CAPA (Clause 10). Corrective-action register with named owners + dated commitments + verification mechanism + closure evidence; preventive-action register; lessons-learned database tied to incident AARs and tabletop AARs; trend analysis (recurring root causes); ISO 42001 Clause 10.2 evidence (continual improvement).
- Tab 8 - A.5 organizational AI policies. Full AI policy stack: board-approved AI policy; acceptable-use policy; data-handling policy; model-development standard; red-team policy; incident-response policy (lesson 049 + 050); third-party AI policy; AI ethics statement; sub-policy distribution log + annual review records + cross-reference table to InfoSec, privacy, ethics policies.
- Tab 9 - A.6 AI system lifecycle. Model lifecycle SOPs covering intake (A.6.1.1), design (A.6.1.2), V&V (A.6.1.3), deployment (A.6.1.4), monitoring (A.6.1.5), technical documentation (A.6.1.6), event logs (A.6.1.7), requirements (A.6.2); change-control procedure for substantial modifications; deprecation/decommissioning procedure; per-Tier-1 system: full lifecycle artifact bundle.
- Tab 10 - A.7 data for AI systems. Datasheets (Gebru et al. 2021 schema or equivalent) per training dataset; data-governance framework; data-quality metrics (completeness, accuracy, consistency, timeliness, validity, uniqueness) measured; provenance metadata via ML-BoM / SBOM-AI (CycloneDX 1.7 + CDXA); Article 10 evidence (training-data governance for high-risk systems); preparation logs; licensing documentation.
- Tab 11 - A.8 information for users. Model cards (Mitchell et al. 2019 schema or equivalent) per system, published externally where applicable; system cards for agentic systems; Annex XII templates (deployer information for high-risk systems) populated; Article 13 instructions-for-use; external transparency cadence record; user documentation per system; A.8.4 stakeholder-communication records.
- Tab 12 - A.9 post-deployment. PMM evidence per system; drift detection dashboards with alert history; performance-decay tracking; concept-drift methodology execution records; responsible-use procedures; deployer-facing guidelines; deviation-detection logs; alignment-with-intended-purpose verification (A.9.4).
- Tab 13 - A.10 third-party. Vendor inventory; vendor DDQ records (lesson 040); SOC 2 + AI attestation reviews (lesson 110); Annex XII receivable verification for GPAI providers; Article 25(2) cooperation evidence, memos documenting cooperation with upstream providers on substantial modifications and incident-affecting changes; supplier-incident workflow records; customer-information procedure evidence.
- Tab 14 - Article 17 QMS cross-walk register (lesson 053). Each of the 10 Article 17 sub-areas (a-k) mapped to AIMS clause + Annex A control + evidence reference + responsible role. Required where the provider is also a high-risk system provider under Article 6.
- Tab 15 - Internal-audit + management-review records. Internal-audit program (Clause 9.2); auditor qualifications (ISACA AI audit certification or IIA AI audit guidance); audit-plan, work papers, reports for prior 12 months; management-review minutes (Clause 9.3) covering documented inputs (audit results, monitoring outputs, incident records, vendor performance, KRI trend, AIRA status, external context changes) and decisions; corrective-action register linkage.
- Tab 16 - Article 73 incident log + tabletop AAR. Article 73 serious-incident reporting log (lesson 089); tabletop drill AARs (lesson 089) for prior 12 months; integrated AI + cyber drill AAR; incident-runbook (lesson 050) version history; competent-authority engagement records (lesson 090); cross-walk to A.6.2.7 + Article 73 + NIST Manage 4.
- Tab 17 - Notified-body coordination (where applicable). Annex VII Module H engagement records (lesson 052); cross-walk between ISO 42001 audit and Module H conformity-assessment audit; dual-framework evidence-sharing memo; notified-body contact + cooperation record. Omit tab if not applicable but document the not-applicable justification in Tab 1 scope statement.
- Tab 18 - EU AI Act Article 71 registration records. EU database registration (Article 71) for each high-risk system; registration update records for substantial modifications; Annex VIII (Section A for providers / Section B for deployers) information current; declaration of conformity (Article 47); CE marking application records; Article 49 EU declaration of conformity records.
- Tab 19 - Regulator engagement log. EU AI Office engagement records (lesson 090); national supervisory authority (MSA) engagement records, BSI U.K., AGID Italy, ACPR France, BNetzA Germany, NCSP Spain; competent-authority readiness exercise participation; AGCM/AGCOM consumer-protection coordination (where applicable); sector-specific regulator engagement (FCA, OCC, FDA, FTC depending on sector); regulator-information-request response log.
Binder discipline: each artifact stored once in a master AIMS portal with explicit cross-references from every tab, the Tab 11 model card is the same artifact referenced in Tab 9 (A.6.1.6), Tab 14 (Article 11 + Annex IV), and Tab 13. Quarterly cross-reference walk-through is required; the T-7 to T-3 integrity check is the audit-window verification. Inconsistent cross-references are a Stage 2 minor finding at every certification body.
Operating-Effectiveness Walks and the 12 High-Frequency Stage 2 Questions
The auditor selects 3-5 controls per day, asks the 1L owner to walk end-to-end execution on a specific recent instance, traces evidence, and looks for the gap between documented procedure and operating reality. Walks run 45-90 minutes. Schellman emphasizes A.6 lifecycle + A.7 data + A.10 third-party; A-LIGN emphasizes regulated-industry cross-walks; BSI emphasizes EU AI Act dual-framing; KPMG emphasizes board-level governance evidence: plus risk-based selection on recent incidents, prior-finding areas, and high-stakes systems.
The 12 high-frequency Stage 2 questions in mid-2026 (drawn from Schellman / A-LIGN / BSI / KPMG Q1-Q2 2026 first-wave audit transcripts and ISACA AI Audit working-group practice data):
- (1) "Show me a Tier-1 system FRIA." Owner pulls the FRIA from Tab 5 FRIA registry; walks methodology (Article 27 alignment); names the affected-population analysis (A.5.4 individual/group); shows the societal-impact section (A.5.5); traces the signature page; demonstrates integration with risk treatment in Tab 3 risk register. Common gap: FRIA exists but is template-filled without evidence of operational analysis behind each section.
- (2) "Show me how a substantial modification was handled." Owner pulls the change record from Tab 5 substantial-modification register; walks Article 43(4) trigger analysis; shows the FRIA refresh, IMV re-execution, model-card revision, deployer notification (Annex XII), Article 71 registration update, board notification (if Tier-1). Common gap: change was processed as routine maintenance with no Article 43(4) analysis documented.
- (3) "Show me Article 73 incident-reporting evidence." Owner pulls Tab 16 incident log; walks a specific incident: detection, Article 3(49) determination within 4 hours, AIGC chair notification within 30 minutes of determination, containment, report drafting start within 2 hours, 15-day (or 2-day widespread) clock tracking, MSA submission, 90-day follow-up. Common gap: log lists incidents but timestamps are aggregated to days, not hours, so clock-discipline evidence is unverifiable.
- (4) "Show me the model-inventory query response." Owner runs the live query in the audit room; pulls back the current inventory; demonstrates the lesson-015 schema fields are populated; shows last-updated timestamps per system; demonstrates the inventory is the source of truth (not a static spreadsheet). Common gap: inventory is a SharePoint file last edited months ago, not a queryable system.
- (5) "Show me red-team finding closure." Owner pulls Tab 6 red-team metrics; walks a Critical-severity finding from detection to closure: CVSS-AI score, owner assignment, MTTR clock, remediation evidence, regression test (lesson 084), closure sign-off. Common gap: findings logged but closure evidence is a Jira ticket marked done without verification artifact.
- (6) "Show me CAPA from last quarter." Owner pulls Tab 7 corrective-action register; walks an action item: root cause, named owner, dated commitment, success criterion, verification mechanism, closure evidence; demonstrates trend analysis (recurring root causes); shows preventive-action distinct from corrective-action. Common gap: register exists but actions have no verification mechanism beyond "owner attests done."
- (7) "Show me Annex XII deployer notification." Owner pulls Tab 11 Annex XII template populated for a specific system; walks the instructions-for-use (Article 13); demonstrates the deployer received the information at deployment + at substantial modification; shows acknowledgment or delivery confirmation. Common gap: template exists but delivery records are absent; can't prove deployers received it.
- (8) "Show me Article 25(2) cooperation memo." Owner pulls Tab 13 third-party records; walks a memo documenting cooperation with an upstream GPAI provider on a substantial modification, incident, or significant change; demonstrates Annex XI receivable verification; shows the cooperation was bilateral (provider response + provider obligations). Common gap: provider relies on upstream attestation as substitute for active cooperation memo.
- (9) "Show me a vendor DDQ." Owner pulls Tab 13 vendor DDQ; walks the DDQ structure (lesson 040); shows the response, the internal-review notes, the SOC 2 + AI evidence (lesson 110), the ISO 42001 certification verification, the AI-specific evidence supplement, the contract-clause incorporation. Common gap: DDQ exists for top tier vendors but is missing for Tier-2/3 vendors that the model-inventory shows are in production use.
- (10) "Show me drift detection." Owner pulls Tab 12 PMM evidence for a specific system; walks the drift methodology, input drift (KS test on feature distributions); prediction drift (PSI on score distributions); concept drift (performance decay against ground truth); fairness drift (subgroup metric stability); shows recent alerts and the disposition of each. Common gap: dashboards exist but no documented threshold, alert history, or review cadence.
- (11) "Show me AIGC minutes." Owner pulls Tab 2 AIGC charter and minutes; walks last-quarter minutes; demonstrates the AIGC reviewed risk-register changes, AIRA breaches, incident log, vendor changes, substantial modifications; shows decisions tracked to closure; demonstrates 1L attendance and 2L independent challenge. Common gap: minutes are present but read as status reports without decisions or challenges documented.
- (12) "Show me independent challenge from 2L." Owner pulls Tab 6 internal-audit reports + Tab 15 management-review minutes; walks a specific 2L challenge, AI Risk Officer formally objected to a 1L decision (e.g., refused to sign a FRIA, escalated a risk decision to AIGC, required additional V&V on a deployment); shows the documented resolution. Common gap: 2L exists organizationally but no record of independent challenge, suggests 2L is captured or under-resourced.
Mock-walkthrough discipline (T-7 to T-3): each 1L owner walks each of the 12 questions with the AI Compliance Officer playing auditor. Owners answering fluently in <5 minutes are ready; owners who need slides, ask for help, or contradict documented procedure get additional rehearsal or de-scoping. Owners get a one-page cue card with the 12 questions and tab-and-document path to each answer, the cue card is permitted in interview; relying on it constantly is the gap signal.
War-Room Composition and On-the-Spot Finding-Management
Too small and capacity is exceeded on Day-2 evidence requests; too large and the audit becomes chaotic. The defensible 2026 composition for mid-market focused-scope audits (3-auditor team, 3-5 days):
- CAF lead (lesson 077) - Conformity Assessment Function head; the single point of accountability for the audit; reports to CAIO. Does not attend interviews unless invited; runs the war-room daily standup; arbitrates evidence-request prioritization; signs the daily memo to the lead auditor.
- AI Compliance Officer, the audit facilitator inside the room; paces sessions; takes notes during interviews; ensures the auditor receives the right document version; coordinates with the document controller; raises objection-to-finding (rare) when factually wrong.
- 1L system owners on-call, the sampled-system owners scheduled for interviews; otherwise in war-room standby; available within 15 minutes for follow-up clarifying questions.
- Document controller, operates the evidence-staging area; receives auditor requests via the chat channel; pulls the named document; verifies version + signature + cross-references; delivers within target response time; logs every delivery for the audit chain-of-custody record.
- Audit facilitator (often the AIMS lead or an external advisor): the in-room session pacer; introduces interviewees; manages clock; takes structured notes on auditor questions, owner answers, evidence shown, side comments; produces the daily memo for the war-room debrief.
- External advisor (typically), the consulting firm or counsel partner who helped build the AIMS or has prior experience hosting Stage 2 audits with this certification body; sits silently in interviews where requested; advises on objection-to-finding posture; reviews findings draft.
Daily standup (0730-0800): CAF lead chairs; reviews previous-day findings catalog from the lead auditor's debrief; previews today's interview slate; confirms 1L owner readiness; flags evidence-request risks; assigns runners for the day. Daily debrief (1730-1830 after audit close): the audit facilitator walks notes; document controller reports delivery exceptions; objections-to-finding considered; T-7 from closing meeting forward, the war-room drafts the closing-meeting talking points.
On-the-spot finding-management: when the auditor verbalizes a potential finding during a walk, the audit facilitator captures it verbatim, asks one clarifying question (multiple questions read as defensiveness), and offers additional evidence if available. The 1L owner does not argue, does not concede classification (only the auditor classifies), and does not promise corrective action in-session. Post-walk, the war-room evaluates whether the finding is factually correct, factually wrong (basis for objection), or factually correct but downgradable with additional evidence. Objections are raised only when factually wrong, in writing, calmly, within 24 hours. Defensive objections (auditor is right and provider disagrees) escalate scrutiny, every additional control gets a tighter look. Mature programs raise one or two on a 4-day audit and win them.
Document control during the audit: every document delivered is logged with timestamp, requester, document name, version, custodian, delivery method. The auditor's request channel is single, chat or audit room with the document controller, never side-channel via interviewee. A 1L owner pulled mid-interview into "I'll email you a copy" is the most common document-control breakdown; the answer is always "the document controller will deliver it to the audit room within 30 minutes."
Finding Classification, 30-60-90 Day Closure, and Seven Common Hosting Mistakes
The closing meeting presents preliminary findings. CAIO, AIMS lead, General Counsel, audit facilitator, and (KPMG-style) board observer attend. The lead auditor walks findings with classification: Critical (Stage 2 conditional; certificate withheld until remediation + supplementary audit within 30-90 days; rare in pre-audited engagements); Major (corrective-action plan within 30 days; remediation evidence within 60-90 days; certificate conditional with verification at first surveillance); Minor (corrective action by next surveillance; does not affect certificate); Observation / OFI (informational). The provider clarifies, presents additional evidence, contests classifications (formally, calmly), and accepts findings as basis for the corrective-action plan.
30-60-90 day closure (lesson 052 + 078 patterns): 30 days: corrective-action plan submitted with named owner, dated commitment, root-cause analysis, remediation steps, verification mechanism, success criterion, AIGC chair sign-off. 60 days, remediation in progress; mid-point evidence package; 2L verification of effectiveness. 90 days, closure evidence package; auditor verification (paper review or supplementary on-site for Critical/Major); closure letter; corrective-action register + lessons-learned database updated. Findings slipped past 90 days escalate to Major classification at next surveillance.
Seven common 2026 hosting mistakes, the recurring failure modes the lesson 058 pre-audit work did not eliminate, surfacing during the host-the-audit window:
- Mistake 1 - Evidence pack ad-hoc-assembled. The 19 tabs are assembled in the final two weeks rather than maintained as the operating artifact; auditor finds inconsistencies (Tab 5 references "model card v3" while Tab 11 has v2; Tab 9 SOP names "AI Review Committee" while Tab 2 charter names "AI Governance Committee"). Fix: the binder is a living artifact maintained quarterly; T-7 integrity check verifies, doesn't construct.
- Mistake 2, 1L owners not prepared. Mock walkthroughs skipped or compressed; owners cannot answer "show me a FRIA" without 10 minutes of searching; contradict documented procedure under pressure. Fix: mock walkthroughs at T-7 and T-3 are non-negotiable; failing owners are de-scoped or backstopped; cue cards prepared and rehearsed.
- Mistake 3 - Too many people in the room. Eight names in war-room becomes 18; auditor walks for a coffee and overhears the AIGC chair criticizing the AIMS lead; chaos shows. Fix: war-room cap at 6-8 named roles; observers (board members, executive sponsors) get a daily 15-minute debrief, not seated presence.
- Mistake 4 - Document control breakdown. Auditor requests "the policy"; three people email three versions; auditor receives the wrong version; the next 30 minutes are spent reconciling. Fix: single document-controller channel; logged delivery; named version + signature + retrieval-path every time.
- Mistake 5 - No pre-audit dry-run. The provider runs the lesson 058 pre-audit work but never runs a host-the-audit dress rehearsal, war-room standup, interview pacing, evidence-request response, so the operational discipline is untested. Fix: at T-14 to T-10, run a half-day dress rehearsal with an external advisor playing lead auditor and the war-room operating as it will during the live audit; surface logistic gaps.
- Mistake 6 - Defensive posture. 1L owner argues with the auditor mid-walk; multiple objections raised on the first day; the auditor escalates scrutiny, every subsequent control gets a tighter look; the 3-minor expected finding profile becomes 1-major-3-minor. Fix: the audit posture is collaborative, not adversarial; objections are raised only when factually wrong, in writing, calmly, after the war-room evaluates.
- Mistake 7 - Skipping the closing meeting. CAIO sends a deputy because of another conflict; closing meeting runs without senior accountability; classification gets accepted without challenge or with reactive over-challenge; no shared understanding of findings. Fix: the CAIO + General Counsel + AIMS lead are present at the closing meeting; the closing meeting is the most operationally consequential 90 minutes of the audit; the corrective-action plan starts from a shared finding-understanding, not a deferred read.
Worked Example - Acme Inc Stage 2 Audit, Q3 2026
Acme Inc - U.S. healthcare-adjacent SaaS, $2.1B revenue, 17 in-scope AI systems (3 Tier-1 Annex III high-risk, 8 Tier-2 limited-risk, 6 Tier-3 minimal-risk per lesson 015 inventory), Schellman as certification body, ANAB accreditation, Stage 2 audit 17-20 August 2026. Audit team: lead auditor (Schellman Director, ISO 42001 Principal Auditor, prior ISO 27001 + SOC 2 history with Acme); two sector auditors (healthcare-AI specialist + ML-platform infrastructure specialist). Four days on-site at Acme SF HQ; half-day remote reserved for supplementary closing review.
The 19-tab evidence pack was finalized 8 August (T-9). Document-controller integrity-check memo to CAIO 10 August (T-7) named two gaps, Tab 11 system-card for ServiceAssist v2.3 last updated 92 days prior (policy: 90); Tab 16 Q1 tabletop AAR had no AIGC chair signature. Both remediated by 13 August (T-4). Pre-disclosure: AIMS lead emailed Schellman 14 August summarizing gaps + remediation; no finding raised on either at audit.
Mock walkthroughs ran T-7 and T-3. Six 1L owners covered three Tier-1 + three Tier-2 systems. Two failed the first mock; one passed the second; one was de-scoped (auditor-named sampling pool included a Tier-2 system whose owner had been promoted out 60 days prior; the new owner was 30 days in, Acme proposed substitution; Schellman accepted).
Audit Day 1 (Mon): opening 0830-0930; Clauses 4-6 review 0930-1200; A.3 + A.5 sampling 1300-1700. Day 1 debrief: lead auditor verbalized two minor findings, A.5 FRIA registry had two entries with template-language identical to the methodology document (copy-paste without operational analysis); A.4 AI-literacy training records showed two recently-hired ML engineers had not completed Article 4 training within the 30-day post-hire window. War-room logged both, did not concede classification, prepared evidence for Day 2.
Audit Day 2 (Tue): Clauses 7-8 + A.6 lifecycle deep sampling. Three Tier-1 systems; eight operating-effectiveness walks (FRIA, substantial-modification, V&V, model-card, change-control). Schellman's 12-question rotation surfaced one major-classifiable finding: model-inventory query for the third Tier-1 system returned a stale record (live database at v2.4 deployed June; inventory at v2.3); AIMS lead acknowledged immediately, traced root cause to manual-inventory-update gap on the model-registry promotion pipeline, committed to fix. Lead auditor classified preliminarily as minor pending remediation; Acme accepted.
Audit Day 3 (Wed): A.7 + A.8 + A.9 + A.10 sampling. Vendor DDQ walk; Annex XII deployer notification walk; Article 25(2) cooperation memo walk. Surfaced: Tab 11 Annex XII templates for two Tier-1 systems had not been re-issued after Q2 substantial-modification (content current; delivery record to named deployers missing). Classified preliminarily as minor pending remediation.
Audit Day 4 (Thu): Clauses 9-10 + internal audit + role-owner interviews + findings consolidation + closing meeting. AIGC chair, AI Risk Officer, AI Internal Audit Lead interviewed. Acme's 2L AI Risk Officer had formally challenged a 1L FRIA in Q2 (documented in management-review minutes; auditor noted approvingly). Closing meeting 1630-1800.
Closing-meeting output: 0 critical; 0 major; 3 minor; 2 OFI. Minor 1, red-team coverage rubric (lesson 083) showed Tier-1 ServiceAssist coverage at 78%, below 85% Tier-1 threshold; remediation = coverage closure plan within 60 days. Minor 2 - Annex XII Q2 substantial-modification delivery record gap; remediation = records reconstituted within 30 days. Minor 3 - AIRA breach log: one Q2 breach logged 7 days after the event (policy: 24 hours); remediation = logging-discipline corrective action verified at next surveillance. OFI 1, model-inventory automation (Day 2 finding remediated in-session manually; OFI for full automation by Year 2). OFI 2 - AI-literacy completion-timing for new hires (Day 1 observation). Corrective-action plan due 17 September; 60-day mid-point 17 October; 90-day closure 17 November.
Acme submitted the corrective-action plan 12 September (5 days ahead); 60-day mid-point 14 October; 90-day closure 11 November. Schellman issued the certificate 21 November 2026, valid through November 2029, with surveillance audits August 2027 + August 2028, recertification August 2029. Total Stage 2 + closure cycle: 96 days. The lesson 077 CAF function logged the closed cycle as Tab 15 evidence for next surveillance.
Cross-Walks, Surveillance Cadence, and Penalty Exposure
The Stage 2 evidence pack supports the full regulatory stack with the same artifacts. Cross-walks: ISO/IEC 42001:2023, Clauses 4-10 + Annex A 38 controls covered by Tabs 1-13 + 15. ISO/IEC 17021-1 + ISO 19011, audit methodology adhered by the certification body, validated by ANAB or UKAS accreditation surveillance. EU AI Act: Article 17 QMS ten sub-areas cross-walked in Tab 14; Article 26 + 27 + 71 + 72 + 73 evidence in Tabs 5, 6, 11, 12, 16, 18. NIST AI RMF: Govern (Tabs 1, 2, 4, 8); Map (Tabs 3, 9, 10); Measure (Tabs 6, 12); Manage (Tabs 5, 7, 12, 16). SR 11-7 + OCC 2011-12 + PRA SS1/23, governance pillar (Tabs 1-4) + validation pillar (Tab 5 IMV) + monitoring pillar (Tab 6) + documentation pillar (Tab 9). SOC 2 + AI (lesson 092) and AICPA AI Trust Services Criteria, shared infrastructure controls plus AI-specific extensions.
Surveillance + recertification cadence: ISO 42001 certificate valid 3 years; annual surveillance audits at 30-50% of Stage 2 effort covering continued operating effectiveness, prior-finding remediation verification, highest-risk areas (A.6 / A.7 / A.8 / A.10), and substantial-modification change-control; recertification at Year 3 at 60-80% of Stage 2 effort covering full clause + Annex A review. Missing a surveillance window invalidates the certificate. The host-the-audit discipline this lesson operationalizes is run again for each surveillance: with reduced scope, the same evidence-pack-integrity, mock-walkthrough, war-room, and finding-management posture.
Penalty exposure: an ISO 42001 finding does not itself trigger EU AI Act penalties, the certification body's role is conformity attestation, not enforcement. But certification status is operationally consequential. Certificate suspension or withdrawal (un-closed Critical findings or systemic Major findings) means market-access loss where ISO 42001 is a procurement gate (most EU public-sector; financial-services third-party-risk for AI vendors; regulated-healthcare AI suppliers). Article 70 MSAs routinely request ISO 42001 audit reports in supervisory engagements (lesson 090), a Major finding becomes a known data point for the competent authority's next-step assessment. Article 99(3) €15M / 3% of global turnover exposure crystallises on the underlying Article 17 QMS, Article 27 FRIA, Article 72 monitoring, Article 73 reporting failures the ISO 42001 finding documents. For U.S. enterprises with SOC 2 + AI attestation (lesson 092), AICPA auditor reuse of ISO 42001 evidence reduces total assurance cost 20-30%.
Key Takeaways
- Stage 2 hosting is the L4 leadership-tier discipline that complements lesson 058 practitioner-tier readiness. The pre-audit work proves the AIMS exists; the host-the-audit work proves the program can operate under regulator-grade scrutiny: pre-audit logistics (T-30 to T-0), evidence-pack final integrity check, war-room composition, operating-effectiveness walk choreography, finding-management posture, post-audit closure.
- The 19-tab evidence pack consolidates the lesson 058 Stage 1 binder with Stage 2 operating-effectiveness extensions: Tabs 1-7 cover Clauses 4-10; Tabs 8-13 cover Annex A.5-A.10; Tabs 14-19 cover Article 17 QMS, internal-audit + management-review, Article 73 incident log, notified-body coordination, Article 71 registration, and regulator-engagement records. Each artifact stored once; quarterly cross-reference walk-through; T-7 integrity check verifies, doesn't construct.
- Operating-effectiveness walks are the core sampling mechanic. Auditor selects 3-5 controls per day; walks 1L owner end-to-end on a specific recent instance; traces evidence; looks for gap between documented procedure and operating reality. Walks run 45-90 minutes per control.
- The 12 high-frequency 2026 Stage 2 questions are rehearsable: show me a Tier-1 FRIA; substantial modification; Article 73 incident; model-inventory query; red-team finding closure; CAPA last quarter; Annex XII deployer notification; Article 25(2) cooperation memo; vendor DDQ; drift detection; AIGC minutes; independent challenge from 2L. Mock walkthroughs at T-7 and T-3 with the AI Compliance Officer playing auditor.
- War-room composition for mid-market focused-scope audits: CAF lead, AI Compliance Officer, 1L system owners on-call, document controller, audit facilitator, external advisor, 6-8 named roles. Daily standup 0730; daily debrief 1730. Single document-controller channel for all evidence delivery. Objections-to-finding raised only when factually wrong, in writing, calmly, within 24 hours.
- Finding classification drives the timeline: Critical (Stage 2 conditional; remediation + supplementary audit within 30-90 days; rare in pre-audited engagements); Major (corrective-action plan 30 days; remediation evidence 60-90 days; certificate conditional); Minor (corrective action by next surveillance); OFI (informational). 30-60-90 day closure cadence with named owner, verification mechanism, AIGC chair sign-off.
- Seven common 2026 hosting mistakes: evidence pack ad-hoc-assembled; 1L owners not prepared; too many people in the room; document control breakdown; no pre-audit dry-run; defensive posture; skipping the closing meeting. Each has a documented fix: living binder, mock walkthroughs, capped war-room, single document-controller channel, T-14 dress rehearsal, collaborative posture, senior closing-meeting attendance.
- Acme Q3 2026 worked example: 3-auditor Schellman team, 4 days on-site, 19-tab evidence pack with two T-7 integrity-check gaps remediated and pre-disclosed, 11 operating-effectiveness walks executed, closing-meeting output 0 critical + 0 major + 3 minor + 2 OFI, 30-60-90 day closure executed 5 days ahead of schedule, certificate issued 21 November 2026 valid through November 2029.
- Cross-walks: ISO/IEC 42001:2023 + ISO/IEC 17021-1 + ISO 19011 methodology; EU AI Act Articles 17 + 26 + 27 + 71 + 72 + 73; NIST AI RMF Govern + Map + Measure + Manage; SR 11-7 + OCC 2011-12 + PRA SS1/23 governance pillar; SOC 2 + AICPA AI TSC. One evidence pack, multiple frameworks of attestation.
- Penalty exposure: certificate suspension = market-access loss in EU public-sector procurement, financial-services AI-vendor requirements, regulated-healthcare AI-supplier requirements. EU AI Act Article 70 designated MSAs request ISO 42001 audit reports in their own supervisory engagements; Major findings become known data points. Article 99(3) €15M / 3% of global turnover exposure crystallises on the underlying Article 17 + 27 + 72 + 73 failures the ISO 42001 finding documents. AICPA AI auditor coordination reduces total assurance cost 20-30% via cross-walk discipline.
Skill.re