ISO/IEC 42001 Annex A Controls - Gap Assessment
In February 2026, a Fortune-500 industrial software vendor decided to pursue ISO/IEC 42001 certification with a Stage 1 audit scheduled for Q3. The program lead, fresh from a successful SOC 2 Type II and ISO 27001 surveillance, proposed skipping a formal pre-Stage-1 gap assessment on the rationale that "we know our control posture from prior audits." Twelve weeks later, three days into Stage 1, the auditor logged seven documentation findings: an AIMS scope statement that excluded two in-scope systems, missing per-control evidence inventory for 14 of the 38 Annex A controls, no documented Statement of Applicability justification for A.5 impact-assessment exclusions, and an internal-audit plan that covered Clauses 4-10 but had not sampled any of the 38 controls in operating terms. Stage 2 slipped a full quarter. The chief AI officer's Q4 customer-commitment calendar, three Fortune-50 procurement deals contingent on certification, re-scheduled. This lesson is the slow, regulator-grade walk through all 38 controls and the per-control evidence inventory the program team should have built before the Stage 1 booking, and the gap-remediation roadmap with target evidence that turns "we think we are ready" into "we are 82% operating evidence and the remaining 18% have named owners and dates."
Why the Pre-Stage-1 Gap Assessment Decides Stage 2 - Findings Cost Quarters
ISO/IEC 17021-1, the conformity-assessment standard governing management-system certification, defines Stage 1 as documentation readiness and Stage 2 as operating effectiveness. The two-stage design is intentional: Stage 1 catches documentation gaps before the more expensive Stage 2 effort. In ISO 27001 and ISO 9001, both decades-mature, most well-prepared organizations pass Stage 1 with minor opportunities and proceed to Stage 2 within 4-8 weeks. ISO 42001 in 2026 does not follow that pattern. First-cycle audits surface significantly higher Stage 1 finding counts because (1) Annex A's 38 controls cover AI-specific evidence territory absent from ISO 27001 / SOC 2 control families, (2) the seven mandatory documents impose specificity (scope statement naming systems by ID; impact-assessment process integrating Article 27 FRIA mechanics; roles with explicit halt authority) the organization's existing documentation typically does not satisfy, and (3) internal-audit programs that worked for ISO 27001 ISMS audit lack the AI-specific competence to meaningfully sample A.6 lifecycle, A.7 data, and A.8 user-information evidence.
The cost of underestimating this is asymmetric. A Stage 1 documentation finding triggers a 30-90 day remediation window before Stage 2 can begin. A clustered set of findings, typical when no pre-Stage-1 assessment is done, slips Stage 2 by a full quarter. For provider organizations whose customer-procurement calendars assume Q-end certification dates, a missed quarter cascades into missed deals, missed insurance underwriting cycles, and missed board commitments. The pre-Stage-1 gap assessment is not optional preparation; it is the artifact that lets the certification body book Stage 1 confidently and the program lead defend the Stage 2 date to the executive team.
A defensible 2026 gap-assessment artifact has four layers: (1) per-control evidence inventory: what evidence exists today, where it lives, who owns it; (2) per-control maturity score (0-5) against an explicit rubric; (3) per-control gap-fill plan with owner, target date, and the specific evidence to be produced; (4) aggregate readiness score against the Stage 1 audit threshold (typically ~80% of controls at maturity 4+ with the remaining 20% having documented remediation plans). This lesson walks all 38 controls with the per-control evidence expectation and the typical maturity-state observed in 2026 first-cycle organizations, the worked Fortune-500 example, the cross-walks to EU AI Act / NIST AI RMF / OWASP-ATLAS that turn one evidence pack into three regulators of value, and the six common methodology mistakes that produce assessments which look defensible on paper but fail the Stage 1 dialogue.
Walking All 38 Controls - Evidence Expectation and Typical Maturity by Control
The walkthrough below covers all 38 controls organized by the nine Annex A areas. For each control: the evidence expectation (what an auditor expects to see at Stage 1 and sample at Stage 2), the typical first-cycle maturity state, and the most common gap.
A.2 Policies Related to AI, 2 Controls
A.2.2 AI policy. Evidence: signed top-management AI policy, version-controlled, intranet-published, acknowledgment records, annual review minutes. Typical state in organizations with prior ISO 27001 maturity: yellow: a draft AI policy exists but lacks board signature, explicit AI-risk-appetite reference, or operational integration into intake forms and ARB workflows. Common gap: policy exists on paper but does not appear in the system-intake form, the FRIA workflow, or the architecture-review-board log, a Stage 2 operating-effectiveness failure even when Stage 1 documentation passes.
A.2.3 Alignment with other organizational policies. Evidence: cross-reference table between the AI policy and information-security policy, privacy policy, ethics policy, risk-management policy, BCM policy, vendor-management policy; gap-analysis memo identifying overlaps and clean delineations. Typical state: yellow, the cross-reference table exists in piece-meal form but has not been formally documented. Common gap: ambiguity between AI policy and existing ethics or vendor-management policy on whose authority approves new AI suppliers.
A.3 Internal Organization, 3 Controls
A.3.2 AI roles and responsibilities. Evidence: RACI matrix; role descriptions for Responsible AI Officer / Chief AI Risk Officer, AI Governance Committee chair, AI Risk Officer, AI Privacy Officer, AI Security Officer, AI Internal Audit Lead, engineering function leads; reporting-line documentation; explicit authority statement (especially halt authority). Typical state: green for organizations that have already named a Chief AI Officer; yellow for organizations where the AI Officer reports through legal or compliance without explicit halt authority. Common gap: the Responsible AI Officer has implicit but not documented authority to halt non-compliant deployments, the auditor at Stage 2 samples halt-or-pause decisions and finds the authority operates ad hoc.
A.3.3 Reporting of concerns. Evidence: AI-specific reporting channel (often integrated with the ethics hotline with AI intake categories), case-management workflow, anonymous reporting option, response-time metrics, sampled cases. Typical state: yellow, the ethics hotline exists but lacks AI-specific intake categories. Common gap: no recorded AI-specific concerns in the prior audit period, which the auditor will probe, silence is not evidence of effectiveness, and the workflow should produce occasional intake even from training and awareness drills.
A.4 Resources for AI Systems, 6 Controls
A.4.2 Resource documentation. Evidence: resource inventory per in-scope system, capacity planning, budget allocation, cross-reference to AIMS scope. Typical state: yellow, engineering has the resource inventory but it is not formally tied to AIMS scope.
A.4.3 Data resources. Evidence: data catalogue with AI-relevant entries, classification, governance integration, training-data inventory keyed to each in-scope system. Typical state: yellow, data catalogue exists but training-data lineage for each system is fragmented. Cross-walks to A.7 and Article 10.
A.4.4 Tooling resources. Evidence: tooling inventory (training, evaluation, monitoring, red-team frameworks), version control, security baseline, sandboxing. Typical state: yellow, tooling exists across engineering teams but the consolidated inventory is missing.
A.4.5 System and computing resources. Evidence: infrastructure inventory, GPU/TPU allocation records, cloud-provider attestations (SOC 2, ISO 27001, ISO 42001 where available), capacity records. Typical state: green for cloud-native organizations with strong CSP evidence, typically the easiest A.4 control to evidence.
A.4.6 Human resources. Evidence: AI-role headcount per function, AI-literacy training records keyed to roles, competency framework, Article 4 deployer-literacy training completion records. Typical state: yellow approaching red, Article 4 deployer-literacy training has been rolled out but completion records and per-role customization are incomplete. The most-scrutinized A.4 control because Article 4 has operational deadlines.
A.5 Assessing Impacts of AI Systems, 5 Controls
A.5.2 AI system impact-assessment process. Evidence: documented process, triggers (new system intake; substantial modification per Article 43; new use context; data-source change), evaluation content, conductor and approver, integration with risk treatment. Maps to mandatory Document 4. Typical state: yellow, the process exists for new high-risk systems but is not consistently applied to substantial-modification triggers.
A.5.3 Documentation of impact assessments. Evidence: completed impact-assessment records, review-and-approval signatures, version history, traceability into risk treatment and AIMS risk register. Typical state: yellow, completed assessments exist but signatures and version history are incomplete.
A.5.4 Impact on individuals or groups of individuals. Evidence: individual/group analysis methodology, identification of affected individuals (cross-walks to Article 3(32) "affected persons"), severity/likelihood analysis, mitigation rationale. Typical state: yellow, the analysis is present but methodology for identifying affected groups is informal. Most-scrutinized A.5 control for employment, credit, education, essential-services systems.
A.5.5 Societal impacts. Evidence: broader societal-impact analysis, methodology referencing NIST Map 3 context-related risks, external-stakeholder engagement records where applicable. Typical state: red, most first-cycle organizations have not addressed societal-impact methodology beyond the immediate-stakeholder analysis. The biggest A.5 gap in first-cycle audits.
A.6 AI System Lifecycle, 8 Controls (The Engineering Core)
A.6 is the largest area and the operational heart of the AIMS. First-cycle audits concentrate Stage 2 nonconformities here.
A.6.1.1 Objectives for development. Evidence: intended-purpose statement per system (cross-walks to Article 3(1)), use-case documentation, success criteria. Typical state: green for organizations with an Article 3(1) scope memo in place, the intended-purpose discipline is established.
A.6.1.2 Responsible design and development. Evidence: secure-development lifecycle for AI (NIST SSDF + existing SDLC), responsible-AI-by-design checklist, model-card and system-card templates, design-review records. Typical state: yellow, SDLC exists but the AI-by-design checklist is informal.
A.6.1.3 Verification and validation. Evidence: V&V plan, evaluation suite results (accuracy, robustness, bias, security), V&V acceptance criteria, sign-off records. Cross-walks to Article 15. Typical state: yellow, V&V results exist for the latest releases but historical evidence and sign-off records are inconsistent.
A.6.1.4 Deployment. Evidence: deployment plan, pre-deployment review, approval, rollback plan, monitoring activation. Typical state: green for organizations with mature CI/CD, the deployment-gate discipline transfers cleanly from non-AI engineering.
A.6.1.5 Operation and monitoring. Evidence: operational monitoring dashboard per system, drift detection, performance monitoring, incident detection, documented review cadence with named reviewer. Cross-walks to Article 72 + NIST Measure 3 + Manage 4. Typical state: red, the most-common single source of Stage 2 nonconformities. Monitoring is stood up at deployment but operates without documented review cadence; drift detection is informal; the reviewer is not designated. Fix: integrate AI monitoring into the existing SRE / observability function with explicit AI dashboards and a designated reviewer with documented monthly review.
A.6.1.6 Technical documentation. Evidence: model card, system card, data card, Annex IV-aligned documentation per in-scope system. Cross-walks to Article 11 + Annex IV. Typical state: yellow, model cards exist for major systems but are not aligned to the nine-section Annex IV structure.
A.6.1.7 Event logs. Evidence: logging architecture, retention policy, integrity controls (cryptographic where applicable), review procedures, log-access controls. Cross-walks to Article 12 automatic logging. Typical state: yellow: operational logs exist but AI-specific log content (prompts, responses, decisions, overrides) and retention discipline are incomplete.
A.6.2 AI system requirements. Evidence: requirements specification per system, traceability matrix to evaluation criteria and design decisions, review records. Cross-walks to EU AI Act tiering memo + NIST Map 2. Typical state: green for organizations with an Article 3(1) scope memo and tiering memo operational, the requirements discipline flows from the upstream tiering work.
A.7 Data for AI Systems, 5 Controls
A.7.2 Data for development and enhancement. Evidence: training-data inventory per system, source documentation, licensing, train/validation/test split, lineage. Cross-walks to Article 10(1)-(4). Typical state: yellow, training-data documentation exists in piece-meal form per engineering team but a consolidated, audit-ready inventory is missing.
A.7.3 Acquisition of data. Evidence: acquisition procedure, vendor licensing, TDM opt-out compliance records (cross-walks to Article 53(1)(c) GPAI copyright and GPAI Code of Practice copyright chapter), approval records. Typical state: yellow, TDM opt-out compliance is recent and historical records are thin.
A.7.4 Quality of data. Evidence: data-quality framework, metrics (completeness, accuracy, consistency, timeliness, validity, uniqueness), monitoring dashboard, remediation backlog. Cross-walks to Article 10(3). Typical state: yellow, data-quality framework exists but AI-specific quality metrics are not yet operationalized.
A.7.5 Data provenance. Evidence: data lineage, provenance metadata, ML-BoM / SBOM-AI integration (CycloneDX 1.7 + CDXA), supplier-supplied provenance attestations. Typical state: red, provenance discipline lags most other A.7 controls. One of the highest-leverage controls for Article 10 + NIST Map 4 once operational.
A.7.6 Data preparation. Evidence: preparation procedures, transformation documentation, bias-mitigation steps, preparation logs. Typical state: yellow, preparation procedures are codified but documentation per dataset is inconsistent.
A.8 Information for Interested Parties, 4 Controls
A.8.2 System documentation and information for users. Evidence: user documentation, instructions for use (Article 13), purpose statement, system limitations, known risks. Typical state: yellow, user documentation exists for the main UI but the formal Article 13 instruction-for-use artifact is missing or thin.
A.8.3 External reporting. Evidence: external transparency cadence (model card publication, public AI usage disclosure), regulator-facing reporting (Article 72 post-market, Article 73 serious-incident). Cross-walks to NIST Measure 2.7 + Govern 4. Typical state: yellow, model cards published for major releases but cadence and content discipline are informal.
A.8.4 Communication of incidents. Evidence: incident-communication procedure, classification rubric, templates, regulator-notification workflow tested via tabletop exercise. Cross-walks to Article 73 (15-day standard / 2-day critical-infrastructure widespread / 10-day serious-and-irreversible critical-infrastructure disruption). Typical state: red, the runbook exists in draft but has not been tested via tabletop exercise; the regulator-notification template is generic. One of the most-common A.8 first-cycle gaps. Fix: tabletop exercise per quarter with regulator-engagement participation; update runbook from each exercise.
A.8.5 Information for interested parties. Evidence: stakeholder communication plan, stakeholder map, communication records, feedback intake. Cross-walks to NIST Govern 5 + Article 86 affected-person right. Typical state: yellow, stakeholder map exists but feedback intake from affected persons is informal.
A.9 Use of AI Systems, 3 Controls (The Deployer Obligations)
A.9.2 Processes for responsible use. Evidence: responsible-use procedures, deployer-facing guidelines, monitoring of actual versus intended use, deviation-detection mechanism. Cross-walks to Article 26 deployer obligations + Article 14 human oversight. Typical state: yellow, procedures exist for the deployer functions but actual-versus-intended-use monitoring is informal.
A.9.3 Objectives for responsible use. Evidence: use-objective per system, alignment with intended purpose, restriction documentation. Typical state: yellow. Use objectives are documented for new systems but legacy systems lack the artifact.
A.9.4 Intended use of AI system. Evidence: intended-use documentation (cross-walks to A.6.1.1), use-context analysis, foreseeable-misuse analysis (cross-walks to Article 9(2)(b)). Typical state: yellow, intended-use is documented but the foreseeable-misuse analysis is shallow.
A.10 Third-Party and Customer Relationships, 3 Controls
A.10 is one of the most consequential areas for organizations with foundation-model upstream exposure.
A.10.2 Allocating responsibilities. Evidence: responsibility-allocation matrix per third-party relationship, contract clauses (substantial-modification reassignment under Article 25(1); upstream-provider cooperation under Article 25(4)), escalation pathways. Cross-walks to Article 25 value chain. Typical state: yellow, relationship matrix exists for the largest suppliers but is patchy for smaller vendors and individual model providers.
A.10.3 Suppliers. Evidence: AI-supplier inventory (foundation-model providers, RAG-infrastructure providers, evaluation tools, monitoring tools), due-diligence framework, Annex XI / XII receivable verification capability for GPAI upstream, supplier-incident workflow. Cross-walks to NIST Govern 6. Typical state: green for the largest suppliers, third-party risk programs exist; yellow for the AI-specific due-diligence content and the Annex XII receivable verification. Common gap: the organization relies on upstream vendor attestations (ISO 42001 cert, SOC 2 report, model card) as A.10 evidence, the auditor will not accept this as substitute for the organization's own documented due-diligence process.
A.10.4 Customers. Evidence: customer-facing obligation documentation (Article 25(4) cooperation when the organization operates as a provider), customer-information procedures, customer-incident coordination, customer-feedback intake. Typical state: green for organizations with mature customer-success and trust-portal functions; yellow for the AI-specific customer-cooperation content where the organization operates as a foundation-model provider or upstream component.
The Gap-Assessment Methodology - Per-Control Evidence Inventory + Maturity Score + Gap-Fill Plan
A defensible 2026 pre-Stage-1 gap-assessment artifact is built from four per-control elements aggregating into a program-level readiness picture.
Per-Control Evidence Inventory
For each of the 38 controls, the inventory captures: what evidence exists today; where it lives (URL / SharePoint path / repository); who owns it; whether it is current (versioned, dated); whether it has been operating in the prior audit period (typical first-cycle target: 6-9 months of operating history); and whether it is accessible to the auditor without ad-hoc retrieval. The inventory is the foundation, without it the gap-assessment is impressionistic.
A common methodology mistake is to merge inventory and assessment into one workshop. They are separate exercises. The inventory must be exhaustive (every control row populated with either evidence or "no evidence") before the maturity scoring begins. Mixing the two produces selection bias, the team scores high for controls where evidence is easy to recall and low for controls where evidence is harder to find, regardless of actual operating effectiveness.
Per-Control Maturity Score (0-5) Against Explicit Rubric
The maturity-scoring rubric must be explicit and applied consistently. A defensible 2026 rubric:
- 0 - Absent: no evidence; control not implemented.
- 1 - Ad hoc: informal practice exists; no documentation; not repeatable.
- 2 - Documented: procedure documented but inconsistently applied; partial evidence; less than 3 months of operating history.
- 3 - Implemented: procedure applied consistently across most in-scope systems; documentation exists; 3-6 months of operating history; minor gaps in evidence accessibility.
- 4 - Operating: procedure applied consistently across all in-scope systems; documentation complete and accessible; 6+ months of operating history; ready for Stage 2 sampling.
- 5 - Optimized: procedure applied consistently; metrics and continuous improvement evidence; integrated with cross-walk frameworks; would withstand external benchmarking.
Stage 1 readiness threshold: most applicable controls at maturity 4 or higher, with the remaining controls at 3 with documented remediation plans showing the path to 4 within the Stage 2 timeline. Maturity 5 across all controls is not the Stage 1 expectation; it is a long-term posture organizations grow into across surveillance cycles.
A common methodology mistake is pseudo-maturity scoring without a rubric, the team self-assesses each control as "we feel this is a 3" or "this is mostly there." The auditor at Stage 1 will probe the rubric; pseudo-scoring collapses under questioning. The rubric must exist before scoring begins, must be the same rubric across all 38 controls, and the scoring must reference specific evidence inventory entries.
Per-Control Gap-Fill Plan with Owner + Target Date + Evidence Reference
For each control scoring below the Stage 1 threshold, the gap-fill plan documents: the specific evidence to be produced; the named owner (individual, not function); the target completion date; the verification approach (who will confirm the evidence is in place); and the cross-reference to the evidence inventory entry that will be created or updated.
Two common methodology mistakes here: (1) no owner, the plan reads "Engineering will produce X by Q3" with no individual on the hook; (2) no target date, the plan reads "We will improve this control" with no date. Both produce roadmaps that never close. The fix is a per-control row with an explicit name and date that can be tracked weekly in the AI Governance Committee's operating cadence.
Aggregate Readiness Score and Stage 1 Audit Threshold
The aggregate readiness score rolls per-control maturity scores into a program-level readiness picture. A defensible 2026 aggregate:
- % of applicable controls at maturity 4+, Stage 1 readiness target: โฅ80%.
- % of applicable controls at maturity 3 with documented remediation plan to 4 within Stage 2 timeline, target: โค20%.
- % of applicable controls below maturity 3, target: 0%. Any below-3 control delays the Stage 1 booking until remediated to 3+.
- Aggregate readiness percentile against peer cohort, informs whether the organization is ahead of or behind comparable first-cycle peers.
The aggregate is the artifact the AI Governance Committee and audit committee see quarterly. The per-control detail is the artifact the program team and certification body see. Both views must reconcile.
Worked Example - Fortune-500 Industrial Software Vendor Pre-Stage-1 Assessment
Consider IndustrialAI Co. (composite anonymized example based on multiple 2026 first-cycle assessments): a Fortune-500 industrial software vendor pursuing ISO 42001 certification with Stage 1 booked for Q3 2026, Stage 2 expected Q4 2026, certification target Q1 2027. The pre-Stage-1 gap assessment runs in May 2026 covering all 38 controls. Aggregate results:
- Maturity 4+ (Stage 1 ready): 27 of 38 controls (~71%). Below the โฅ80% threshold, Stage 1 booking should not proceed without remediation.
- Maturity 3 with remediation plan: 8 of 38 controls (~21%). Approximately at the โค20% threshold.
- Below maturity 3: 3 of 38 controls (~8%). All three must be remediated to 3+ before Stage 1 booking is defensible.
The three below-3 controls: A.5.5 societal impacts (no documented methodology); A.6.1.5 operation and monitoring (monitoring dashboards exist but documented review cadence absent for half the in-scope systems); A.8.4 communication of incidents (runbook in draft, never tabletop-tested). The remediation plan: A.5.5: write the societal-impact methodology document and apply retrospectively to the three highest-impact in-scope systems, owner Risk Officer, target +8 weeks; A.6.1.5: designate the named monthly reviewer per system, document the review cadence, log the next two months of reviews, owner SRE Lead, target +10 weeks; A.8.4: first tabletop exercise, runbook update, second tabletop exercise, owner Privacy Officer, target +12 weeks. Stage 1 booking moves from Q3 to mid-Q4 2026 to accommodate the remediation. Stage 2 moves to Q1 2027. Certification target moves to Q2 2027, one quarter delay.
The alternative, pushing through to the original Q3 Stage 1 booking, would have produced an estimated 5-7 documentation findings (the three below-3 controls plus likely 2-4 of the maturity-3 controls discovered to be weaker than scored). Stage 1 remediation plus rebooked Stage 2 would have slipped certification to Q3 2027, two quarters of delay versus the one-quarter delay produced by an honest pre-Stage-1 assessment. The math always favors the assessment.
Cross-Walks - One Gap-Assessment Roadmap, Multiple Regulator Coverage
The gap-assessment roadmap should not be a single-framework artifact. The L3 design principle is multi-framework evidence efficiency: each control row references the corresponding evidence in EU AI Act, NIST AI RMF, OWASP / ATLAS (for user-facing security controls), and applicable sectoral frameworks. Per-control evidence then serves multiple regulator-facing artifacts.
ISO 42001 Clauses 4-10 + Annex A 38 Controls
The base structure. Clauses 4-10 are the management-system shape; Annex A is the 38 AI-specific controls. The gap-assessment matrix rows are the 38 Annex A controls; the columns include cross-reference to the Clauses 4-10 evidence and to other frameworks.
EU AI Act Articles
Per-control cross-walks (per lesson 012 mapping): A.2.2 โ Article 17(1)(a); A.3.2 โ Article 17(1)(k); A.4.6 โ Article 4 deployer literacy; A.5.2/5.3/5.4 โ Article 27 FRIA; A.6.1.1 โ Article 3(1); A.6.1.3 โ Article 15; A.6.1.5 โ Article 72; A.6.1.6 โ Article 11 + Annex IV; A.6.1.7 โ Article 12; A.6.2 โ tiering memo; A.7.2 โ Article 10(1)-(4); A.7.3 โ Article 53(1)(c); A.7.4 โ Article 10(3); A.8.2 โ Article 13; A.8.3 โ Articles 50, 72; A.8.4 โ Article 73; A.8.5 โ Article 86; A.9.2 โ Articles 14, 26; A.9.4 โ Article 9(2)(b); A.10.2 โ Article 25; A.10.3 โ Articles 25(4), Annex XI/XII; A.10.4 โ Article 25(4). The cross-walks make ISO 42001 the structural backbone of the Article 17 QMS, ~70% of high-risk documentation coverage from one AIMS implementation.
NIST AI RMF Govern / Map / Measure / Manage
A.2-A.3 โ Govern 1, 2; A.4 โ Govern 3; A.5 โ Map 1, 3; A.6.1.1-1.2 โ Map 2, 4; A.6.1.3 โ Measure 2; A.6.1.5 โ Measure 3 + Manage 4; A.6.1.6 โ Govern 4; A.6.1.7 โ Measure 2.9; A.7 โ Map 4; A.8 โ Measure 2.7 + Govern 4, 5; A.9 โ Govern 4 + Manage 1; A.10 โ Govern 6. The cross-walk lets one AIMS satisfy NIST AI RMF maturity reporting expected by U.S. federal procurement, U.S. state programs, and many board-reporting frameworks.
OWASP LLM / Agentic Top 10 + MITRE ATLAS
For user-facing security controls, A.6.1.3 V&V (LLM01 prompt injection, LLM05 supply-chain), A.6.1.5 monitoring (LLM06 sensitive-info disclosure, LLM09 misinformation), A.6.1.7 logs (ATLAS reconnaissance + initial-access detection), A.8.4 incidents (ATLAS impact + exfiltration classification). The OWASP/ATLAS cross-walk turns the AIMS A.6/A.8 evidence into Red Team and SOC-facing artifacts.
Sectoral Overlays
Per industry: SR 11-7 (banking model risk management) โ A.5 + A.6 for credit-decisioning; HIPAA + 21 CFR 820 (medical-device AI) โ A.5 + A.6 + A.7; Colorado SB 24-205 + NYC LL 144 + Texas TRAIGA (U.S. state AI) โ A.5 impact assessment; ECOA + CFPB (fair lending) โ A.5.4 individual/group analysis. The sectoral cross-walks let one AIMS evidence pack satisfy industry regulator inquiries without redundant assessment.
The Gap-Remediation Roadmap - The L3 Artifact
The gap-remediation roadmap is the deliverable the AI Governance Committee, audit committee, and program team operate from. Structure:
Per-Control Row
- Control ID (e.g., A.6.1.5).
- Control title.
- Applicable / Not applicable with SoA justification reference.
- Current state: maturity score 0-5 against rubric.
- Target state: maturity 4 (Stage 1 ready) or 5 (optimized).
- Gap description: specific evidence missing or insufficient.
- Remediation actions: specific deliverables (procedure, dashboard, runbook, training).
- Owner: named individual.
- Target date: specific date.
- Evidence reference: pointer to evidence inventory entry where the deliverable will live.
- Cross-walk references: EU AI Act article, NIST AI RMF category, OWASP/ATLAS, sectoral overlay.
- Status: not started / in progress / blocked / complete / verified.
Aggregated Dashboard for AI Governance Committee
The committee sees the aggregate readiness percentage, the count of controls below threshold, the count of controls with overdue owners, the burn-down trend over the prior 90 days, and the projected Stage 1 booking date based on current remediation pace. The dashboard surfaces blocked items and pace-of-burn-down concerns for committee decision.
Quarterly Refresh and Operating Cadence
The roadmap is not a one-time deliverable. It refreshes quarterly with re-scoring of all 38 controls, updated evidence inventory, new gap-fill rows for surveillance findings, and burn-down tracking. The quarterly refresh continues through surveillance audits across the 3-year certification cycle. Static roadmaps, produced once at pre-Stage-1, never updated, fail at Year 1 surveillance when control maturity that was 4 at certification drops to 3 from staffing turnover or process drift.
Six Common Methodology Mistakes
Mistake 1 - Skipping the Pre-Stage-1 Assessment Entirely
The most expensive mistake. Booking Stage 1 on the assumption that prior ISO 27001 / SOC 2 maturity is sufficient. Stage 1 surfaces 5-10 documentation findings, Stage 2 slips a quarter, customer commitments slip, board credibility takes a hit. The fix is non-negotiable: run a formal pre-Stage-1 assessment 12-16 weeks before the booking and use the results to set the booking date defensibly.
Mistake 2 - Weak Per-Control Evidence Inventory
The team builds the assessment from memory or a partial document census, not from a row-by-row inventory of all 38 controls. The result is selection bias: high scores for controls where evidence is recent and accessible, low scores for controls where evidence is older or fragmented, regardless of operating effectiveness. The fix: the evidence inventory is a standalone exercise completed before maturity scoring begins, with every control row populated (even if the entry is "no evidence").
Mistake 3 - Pseudo-Maturity Scoring Without Explicit Rubric
The team assigns scores from gut feel, "we think this is a 3", without an explicit 0-5 rubric applied consistently. The Stage 1 auditor will probe the scoring methodology; pseudo-scoring collapses. The fix: document the rubric, apply it consistently, reference specific evidence inventory entries in each score justification.
Mistake 4 - Gap-Fill Plans with No Owner or No Target Date
"Engineering will improve A.6.1.5 monitoring." With what deliverable, by whom, by when? Without a named individual and a specific date, the plan does not close. The fix: every gap-fill row has an individual name and a date in the next 16 weeks. Items beyond 16 weeks need an intermediate milestone and a check-in date.
Mistake 5 - Treating the Roadmap as a One-Time Deliverable
The roadmap is built for pre-Stage-1, then shelved. Year 1 surveillance audit pulls evidence; control maturity that was 4 at certification has drifted to 3 from staffing changes, process drift, or new in-scope systems. The fix: quarterly refresh of all 38 control scores; the roadmap is a living artifact, updated for every surveillance finding and every new in-scope system.
Mistake 6 - No Cross-Walk to EU AI Act / NIST / Sectoral for Evidence Efficiency
The gap-assessment treats ISO 42001 as a stand-alone framework rather than the structural backbone for multi-framework evidence. The result is parallel evidence work, separate FRIA evidence, separate NIST AI RMF mapping, separate sectoral assessment, instead of one cross-walked evidence pack. The ~70% EU AI Act coverage and the parallel NIST coverage are lost. The fix: design the gap-assessment matrix with cross-walk columns from day one; every control row references the corresponding EU AI Act article, NIST AI RMF category, OWASP/ATLAS, and sectoral overlay; one piece of evidence serves three or four regulator-facing artifacts.
Key Takeaways
- The pre-Stage-1 gap assessment is non-negotiable. Skipping it on the rationale of prior ISO 27001 / SOC 2 maturity produces 5-10 Stage 1 findings, a slipped Stage 2 quarter, missed customer commitments, and damaged board credibility. The assessment runs 12-16 weeks before the Stage 1 booking.
- Four-layer methodology: per-control evidence inventory; per-control maturity score (0-5 against explicit rubric); per-control gap-fill plan with owner and target date; aggregate readiness score against Stage 1 threshold (โฅ80% controls at maturity 4+).
- All 38 controls walked per evidence expectation and typical maturity: A.2-A.4 governance (typically yellow-green with prior ISO 27001 maturity); A.5 impact assessment (yellow with A.5.5 societal impact often red); A.6 lifecycle (mixed, A.6.1.1/1.4/6.2 green; A.6.1.5 monitoring typically red; A.6.1.3/1.6/1.7 yellow); A.7 data (yellow with A.7.5 provenance often red); A.8 information (yellow with A.8.4 incidents often red); A.9 use (yellow); A.10 third-party (mixed, A.10.4 green; A.10.3 supplier-AI-specific content yellow).
- Maturity rubric (0-5): 0 absent; 1 ad hoc; 2 documented; 3 implemented (3-6 months operating); 4 operating (6+ months, Stage 1 ready); 5 optimized. Stage 1 readiness: most controls at 4+, remaining at 3 with remediation plan; no controls below 3.
- Worked Fortune-500 example: 27/38 at maturity 4+ (~71%, below threshold); 8/38 at 3 with plan (~21%); 3/38 below 3 (A.5.5, A.6.1.5, A.8.4). Three remediation tracks (8/10/12 weeks) move Stage 1 booking from Q3 to mid-Q4 2026; certification slips one quarter versus an estimated two-quarter slip from pushing through unprepared.
- Cross-walks turn one gap-assessment into multi-framework evidence: EU AI Act articles per control (~70% of high-risk documentation coverage); NIST AI RMF Govern/Map/Measure/Manage; OWASP LLM / Agentic Top 10 + MITRE ATLAS for user-facing security; sectoral overlays (SR 11-7, HIPAA, Colorado SB 24-205, NYC LL 144, Texas TRAIGA, ECOA).
- Gap-remediation roadmap structure: per-control row with control ID, applicability, current/target maturity, gap description, remediation actions, named owner, target date, evidence reference, cross-walk references, status. Aggregated dashboard for the AI Governance Committee. Quarterly refresh across the 3-year certification cycle.
- Stage 1 readiness threshold: โฅ80% of applicable controls at maturity 4+ with documented evidence; โค20% at maturity 3 with documented remediation plan; 0% below 3. Below-threshold posture delays the Stage 1 booking until remediated.
- Six common methodology mistakes: (1) skipping the pre-Stage-1 assessment entirely; (2) weak per-control evidence inventory (selection bias); (3) pseudo-maturity scoring without explicit rubric; (4) gap-fill plans with no owner or no target date; (5) treating the roadmap as a one-time deliverable (Year 1 surveillance gaps); (6) no cross-walk to EU AI Act / NIST / sectoral (lost evidence efficiency).
- The L3 artifact: the gap-remediation roadmap as a living, cross-walked, owner-and-date-bearing instrument that turns "we think we are ready" into "we are 82% operating evidence and the remaining 18% have named owners and dates": defensible to the certification body, the audit committee, the board, and the customer trust portal.
Skill.re