ISO/IEC 23894:2023 AI Risk Management - Process Application
In March 2026, a €1.8B EU e-commerce platform stood up a refreshed product-recommendation system across 14 Member State storefronts. The launch slide deck called the system "low-risk" because it did not personalize credit terms, did not gate access to public services, and did not surface in any Annex III category. By June, a Spanish consumer-protection regulator opened an inquiry, not under the AI Act, but under unfair-commercial-practices law, citing systematic over-recommendation of premium-tier products to demographic segments that the platform's own A/B tests had not analyzed for disparate impact. The Responsible AI Officer had a generic risk assessment from intake. What she did not have was a documented application of ISO/IEC 23894:2023, scope, context, criteria; identification with source analysis; analysis with likelihood × impact; evaluation against pre-set criteria; treatment with owners and dates; monitoring with triggers, and a risk treatment plan with acceptance-authority signatures. This lesson is the slow walk through 23894 process application on that recommendation system, the cross-walks to ISO 42001 A.6 (AI impact assessment) and ISO 31000 risk-management process, and the L3 artifact: scope-context-criteria document + risk register + risk treatment plan + monitoring schedule + management-review evidence.
Why ISO/IEC 23894:2023 Matters in 2026 - The AI Overlay on ISO 31000
ISO/IEC 23894:2023, "Information technology, Artificial intelligence, Guidance on risk management", was published February 2023 by ISO/IEC JTC 1/SC 42. It is guidance (not certifiable on its own) on how to apply the ISO 31000:2018 risk-management process to AI-specific risk sources, events, and consequences. In 2026 it sits at the operational heart of three concurrent compliance regimes.
First, ISO/IEC 42001:2023 Clauses 6.1.2 (AI risk assessment) and 6.1.3 (AI risk treatment) require a risk-management methodology, and ISO 23894 is the standard the auditor expects to see referenced. The defensible 2026 answer to "what methodology underpins your 6.1.2/6.1.3 work?" is "ISO/IEC 23894:2023 applied through the ISO 31000:2018 process." Adopting 23894 satisfies both 42001 clauses and cross-walks cleanly to EU AI Act Article 9 for high-risk providers.
Second, EU AI Act Article 9 requires high-risk providers to establish, implement, document, and maintain a risk-management system "as a continuous iterative process planned and run throughout the entire lifecycle". Article 9(2) prescribes content (identification and analysis of known and reasonably foreseeable risks; estimation and evaluation; evaluation of post-market monitoring; adoption of appropriate and targeted measures). The Article 9 RMS is the ISO 23894 process applied to the lifecycle, 23894 produces Article 9 evidence as a byproduct.
Third, NIST AI RMF 1.0 Govern/Map/Measure/Manage maps function-by-function to 23894 process steps. Map ≈ scope/context/criteria + identification + analysis; Measure ≈ analysis (measurement focus); Manage ≈ evaluation + treatment + monitoring; Govern ≈ communication/consultation + recording/reporting.
The L3 question is how to apply the eight steps to a specific AI system, produce the risk treatment plan, and integrate with the AIMS, Article 9 RMS, and NIST AI RMF profile so one analysis serves four obligations. The L3 artifact: scope-context-criteria + risk register + RTP + monitoring schedule + management-review evidence.
The Eight Process Steps - ISO 31000 Spine, ISO 23894 AI Overlay
ISO 23894 follows the ISO 31000:2018 risk-management process verbatim. ISO 31000 establishes the universal framework; ISO 23894 adds AI-specific guidance at each step. Eight steps in operational order:
Step 1 - Scope, Context, Criteria (ISO 31000 §6.3 / ISO 23894 §6.3)
Scope defines the boundaries of the risk-management exercise: the AI systems in scope, the lifecycle phases covered, the organizational units involved, the deployment geographies, and the time horizon. Context covers internal context (mission, AI strategy, risk appetite, existing controls, culture) and external context (regulatory landscape: EU AI Act, GDPR, sector law; competitive pressure; technological evolution; societal expectations). Criteria specify the thresholds against which risks will be evaluated: acceptable likelihood, acceptable consequence severity, acceptable residual risk, escalation thresholds.
This is the most-skipped step in first-time 23894 applications. Practitioners jump straight to identification because "we already know what the risks are." The auditor finds the analysis lacks a defensible threshold structure. There is no documented basis for declaring a risk "acceptable" or "requiring treatment." L3 fix: produce a 4-8 page scope-context-criteria document signed by the Responsible AI Officer and the system business owner before any identification work begins.
Step 2 - Risk Identification (ISO 31000 §6.4.2 / ISO 23894 §6.4.2)
Identification produces the inventory of risk sources, risk events, causes, and consequences across the AI lifecycle. ISO 23894 Annex A and Annex B catalog AI-specific risk sources organized into five families: data-related (quality, representativeness, bias, drift); model-related (accuracy, robustness, interpretability, security); human-AI-interaction (over-reliance, under-reliance, accessibility, transparency); lifecycle (development, deployment, post-deployment monitoring, retirement); third-party (foundation models, datasets, tools, infrastructure).
The defensible identification process combines top-down (drawing from the 23894 Annex A/B catalog plus the NIST AI 600-1 GenAI Profile twelve-risk list plus the OWASP LLM Top 10 plus the OWASP Agentic Top 10) with bottom-up (system-specific brainstorming with engineering, data, security, legal, and the business owner). The output is a risk register entry per identified risk with risk ID, description, source family, lifecycle phase, and initial hypothesis on likelihood and consequence.
Step 3 - Risk Analysis (ISO 31000 §6.4.3 / ISO 23894 §6.4.3)
Analysis assesses each identified risk's likelihood × consequence against the criteria set in Step 1. ISO 23894 explicitly endorses both qualitative methods (likelihood bands defined in plain language; consequence severity bands tied to organizational impact thresholds) and quantitative methods (Bayesian networks, Monte Carlo on specific failure modes, statistical analysis on observed metrics). Most 2026 implementations use a hybrid, qualitative scoring for risks where quantitative data is sparse, quantitative analysis for risks where measurement is feasible (e.g., bias-disparity testing produces quantitative inputs to a fairness risk).
A common L3 pitfall is pseudo-quantification, putting a "0.05 likelihood" on a risk with no statistical basis. The defensible 2026 practice: use a 5-band qualitative scale (VL/L/M/H/VH) with verbal anchors for both likelihood and consequence; reserve quantitative scoring for risks with empirical measurement; document the basis for each score.
Step 4 - Risk Evaluation (ISO 31000 §6.4.4 / ISO 23894 §6.4.4)
Evaluation compares each analyzed risk to Step-1 criteria. Output: a prioritization, risks exceeding the "treat" threshold go to Step 5; risks within tolerance go to monitoring (Step 6) with documented acceptance. If criteria are weak, evaluation is arbitrary.
ISO 23894 emphasizes AI-specific factors beyond severity-likelihood: irreversibility (an over-recommendation reverses in a session; a denied credit decision is durable); cascading impacts (recommendation bias → vendor-revenue inequality → platform-trust erosion); detectability (CTR drift is detectable; subtle homogenization is not); universe of affected persons (Article 3(32)).
Step 5 - Risk Treatment (ISO 31000 §6.5 / ISO 23894 §6.5)
Treatment selects options per "treat" risk: avoid (do not deploy or change scope), mitigate (reduce likelihood/consequence), transfer (insure, contractually allocate), or accept (document and monitor). Treatment plans must specify control type (preventive/detective/corrective), implementation owner, target date, post-treatment residual rating, and residual-risk acceptance authority.
The risk treatment plan (RTP) is the central Step-5 artifact. A 2026-defensible RTP per risk includes: risk ID + description + source; pre-treatment likelihood/consequence/overall rating; treatment option(s) with justification; controls implemented; owner; target and actual completion dates; post-treatment residual rating; residual-risk acceptance signature; review cadence; trigger events.
Step 6 - Monitoring and Review (ISO 31000 §6.6 / ISO 23894 §6.6)
Monitoring and review is continuous through the lifecycle and trigger-based on specific events (model update; data-source change; deployment-context change; incident; regulator inquiry; bias-alert; drift-alert; vendor model change). ISO 23894 emphasizes monitoring includes both control-effectiveness monitoring (are the implemented controls working as designed?) and risk-environment monitoring (has the risk landscape changed?).
A defensible monitoring schedule specifies cadence per risk (some risks reviewed quarterly, some annually); KPIs and KRIs tied to each material risk; trigger events that force ad-hoc review; reporting cadence to the AI Governance Committee; escalation thresholds.
Step 7 - Communication and Consultation (ISO 31000 §5.2 / ISO 23894 §5.2)
Communication and consultation runs in parallel through all seven other steps. Stakeholders include internal (executive sponsors, engineering, data, security, legal, compliance, internal audit, business owners) and external (regulators, customers, affected persons, vendors, civil society). Engagement focuses on scope/criteria-setting, identification, evaluation, and treatment review.
This step bridges EU AI Act Article 86 (affected-person right to explanation), Article 27(1)(f) FRIA complaint mechanisms, NIST Govern 5 (stakeholder engagement), and ISO 42001 A.8.5 (information for interested parties). One program serves all four obligations.
Step 8 - Recording and Reporting (ISO 31000 §6.7 / ISO 23894 §6.7)
Recording produces the documented evidence consumed by the ISO 42001 Stage 2 auditor, EU AI Act notified body, NIST AI RMF maturity assessor, SR 11-7 validator, and customer trust-portal reviewer. Records: scope-context-criteria document; risk register; risk treatment plan; monitoring schedule and execution evidence; communication/consultation log; management-review evidence.
This step is the most underbuilt. Practitioners document informally (slide decks, ad-hoc spreadsheets) and discover at audit time the evidence is not traceable, signed, or versioned. L3 fix: build the templates at Step 1 so artifacts populate as the process runs.
AI-Specific Risk Sources - The 23894 Annex Catalog Applied
ISO 23894 Annexes A and B catalog AI-specific risk sources that distinguish AI risk management from generic IT risk management. Internalize the five families and principal sub-risks within each, the catalog drives identification and prevents the "we missed it" finding at evaluation.
Data-Related Risks
Data quality (the six classical dimensions, completeness, accuracy, consistency, timeliness, validity, uniqueness, applied to training and operational data; Article 10(3) anchor). Representativeness (whether training data reflects the deployment population; Article 10(2)-(3)). Bias (systematic group-level disparate treatment; multiple operationally-incompatible definitions). Drift (concept drift + data drift; the most operationally important data-related risk in continuous deployments).
Model-Related Risks
Accuracy (aggregate vs per-group). Robustness (distribution shift, adversarial inputs, edge cases; Article 15 + NIST Measure 2.5). Interpretability (drives Article 14 oversight and Article 86 explanation right). Security (prompt injection, jailbreaks, prompt leakage, data poisoning, model extraction, membership inference; OWASP LLM Top 10 + MITRE ATLAS). Cost / latency / energy (operational economics especially for foundation-model deployments).
Human-AI Interaction Risks
Over-reliance (automation bias) (Article 14 oversight challenge). Under-reliance (often after a high-profile failure; safety-critical settings). Accessibility (Article 16 provider obligation; AI Office accessibility guidance). Transparency (Article 50(1) AI-interaction disclosure + Article 13 instructions for use).
Lifecycle Risks
Development (specification errors, requirements drift, training-pipeline defects, evaluation gaps). Deployment (configuration errors, integration failures, scope creep, missing rollback). Post-deployment monitoring (absent monitoring, monitoring without a reviewer, monitoring without trigger workflows, the most common Stage 2 nonconformity area for ISO 42001 A.6.1.5). Retirement (decommissioning without affected-user notice, data-deletion gaps, dependency-residue defects).
Third-Party Risks
Foundation-model (upstream updates, deprecations, behavior changes, training-data unknowns, copyright exposure, GPAI CoP signatory status). Dataset (licensing changes, TDM opt-out exercise, provenance ambiguity, contamination). Tool (vendor-tool deprecation, evaluation-framework defects). Infrastructure (cloud-provider outages, region-specific availability, data-residency changes).
A defensible identification artifact covers all five families per system. The L3 register typically has 20-40 risks per material AI system.
Worked Example - Applying ISO 23894 to a Recommendation System
Consider a product-recommendation system at an EU e-commerce platform deployed across 14 Member States with a hybrid architecture: foundation-model NL understanding for query interpretation, collaborative-filtering recommendation engine, re-ranking layer with promotional logic.
Step 1 - Scope, Context, Criteria
Scope. The product-recommendation AI system serving EU customers across 14 Member State storefronts; lifecycle phases from training-data acquisition through post-deployment monitoring; in scope are the engineering, data, security, legal, compliance, business-owner, and customer-experience functions; time horizon is 12 months with quarterly review.
Context, internal. Organizational mission of customer trust and revenue growth; AI strategy emphasizes personalization within fundamental-rights guardrails; risk appetite is "moderate" for revenue impact and "low" for fundamental-rights impact; existing controls include the platform's general A/B testing infrastructure, the engineering peer-review process, the customer-trust portal; culture is engineering-led with a building-out compliance function.
Context, external. EU AI Act applies, not in Annex III categories; Article 3(1) intended-purpose scope hits; Article 50(1) interactive-disclosure not triggered; Article 4 literacy obligation applies to deployer-staff handling outputs; GDPR Article 35 DPIA likely required; Article 22 GDPR likely not triggered but must be documented; competitive pressure; technological evolution toward generative recommendations.
Risk criteria. Acceptable thresholds set by AI Governance Committee:
- Drift (KL divergence on output distribution): <0.05 tolerated; 0.05-0.10 monitor; >0.10 trigger review.
- Bias disparity (demographic parity on premium-tier recommendation): max 1.25× ratio across protected-attribute groups; ratios >1.25× trigger remediation.
- Recommendation diversity (intra-list diversity score): minimum 0.6; below 0.6 triggers review for filter-bubble risk.
- Revenue intervention: any treatment expected to reduce per-session revenue >3% requires CFO sign-off.
- IP infringement: zero tolerance for confirmed third-party IP in generative components; opt-out responsiveness within 14 calendar days.
This 6-page document, signed by the Responsible AI Officer and head of e-commerce, is the foundation for everything that follows. Without it, Step-4 evaluation has no defensible basis.
Step 2 - Risk Identification
Combine top-down (23894 Annex A/B, NIST AI 600-1, OWASP LLM Top 10) with bottom-up (engineering + data + business workshop). The risk register for this system includes 28 risks across the five families. Selected entries:
- R-01 Recommendation bias (data-related). Training data skews toward historical purchase patterns; women-coded users may receive fewer recommendations for high-value categories. Phase: training. Hypothesis: medium L × high C.
- R-02 Filter bubble (homogenization) (model + human-AI). Collaborative filtering converges on dominant patterns, trapping users in narrow sets. Phase: deployment + operation. Hypothesis: high L × medium C.
- R-03 Cold-start failure (model). New users with no purchase history get low-quality recommendations. Phase: deployment. Hypothesis: high L × low C.
- R-04 Drift on category mix (data, drift). Seasonal and macro shifts change customer interest distribution faster than retraining cadence. Phase: operation. Hypothesis: medium L × medium C.
- R-05 Cross-vendor data poisoning (third-party). Vendors providing product feeds may inject manipulated metadata. Phase: training + operation. Hypothesis: low L × medium C.
- R-06 Prompt-injection in NL component (model + human-AI). Users craft queries that manipulate the foundation-model NL component. Phase: operation. Hypothesis: medium L × medium C.
- R-07 IP infringement in generative explanations (model + third-party). NL component may generate text reproducing copyrighted material; Article 53(1)(c) opt-out signals must be honored. Phase: operation. Hypothesis: low-medium L × high C.
- R-08 Over-reliance by merchandising staff (human-AI). Staff use recommendation rankings for inventory priorities without challenging the model. Phase: operation. Hypothesis: medium L × medium C.
- R-09 Foundation-model upstream change (third-party). NL component vendor pushes model updates mid-quarter without notice. Phase: operation. Hypothesis: medium-high L × medium C.
- R-10 GDPR Article 22 boundary (lifecycle + legal). If recommendations become so personalized they constitute automated decision-making with significant effect, Article 22 triggers. Phase: operation. Hypothesis: low L × high C.
The remaining 18 risks cover finer-grained data quality, model security, deployment configuration, retirement, and infrastructure risks.
Step 3 - Risk Analysis
Apply the 5-band scale (VL/L/M/H/VH) per risk; use quantitative data as the basis where available. For the recommendation system:
- R-01 bias: M × H (preliminary 1.4× ratio in premium-tier exposure across proxy groups + 14-Member-State exposure) → High.
- R-02 filter bubble: H × M (well-documented CF tendency + revenue impact) → Medium-High.
- R-03 cold-start: H × L (known issue, recoverable bounce) → Low-Medium.
- R-04 drift: M × M (observed seasonal variation) → Medium.
- R-05 cross-vendor poisoning: L × M (existing vendor controls) → Low.
- R-06 prompt-injection: M × M (inherent foundation-model exposure) → Medium.
- R-07 IP infringement: L-M × H (litigation + brand) → Medium-High.
- R-08 over-reliance: M × M (inventory mis-allocation cost) → Medium.
- R-09 foundation-model upstream: M-H × M → Medium-High.
- R-10 Article 22 boundary: L × H (current non-binding architecture) → Medium.
The analysis documents the basis per score, quantitative data where available, expert judgment otherwise. Pseudo-quantification avoided.
Step 4 - Risk Evaluation
Risks rated High or Medium-High enter the committee treatment workflow; Medium risks treated within engineering with periodic committee review; Low risks documented and monitored. Applied:
- Treat (committee-level): R-01 bias (exceeds 1.25× criterion); R-02 filter bubble (likely to violate the 0.6 intra-list-diversity floor); R-07 IP infringement (zero-tolerance criterion).
- Treat (engineering with committee visibility): R-04 drift; R-06 prompt-injection; R-08 over-reliance; R-09 foundation-model upstream; R-10 Article 22 boundary.
- Monitor with documented acceptance: R-03 cold-start; R-05 cross-vendor poisoning; remaining 18 lower-rated risks.
Evaluation records the prioritization decision and rationale per risk.
Step 5 - Risk Treatment
The risk treatment plan (RTP) per treat-tier risk specifies treatment option, controls, owner, dates, residual rating, acceptance authority. Selected entries:
- R-01 Recommendation bias. Mitigate. Controls: bias-monitoring dashboard with weekly KPI on premium-tier recommendation ratio across proxy demographic groups (alert at >1.20× to allow remediation lead-time); training-data re-weighting; post-ranking calibration enforcing demographic-parity on premium-tier exposure. Owner: ML platform lead. Target: Q3 2026. Residual: Low-Medium. Acceptance: Chief AI Risk Officer (CAIRO) + Head of E-commerce. Review: monthly during Q3, quarterly thereafter.
- R-02 Filter bubble. Mitigate. Controls: intra-list-diversity score per session (target ≥0.65); 10% exploration injection within category constraints; cross-category surfacing for long-history users. Owner: Recommendation engine team. Target: Q3 2026. Residual: Low-Medium. Acceptance: CAIRO. Review: quarterly.
- R-07 IP infringement. Mitigate + transfer. Controls: IP-attribution mechanism in generative explanations; opt-out registry honored within 14 calendar days; periodic IP-bleed red-team; contractual IP-warranty from foundation-model vendor; legal-led claim-handling workflow. Owner: Legal + Engineering. Target: Q3 2026. Residual: Low. Acceptance: General Counsel + CAIRO. Review: quarterly + triggered on claim.
- R-04 Drift on category mix. Mitigate. Controls: daily KL-divergence drift pipeline (alert 0.05; review trigger 0.10); accelerated retraining during high-drift periods. Owner: ML platform. Target: Q2 2026 (in flight). Residual: Low. Acceptance: CAIRO. Review: quarterly.
- R-06 Prompt-injection. Mitigate. Controls: input sanitization on NL component; OWASP LLM01 eval suite pre-deployment and on each NL-component update; periodic red-team. Owner: Application security. Target: Q3 2026. Residual: Low-Medium. Acceptance: CAIRO + CISO. Review: quarterly + triggered on NL-component update.
- R-09 Foundation-model upstream change. Mitigate. Controls: vendor SLA requiring 30-day notice of material changes; 2-week shadow-deployment of new versions; drift-and-regression eval on switch. Owner: Engineering + Vendor Management. Target: Q3 2026. Residual: Low. Acceptance: CAIRO. Review: triggered on vendor change.
- R-10 GDPR Article 22 boundary. Monitor + control architecture. Controls: explicit architectural commitment that recommendations are non-binding (no auto-add-to-cart based solely on model output); quarterly legal review of personalization depth; documented analysis on every material personalization change. Owner: Legal + Product. Target: Ongoing. Residual: Low. Acceptance: General Counsel. Review: quarterly + triggered.
The RTP shows the auditor each material risk has documented treatment with owner, date, and residual-acceptance signature. Without it, treatment is informal and not auditable.
Step 6 - Monitoring and Review
The schedule: quarterly review of all material risks; monthly review of R-01 and R-07 during active remediation; daily automated KPI reporting on bias-disparity, drift, intra-list-diversity, prompt-injection detections; trigger-based review on model update, vendor change, regulator inquiry, incident, customer complaint pattern. Monthly KRIs to the AI Governance Committee cover the top-5 risks and residual-risk trend.
Step 7 - Communication and Consultation
The communication plan covers: monthly AI Governance Committee briefings; quarterly board AI risk dashboard inclusion; customer-facing transparency on the trust portal (recommendation methodology and diversity commitments); regulator-facing engagement (Spanish consumer-protection inquiry response; AI Office Article 13 information-sharing); merchandising-staff Article 4 literacy training; vendor consultation on data quality and model-change cadence; affected-person feedback intake channel.
Step 8 - Recording and Reporting
Recording artifacts: signed scope-context-criteria document; versioned risk register; versioned RTP with owner/date evidence; monitoring schedule plus execution evidence (dashboards, alert logs, review minutes); communication/consultation log; management-review evidence (quarterly committee minutes; annual board review). These populate the ISO 42001 Stage 2 evidence pack, Article 9 RMS evidence, NIST AI RMF maturity assessment, SOC 2 AI evidence, and customer trust portal, one analysis, five purposes.
Cross-Walks - ISO 42001 A.6, ISO 31000, NIST AI RMF, EU AI Act Article 9
L3 leverage comes from cross-walk efficiency. The 23894 application produces evidence that maps directly onto four other frameworks.
Cross-Walk to ISO/IEC 42001:2023
- 23894 scope/context/criteria ↔ ISO 42001 Clause 6.1.2 (AI risk assessment establishment) + A.6.2 (AI system requirements).
- 23894 identification + analysis ↔ ISO 42001 A.6.1.1 (AI impact assessment), the central cross-walk for this lesson. The AIMS impact-assessment process operationalizes 23894 identification and analysis.
- 23894 evaluation + treatment ↔ ISO 42001 Clause 6.1.3 (AI risk treatment) + A.6.2 (system requirements).
- 23894 monitoring and review ↔ ISO 42001 A.6.4 (post-deployment monitoring) + A.6.1.5 (operation and monitoring).
- 23894 communication and consultation ↔ ISO 42001 A.5 (impact assessment communication) + A.8 (information for interested parties).
- 23894 recording and reporting ↔ ISO 42001 Clause 7.5 (documented information) + Clause 9.2 (internal audit evidence).
Name ISO 23894 as the underlying methodology in the AIMS document (Clause 4.4), answers the Stage 1 auditor's methodology question and produces cross-walk evidence in one go.
Cross-Walk to ISO 31000:2018
ISO 23894 is the AI overlay on ISO 31000:2018. The eight steps map verbatim to ISO 31000 §6.3-§6.7 plus §5.2. Organizations with existing ISO 31000 ERM extend to AI by adopting the 23894 Annex A/B catalog and AI-specific guidance at each step, unified ERM + AI-RM evidence, no parallel processes.
Cross-Walk to NIST AI RMF 1.0
- 23894 scope/context/criteria + identification ↔ NIST Map (1 context, 2 categorization, 3 capabilities/limits, 4 risks/benefits, 5 impacts).
- 23894 analysis ↔ NIST Measure (1 identify methods, 2 evaluate, 3 mechanisms for tracking, 4 feedback).
- 23894 evaluation + treatment + monitoring ↔ NIST Manage (1 prioritize, 2 strategies, 3 third-party, 4 ongoing monitor and improve).
- 23894 communication + consultation + recording + reporting ↔ NIST Govern (1 policies, 2 accountability, 3 workforce, 4 culture, 5 engagement, 6 third-party).
Lesson 043 walks NIST AI RMF profile-building for a sector use case; this lesson establishes the underlying 23894 process the profile sits on.
Cross-Walk to EU AI Act Article 9 (Risk-Management System for High-Risk Providers)
Article 9 requires a continuous iterative RMS across the lifecycle. Article 9(2)(a)-(d):
- (a) identification and analysis of known and reasonably foreseeable risks ↔ 23894 Steps 2-3.
- (b) estimation and evaluation when used in accordance with intended purpose and reasonably foreseeable misuse ↔ 23894 Step 3 (analysis) + Step 4 (evaluation).
- (c) evaluation of risks emerging from post-market monitoring data ↔ 23894 Step 6 (monitoring and review).
- (d) adoption of appropriate and targeted risk-management measures ↔ 23894 Step 5 (treatment).
Article 9(3) requires measures to consider combined application of Section 2 (Articles 8-15): risk-management, data, technical documentation, logging, transparency, oversight, accuracy/robustness/cybersecurity. Structure RTP controls against Article 9(3) to produce one-document Article 9 evidence.
Article 9(4) requires residual risk be judged acceptable, with Article 13 user information. The 23894 residual-risk acceptance signature satisfies this.
Sectoral and Supervisory Overlays
The same analysis informs SR 11-7 / OCC 2011-12 model-risk (lessons 067-069); EBA credit-scoring guidelines; NYC LL 144 bias audits (lesson 015); Texas TRAIGA disparate-impact (lesson 014); Colorado SB 24-205 (lesson 013); Article 27 FRIA (lessons 041, 044-049). Cross-walks in the 23894 artifacts reduce duplicative work across overlays.
Risk Treatment Plan (RTP) Template - The L3 Deliverable
The RTP is the central L3 artifact from Step 5. A defensible 2026 template has the following columns per risk row:
- Risk ID. Unique identifier tied to the risk register (e.g., R-01).
- Risk description and source family. Plain-language description and the 23894 source-family classification (data / model / human-AI / lifecycle / third-party).
- Pre-treatment likelihood × consequence × overall rating. The Step-3 analysis result.
- Treatment option(s). Avoid / mitigate / transfer / accept; rationale for selection.
- Controls implemented. Specific preventive / detective / corrective controls; reference to control documents and SOPs.
- Owner. Named role (not just function) responsible for implementation.
- Target completion date. Planned date; actual completion date once achieved.
- Post-treatment residual likelihood × consequence × overall rating. The expected residual once controls are in place; updated to actual after implementation.
- Acceptance authority. Named role with authority to accept residual risk (typically CAIRO + business owner + General Counsel for legal risks + CISO for security risks); signature evidence.
- Review cadence. Monthly / quarterly / annually / triggered; basis for the cadence selection.
- Trigger events. Specific events forcing re-evaluation (model update / vendor change / regulator inquiry / incident / metric breach).
- Cross-walk references. ISO 42001 Annex A controls; EU AI Act Article 9(3) reference; NIST AI RMF Manage function; sectoral overlay.
- Status. Open / In Progress / Implemented / Verified / Accepted / Closed.
The RTP is a living document: reviewed monthly by the AI Governance Committee, included in the quarterly board AI risk dashboard, sampled by ISO 42001 internal audit, inspected by external auditors at Stage 2. Template quality compounds across the multi-framework audit set.
Common ISO 23894 Application Mistakes
Mistake 1 - Skipping the Scope-Context-Criteria Step
The most common 23894 failure. Practitioners jump to identification because "we know the risks." Evaluation then lacks a defensible threshold structure. Auditor finds: arbitrary evaluation; inconsistent treatment decisions; no traceable basis for residual-risk acceptance. Fix: produce the scope-context-criteria document first, signed by Responsible AI Officer and business owner; reference it in every subsequent step.
Mistake 2 - Identification Without Source Analysis
Practitioners list "risks" as undifferentiated brainstorm output without source-family classification. Auditor finds: gaps in entire source families (typically third-party and lifecycle); inability to demonstrate completeness. Fix: use the 23894 five-family taxonomy as the scaffold; document coverage per family; cross-reference NIST AI 600-1, OWASP LLM Top 10, OWASP Agentic Top 10, MITRE ATLAS where relevant.
Mistake 3 - Pseudo-Quantification on Likelihood Without Data
Practitioners attach numerical probabilities (0.05, 0.12) with no statistical basis. Numbers create false confidence; audit cannot trace the analytical basis. Fix: 5-band qualitative scales with verbal anchors as default; reserve quantitative scoring for empirically measured risks (bias-disparity, drift KL divergence, prompt-injection rates); document the basis per score.
Mistake 4 - Treatment Without Monitoring
The RTP lists controls and owners but no control-effectiveness monitoring. Six months later, the auditor asks "is the bias-monitoring dashboard being reviewed?" and the answer is "we built it but stopped reviewing in May." Fix: monitoring schedule with named reviewers and cadence per control; AI Governance Committee monthly briefing; quarterly internal-audit sampling of control execution evidence.
Mistake 5 - Static RTP
The RTP is produced at intake and never updated. New operational risks, incident-response controls, vendor changes, regulator guidance, none reaches the RTP. Auditor finds: RTP 6 months stale; incident-report risks not in the register; operational controls not in the RTP. Fix: monthly RTP refresh as part of AI Governance Committee cadence; explicit trigger events; versioning with change-tracking.
Mistake 6 - No Recording and Reporting Evidence
The analysis exists in slide decks, spreadsheets, and engineering tickets, not in audit-traceable artifacts. Auditor cannot trace decisions, owners, dates, signatures. Fix: build the recording templates at Step 1; populate as the process runs; integrate with AIMS Clause 7.5 documented information; treat the artifact set as the primary deliverable.
Key Takeaways
- ISO/IEC 23894:2023 is the AI-specific overlay on ISO 31000:2018, published February 2023; guidance (not certifiable); the standard ISO 42001 Clauses 6.1.2/6.1.3 auditors expect to see referenced.
- Eight process steps mirror ISO 31000: (1) scope, context, criteria; (2) risk identification; (3) risk analysis; (4) risk evaluation; (5) risk treatment; (6) monitoring and review; (7) communication and consultation (parallel); (8) recording and reporting.
- Five AI-specific risk source families (from 23894 Annex A/B): data-related (quality, representativeness, bias, drift); model-related (accuracy, robustness, interpretability, security); human-AI-interaction (over-reliance, under-reliance, accessibility, transparency); lifecycle (development, deployment, post-deployment monitoring, retirement); third-party (foundation models, datasets, tools, infrastructure).
- Worked example: EU e-commerce recommendation system across 14 Member States: scope/context/criteria producing 6-page signed document; identification of 28 risks with bias, filter bubble, drift, prompt-injection, IP, foundation-model-upstream as top tier; analysis with 5-band qualitative + quantitative where data exists; evaluation prioritizing 3 risks for committee-level treatment; RTP with named owners, target dates, residual ratings, and acceptance signatures; monitoring schedule with quarterly + triggered refresh.
- Risk treatment plan (RTP) template: the L3 deliverable: risk ID, description, source family, pre-treatment rating, treatment option, controls, owner, dates, residual rating, acceptance authority, review cadence, trigger events, cross-walk references, status.
- Cross-walk to ISO 42001: A.6.1.1 AI impact assessment ↔ 23894 identification + analysis; A.6.4 post-deployment monitoring ↔ 23894 monitoring; A.5 impact-assessment communication ↔ 23894 communication and consultation; A.6.2 system requirements ↔ 23894 scope + treatment.
- Cross-walk to ISO 31000: 23894 IS the AI-specific overlay on 31000; same eight steps; organizations with existing ISO 31000 ERM extend to AI by adopting the 23894 Annex A/B catalog and AI guidance at each step.
- Cross-walk to NIST AI RMF 1.0: Map ≈ scope/context/criteria + identification; Measure ≈ analysis; Manage ≈ evaluation + treatment + monitoring; Govern ≈ communication + consultation + recording + reporting.
- Cross-walk to EU AI Act Article 9 RMS: 9(2)(a) ↔ identification + analysis; 9(2)(b) ↔ evaluation; 9(2)(c) ↔ monitoring; 9(2)(d) ↔ treatment; 9(3) ↔ controls structured against Articles 8-15 reference; 9(4) ↔ residual-risk acceptance signature.
- Six common mistakes: skipping scope/context/criteria; identification without source analysis; pseudo-quantification on likelihood without data; treatment without monitoring; static RTP; no recording/reporting evidence.
- The L3 artifact set: scope-context-criteria document + risk register + risk treatment plan + monitoring schedule + management-review evidence: produced once, cited four times across ISO 42001, ISO 31000, NIST AI RMF, and EU AI Act Article 9.
Skill.re