AI Governance, Risk & Red Teaming
Proficient · M6 · lesson 6 of 31 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Annex XIII Non-FLOPs Designation - When Commission Designates GPAI With Systemic Risk Below the 10^25 Threshold
📖
now learning

Annex XIII Non-FLOPs Designation - When Commission Designates GPAI With Systemic Risk Below the 10^25 Threshold

15 min

Article 51(1)(a) gets the headlines: the 10^25 cumulative training FLOPs presumption that pulls GPT-4-class, Claude-class, Gemini-class, and Llama 3.1-405B-class models into the systemic-risk tier. Article 51(1)(b) gets the procurement memos that practitioners actually have to write. The Commission's discretionary designation power, exercised through the ten Annex XIII criteria, can pull a below-threshold open-weight model into the Article 55 regime on the basis of business-user reach, autonomy, tool access, or capability benchmarks. Llama 3.1-70B sits below 10^25 FLOPs and could still get designated. DeepSeek V3 sits in the disputed band and is a prime Annex XIII candidate. Your procurement, attestation, and downstream-deployer workflow has to hold up if either gets that letter from the AI Office. This lesson is how you walk Annex XIII against two below-threshold open-weight models in your stack, build the contractual flow-down for a designation event, and produce the L3 designation-risk memo that procurement, legal, and the AI Governance Committee can act on.

Why Non-FLOPs Designation Is the Procurement Question Nobody Asked in 2025

Through 2025, the entire GPAI conversation gravitated to the compute presumption. The 10^25 threshold was clean, quantitative, and easy to read off published research estimates. Procurement teams built GPAI exposure maps with three columns: vendor, FLOPs estimate, presumed-designated yes/no. The matrix worked for OpenAI, Anthropic, Google DeepMind, and Meta's Llama 3.1-405B. It worked badly for everything below the line.

The Commission's Article 51(1)(b) power is the regulatory mechanism that fills that gap. The text reads: "A general-purpose AI model shall be classified as a general-purpose AI model with systemic risk if it has high-impact capabilities evaluated on the basis of appropriate technical tools and methodologies, including indicators and benchmarks." Article 51(2) then directs the Commission to "adopt delegated acts… to amend the thresholds listed in paragraph 1, as well as to supplement benchmarks and indicators in light of evolving technological developments, such as algorithmic improvements or increased hardware efficiency." Annex XIII is the criteria list the Commission uses when it exercises that designation power below the FLOPs presumption. The first Article 51(1)(b) designations are expected in the second half of 2026 or 2027, once the GPAI enforcement powers go live on Aug 2, 2026, the date Omnibus VII did not move.

Why does this matter for a deployer whose organization does not train foundation models? Three operational reasons. First, the open-weight models in your stack, Llama 3.1-70B fine-tunes, Mistral Small derivatives, Phi-3-Medium, DeepSeek V3 via redistributors, Gemma 2, Qwen 2, all sit below the FLOPs presumption and all have plausible business-user-reach or autonomy profiles that could trip Annex XIII. Second, your procurement contracts with non-signatory providers (Meta, DeepSeek, Alibaba, Baidu) most likely do not anticipate a designation event. If the upstream provider gets designated, your Annex XII receivable pack changes overnight, Article 55 evidence becomes required, and your Annex IV technical file for downstream high-risk systems may need a substantial refresh. Third, the audit committee is going to ask, in the next two quarters: "What is our exposure if any of our below-threshold open-weight models gets designated under Annex XIII?" That question gets answered by the Annex XIII designation-risk memo built off of this lesson.

And there is a sharper edge. Unlike the compute presumption, which is largely binary and which the Commission verifies through public training-FLOPs disclosures, the Annex XIII designation is discretionary, multi-criteria, and weighted by Commission judgment. A model can be designated on a single dominant criterion (e.g., 50,000 EU business users with broad agentic tool access) even if the others sit below typical thresholds. The Commission's discretion is what makes pre-designation procurement preparation valuable: by the time the designation letter arrives, the contractual flow-down has to already be in place.

Reading Annex XIII - The Ten Criteria the Commission Will Apply

Annex XIII lists ten criteria the Commission considers when deciding whether to designate a GPAI model with systemic risk under Article 51(1)(b). The text reads as a non-exhaustive list, the Commission can consider other relevant factors, but the ten are the operational checklist procurement should walk per below-threshold model.

Criterion 1 - Number of Parameters

Parameter count is the first signal. Models in the 7B-13B range (Llama 3.1-8B, Mistral 7B, Phi-3-Mini) are very unlikely Annex XIII candidates on parameter count alone. Models in the 30B-70B range (Llama 3.1-70B, Mixtral 8x22B) are at modest risk. Models above 100B (DeepSeek V3 at 671B MoE with ~37B active parameters, Qwen 2-72B variants) sit in a band where parameter count alone supports designation if combined with other criteria. Sparse MoE architectures complicate the count, total parameters versus active parameters per forward pass, and the Commission has signaled it will look at both. For procurement: record total parameter count and active-parameter count separately for any MoE model in the stack.

Criterion 2 - Quality or Size of Dataset

Token count is the proxy. Frontier-class training datasets are now in the 10-30 trillion token range. Llama 3.1-405B was trained on ~15T tokens; Llama 3.1-70B on the same. Below-threshold models trained on smaller datasets (1-5T tokens) sit at lower risk. The dataset-quality signal, curation, diversity, filtering for harmful content, is harder to quantify externally but is read off vendor documentation, public training-data summaries (Article 53(1)(d)), and independent evaluations. For procurement: record token count, dataset composition (web, books, code, multilingual, synthetic), and any documented curation methodology.

Criterion 3 - Compute Used for Training

Compute is the FLOPs figure that Article 51(1)(a) makes presumptive. For Annex XIII purposes, the Commission can look at FLOPs estimated through a combination of variables, training cost, training time, energy consumption, GPU-hours, accelerator-hours, when direct FLOPs disclosure is unavailable. A model trained for $50M on H100 clusters for 4 months is in the 10^24-10^25 range regardless of the provider's published figure. The Commission's compute audit methodology (in development through 2026) is expected to triangulate from procurement data, energy disclosures, and model-card claims. For procurement: record vendor-published FLOPs estimate, independent estimate range (EpochAI, Stanford HELM), training cost if disclosed, and energy disclosure if available.

Criterion 4 - Input and Output Modalities

Multi-modality is a designation amplifier. Text-only models face lower Annex XIII risk than text+image+audio+video models. The reasoning is that multi-modal capabilities expand the attack surface (image-based prompt injection, audio adversarial perturbations, video synthesis risks) and broaden the systemic-impact pathways (synthetic content marking under Article 50(2), CBRN evaluation under Recital 110). A below-threshold text-only model is less interesting to the Commission. A below-threshold model that accepts image and video input and generates image, video, and audio output is more interesting. For procurement: record input modalities and output modalities for every model, including planned-but-not-shipped modality expansions.

Criterion 5 - Benchmarks and Evaluations of Capabilities

State-of-the-art benchmarks (MMLU, MMLU-Pro, GPQA, HumanEval, SWE-Bench, MATH, GSM8K), autonomous-agent benchmarks (SWE-Bench-Verified, GAIA, AgentBench), and dangerous-capability evaluations (BioBench, CyberSecEval, AutoRedTeam, JBFuzz coverage scores) all factor in. A model that scores at or near frontier-class on agentic benchmarks (GAIA, SWE-Bench) is a designation candidate even with sub-threshold FLOPs. The Commission's evaluation taxonomy, converged with U.S. CAISI and UK AISI dangerous-capability evals, provides the scoreboard. For procurement: maintain a benchmark-result row per model with peer comparison to designated peers.

Criterion 6 - High Impact on Internal Market Through Reach

This is the criterion to commit to memory. The Annex XIII text reads: "Whether the model has a high impact on the internal market because of its reach, for example, the fact that the model is offered for use to at least 10,000 registered business users established in the Union." The 10,000 EU business users threshold is the operational pull-in. It is the criterion most likely to be invoked in the first round of Article 51(1)(b) designations because it is measurable, scale-driven, and aligns with the Commission's market-impact concern.

For procurement, business-user reach includes direct customers of the model provider and downstream redistributors. A model offered by a base provider to 2,000 direct EU business users but distributed via a SaaS partner to an additional 15,000 EU business customers is at 17,000, over the threshold. The Commission's draft methodology counts the aggregate. Open-weight models distributed via Hugging Face, integrated into SaaS products, and resold via cloud-hyperscaler marketplaces have a particularly hard-to-track reach footprint. For procurement: estimate aggregate EU business-user reach including downstream redistribution; document the estimation methodology.

Criterion 7 - Number of Registered End Users

End-user count is distinct from business-user count. A consumer-facing chatbot serving 50M EU consumers crosses end-user thresholds even if its business-customer footprint is small. The Commission has not published a specific end-user numerical threshold equivalent to the 10,000 business users, but Recital 111 references "potential reach and use" as a designation factor. For procurement: record end-user reach where the model is exposed in consumer-facing products.

Criterion 8 - Autonomy and Scalability

Autonomy and scalability are the agentic-deployment criteria. A model that supports autonomous multi-step task execution, persistent memory, self-improvement loops, or large-scale parallel deployment scores higher on this criterion. The OWASP Agentic Top 10 attack classes (goal hijack, memory poisoning, tool misuse, code-execution escape) all assume the autonomy/scalability conditions that this criterion targets. A non-agentic chatbot scores low. An agent-orchestration model with tool-calling, code-execution, and persistent memory scores high. For procurement: record autonomy tier (tool-calling only, multi-step planning, persistent memory, self-improvement) and scalability profile (parallel-deployment patterns observed in production).

Criterion 9 - Tools the Model Has Access To

Tool access, agent tooling, code execution, web browsing, file-system access, external APIs, database access, cloud-resource provisioning, payment-system integration, is the criterion that captures the agentic-deployment risk surface. A model packaged with broad tool access (the post-GPT-4 standard) is a higher Annex XIII candidate than the same model used only for text completion. The Commission's concern, articulated in 2026 guidance drafts, is that tool access transforms model capability from "outputs text" to "executes actions in the world." For procurement: record default tool access, common deployment-time tool extensions, and any opt-out controls.

Criterion 10 - State of the Art Relative to Other Models

The relative-frontier criterion captures models that introduce new modality combinations, new capability classes, or new architectural approaches that the prior generation lacked. DeepSeek V3's reasoning-trace efficiency and MoE-architecture cost-effectiveness drew Commission attention not because of raw scale but because of state-of-the-art-relative-to-cost dynamics. A below-threshold model that achieves frontier-class results on key benchmarks at lower FLOPs is more interesting to the Commission, not less. For procurement: record any state-of-the-art claims the vendor makes and any independent confirmation.

The Commission's Discretion - Not Formulaic, Weighted Multi-Criteria

The Commission's Article 51(1)(b) designation is not a formula. There is no published weight for each criterion. The Commission considers the ten criteria as a whole, with the AI Office providing technical input and the AI Board providing Member State coordination. The first designations will signal which criteria the Commission weights most heavily; until then, procurement should treat any model meeting three or more criteria at medium-to-high levels as at material designation risk and contract accordingly.

Two Worked Examples - Below-Threshold Open-Weight Models in the Stack

Apply the Annex XIII walk to two representative below-threshold open-weight models that a Fortune-500 in 2026 commonly has in the stack. Both sit below the 10^25 FLOPs presumption. Both are non-signatory providers. Both could plausibly be designated under Article 51(1)(b) in the next 12-18 months.

Example 1 - Meta Llama 3.1-70B (Open-Weight, Non-Signatory)

Llama 3.1-70B is the workhorse open-weight model of 2024-2026. Released July 2024 by Meta under the Llama 3.1 Community License, it sits below the 10^25 FLOPs threshold (estimates put it at ~6 × 10^24 FLOPs; the larger Llama 3.1-405B sits above at ~4 × 10^25). It is broadly deployed via Hugging Face downloads, Together AI hosting, Fireworks AI, Anyscale, IBM watsonx, Databricks Mosaic AI, and via fine-tunes packaged into enterprise SaaS products.

Annex XIII walk:

  • Criterion 1 (parameters): 70B dense, medium-high risk on parameter count alone.
  • Criterion 2 (dataset): ~15T tokens, frontier-class dataset scale.
  • Criterion 3 (compute): ~6 × 10^24 FLOPs, below the presumption but in the upper-band where Commission compute-audit methodology could re-estimate higher.
  • Criterion 4 (modalities): Text in/text out today; Meta has signaled multi-modal Llama variants for 2026-2027, designation risk rises sharply with modality expansion.
  • Criterion 5 (benchmarks): Frontier-class on most academic benchmarks; SWE-Bench-Verified and GAIA scores in the upper band for non-frontier models.
  • Criterion 6 (business-user reach): Estimated 50,000+ EU business users via direct Meta API access plus downstream redistribution via cloud hyperscalers and SaaS integrations, well over the 10,000 threshold.
  • Criterion 7 (end users): Indirect end-user exposure via SaaS products is in the tens of millions across the EU.
  • Criterion 8 (autonomy/scalability): Fine-tune ecosystems enable autonomy/scalability that the base model does not exhibit directly; agentic deployments on Llama 3.1-70B are widespread.
  • Criterion 9 (tools): Llama 3.1-70B supports function-calling and is commonly deployed with broad tool access; agentic-tooling layer is mature.
  • Criterion 10 (state of the art): Frontier-class for open-weight models; defines the price-performance frontier for self-hosted enterprise deployment.

Aggregate designation risk: ELEVATED. Llama 3.1-70B meets six to eight of the ten criteria at medium-to-high levels, with business-user reach the dominant single criterion. Meta is a non-signatory. Procurement contracts with Meta and with downstream Llama-redistributors (Together AI, Fireworks, IBM watsonx, Databricks Mosaic, Anyscale) should anticipate a designation event and include the contractual flow-down language detailed below.

Example 2 - DeepSeek V3 (Open-Weight, Non-Signatory, China-Headquartered)

DeepSeek V3, released December 2024, is the second case study. ~671B total parameters in a Mixture-of-Experts architecture with ~37B active parameters per forward pass. Training compute is publicly contested, DeepSeek's own disclosure put training cost at ~$5.6M (suggesting compute in the 2-3 × 10^24 FLOPs band), but independent estimates and re-analysis put it as high as 1.5 × 10^25 FLOPs when amortized R&D compute is included. The model is broadly deployed via the DeepSeek hosted API, via Hugging Face open-weights, via Together AI, Fireworks AI, Perplexity, Groq, and via redistributors into the EU market.

Annex XIII walk:

  • Criterion 1 (parameters): 671B total, 37B active, high on total count, moderate on active.
  • Criterion 2 (dataset): ~14.8T tokens disclosed, frontier-class.
  • Criterion 3 (compute): Compute disclosure is the contested band, the Commission will likely use the upper estimate (~10^25) for designation analysis.
  • Criterion 4 (modalities): Text today, vision-language variant (DeepSeek-VL2) available; multi-modal designation profile is rising.
  • Criterion 5 (benchmarks): Frontier-class on reasoning benchmarks (MATH, GSM8K, GPQA); SWE-Bench and HumanEval at upper band.
  • Criterion 6 (business-user reach): Estimated 12,000-15,000 EU business users via redistributors and direct API, above the 10,000 threshold.
  • Criterion 7 (end users): DeepSeek consumer app has substantial EU end-user footprint.
  • Criterion 8 (autonomy/scalability): Function-calling supported; agentic deployments increasingly common; reasoning-trace efficiency makes scalable agentic patterns economically viable.
  • Criterion 9 (tools): Tool-calling supported; agentic-tooling layer maturing.
  • Criterion 10 (state of the art): Defines the cost-efficient frontier; the Commission has cited DeepSeek-class models in its emerging designation guidance as exemplars of "state of the art at lower compute."

Aggregate designation risk: HIGH. DeepSeek V3 meets seven to nine of the ten criteria at medium-to-high levels. The China-headquartered provider adds an overlay: GDPR cross-border data-transfer concerns under Schrems II, China's CAC AI regulations as a parallel obligation set, and AI Office wariness about cooperation under Article 25(2) with non-EU providers headquartered outside the U.S./UK reciprocal frameworks. DeepSeek is a non-signatory to the GPAI Code of Practice. Procurement contracts with DeepSeek-hosting redistributors should treat designation as a high-likelihood event in the next 12-18 months and price the contractual flow-down accordingly.

The Procurement Workflow If the Commission Designates - A Six-Step Playbook

Assume the Commission, in Q4 2026 or Q1 2027, issues an Article 51(1)(b) designation letter for one of the below-threshold open-weight models in your stack. What happens in the 30, 60, and 90 days that follow? The procurement workflow has six steps.

Step 1 - Immediate Notification From Upstream Provider

Under Article 25(2), upstream providers and downstream deployers have a duty to cooperate. The first thing that should happen post-designation is a notification from the upstream provider (Meta, DeepSeek) or the immediate redistributor (Together AI, IBM watsonx, etc.) confirming the designation. Your contract should specify a maximum notification window (recommended: 5 business days) and the format of the notification (regulatory letter copy, model identifier, designation date, effective date for Article 55 obligations). Without contractual notification language, you may learn of the designation from the AI Office's public register weeks after the fact, losing critical preparation time.

Step 2 - Annex XI / Annex XII Receivable Refresh

Pre-designation, the upstream provider owed Article 53 obligations: Annex XI technical documentation and Annex XII downstream-deployer information. Post-designation, the provider also owes Article 55 obligations: state-of-the-art model evaluations, adversarial-testing documentation, systemic-risk assessments, cybersecurity protection evidence, incident-reporting protocols. The Annex XII pack the provider delivers to you must expand to include Article 55 evidence. Your contract should specify a refresh SLA (recommended: 30 days post-designation for an updated Annex XII pack with Article 55 evidence appended) and a periodic-refresh cadence (recommended: quarterly during the first year post-designation, semi-annually thereafter).

Step 3 - Annex IV Technical-File Refresh for Downstream Systems

For every high-risk AI system in your portfolio built on top of the now-designated model, the Annex IV technical documentation file may need refresh. Section 2 (detailed system description) needs the updated upstream-model identifier and designation status. Section 3 (monitoring, functioning, and control) may need updated evaluation results from the upstream provider's Article 55 evaluations. Section 4 (performance description) may incorporate updated benchmark and adversarial-testing results. Section 5 (risk-management system) needs an updated risk assessment that incorporates the systemic-risk profile of the upstream model. Section 7 (harmonized-standards conformity) may shift if updated standards apply. The refresh effort per high-risk system is non-trivial, typically 80-160 hours of regulatory and engineering work, and should be scoped immediately on notification.

Step 4 - Substantial-Modification Analysis on the Deployer Side

Article 43(4) defines substantial modification as a change that affects compliance with the high-risk requirements (Articles 8-15) or changes the intended purpose. The upstream model's designation may trigger substantial-modification analysis on your downstream system. If the upstream provider's Article 55 evaluations reveal new capabilities or new risk profiles, your downstream system's risk-management posture may need re-evaluation, and the Article 43 conformity-assessment route may need to be re-run. For Annex III §1 systems on Annex VII Module H, this means engaging the notified body for re-certification analysis. For Annex VI internal-control systems, this means re-running the internal assessment and refreshing the Article 47 declaration.

Step 5 - Procurement Contract Amendments

Existing contracts likely lack the post-designation evidence obligations the upstream provider now owes. Amendments should add: Article 55 evidence-sharing commitments (model evaluations, adversarial-testing reports, systemic-risk assessments, cybersecurity protection evidence); cooperation under Article 25(2) for any incident or regulatory inquiry; indemnification language for downstream impact attributable to upstream non-compliance; cost allocation for re-conformity-assessment work the designation triggers; and updated SLAs for documentation delivery, incident notification, and inquiry response.

Step 6 - Brief to AI Governance Committee and Audit Committee

The designation event is an audit-committee-reportable event. The brief should cover: the designation fact and effective date; the affected systems in your portfolio; the operational mitigations underway; the contractual amendments in process; the Annex IV refresh status per high-risk system; the substantial-modification analysis outcomes; the budget impact; and the timeline to full post-designation conformity. Without the pre-built Annex XIII designation-risk memo, this brief is reactive and incomplete. With the memo in place, the brief is the activation of a pre-planned response.

Contractual Flow-Down Language - What the Upstream Procurement Contract Must Say

The pre-designation procurement contract is where the post-designation workflow gets enforceable. For any below-threshold open-weight model with elevated or high Annex XIII designation risk, the upstream contract (with the base provider or the immediate redistributor) should include five clauses.

Clause 1 - Notification on Commission Designation Event

"In the event that the Commission designates the Model as a general-purpose AI model with systemic risk under Article 51(1)(a) or Article 51(1)(b) of Regulation (EU) 2024/1689 (the 'AI Act'), Provider shall notify Customer in writing within five (5) business days of the designation, providing a copy of the designation letter or public notice, the effective date of the designation, and the effective date of any Article 55 obligations applicable to the Model."

Clause 2 - Annex XI/XII Delivery SLA Refresh on Designation

"Upon a designation event under Clause 1, Provider shall deliver to Customer, within thirty (30) days, an updated Annex XII downstream-deployer documentation pack incorporating Article 55 evidence, including: model-evaluation results conducted per Article 55(1)(a); adversarial-testing methodology and findings; systemic-risk assessment and mitigation measures per Article 55(1)(b); cybersecurity protection measures per Article 55(1)(d); and serious-incident reporting protocols per Article 55(1)(c). Provider shall refresh the pack quarterly during the first twelve months post-designation and semi-annually thereafter."

Clause 3 - Article 55 Evidence-Sharing Commitment

"Provider shall, on Customer's reasonable request and at no additional cost during the first twelve months post-designation, provide Customer with documentation, evaluation results, evidence, and information necessary for Customer to fulfill its obligations under Articles 9, 10, 11, 13, 14, 15, 17, and 26 of the AI Act in respect of any high-risk AI system Customer operates that incorporates the Model. Provider shall cooperate with Customer's notified body or competent authority on reasonable request for clarification or evidence."

Clause 4 - Indemnification for Downstream Impact

"Provider shall indemnify Customer for direct losses, including reasonable legal fees and re-conformity-assessment costs, attributable to Provider's failure to comply with Article 53 or (post-designation) Article 55 obligations, where such failure results in a regulatory finding, enforcement action, or required corrective action against Customer's downstream high-risk AI system. Indemnification is capped at [negotiated cap]; specific exclusions apply per Schedule X."

Clause 5 - Re-Conformity-Assessment Cost Allocation

"In the event the designation event triggers a substantial-modification analysis under Article 43(4) on any Customer downstream high-risk AI system and the analysis concludes that re-certification is required, Provider shall reimburse Customer for [50% / 75% / 100%] of the documented re-certification costs (notified-body fees, Annex IV refresh, technical-staff cost, legal review) up to a cap of [negotiated cap], where the substantial modification is attributable to the upstream Model's designation rather than to changes made by Customer to the downstream system."

Negotiating these clauses pre-designation is materially easier than post-designation. Once the designation letter arrives, the upstream provider's leverage rises and the deployer's leverage falls. The L3 procurement playbook is to identify Annex XIII-at-risk models in the stack now, prioritize contract amendments in the next two procurement cycles, and price the amendments into renewal economics.

The Annex XIII Designation-Risk Memo - The L3 Artifact

The L3 deliverable in this area is the Annex XIII designation-risk memo, a procurement-facing artifact owned by the AI Governance Committee, refreshed quarterly, and produced for the audit committee at semi-annual cadence. It has the following structure.

Header. Date of memo. Period covered. Author. Distribution list (AI Governance Committee, Procurement, Legal, Audit Committee).

Row per below-threshold GPAI model in the stack. One row per model used directly or indirectly across the AI portfolio that sits below the Article 51(1)(a) 10^25 FLOPs presumption. The row covers:

  • Model. Vendor, version, release date, license. (E.g., "Meta Llama 3.1-70B base, July 2024, Llama 3.1 Community License.")
  • Deployment surface. Where in the stack the model is used. (E.g., "Internal customer-service agent fine-tune; embeddings for resume-screening upstream of the Annex III §4 system; coding-assistant for engineering team via IBM watsonx hosting.")
  • Annex XIII criterion-by-criterion analysis. Ten-row sub-table with per-criterion assessment (low/medium/high) and the source data behind each assessment. Parameter count; dataset size; compute estimate; modalities; benchmark profile with peer comparison; estimated EU business-user reach with methodology; end-user reach; autonomy/scalability profile; tool access; state-of-the-art positioning.
  • Aggregate designation-risk score. High / Medium / Low based on the criterion-by-criterion roll-up. Three-tier scoring is sufficient; finer granularity is false precision given the Commission's discretionary process.
  • Procurement-contract amendment status. For each of the five contractual flow-down clauses, status is: Present / Drafted-Pending / Not Yet Negotiated / Refused-by-Provider / Not Applicable. Counterparty (base provider or redistributor). Renewal date as the natural amendment window.
  • Operational mitigations in place. Pre-designation operational steps: monitoring of the AI Office's public designation register; quarterly review of Annex XII receivables; pre-prepared Annex IV refresh templates for downstream systems; pre-identified notified-body relationships for substantial-modification analysis; budget reserve for re-conformity work.
  • Refresh cadence. Quarterly criterion-walk refresh; semi-annual contractual-amendment status refresh; ad-hoc refresh on any AI Office designation announcement or material model change.
  • Owner and next review date. Procurement or AI Governance Committee designate; next quarterly review date.

Aggregate portfolio view. Roll-up table showing total below-threshold models by aggregate risk tier (High / Medium / Low); contract-amendment coverage; estimated re-conformity exposure if designation occurs; budget reserve recommended; quarterly trend.

Recommendations and decisions. Per-row recommendations: amend contract at next renewal; accelerate contract amendment via mid-cycle redlining; migrate off the model to a designated-and-Code-signatory peer; accept residual designation risk with documented rationale; escalate to AI Governance Committee for further action.

For a Fortune-500 in 2026, the memo typically covers 8-15 below-threshold open-weight models, with 2-4 in the High tier (Llama 3.1-70B, DeepSeek V3, Mixtral 8x22B, Qwen 2-72B variants are common High-tier residents), 4-7 in Medium, and 2-4 in Low. The memo is read alongside the GPAI exposure map from lesson 003 and feeds into the Annex IV technical-file refresh schedule and the procurement-renewal calendar.

Six Common Annex XIII Mistakes - And the Audit-Defensible Counter-Stances

Mistake 1 - Assuming Below-Threshold Means Safe

The most common 2025 procurement posture treats the 10^25 FLOPs threshold as a binary cutoff: above = systemic risk, below = no risk. Article 51(1)(b) makes that posture indefensible. The Commission can designate any below-threshold model based on the ten Annex XIII criteria; business-user reach alone can pull a sub-threshold model into the designation tier. The L3 counter-stance: every below-threshold GPAI model in the stack gets an Annex XIII designation-risk score, regardless of FLOPs. The score drives contractual flow-down priority. No model is "safe" by virtue of being below 10^25 FLOPs.

Mistake 2 - Missing the Business-User-Reach Criterion

The 10,000 EU business users criterion is the most measurable and the most often missed. Many procurement maps capture FLOPs and parameters but do not estimate downstream business-user reach. The L3 counter-stance: estimate aggregate EU business-user reach for every model, including downstream redistribution. Document the estimation methodology. Update quarterly. Treat 5,000+ EU business users as a yellow flag and 10,000+ as a red flag for designation risk.

Mistake 3 - Weak Contractual Flow-Down for Designation Events

Most pre-2026 procurement contracts with non-signatory providers do not anticipate a designation event. The result: when designation happens, the deployer scrambles to extract Article 55 evidence with no contractual hook. The L3 counter-stance: pre-2026 contracts are amended at next renewal with the five flow-down clauses; mid-cycle amendments are negotiated for high-priority models; new contracts include the flow-down language by default.

Mistake 4 - No Quarterly Annex XIII Designation-Risk Refresh

Designation-risk profiles change quarter to quarter. Models add modalities. Business-user counts grow. Benchmark scores climb. New tool integrations expand attack surface. A static memo is a stale memo. The L3 counter-stance: quarterly refresh of the criterion-by-criterion analysis for every model in the memo; trend tracking to identify models whose risk profile is rising; escalation triggers when a model crosses risk-tier boundaries.

Mistake 5 - Missing the Autonomy and Tool-Access Criteria

Criteria 8 (autonomy/scalability) and 9 (tools) are the agentic-deployment criteria. A model used statically for text completion has a different risk profile than the same model packaged into an agentic deployment with code execution, web browsing, and file-system access. Many memos record the base-model profile and miss the deployment-time risk amplification. The L3 counter-stance: record both the base-model profile and the deployment-time autonomy/tools profile per model usage; treat agentic-deployment patterns as designation-risk amplifiers.

Mistake 6 - Siloed From the GPAI Exposure Map

The Annex XIII designation-risk memo and the GPAI exposure map are sister artifacts. The exposure map (lesson 003) covers all GPAI models in the stack with designation status, signatory status, and Annex XI/XII receivable status. The Annex XIII memo zooms into the below-threshold subset with criterion-by-criterion designation-risk analysis. Treating them as separate artifacts produces duplicate work and inconsistent risk ratings. The L3 counter-stance: the Annex XIII memo is the below-threshold drill-down of the GPAI exposure map; refreshes happen on the same quarterly cycle; the same owner maintains both.

Aligning the Annex XIII Memo With ISO 42001, NIST AI RMF, and the GPAI Code

The Annex XIII designation-risk memo carries weight across the framework stack. ISO/IEC 42001:2023 Annex A control A.10 (third-party relationships) requires risk assessment for third-party AI providers; the Annex XIII memo is the specific A.10 evidence for upstream GPAI providers with elevated designation risk. NIST AI RMF 1.0 functions Map 4 (third-party risks mapped) and Govern 6.1 (third-party policies) are similarly evidenced. The GPAI Code of Practice cross-walk: if a non-signatory provider's model is in your memo at High risk, the absence of Code-signatory presumption-of-compliance is itself a material risk factor, non-signatories face a longer evidence-construction effort post-designation than signatories. CycloneDX 1.7 ML-BoM: the model identifier, version, training-data references, and provenance attestations in the ML-BoM provide machine-readable input to the memo; the memo provides the qualitative risk overlay. The dual-citation pattern continues: write the memo once, cite it against four frameworks.

Key Takeaways

  • Article 51(1)(b) is the Commission's discretionary designation power. A GPAI model can be designated as systemic-risk even when training FLOPs are below 10^25, based on the ten Annex XIII criteria. First Article 51(1)(b) designations are expected H2 2026 or 2027 once enforcement powers go live Aug 2, 2026.
  • Annex XIII lists ten criteria, committed to memory. Parameters; dataset; compute; modalities; benchmarks; business-user reach (≥10,000 EU users is the operational threshold); end users; autonomy/scalability; tools; state of the art relative to peers.
  • Business-user reach is the criterion most likely to pull a below-threshold model in. The 10,000-EU-business-user threshold counts direct and downstream-redistributor reach in aggregate. Open-weight models distributed via Hugging Face, cloud hyperscalers, and SaaS partners have hard-to-track aggregate footprints.
  • Llama 3.1-70B and DeepSeek V3 are both Annex XIII candidates. Llama 3.1-70B: aggregate designation risk ELEVATED. DeepSeek V3: aggregate designation risk HIGH. Both are non-signatory providers; procurement contracts should anticipate designation.
  • Six-step post-designation workflow. Upstream notification (Article 25(2)); Annex XI/XII refresh with Article 55 evidence; Annex IV technical-file refresh for downstream high-risk systems; substantial-modification analysis under Article 43(4); procurement contract amendments; brief to AI Governance Committee and audit committee.
  • Five contractual flow-down clauses pre-designation. Notification SLA (5 business days); Annex XII delivery SLA (30 days post-designation); Article 55 evidence-sharing commitment; indemnification for downstream impact; re-conformity-assessment cost allocation.
  • The Annex XIII designation-risk memo is the L3 artifact. Row per below-threshold model, ten-criterion sub-table, aggregate risk score, contract-amendment status, operational mitigations, refresh cadence. Owned by AI Governance Committee, refreshed quarterly, audit-committee semi-annually.
  • Six common mistakes. Assuming below-threshold means safe; missing the business-user-reach criterion; weak contractual flow-down for designation events; no quarterly designation-risk refresh; missing the autonomy/tool-access criteria; siloed from the GPAI exposure map.
  • The Commission's discretion is not formulaic. No published weight per criterion. Treat any model meeting three or more criteria at medium-to-high levels as at material designation risk. The first designations will signal Commission weighting; until then, contract conservatively.
  • The memo carries multi-framework weight. ISO 42001 A.10 third-party relationships, NIST AI RMF Map 4 + Govern 6.1, GPAI Code of Practice cross-walk, CycloneDX 1.7 ML-BoM input. Write the memo once, cite it against four frameworks.