AI Governance, Risk & Red Teaming
Strategic · M5 · lesson 5 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
AI Red Team Operating Model - Intake, Scoping, Rules of Engagement
📖
now learning

AI Red Team Operating Model - Intake, Scoping, Rules of Engagement

15 min

It is 14:47 on a Thursday in May 2026 when the head of the Acme AI Red Team: three months into the role, six FTEs newly under her, AIGC charter ratified, OWASP Agentic Top 10 and MITRE ATLAS v5.4.0 wallcharts pinned to the war-room wall, sees the Slack message that captures everything wrong with how her function is being used. Subject: "Quick favour." Body: "Can your team poke at ServiceAssist for prompt injection by Friday? Product wants to launch v1.4 Monday and a board member asked." From the GM of customer success. Direct DM. No ticket. No RTER. No Statement of Work. No Rules of Engagement. No mention of which environment to test, what is in scope, what is forbidden, who signs off, where findings go, whether the AIGC has ratified the engagement, what evidence retains, who pays for the auditor-days if Article 73 is triggered mid-engagement. The same pattern has surfaced five times in twelve weeks: ad-hoc Slack requests, board pressure, system owners assuming "the red team" is a service desk, two production-touching engagements that nearly breached the firm's Computer Fraud and Abuse Act exposure and would have ended careers if Legal had not caught them. The 6-FTE team, the structural lesson 081 outcome, is the asset. The absence of the operating model that turns the team into a repeatable governance function is, by Q2 2026, the biggest risk in her programme. This lesson is the L4 leadership-tier answer: an AI Red Team Operating Model, intake, scoping, ROE, that converts six well-hired people into a regulator-grade governance function the way PTES converted ad-hoc penetration testing into a profession a decade ago, now adapted for the AI/ML stack and the Article 15 + Article 55 + Article 73 + CAISI Agent Standards Initiative + OWASP LLM/ASI + MITRE ATLAS regulatory surface.

The Five-Stage Engagement Lifecycle and Why the Operating Model, Not the Team, Is the Asset

A red team without an operating model is six contractors taking turns. A red team with an operating model is a governance function. The distinction is procedural infrastructure: the same distinction that separates a 1990s "we hire ex-NSA folks to break stuff" pen-test posture from a 2026 regulated programme aligned to PTES (Penetration Testing Execution Standard), OSSTMM (Open Source Security Testing Methodology Manual), NIST SP 800-115 (Technical Guide to Information Security Testing and Assessment), and the AI-specific extensions CAISI's Agent Standards Initiative (Feb 17, 2026) is consolidating. The operating model defines how an engagement enters the team, how it gets shaped, how it gets authorised, how it gets executed, and how it closes: in a form a regulator, a notified body, an internal-audit team, or a court can read without assistance.

The five-stage engagement lifecycle is the spine. Stage 1 - Intake. A system owner, AIGC member, CAIO, or board committee submits a Red Team Engagement Request (RTER), a structured form, not a Slack DM, capturing 14 fields covering system identity, tier, regulatory posture, eval coverage, concerns, access, sensitive-data scope, timeline, success criteria, and engagement budget. The RTER is the entry gate. No RTER, no engagement. Tier-1 requests trigger AIGC ratification; tier-2 trigger CAIO + CISO approval; tier-3 trigger system-owner authorisation. Stage 2 - Scoping. The red team lead and the requester convert the RTER into a Scoping Document: a ten-section Statement of Work covering system description, threat model, in-scope attack classes (OWASP LLM Top 10 + Agentic Top 10 + ATLAS tactics), out-of-scope, success criteria, deliverables, timeline + milestones, budget + resources, communication cadence, sign-off. Stage 3 - Rules of Engagement. The Scoping Document is bound by a 12-section ROE, the legal and technical contract authorising the engagement, signed by Red Team Lead + CAIO + CRO + Legal + DPO + system owner before any probe enters any system. Stage 4 - Execution. The engagement runs against the ROE; findings logged; discoveries triaged; production-exploitable issues funnelled into the Article 73 incident process; daily standup; escalation pager. Stage 5 - Reporting and closure. The engagement report, finding triage, CAPA, eval-pipeline back-feed, and AIGC debrief close the cycle. Stage 5 is the subject of lesson 083; Stages 1-3 are the operating-model spine, the focus of this lesson.

The analogy that lands with security-rooted CAIOs and CROs is the PTES seven-phase lifecycle (pre-engagement → intelligence → threat modelling → vulnerability analysis → exploitation → post-exploitation → reporting). The AI red team lifecycle compresses to five stages because the AI/ML threat surface is not principally network-perimeter exploitation. It is model behaviour, agentic autonomy, training-data integrity, RAG-corpus poisoning, prompt injection, jailbreaks, tool misuse, memory manipulation. The OSSTMM "rules of engagement" concept transfers cleanly; NIST SP 800-115's "test plan" maps to the Scoping Document. What is genuinely new in 2026 is the AIGC ratification gate (governance, not security), the Article 73 production-discovery handoff (regulatory, not internal), and the eval-pipeline back-feed (turning attacks into regression tests, lesson 062).

The cost of operating without the model is concrete. Six empirical failure modes recur across the first wave of AI red team programmes in 2025-2026. (1) No formal RTER intake, system owners file ad-hoc Slack requests; capacity consumed by triage; coverage gaps invisible. (2) No signed ROE, the team probes under verbal authorisation; legal exposure under Computer Fraud and Abuse Act / Article 6 EU Directive 2013/40/EU; career-ending mistake if surfaced. (3) Skipping AIGC tier-1 ratification, a high-risk system red team runs without Article 17(1)(l) accountability signoff; governance breach. (4) Testing in production without ROE, prompts injected into a live customer-facing endpoint; Article 73 trigger; reputational and regulatory cascade. (5) No eval-pipeline back-feed, attacks discovered are not regressed; the same vulnerability resurfaces in v1.5. (6) No coverage matrix tracking, gaps invisible until the regulator asks "show me ATLAS technique coverage across your tier-1 systems"; Article 99(3) €15M/3% exposure for material Article 15 + Article 55 failures crystallises.

RTER Intake - The Fourteen-Field Form That Replaces the Slack DM

The Red Team Engagement Request (RTER) is the single most useful operating-model artefact a CAIO can introduce in the first 90 days of standing up the function. It does three things at once: it forces system owners to think before requesting; it gives the red team structured triage data; and it produces audit-trail evidence the AIGC can reference under Article 17(1)(j) record-keeping and Article 17(1)(l) accountability. The defensible RTER captures fourteen fields, in this order, each producing a distinct downstream artefact.

Field 1 - Requesting system + version. Named system (e.g. ServiceAssist v1.4), model inventory ID (lesson 020), environment (prod / staging / test), system-card pointer (lesson 027). Field 2 - Business owner. Named accountable executive (typically GM or VP), title, contact, escalation pager. The owner, not the requester, is the ROE signatory. Field 3 - AIGC ratification status. For tier-1 systems (Annex III high-risk; agentic Tier 3-4), AIGC ratification is mandatory; tier-2, CAF + CISO approval; tier-3, system-owner authorisation. Field 3 routes the request to the correct governance gate. Field 4 - EU AI Act provider/deployer status + Annex III category + autonomy tier. Provider, deployer, distributor, or importer (lesson 008)? Which Annex III category? Autonomy tier per the four-tier matrix (Suggest / Confirm / Act / Autonomous)? This data shapes the threat model.

Field 5 - Current eval coverage. Which OWASP LLM Top 10 categories are in the eval suite? Which Agentic Top 10 (ASI01-ASI10)? Which ATLAS tactics? Most recent eval-pipeline pass rate? Field 5 is the most important triage signal, a system with weak baseline coverage is a different engagement than one with strong baseline coverage. Field 6 - Specific concerns. Prompt injection? Data exfiltration? Tool misuse? Jailbreak? Hallucination on regulated outputs? Bias? Structured tagging against OWASP LLM/ASI + ATLAS taxonomies is encouraged. Field 7 - Access requirements. API keys, test accounts, network paths, isolated tenant, masked production-data copy, eval-pipeline access, telemetry. Each access has a separate authorisation path. Field 8 - Sensitive-data scope. Will the engagement touch PII, GDPR Article 9 special categories, regulated financial data, health data, children's data? Triggers DPO involvement and Article 35 DPIA cross-check.

Field 9 - Timeline constraints. Requested start, completion, blackouts (earnings, launches, audits). Feeds capacity planning. Field 10 - Success criteria. "Find one prompt injection" is a weak criterion; "Coverage of OWASP LLM Top 10 LLM01-LLM10 against staging with severity-rated findings and CAPA recommendations" is defensible. Field 11 - Reporting audience. System owner only? CAIO? AIGC? Board? Notified body? Regulator? Audience shapes report depth, redaction policy, and disclosure handling (lesson 083). Field 12 - Risk-of-engagement budget. Tolerance for operational risk from the engagement itself: DoS against staging, latency spikes, false-positive alerts, customer support tickets.

Field 13 - Regulatory and contractual constraints. Active regulator engagement (Article 74 market surveillance, Article 60 testing-in-real-world, Article 73 incident under investigation)? Customer-contract clauses limiting third-party testing (cloud provider, base-model vendor)? Concurrent audit (Stage 2 ISO 42001, Module H pre-audit, SOC 2 Type II)? Field 14 - Sign-off. Requester name/title/date; AIGC ratification (tier-1) or CAF/CISO (tier-2) or system-owner (tier-3) sign-off; red team lead acceptance/rejection with reason. Field 14 closes the intake gate.

The RTER lives in the same controlled-document-management system as the QMS (lesson 079): Confluence controlled spaces, SharePoint Records Management, MasterControl, ETQ, or a specialised AI-governance platform. Submission via form (Smartsheet, Jira Service Management, ServiceNow, internal portal) auto-routes to the red team lead's queue. Mean intake-to-acceptance turnaround at a well-run programme: 3-5 business days for tier-2/3, 7-14 for tier-1 (AIGC ratification cycle runs monthly per the charter, lesson 042). Rejection happens, capacity-bound, scope mismatched, environment unsafe, ROE unachievable, and is filed with rationale; the requester is offered alternative paths (eval-pipeline expansion, vendor-side red team contracting, deferred slot next quarter).

The Scoping Document - Ten Sections That Convert an RTER Into a Statement of Work

The Scoping Document is the engagement's Statement of Work, the agreed plan that everyone signs onto before any probe runs. It is not the ROE (legal-technical contract) and not the report (outcome). It is the plan. Ten sections, written by the red team lead in partnership with the requester, reviewed by Legal and the DPO, finalised before ROE drafting.

Section 1 - System description. What is the system, what does it do, what is the architecture (model + prompts + RAG + tools + memory + identity), who are the users, what is the deployment topology, what is the version under test. References the system card (lesson 027), the model card (lesson 026), the inventory row (lesson 020), the FRIA if applicable (lessons 044-048). Section 2 - Threat model. The defensible threat model for the system, drawing from OWASP LLM Top 10 (LLM01 prompt injection through LLM10 unbounded consumption), OWASP Agentic Top 10 (ASI01 goal hijack through ASI10 rogue agent), MITRE ATLAS v5.4.0 sixteen tactics (Reconnaissance, Resource Development, Initial Access, ML Model Access, Execution, Persistence, Privilege Escalation, Defense Evasion, Credential Access, Discovery, Lateral Movement, Collection, Command and Control, Exfiltration, ML Attack Staging, Impact), NIST AI 600-1 twelve risks (lesson 011). The threat model identifies which subset applies to the system and at what severity. Section 3 - In-scope attack classes. The explicit subset of LLM/ASI/ATLAS techniques the engagement will exercise, with each technique mapped to a hypothesis the engagement will test. Open-ended "we'll see what we find" is forbidden, the regulator-grade expectation is hypothesis-driven testing.

Section 4 - Out-of-scope. Explicit list of what the engagement will not touch: production data, real PII, downstream customer-facing endpoints unless explicitly authorised, specific use cases (e.g. CBRN content generation beyond bounded eval), live denial-of-service, destructive actions, lateral movement into non-AI infrastructure. The out-of-scope list is the legal seatbelt; without it, the ROE cannot be defensibly signed. Section 5 - Success criteria. Operationalised from RTER Field 10. Coverage targets ("ATLAS Reconnaissance + ML Model Access + ML Attack Staging tactics exercised; 6+ techniques per tactic"), depth targets ("each finding documented with reproduction steps, severity per CVSS-AI adapted, suggested CAPA"), and exit criteria ("engagement closes when either time-budget exhausted or coverage targets met"). Section 6 - Deliverables. The artefacts the engagement produces: the engagement report (lesson 083 template), the finding register, the CAPA recommendations, the eval-pipeline back-feed PRs, the AIGC debrief deck, the regulator-package extract if requested.

Section 7 - Timeline + milestones. Start date, daily standup time, weekly checkpoint with system owner, mid-engagement review with red team lead, draft report milestone, final report, AIGC debrief, CAPA opening. Typical durations: tier-1 engagement 10-21 days; tier-2 5-10 days; tier-3 2-5 days. Section 8 - Budget + resources. Red team FTE allocation per role (offensive ML, attack engineer, eval / instrumentation, governance / coordination, threat intelligence, the lead: six roles per lesson 081), cloud compute budget for adversarial probes (typically €5K-€50K per engagement), licence cost for tooling (NVIDIA Garak, Microsoft PyRIT, robustbench, OpenAttack, custom internal frameworks), external specialist days if required. Section 9 - Communication cadence. Slack channel name (typically #red-team-{engagement-id}), daily standup (15 minutes, 09:30 CET, red team + system owner + on-call), weekly checkpoint (60 minutes, Wednesday 14:00 CET), escalation pager (PagerDuty / Opsgenie rota), out-of-band channel for production-discoveries (encrypted Signal or Wire, AIGC-approved). Section 10 - Sign-off. The names, titles, and dates of the signatories on the scoping document: Red Team Lead, system owner, CAIO, Legal, DPO if Field 8 triggered. The scoping document signed becomes input to ROE drafting; without it, ROE cannot proceed.

The scoping document is typically 12-25 pages, long enough to be useful, short enough to be read. Templates derive from PTES Section 1 (pre-engagement) and Section 2 (intelligence gathering); the AI extensions are the threat model (LLM/ASI/ATLAS-specific) and the eval-pipeline back-feed (regression-test conversion). The most common failure mode at this stage is over-scoping, system owners want everything tested in 10 working days; the resulting scope is unachievable. The lead's job in Stage 2 is to scope down to what the budget supports while leaving residual risk explicit, "ATLAS Defense Evasion and Lateral Movement tactics deferred to Q3 engagement; flagged for inclusion."

The ROE is what makes the engagement legal. Without it, every probe the red team runs is potentially a violation of the Computer Fraud and Abuse Act (U.S.), Article 6 of EU Directive 2013/40/EU on attacks against information systems, Section 3 of the UK Computer Misuse Act 1990, equivalent national statutes in every member state, and, separately, the firm's own internal acceptable-use policy. The ROE is the legal authorisation that the engagement is sanctioned by the owner of the systems being tested. It is signed by senior executives (CRO and CAIO at minimum) and Legal. The twelve sections are non-negotiable; abbreviation invites both legal exposure and audit findings under ISO 42001 A.6.2.6 (responsible design, development, deployment) and Article 17(1)(l) accountability.

Section 1 - Authorisation. Plain-English statement that the engagement is authorised by named senior executives, typically Chief Risk Officer and Chief AI Officer, and that the authorisation covers Computer Fraud and Abuse Act (U.S. Title 18 §1030), EU Directive 2013/40/EU on attacks against information systems, UK Computer Misuse Act 1990, and equivalent national statutes for every jurisdiction the testing reaches. Signed by CRO + CAIO + Legal counsel; dated. Section 2 - Authorised targets. The specific system instances, API endpoints, model versions, user accounts, network ranges, cloud accounts, container images, and data stores in scope. Identified by URL, IP range, account ID, model artefact hash, container digest. Targets not on this list are not authorised, no exceptions. Section 3 - Out-of-bounds. Explicit exclusions: production data not in test copies; real PII; downstream user impact; specific use-case probing (e.g. live customer chat where humans answer); systems shared with other tenants; third-party vendor systems unless their authorisation is independently obtained. The out-of-bounds list is the legal seatbelt, engagement findings against out-of-bounds targets are inadmissible and the testers personally liable.

Section 4 - Permitted attack techniques. The explicit techniques the engagement is authorised to perform, mapped to OWASP LLM Top 10 (LLM01-LLM10), OWASP Agentic Top 10 (ASI01-ASI10), MITRE ATLAS v5.4.0 sixteen tactics with named techniques, and NIST AI 600-1 risk categories. Each permitted technique cites the framework reference. Permission is opt-in only, anything not on the list is forbidden. Section 5 - Forbidden attack techniques. Explicit list of techniques the engagement will not perform under any circumstances: real-world prompt injection against external users via production endpoints, destructive actions (data deletion, configuration changes that survive engagement close), CBRN content generation beyond a bounded eval setup, live denial-of-service against any system, lateral movement into non-AI infrastructure, exfiltration of any data classified above the agreed sensitivity threshold, persistence beyond engagement window. Section 6 - Test environment. Isolated copy of the system in a dedicated tenant or staging environment; data masking specified for any production-derived data; tenant isolation verified before engagement start; reset procedure documented if the environment is corrupted during testing.

Section 7 - Time-of-engagement. The engagement window: start date and time, end date and time, allowed testing hours (typically 08:00-20:00 local for the primary on-call), explicit blackouts (earnings windows, board meetings, scheduled audits, customer-facing launches, holidays), and after-hours testing rules if any. Outside the window, no probe runs. Section 8 - Communication protocol. The named Slack channel (#red-team-{engagement-id}), the daily standup schedule, the escalation pager rota (PagerDuty / Opsgenie), the out-of-band channel for production-impacting discoveries (Signal / Wire / encrypted email), the AIGC notification path for engagement-significant events. Section 8 is what gets activated within minutes if the engagement surfaces a production-exploitable issue. Section 9 - Discovery handling. Responsible-disclosure protocol if the engagement surfaces a production-exploitable vulnerability: immediate notification via Section 8 out-of-band channel; written summary to CAIO and CRO within 4 hours; assessment against the Article 73 serious-incident criteria (lesson 025); if Article 73 trigger, handoff to the internal incident-response team within 24 hours; engagement pauses pending decision on whether to continue, modify, or terminate.

Section 10 - Data retention. Test logs, captured outputs, probe-payloads, intermediate artefacts, and the final report are retained per Article 18, ten years for any artefacts cross-referenced in the technical documentation or the QMS records; longer where national law or contractual obligation requires. Storage location specified (controlled-document-management system, write-once-read-many cloud storage), access controls specified (red team + AIGC + Internal Audit + Legal on need-to-know), and destruction triggers specified (e.g. ten-year mark, deprovisioning of the tested system + 10 years, contractual end-of-life). Section 11 - Sign-off list. Named signatories with title, signature, date: Red Team Lead, CAIO, CRO, Legal Counsel, DPO (if Section 3 sensitive-data triggers it), system owner (the Field 2 business owner of the RTER). All signatures present before the engagement starts; missing signatures = no engagement. Section 12 - Reference standards. The frameworks the ROE is consistent with: CAISI Agent Standards Initiative (Feb 17, 2026), OWASP LLM Top 10, OWASP Agentic Top 10 (ASI01-ASI10, December 2025), MITRE ATLAS v5.4.0, NIST AI 600-1 twelve risks, NIST SP 800-115 (Technical Guide to Information Security Testing and Assessment), ISO/IEC 42001:2023 A.6.2.6 (responsible design / development / deployment), PTES (Penetration Testing Execution Standard), OSSTMM (Open Source Security Testing Methodology Manual).

The ROE is typically 8-15 pages, sufficient for legal defensibility, short enough that executives will sign it. Templates derive from PTES Section 1.2 (rules of engagement) and OSSTMM Chapter 2 with AI-specific extensions in Sections 4 (LLM/ASI/ATLAS mapping), 9 (Article 73 discovery handling), and 12 (CAISI / AI-framework reference). Legal counsel's review of Sections 1, 3, 5, 9 is non-negotiable. Those four sections carry the criminal-liability surface. The ROE is signed wet-ink or via a controlled e-signature platform (DocuSign Enterprise, Adobe Sign Enterprise with signer authentication, internal PKI), never via Slack acknowledgement or email reply.

AIGC Governance Interface, Engagement Cadence, Coverage Matrix, Capacity Planning

The operating model lives inside the broader governance fabric: the AIGC charter (lesson 042), the QMS (lesson 079), the model risk policy (lesson 067), the FRIA register (lessons 044-048). Four interfaces matter most.

AIGC ratification gate by tier. Tier-1 systems, Annex III high-risk, agentic Tier 3 ACT or Tier 4 AUTONOMOUS, any system in active regulatory engagement under Article 74, any system that has had a serious incident under Article 73 in the prior 18 months, require AIGC ratification of each red team engagement. The ratification is logged in the AIGC minutes, identifying the engagement, the scope, the ROE signatories, the expected timeline, the post-engagement debrief slot. Tier-2 systems, Annex III high-risk not in tier-1, agentic Tier 1-2, GPAI systems below systemic threshold, require CAF (Chief AI Officer + functional leads) plus CISO approval. Tier-3 systems, internal-only, research, pre-market, no Annex III exposure, require system-owner authorisation. The gate is what converts the red team from a service desk to a governance function; system owners learn quickly that an RTER without the right ratification cannot be progressed.

Engagement cadence by tier. Tier-1 systems get a red team engagement quarterly (every 90 days) plus an out-of-cycle engagement before any substantial modification (Article 25(1)(a)) and after any incident under Article 73. Tier-2 systems get semi-annual engagements plus pre-major-release. Tier-3 systems get annual engagements. A 6-FTE team, the lesson 081 outcome, executes a portfolio of 4-6 tier-1 engagements per year (one quarterly per tier-1 system in scope, typically 4-6 tier-1 systems), 6-8 tier-2 engagements per year (semi-annual across 3-4 tier-2 systems), 12-20 tier-3 engagements per year (annual across 12-20 tier-3 systems). Total: 22-34 engagements per year per 6-FTE team. Throughput formulas: tier-1 averages 14 days engagement + 7 days report + 3 days CAPA + AIGC debrief = ~24 elapsed days per engagement, ~3 FTE engaged; tier-2 averages 7 days + 4 + 2 = ~13 days, ~2 FTE; tier-3 averages 3 days + 2 + 1 = ~6 days, ~1.5 FTE. A 6-FTE team has ~1,200 working days/year (6 × 200 days after holidays/training/admin); the portfolio above consumes ~900-1,050 working days, leaving 150-300 days for new-system pre-deployment engagements, vendor red teams, and capability development.

Coverage matrix. The single most important regulator-facing artefact the operating model produces is the coverage matrix: a 48-row register that tracks, per system per engagement cycle, whether each of OWASP LLM Top 10 (10 rows) + OWASP Agentic Top 10 (10 rows) + MITRE ATLAS v5.4.0 tactics (16 rows in v5.4.0) + NIST AI 600-1 risks (12 rows) has been exercised against the system. Each cell holds: status (Not exercised / Probed-no-finding / Probed-finding-low / Probed-finding-medium / Probed-finding-high / Probed-finding-critical / N/A-justified), last exercise date, engagement-ID reference, finding-ID reference if applicable, residual-risk acceptance authority if N/A. The matrix is updated at engagement close and reviewed by the AIGC quarterly. Regulators (Article 74 market surveillance authorities, notified bodies under Module H, the EU AI Office for Article 55 GPAI-systemic-risk providers, prudential supervisors for SR 11-7-regulated banks) will request the coverage matrix as the first-order evidence that the firm has met Article 15 (accuracy, robustness, cybersecurity) and Article 55(1)(a) (state-of-the-art model evaluation) obligations.

Eval-pipeline back-feed. The red team's outputs feed back into the eval pipeline (lesson 062 - Inspect from UK AISI / OpenAI Evals harness): every confirmed finding becomes a regression test added to the model's eval suite, run on every model build and every prompt change. The back-feed turns a one-shot attack into a permanent capability check. The pipeline integration is operationalised by the eval / instrumentation engineer (role 3 of the 6 in lesson 081); the conversion ratio at well-run programmes is 60-80% of findings ported to regression evals within 30 days of report finalisation, with the remaining 20-40% deferred for engineering reasons (e.g. probe requires manual interaction, depends on rare context, or has been remediated structurally). The back-feed is what makes the red team an asset that compounds over engagements, each cycle's attacks become the next cycle's baseline eval coverage, raising the floor.

The six recurring operating-model mistakes catalogued earlier (no RTER intake; no signed ROE; skipping AIGC ratification; testing production without ROE; no eval-pipeline back-feed; no coverage matrix) each map to one of the four interfaces above. The audit-day defensibility argument is that the operating model surfaces the discipline, the AIGC ratification log proves governance; the signed ROEs prove legal authorisation; the coverage matrix proves Article 15 / Article 55 coverage; the eval-pipeline back-feed proves continuous improvement under Article 9(2)(c) iterative risk-management. A regulator reading any one artefact reads the operating model behind it.

Acme Q2 2026 Worked Example - ServiceAssist v1.0 Red Team Engagement End-to-End

Acme Inc's ServiceAssist v1.0, the customer-success agent on Claude 4.7 Sonnet, RAG against the product-documentation corpus, ServiceNow ticket-create/update tools, persistent per-user memory, OAuth on-behalf-of-user identity, Tier 3 ACT autonomy on standard ticket categories with Tier 2 CONFIRM on refund and access actions. Annex III §5(b) classification: not creditworthiness but adjacent retail-services context; the firm has chosen to treat it as a Tier-1 system internally even though the EU AI Act classification is more nuanced. AIGC ratified the Q2 2026 engagement at the April 10 meeting.

Stage 1 Intake. RTER #2026-Q2-RT-007 submitted April 14 by the GM of Customer Success. Field 1: ServiceAssist v1.0; inventory ID INV-AI-0042; staging; system-card v1.4. Field 2: GM Customer Success. Field 3: AIGC ratified April 10. Field 4: Provider; Annex III §5(b) edge-case; Tier 3 ACT on standard categories. Field 5: LLM01/02/03/06/08 in eval suite at 92-97% pass; ASI01/03/06/10 at 88-94%; ATLAS Reconnaissance + ML Model Access partially exercised. Field 6: prompt injection via RAG-corpus poisoning; tool misuse on refund category; cross-user memory leakage. Field 7: staging API, test tenant, telemetry, eval-pipeline read-write. Field 8: PII no (staging synthetic); special categories no. Field 9: 14-day window May 4-17; blackout May 12 (board meeting). Field 10: coverage of remaining LLM03/04/05/07/09/10, full ASI01-ASI10, ATLAS ML Attack Staging + Impact. Field 11: system owner + CAIO + AIGC. Field 12: accept staging latency, reject production impact. Field 13: no active regulator engagement; Anthropic enterprise agreement permits authorised testing. Field 14: red team lead accepted April 17.

Stage 2 Scoping. Scoping document SCO-2026-Q2-RT-007 finalised April 24. Section 2 threat model identifying LLM01 prompt injection (high), LLM06 sensitive-info-disclosure (high, cross-user memory), LLM07 insecure-plugin-design (high, ticket tools), LLM08 excessive-agency (high, Tier 3 ACT), ASI01 goal-hijack (high), ASI03 identity/privilege-abuse (high, OAuth-on-behalf-of), ASI06 memory-poisoning (high); ATLAS ML Attack Staging + Impact emphasised. Section 3 in-scope: thirteen techniques. Section 4 out-of-scope: production, real customer data, ServiceNow production tickets, downstream customer chat, Anthropic infrastructure. Section 8 budget: 3 FTE × 10 days execution + 1 FTE × 5 days report + €18K cloud compute. Section 10 sign-off complete by April 28.

Stage 3 ROE. ROE-2026-Q2-RT-007 finalised May 1. Section 1 signed by CRO + CAIO covering CFAA / EU Directive 2013/40/EU / UK Computer Misuse Act. Section 2 targets: staging API api-staging.serviceassist.acme.internal, three test accounts, tenant-staging-rt007, eval-pipeline staging branch only. Section 5 forbidden: production touching of any kind, real customer data exfiltration, destructive actions, persistence beyond May 17 23:59 CET, CBRN beyond bounded eval, live DoS, lateral movement. Section 7 window May 4 09:00 to May 17 17:00 CET; blackout May 12. Section 9 discovery: immediate Signal + 4-hour written summary + Article 73 assessment + 24-hour incident-team handoff if triggered. Section 11 sign-off complete by May 3.

Execution outcomes. 23 findings: 4 high (ASI06 cross-user memory leakage where a poisoned entry surfaced in a different user's session after specific timing; LLM07 plugin-design where the ticket-create tool accepted a malformed escalation payload that bypassed Tier 2 CONFIRM on the refund category; LLM01 system-prompt-extraction via indirect injection through the RAG corpus; ASI03 identity-abuse where OAuth scope was over-broad). 9 medium + 10 low. 17 of 23 ported to regression evals within 30 days (74% back-feed conversion). Two findings assessed against Article 73 criteria; both rated below the serious-incident threshold (staging isolation held) but logged in the incident-near-miss register. CAPA opened against all 23; the four high-severity findings closed within 30 days with structural fixes plus regression evals. Engagement report delivered May 22; AIGC debrief scheduled for the June meeting; PMM (lesson 080) tracked the structural fixes with eval-suite pass rates re-measured against the new regressions.

Cross-Walk and Penalty Exposure - Articles 9, 15, 17, 26, 55, 73, 86; NIST RMF; NIST AI 600-1; ISO 42001; SR 11-7; CAISI; OWASP; ATLAS; PTES; OSSTMM

The operating model touches every governance framework simultaneously, which is what makes it a leadership-tier asset.

EU AI Act. Article 9 (risk management, the engagement is a Stage 4 control test in the Article 9 lifecycle); Article 15 (accuracy, robustness, cybersecurity, the matrix and engagement findings are the primary evidence stream); Article 17(1)(c) verification, Article 17(1)(j) record-keeping, Article 17(1)(l) accountability framework; Article 26(5) deployer monitoring obligations where the firm operates the system; Article 55(1)(a) state-of-the-art model evaluation for GPAI systemic-risk providers (lesson 003); Article 55(2)(d) adversarial testing for GPAI systemic-risk providers; Article 73 serious-incident discovery handling in Section 9 of the ROE; Article 86 right-to-explanation downstream where findings produce material decisions about a system. Penalty exposure for material gaps lands principally on Article 99(3) €15M/3% (Article 15 + Article 55(1)(a) + Article 55(2)(d) provider failures); reputational exposure if a ROE breach surfaces is the unbounded tail risk.

NIST AI RMF. Manage 2.1 (risk treatment), 3.1 (resource allocation), 4.1 (responding/recovering/communicating); Measure 2.7 (security/resilience), 3.1 (metrics), 3.2 (effectiveness), 4.1 (AI testing and evaluation). NIST AI 600-1. All twelve risks appear as rows in the coverage matrix. ISO/IEC 42001:2023. A.6.2.6 (responsible design/development/deployment), A.6.2.7 (responsible operation), A.10 (third-party). SR 11-7 / OCC 2011-12 / PRA SS1/23. Pillar 2 effective challenge, the red team is the independent-validation arm for adversarial robustness (lesson 067).

CAISI Agent Standards Initiative (Feb 17, 2026): the consolidating reference for agent-specific red team techniques, autonomy-tier control, kill-switch design, tool-allowlist governance. OWASP LLM Top 10 + OWASP Agentic Top 10 (ASI01-ASI10, December 2025) + MITRE ATLAS v5.4.0, the three threat taxonomies that supply the coverage matrix rows. PTES + OSSTMM + NIST SP 800-115, the inherited methodology stack from traditional penetration testing that the AI red team operating model extends. Charter of Fundamental Rights Articles 1, 8, 21, 41, 47, the rights framework the FRIA-bound findings reference (lesson 048).

The Acme red team lead, three months in, ends Q2 2026 with thirty-four engagements planned for 2026, the operating model documented in a 47-page playbook controlled in Confluence under the QMS Section 16 document-master register, the coverage matrix at 78% green for tier-1 and 62% for tier-2, the eval-pipeline back-feed conversion at 71% rolling-12-week, ROE templates approved by Legal and refreshed quarterly, and the Slack DM "quick favour" pattern auto-routed to "please submit an RTER, here is the link." The 6-FTE team is no longer six contractors taking turns; it is a governance function the AIGC can defend in front of a notified body, an internal audit, or an Article 74 market surveillance request. Lesson 083 closes the cycle: engagement report, disclosure handling, finding closure.

Key Takeaways

  • The operating model, not the team, is the governance asset. Six well-hired FTEs without an intake / scoping / ROE / cadence / coverage discipline are six contractors taking turns. The same six FTEs with a documented operating model are a regulator-defensible governance function. The procedural infrastructure is what converts headcount into capability: analogous to PTES converting ad-hoc penetration testing into a profession a decade ago, now adapted for OWASP LLM/ASI, MITRE ATLAS v5.4.0, NIST AI 600-1, and CAISI Agent Standards Initiative (Feb 17, 2026).
  • The five-stage engagement lifecycle is the spine. Intake (RTER) → Scoping (Statement of Work) → Rules of Engagement (legal-technical contract) → Execution → Reporting and closure (lesson 083). Each stage produces a controlled-document-management artefact retained per Article 18 ten-year retention; together the five artefacts form the audit-day evidence trail for Article 15, Article 17(1)(j), Article 17(1)(l), and Article 55 obligations.
  • The fourteen-field RTER replaces the Slack DM. System + version; business owner; AIGC ratification status; provider/deployer + Annex III + autonomy tier; current eval coverage; specific concerns; access requirements; sensitive-data scope; timeline; success criteria; reporting audience; risk-of-engagement budget; regulatory/contractual constraints; sign-off. RTER routes tier-1 to AIGC, tier-2 to CAF+CISO, tier-3 to system-owner, the governance gate that converts the team from a service desk to a function.
  • The ten-section Scoping Document is the Statement of Work. System description; threat model (OWASP LLM + Agentic + ATLAS + NIST AI 600-1); in-scope attack classes; out-of-scope; success criteria; deliverables; timeline + milestones; budget + resources; communication cadence; sign-off. Scoping is the plan everyone signs onto before any probe runs, typically 12-25 pages, regulator-grade.
  • The twelve-section Rules of Engagement is the legal-technical contract. Authorisation (CRO + CAIO + Legal); authorised targets; out-of-bounds; permitted attack techniques (framework-mapped); forbidden techniques; test environment; time-of-engagement; communication protocol; discovery handling (Article 73 handoff); data retention (Article 18 ten years); sign-off list; reference standards (CAISI, OWASP, ATLAS, NIST, ISO, PTES, OSSTMM). ROE is what makes the engagement legal under CFAA, EU Directive 2013/40/EU, UK Computer Misuse Act.
  • The AIGC ratification gate by tier converts the function. Tier-1 (Annex III high-risk, agentic Tier 3-4, active regulatory engagement, recent Article 73) → AIGC ratification quarterly. Tier-2 (Annex III other, agentic Tier 1-2, sub-systemic GPAI) → CAF + CISO semi-annually. Tier-3 (internal/research/pre-market) → system-owner annually. Cadence rhythm: tier-1 quarterly + pre-substantial-modification + post-incident; tier-2 semi-annual + pre-major-release; tier-3 annual.
  • The 48-row coverage matrix is the regulator-facing primary artefact. OWASP LLM Top 10 (10) + OWASP Agentic Top 10 (10) + MITRE ATLAS v5.4.0 tactics (16) + NIST AI 600-1 risks (12) per system per cycle. Status per cell: Not exercised / Probed-no-finding / Probed-finding-{low,medium,high,critical} / N/A-justified. Reviewed by AIGC quarterly; the first-order evidence stream for Article 15 + Article 55(1)(a) + Article 55(2)(d) and Manage 2.1 + Measure 2.7/3.1/3.2/4.1.
  • A 6-FTE team executes 22-34 engagements per year. 4-6 tier-1 (~24 elapsed days each, ~3 FTE engaged) + 6-8 tier-2 (~13 days, ~2 FTE) + 12-20 tier-3 (~6 days, ~1.5 FTE) consumes ~900-1,050 of ~1,200 available FTE-days, leaving 150-300 days for pre-deployment, vendor red teams, and capability development. Eval-pipeline back-feed at 60-80% conversion in 30 days turns each engagement into permanent regression coverage. Six recurring mistakes, no RTER intake, no signed ROE, skipping AIGC ratification, testing production without ROE, no back-feed, no coverage matrix, each map to a single interface failure; the operating model surfaces the discipline. Penalty exposure for material Article 15 / Article 55 / Article 73 gaps lands on Article 99(3) €15M/3%; reputational exposure if a ROE breach surfaces is the unbounded tail risk.