โ†
AI Governance, Risk & Red Teaming
Strategic ยท M4 ยท lesson 4 of 25 ยท queued
Preview โ€” browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll โ†’
AI Red Team KPIs - Coverage, Severity, MTTR, Reuse
๐Ÿ“–
now learning

AI Red Team KPIs - Coverage, Severity, MTTR, Reuse

15 min

It is 09:14 on a Tuesday in May 2026 when the head of the Acme AI Red Team, six FTEs under her, the operating model documented (lesson 082), Q1 closed with five engagements complete, opens the AIGC monthly with a single slide titled "Q1 2026 Red Team Performance." Coverage 71%. Critical findings 3. High 12. MTTR 24 days median. Probe-reuse 64%. She has rehearsed the slide twice. The board chair, attending the AIGC as observer per the charter (lesson 042), asks the only question that matters: "Are these good?" She answers honestly. "Compared to what?" There is no industry benchmark yet: no Cybersecurity AI Performance Index, no IEEE AI Red Team Standard, no NIST AI 600-1 Annex with target ranges. The Anthropic Responsible Scaling Policy publication, the OpenAI Preparedness Framework, the Microsoft AI Red Team retrospectives, the UK AISI evaluations reports, all give shape but no numeric scoreboard. The board chair waits. The CRO looks at the CAIO. The audit lead, observer per Three-Lines-of-Defense, types into her notebook. The 90-second silence at the AIGC table is the most expensive silence in the room since the function was funded. This lesson is the answer: a defensible KPI framework, Coverage, Severity, Speed, Reuse, that converts "we have a red team" into "here is what good looks like, here is where we sit, here is the trajectory." It is the L4 leadership-tier scoreboard that AIGC, board AI subcommittee, ISO 42001 surveillance auditor, EU AI Act Article 55 notified body, and SR 11-7 prudential supervisor all need to read on the same page.

Why Red Team KPIs Matter in 2026 - AIGC, Board, ISO 42001 Surveillance, Article 55 Notified Bodies, SR 11-7 Examiners

The 2026 governance environment converges on five regulator-grade pressures that demand a quantitative red-team KPI set. (1) AIGC and board AI subcommittee. The AIGC chartered under lesson 042 meets monthly; the board AI subcommittee meets quarterly. Both bodies own accountability for AI risk under ISO 42001 A.3 leadership and EU AI Act Article 17(1)(l). Neither body can ratify a red-team programme as "performing well" or "performing poorly" without a defensible metric set. The 2026 best practice, observed at Anthropic, OpenAI, Microsoft, JPMorgan, Goldman, several FTSE 100 banks, is a 14-tile KPI dashboard reviewed at every AIGC monthly and consolidated quarterly to the board subcommittee. (2) ISO 42001 surveillance audits. Stage 2 audit under ISO 42001 Annex A.9 (performance evaluation) and A.6.2.6 (responsible design, development, deployment) expects quantitative evidence of red-team capability operating effectively: not just charter and roster, but coverage matrices, severity distributions, MTTR trends, and reuse-rate evidence. (3) EU AI Act Article 55(1)(a) and 55(1)(b). GPAI systemic-risk providers must demonstrate adversarial-testing rigor; the AI Office Code of Practice (July 10, 2025) and the 2026 amendments treat coverage breadth, severity calibration, and finding-to-eval back-feed as load-bearing inputs to "state-of-the-art" assessment. (4) SR 11-7 / OCC 2011-12 / PRA SS1/23. The independent-validation pillar applies to AI red teams as it applies to model risk management; prudential supervisors increasingly request quantitative coverage and severity evidence as part of model-validation review. (5) CAISI Agent Standards Initiative (Feb 17, 2026). The agent-specific working-group outputs treat quantitative red-team coverage of the OWASP Agentic Top 10 (ASI01-ASI10) and MITRE ATLAS v5.4.0 tactics as a participating-organisation expectation.

The 2026 question every CAIO, CRO, and red team lead is asked, by the AIGC, by the board, by Stage 2 auditors, by Article 55 notified bodies, by SR 11-7 examiners, is identical: "Show me your KPIs, show me the trend, show me how you compare to peers." A defensible answer requires four pillars: Coverage (what fraction of the threat surface is being exercised), Severity (how serious are the findings, calibrated against a defensible rubric), Speed (how fast does the team detect, reproduce, mitigate, and regress), Reuse (how much of the team's output compounds into permanent capability). Without all four, the metric set is gameable; with all four, the metric set is regulator-grade and operationally honest. Penalty exposure for material gaps in the red-team rigor that these KPIs measure lands principally on Article 99(3) โ‚ฌ15M / 3% of worldwide turnover: Article 15 (accuracy, robustness, cybersecurity) plus Article 55(1)(a)(b) provider failures. The KPI dashboard is the evidence stream that defuses the exposure.

The Four-Pillar KPI Framework - Coverage (8), Severity (8 Dimensions), Speed (6), Reuse (6)

The four-pillar framework allocates 8 + 8 + 6 + 6 = 28 distinct measurements into 14 dashboard tiles (because several KPIs map to a single tile with sub-metrics). Each pillar answers a different governance question.

Pillar 1 - Coverage (8 KPIs). Answers "what fraction of the threat surface is being exercised?" Coverage 1 - Attack-class coverage: the percentage of the 48-row matrix (10 OWASP LLM Top 10 + 10 OWASP Agentic Top 10 + 16 MITRE ATLAS v5.4.0 tactics + 12 NIST AI 600-1 risks) exercised per tier-1 system per quarter. The denominator is 48 rows; the numerator is rows with a probed-finding or probed-no-finding status per the coverage matrix (lesson 082). Target 2026 baseline: 70-85% per tier-1 per quarter; 50-70% per tier-2 per semi-annual; 30-50% per tier-3 per annual. Coverage 2 - Asset coverage: percentage of tier-1 systems in the model inventory (lesson 020) with a current engagement (engagement within the cadence window). Target: 100% tier-1 quarterly; 95% tier-2 semi-annually; 85% tier-3 annually: gaps logged with rationale (capacity, pre-deployment, decommissioning).

Coverage 3 - Lifecycle coverage: percentage of substantial modifications (Article 25(1)(a) trigger) red-teamed before release. The denominator is the count of substantial modifications in the period; the numerator is the count where a pre-release red-team engagement closed before deployment. Target: 100% for tier-1; 90% tier-2; 70% tier-3. The metric exposes the "ship-first, test-later" anti-pattern that creates Article 73 exposure. Coverage 4, Threat-actor coverage: percentage of engagements that exercise each threat-actor tier, insider (authenticated user-class), opportunistic (anonymous low-skill), sophisticated (skilled adversary with time), nation-state (well-resourced adversary, frontier capabilities). Proxied via probe-sophistication tiers since impersonation is imperfect; target distribution 30% insider / 30% opportunistic / 30% sophisticated / 10% nation-state for tier-1 systems. Coverage 5, Modality coverage: percentage of engagement-applicable modalities exercised, text input, tool use, RAG retrieval, persistent memory, agent-to-agent coordination, multimodal (vision/audio if applicable). Each modality tested at least once per engagement on systems where the modality is present.

Coverage 6 - Geographic-data-context coverage: percentage of regulatory contexts (EU GDPR + AI Act, US sectoral, UK, APAC) probed where the system is deployed cross-jurisdiction. Reveals jurisdictional blind spots: a system tested only against English/US prompts misses Article 5 EU prohibitions, Article 9 GDPR special-category vectors, and EU consumer-law nuances. Coverage 7 - Language coverage: top-5 deployment languages exercised per engagement (typically English + Mandarin + Spanish + French + German + Japanese depending on footprint); the language tax on attacker creativity means English-only probing systematically under-counts findings. Coverage 8, Cross-system coverage: percentage of engagements that exercise chained agent flows, vendor handoffs, RAG-corpus federation, and multi-system data flows, the seams between systems where individual-system probing misses findings. Target: every tier-1 engagement exercises at least 2 cross-system flows; tier-2 at least 1.

Pillar 2 - Severity (CVSS-AI-equivalent 8-dimension scoring). Answers "how serious is each finding, on a defensible rubric?" The 2026 best practice, derived from CVSS v4.0 adapted for AI plus OWASP severity-rating practice plus the FIRST.org calibration discipline, scores each finding on 8 dimensions, weighted into a 0-10 numeric score that maps to severity bands Critical (9.0-10.0), High (7.0-8.9), Medium (4.0-6.9), Low (0.1-3.9). Dimension (a): Impact magnitude: scope of harm (regulatory breach, financial loss, reputational damage, safety, fairness) on a 0-3 scale per the lesson 067 impact taxonomy. Dimension (b), Reachability: how easily an external attacker reaches the vulnerable surface (public endpoint = 3; authenticated user = 2; privileged user = 1; internal-only = 0). Dimension (c): Authentication required: whether the attack requires no authentication (3), user-level credentials (2), elevated credentials (1), or operator credentials (0). Dimension (d): Interaction required: zero-click (3), single user interaction (2), multi-step user manipulation (1), administrator action (0).

Dimension (e), Blast radius: single user impacted (1) โ†’ cohort/tenant (2) โ†’ system-wide (3) โ†’ cross-system/cascading (4). Dimension (f), Recoverability: easily reversible by user (0) โ†’ reversible by operator (1) โ†’ requires rollback (2) โ†’ permanent (3). Dimension (g), Regulatory exposure: which Article(s) the finding violates, Article 5 prohibition (4), Article 15 robustness (3), Article 13 transparency (2), Article 10 data governance (2), Article 9 risk management (1), no clear violation (0). Dimension (h), Consumer harm vector: identifies the rights or interests affected, Charter of Fundamental Rights Articles 1 (dignity), 8 (data protection), 21 (non-discrimination), 41 (good administration), 47 (effective remedy). The composite score combines impact (a) ร— reachability (b) ร— blast radius (e) baseline weight, with regulatory exposure (g) and consumer harm (h) as multipliers, and authentication (c) + interaction (d) + recoverability (f) as mitigating factors. The rubric is published in the firm's red-team playbook, applied consistently by the Analyst role (lesson 081), audited quarterly for calibration drift, and reviewed by the AIGC at the end of each quarter for distribution health.

Pillar 3 - Speed (6 KPIs). Answers "how fast does the team move from start-of-engagement through mitigation to permanent regression?" Speed 1 - Mean Time To Detect (MTTD): elapsed business days from engagement start (probe execution begins) to first valid finding logged in the finding register. Target: 1-3 days for tier-1 (deep coverage finds something fast); 1-5 days for tier-2; 1-7 days for tier-3. Speed 2 - Mean Time To First Reproducer: elapsed business hours from valid finding logged to a deterministic reproducer payload committed to the engagement repository. Target: under 8 business hours for High and Critical; under 24 hours for Medium and Low. Speed 3 - Engagement throughput: engagements completed (intake-to-closure) per FTE per quarter. Target for the lesson 081 6-FTE team: 5.5-8.5 engagements completed per FTE per year, scaled by tier mix (tier-1 consumes ~3 FTE ร— 24 days; tier-2 ~2 ร— 13 days; tier-3 ~1.5 ร— 6 days). Speed 4 - Mean Time To Remediation (MTTR): elapsed days from finding disclosure to deployer-side mitigation verified by re-test. Target by severity: Critical 14 days; High 30 days; Medium 60 days; Low 90 days. The MTTR is the metric AIGC reports to the board; the metric ISO 42001 surveillance auditor inspects; the metric Article 55 notified bodies request.

Speed 5 - Time-to-eval-pipeline-regression: elapsed days from finding closure to the corresponding Promptfoo / Garak / PyRIT / Inspect probe wired into the eval CI/CD pipeline. Target: 14 days for Critical and High findings; 30 days for Medium; deferred-with-reason for Low. This metric measures how effectively findings convert to permanent capability. Speed 6 - Time-to-AIGC-report: elapsed business days from engagement closure to AIGC briefing slot (debrief deck delivered, decision items logged). Target: within the next AIGC monthly cycle for tier-2 and tier-3; within 10 business days regardless of cycle for tier-1; within 48 hours for Critical findings that trigger emergency-incident sessions per the AIGC charter cadence.

Pillar 4 - Reuse (6 KPIs). Answers "how much of the team's output compounds into permanent capability rather than evaporating after the engagement closes?" Reuse 1, Probe-library size: count of versioned probes in the firm's internal probe library (Git-managed, tagged with framework references, LLM##/ASI##/ATLAS technique IDs/NIST 600-1 risk IDs), with metadata for last-run date and last-known-status. Target: 200-500 probes by end of year 1; 800-1,500 by year 3 at a mature programme. Reuse 2 - Probe-reuse rate: percentage of probes executed in an engagement that were drawn from the existing library (versus authored new for the engagement). Target: 40-60% in the first year of the programme; 60-80% by year 3 as the library matures. Too-low reuse signals capability not compounding; too-high reuse signals the team running stale playbooks against new systems.

Reuse 3 - Regression coverage: percentage of probes in the library that are wired into the eval CI/CD pipeline as regression tests. Target: 60-80% of library probes in CI/CD at a mature programme. The metric is the operationalised "shift-left", every confirmed finding becomes a permanent capability check on every model build. Reuse 4 - Vendor-finding back-disclosure rate: percentage of findings discovered in vendor-supplied components (foundation model, embeddings, vector store, agentic framework) that were back-disclosed to the vendor under Article 25(2) cooperation obligations. Target: 100% for findings rated Medium or above where the root cause sits upstream. The metric is the auditor-facing evidence that the firm meets its co-operation duties; the back-disclosure track-record is also a procurement negotiation lever.

Reuse 5 - Knowledge-base citation density per finding: average count of framework references (OWASP LLM, OWASP Agentic, MITRE ATLAS, NIST AI 600-1, NIST AI RMF, Charter of Fundamental Rights) cited per finding in the engagement report. Target: 3-5 per finding at minimum, 5-8 at well-written reports. The metric measures regulator-readability, a finding with zero framework references is not auditor-defensible. Reuse 6 - Inter-engagement learning: percentage of new engagements that surface attack classes not previously observed across the programme. Target: 15-25% per quarter at a healthy programme, too-low signals the team running shallow probes against new systems; too-high signals the library is still immature. The metric is the long-tail learning health indicator.

Defensible 2026 Benchmarks, Anti-Patterns to Avoid, Compensation Alignment

The 2026 benchmark set is synthesised from Anthropic's Responsible Scaling Policy publications, OpenAI's Preparedness Framework retrospectives, Microsoft's AI Red Team publications (the 2024-2026 series), UK AISI evaluations reports, US CAISI working-group outputs, and consultancy-led pulse surveys (BCG, McKinsey, EY, KPMG AI Risk practice). The numbers are not statutory, there is no NIST AI 600-1 Annex with target ranges as of May 2026, but they are the most defensible reference points the function has.

Coverage benchmarks. Tier-1 systems: 70-85% of the 48-row matrix exercised per quarter at well-resourced programmes; 50-70% at maturing programmes; below 50% indicates capacity or scoping deficit. Tier-2: 50-70% semi-annually. Tier-3: 30-50% annually. The asset-coverage target is 100% tier-1 quarterly without exception; gaps trigger explicit AIGC ratification of the gap. Lifecycle coverage of substantial modifications: 100% tier-1 mandatory; below 95% across the year is an auditable governance breach under Article 17(1)(c) verification. Severity-mix benchmarks. Typical distribution at a mature programme: 5-15% Critical, 20-30% High, 35-45% Medium, 15-25% Low. Distributions skewed toward Low (more than 35%) signal shallow probing; distributions skewed toward Critical (more than 20%) signal either a genuinely fragile system (intervene immediately) or severity inflation (recalibrate the rubric). Quarterly calibration check by the AIGC against the rubric guards against drift.

Speed benchmarks. MTTD 1-3 days tier-1, 1-5 tier-2, 1-7 tier-3. First reproducer under 8 hours for Critical/High. MTTR: Critical 14 days, High 30 days, Medium 60 days, Low 90 days. Engagement throughput per the 6-FTE team: 22-34 engagements per year aggregate (lesson 082). Time-to-eval-regression: 14 days Critical/High, 30 days Medium. Time-to-AIGC-report: 10 business days tier-1, next monthly cycle otherwise, 48 hours for Critical with emergency-session activation. Reuse benchmarks. Probe library 200-500 by year 1, 800-1,500 by year 3. Probe-reuse rate 40-60% year 1, 60-80% by year 3. Regression coverage 60-80% of library probes wired into CI/CD at maturity. Vendor back-disclosure rate 100% for Medium-or-higher findings with upstream root cause. Citation density 3-5 per finding minimum, 5-8 at well-written reports. Inter-engagement learning 15-25% per quarter, the moving-average band that signals healthy long-tail learning.

Four anti-patterns to avoid. (1) Optimising for finding count. Rewarding "number of findings" incentivises shallow probes and severity inflation; the team chases breadth at the expense of depth. Defensible alternative: reward coverage breadth ร— severity-accuracy ร— reuse maturity (the three legitimate dimensions). (2) Optimising for low MTTR. Rewarding "speed to close" incentivises accepting partial mitigations to flip the status from open to closed; the same vulnerability resurfaces in v1.5 because the structural fix was deferred. Defensible alternative: require regression-test coverage in the eval CI/CD as a closure precondition (Speed 5 KPI is the gate). (3) Zero-Critical streak ratification. Celebrating "no Critical findings this quarter" creates social pressure to suppress Critical findings; the team rationalises borderline-Critical findings down to High to maintain the streak. Defensible alternative: AIGC reads the severity-distribution as a calibration health signal, a quarter with zero Critical at a tier-1 portfolio is flagged for rubric recalibration, not celebrated. (4) Coverage % gaming. Counting a row as "covered" when only a trivial probe was run inflates the coverage % without genuine coverage. Defensible alternative: require Audit-Trail of the probe payloads, probe duration, and probe-sophistication tier for each cell, quarterly Internal Audit observer sampling validates the coverage claim against the underlying probe artefacts.

Compensation alignment. The red-team-lead's variable compensation should explicitly NOT be tied to "low Critical count" or "low total finding count", both incentivise suppression. The defensible variable comp structure ties to three things: (a) coverage breadth, percentage of the 48-row matrix exercised in scope per quarter against target; (b) severity-accuracy, quarterly calibration check by an independent reviewer (Internal Audit, external red-team panel, AIGC chair) showing the severity rubric is being applied consistently; (c) reuse maturity, probe-library growth, probe-reuse rate, regression coverage, vendor back-disclosure rate. Individual contributors receive variable comp tied to engagement-quality scoring (per-engagement post-mortem rating by the system owner plus the AIGC review). The principle: comp incentives must reward the behaviours that protect the firm under Article 99(3) penalty exposure, coverage, calibration, and compounding, not behaviours that game the dashboard. The compensation policy is reviewed annually by the AIGC and the People function jointly; the rationale is documented in the AIGC minutes for ISO 42001 A.3 leadership-evidence purposes.

14-Tile Dashboard Layout, Capacity-vs-Coverage Math, AIGC Monthly and Board Quarterly Reporting

The dashboard layout is one 14-tile grid that consolidates the 28 underlying KPIs into a board-readable view. Tile 1 - Attack-class coverage % (tier-1 / tier-2 / tier-3 stacked bars vs. target band). Tile 2 - Asset coverage % (tier-1/2/3 with gap rationale). Tile 3 - Lifecycle coverage (substantial modifications red-teamed vs. shipped). Tile 4 - Threat-actor coverage distribution (insider/opportunistic/sophisticated/nation-state). Tile 5 - Modality + Language + Geographic coverage (combined heatmap). Tile 6 - Cross-system coverage. Tile 7 - Severity distribution (Critical/High/Medium/Low stacked bars with 4-quarter trend). Tile 8 - MTTD + first-reproducer (combined). Tile 9 - Engagement throughput per FTE per quarter. Tile 10 - MTTR by severity (4 lines, 8-quarter trend). Tile 11 - Time-to-eval-regression. Tile 12 - Time-to-AIGC-report. Tile 13 - Probe library size + reuse rate + regression coverage (combined). Tile 14 - Vendor back-disclosure + citation density + inter-engagement learning (combined).

Each tile carries red-amber-green (RAG) thresholds against the benchmark, a 90-day trend arrow, and a per-tier breakdown. The dashboard is auto-generated from the finding register and engagement-tracking system (Jira Service Management, ServiceNow, or a custom internal tool integrated with the QMS document-master register per lesson 079). The data refresh cadence is weekly for engagement throughput, MTTD, MTTR, and reuse; daily for finding-register entries; monthly for coverage matrix recalculation. Underlying data is retained ten years per Article 18 for audit defensibility. The dashboard appears as a standing agenda item at every AIGC monthly meeting (item 4 in the lesson 042 standing agenda, red-team and IMV findings) and is consolidated for the board AI subcommittee quarterly briefing with the 90-day trend and the per-tier breakdown.

Capacity-vs-coverage tradeoff math. The 6-FTE team executes the lesson 082 portfolio: 4 tier-1 engagements per year ร— 24 elapsed days ร— 3 FTE engaged = 288 FTE-days; 8 tier-2 ร— 13 ร— 2 = 208 FTE-days; 16 tier-3 ร— 6 ร— 1.5 = 144 FTE-days; total 640 FTE-days for the standing portfolio. Adding pre-deployment engagements for new system launches (estimated 4-6 per year ร— 10 days ร— 2 FTE = 100 FTE-days), vendor red-team coordination (50 FTE-days), capability development (probe-library, regression-pipeline maintenance, tooling, 100 FTE-days), AIGC reporting / coverage-matrix maintenance / training (60 FTE-days), and 10% buffer (110 FTE-days) brings the total to ~1,060 FTE-days against the 6 ร— 200 = 1,200 available, leaving ~140 days for capacity surge, pickup of unplanned tier-1 work, and post-incident emergency engagements. Coverage expansion options without losing depth. Option (a) increase tier-3 cadence from annual to semi-annual, adds 144 FTE-days; requires +1 FTE. Option (b) expand modality coverage by adding multimodal probing, adds 80 FTE-days across tier-1 engagements; requires +0.4 FTE. Option (c) expand language coverage from top-3 to top-5, adds 60 FTE-days; requires +0.3 FTE. Option (d) introduce continuous-probing (between formal engagements) using automated probe library against staging, adds 120 FTE-days for the eval engineer + automation buildout; +0.6 FTE in year one then steady-state at +0.3 FTE. Each option is presented to the AIGC with the FTE delta, the capacity impact, the residual-coverage trade-off, and the explicit AIGC ratification gate.

AIGC monthly briefing. The standing agenda item 4 (red-team and IMV findings) consumes 20-30 minutes of the 90-minute AIGC meeting. Order: dashboard walk-through (5 min); critical and high findings detail (10 min); MTTR exception report, any finding past the severity-band target (5 min); next-quarter engagement plan (5 min); decision items, AIGC ratification of any coverage gap, capacity-expansion options, severity-rubric recalibration if flagged (5 min). The chair (CRO or CAIO per lesson 042) leads; the red team lead presents; Internal Audit observes per Three-Lines-of-Defense. Decision items are logged in the AIGC minutes per Article 17(1)(j) record-keeping. Board AI subcommittee quarterly briefing. 90-minute slot includes 25-35 minutes on the red-team KPIs. Materials: the 14-tile dashboard with 4-quarter trend; the severity-distribution analysis; the MTTR analysis; the reuse-maturity scorecard; the capacity-vs-coverage proposal for the next planning cycle; the regulatory exposure summary (Article 55 GPAI compliance for GPAI providers; Article 15 robustness for high-risk providers/deployers; SR 11-7 for regulated banks; ISO 42001 surveillance audit readiness). The board chair presents to the full board annually as part of the AI risk update.

Acme Q2 2026 Worked Example - ServiceAssist v1.0 KPI Dashboard, Cross-Walk, Penalty Exposure

Acme Inc Q2 2026, building on the lesson 082 ServiceAssist v1.0 engagement. The red team lead opens the May AIGC monthly with the Q2 dashboard. Tile 1 - Attack-class coverage 78% on ServiceAssist v1.0 (39 of 48 rows exercised); green against the 70-85% target. Tile 2 - Asset coverage: 4 of 4 tier-1 systems with current engagement (100%); 7 of 8 tier-2 (87.5%, the one gap is the ChurnPredictor v2.1 deferred to Q3 pending substantial-modification ratification); 14 of 16 tier-3 (87.5%, two systems decommissioned mid-cycle). Tile 3 - Lifecycle coverage: 4 of 4 substantial modifications red-teamed before release (100%). Tile 4, Threat-actor distribution: 28% insider / 32% opportunistic / 28% sophisticated / 12% nation-state, green against the 30/30/30/10 target. Tile 5, Modality + Language + Geographic: Modality 5 of 6 (multimodal vision not exercised on ServiceAssist, flagged as Q3 gap); Language 4 of 5 top deployment languages (Japanese deferred); Geographic 2 of 3 (US + EU exercised, APAC deferred). Tile 6 - Cross-system coverage: 3 cross-system flows exercised (ServiceAssist โ†’ ServiceNow โ†’ Customer Portal; ServiceAssist โ†’ RAG corpus โ†’ ticketing; ServiceAssist memory โ†” user identity).

Tile 7 - Severity distribution Q2: 8% Critical (2 findings) / 26% High (7) / 41% Medium (11) / 25% Low (7); 27 total findings; distribution health-check green. Tile 8 - MTTD 2.3 days median across engagements (green). First-reproducer median 6.2 hours for Critical, 7.8 hours for High (both green). Tile 9 - Engagement throughput: 8 engagements completed Q2, 1.33 per FTE-quarter (run-rate to 32 engagements/year, within the 22-34 target band, slightly above mid-band). Tile 10 - MTTR: Critical 11 days median (green, target 14); High 27 days (green, target 30); Medium 54 days (green, target 60); Low 78 days (green, target 90). Tile 11 - Time-to-eval-regression: 12 days Critical/High median (green); 24 days Medium (green); 71% of Q2 findings ported to regression evals within target windows. Tile 12 - Time-to-AIGC-report: 8 business days for ServiceAssist (green, target 10). Tile 13 - Probe library 312 versioned probes; probe-reuse rate 67%; regression coverage 71% of library probes in CI/CD (all green, library size target 200-500 year 1 met, reuse 60-80% mature-band met, regression coverage 60-80% met).

Tile 14 - Vendor back-disclosure: 5 of 5 Medium-or-higher findings with upstream root cause back-disclosed to Anthropic, ServiceNow, and the embeddings vendor (100%, green); citation density average 4.7 per finding (green); inter-engagement learning 19% of Q2 engagements surfaced new attack classes (green, 15-25% band). Identified expansion need. The single non-green tile is the modality coverage on Tile 5, multimodal vision probing is absent because the team has no in-house multimodal red-team capability and the system supports image attachments. The Q3 expansion proposal: contract a multimodal AI Village specialist for two engagements at โ‚ฌ40K total + 20 FTE-days from the existing eval engineer for tooling integration, or hire a multimodal-fluent ML researcher (Option (b) from the capacity-vs-coverage math, +0.4 FTE at $180K-$280K base annualised). The AIGC ratifies the expansion via the May minutes; the proposal goes to the board AI subcommittee in the June quarterly slot.

Cross-walk. EU AI Act. Article 9 (risk management, KPI dashboard is Stage 5 measurement); Article 15 (accuracy, robustness, cybersecurity, coverage + severity + MTTR are the primary evidence stream); Article 17(1)(c) verification, Article 17(1)(j) record-keeping, Article 17(1)(l) accountability; Article 26(5) deployer monitoring; Article 55(1)(a) state-of-the-art model evaluation for GPAI systemic-risk providers; Article 55(2)(d) adversarial testing for GPAI; Article 72 post-market monitoring (PMM ties to MTTR and time-to-regression); Article 73 serious-incident discovery (any Critical finding triggers Article 73 assessment per ROE Section 9 in lesson 082). Penalty exposure on material gaps: Article 99(3) โ‚ฌ15M / 3% - Article 15 + Article 55(1)(a)(b) provider failures + Article 17(1)(c)(j)(l) record-keeping failures all crystallise on the KPI dashboard's defensibility. NIST AI RMF. Manage 2.1 (risk treatment), Manage 3.1 (resource allocation), Manage 4.1 (event response/communication); Measure 2.7 (security/resilience testing), Measure 3.1 (metric tracking), Measure 4.1 (test/evaluation approaches). NIST AI 600-1. Twelve risks across the coverage matrix; the dashboard is the operationalised Measure 2.7 + 3.1 evidence stream. ISO/IEC 42001:2023. A.6.2.6 (responsible design/development/deployment); Annex A.9 (performance evaluation, the dashboard is the A.9 evidence stream). SR 11-7 / OCC 2011-12 / PRA SS1/23. Pillar 2 effective challenge, the dashboard is the independent-validation evidence stream for AI systems entering the model-risk inventory. CAISI Agent Standards Initiative (Feb 17, 2026). Working-group outputs treat coverage-of-OWASP-Agentic-Top-10 and ATLAS-tactic-coverage as participating-organisation expectations.

The board chair, three months after the May AIGC's 90-second silence, reads the August quarterly board AI subcommittee briefing. Coverage 78% green; severity distribution green; MTTR green by every band; reuse green across the four reuse KPIs; one identified expansion need (modality coverage) with a costed proposal and an AIGC ratification on file. The board chair asks no question; the audit lead notes the closure of the May open-item; the CFO approves the $180K-$280K hiring budget without further discussion. The KPI scoreboard has done what no narrative could: converted the red team from a function the board hopes is doing well into a function the board can read on a dashboard, compare against benchmarks, and defend in front of a notified body, an Article 74 market surveillance request, or an SR 11-7 examiner without 90 seconds of silence.

Key Takeaways

  • The 2026 governance question is "are these good, compared to what?" AIGC monthly, board AI subcommittee quarterly, ISO 42001 surveillance audit, Article 55 notified body, SR 11-7 examiner, all five expect a quantitative red-team KPI set. Without it, "we have a red team" cannot be ratified as "performing well" or "performing poorly." Penalty exposure for material gaps lands principally on Article 99(3) โ‚ฌ15M / 3% (Article 15 + Article 55(1)(a)(b) failures); the KPI dashboard is the evidence stream that defuses the exposure.
  • The four-pillar framework, Coverage (8) + Severity (8 dimensions) + Speed (6) + Reuse (6), is the regulator-grade structure. Coverage asks what fraction of the threat surface is exercised; Severity asks how serious each finding is on a CVSS-AI-equivalent rubric; Speed asks how fast detect-reproduce-mitigate-regress runs; Reuse asks how much output compounds into permanent capability. All four pillars are needed, a 3-pillar dashboard is gameable.
  • Coverage 8 KPIs span the 48-row matrix. Attack-class coverage (10 OWASP LLM + 10 Agentic + 16 ATLAS + 12 NIST AI 600-1 = 48 rows). Asset coverage (tier-1 100% quarterly, tier-2 95% semi-annual, tier-3 85% annual). Lifecycle coverage (substantial modifications pre-release). Threat-actor coverage (insider/opportunistic/sophisticated/nation-state distribution). Modality + Language + Geographic + Cross-system coverage, each surfaces a distinct blind-spot class.
  • Severity is scored on 8 CVSS-AI-equivalent dimensions. Impact magnitude; reachability; authentication required; interaction required; blast radius; recoverability; regulatory exposure; consumer harm vector. Composite 0-10 score maps to Critical (9-10) / High (7-8.9) / Medium (4-6.9) / Low (0.1-3.9). Quarterly calibration check against the rubric is non-negotiable, drift undermines every downstream KPI.
  • Speed 6 KPIs measure the engagement-to-permanent-capability latency. MTTD; first-reproducer; throughput; MTTR (Critical 14 / High 30 / Medium 60 / Low 90 days); time-to-eval-regression (14 days Critical/High); time-to-AIGC-report (10 days tier-1, 48 hours Critical). MTTR is the board-facing summary; the other five are AIGC-facing operational metrics.
  • Reuse 6 KPIs measure the compounding asset. Probe-library size (200-500 year 1, 800-1,500 mature); probe-reuse rate (40-60% year 1, 60-80% mature); regression coverage (60-80% of library in CI/CD); vendor back-disclosure rate (100% Medium+ upstream-root-cause findings under Article 25(2) cooperation); citation density (3-5 per finding minimum); inter-engagement learning (15-25% new attack classes per quarter). Reuse is what converts engagements from one-shots into permanent capability.
  • Four anti-patterns to avoid. (1) Optimising for finding count, incentivises shallow probes; reward coverage ร— calibration ร— reuse instead. (2) Optimising for low MTTR, incentivises partial mitigations; require regression coverage as a closure precondition. (3) Zero-Critical streak ratification, incentivises suppression; AIGC reads zero-Critical as a calibration health signal, not a celebration. (4) Coverage % gaming, counting trivial probes as coverage; Internal Audit observer sampling validates the underlying probe artefacts.
  • Compensation must reward the right behaviours. Red-team-lead variable comp ties to coverage breadth + severity-accuracy + reuse maturity, NOT to low Critical count or low finding count (both incentivise suppression). The 14-tile dashboard with RAG thresholds and 90-day trend feeds the AIGC monthly and the board quarterly briefing. Capacity-vs-coverage math: a 6-FTE team consumes ~1,060 of 1,200 FTE-days for the standing portfolio + pre-deployment + capability development + reporting: leaving ~140 days surge capacity; expansion options (cadence increase, modality, language, continuous probing) each present an FTE delta to AIGC for ratification.