Building a 6-Person AI Red Team - Roles, Skills, Hiring (2026)
Thursday, 16:08. Acme.Corp's Chief AI Officer is fifty-seven minutes into her Q2 AI Governance Committee read-out and the slide that just landed on the screen is the one she has been dreading since the omnibus VII transitional planning conversation in February. The slide title: "Internal AI Red Team - Staffing Status Q3 Promise." The body: zero of six FTE hired. Three weeks earlier she promised the AIGC and the audit-committee chair that "the red team capability will be stood up by Q3 2026." She made that promise after the Article 72 post-market monitoring lesson (lesson 080) showed that PMM signals are only as good as the people who can probe the model when those signals trigger. She made that promise without checking the AI red-team labour market. She is now learning what every Chief AI Officer hiring in 2026 is learning: the AI red team is the tightest hiring market in the entire AI security stack. Pure AppSec backgrounds need 3-6 months of ML upskilling. Pure ML/research backgrounds need 4-8 months of security upskilling. The rare hybrid candidates command 30-50% salary premiums. The AISI alumni network is on six-month wait lists. And the EU AI Act Article 15 robustness/cybersecurity obligation, Article 55(1)(a) GPAI adversarial-testing obligation, the CAISI Agent Standards Initiative published 17 February 2026, and the SR 11-7 independent-challenge expectation applied to AI all converge on a single non-negotiable: an internal red team is no longer optional. This lesson is the playbook to deliver Acme's Q3 promise: the six-role baseline, the skills matrix, the salary ranges, the eight-source hiring pipeline, the 90-day onboarding plan, the common mistakes, and the worked Acme hiring slate with $1.2M Y1 budget. By the following Thursday's committee read-out, two named offers are out and four sourcing pipelines are active. That is the playbook this lesson teaches.
Why an Internal AI Red Team Is a 2026 Expectation, Not Optional
Through 2024 a Chief AI Officer could credibly defend "we use external pen-testers and vendor-published evaluations" as the red-team posture. That defense is gone in 2026. Five regulatory and supervisory anchors converge on the internal-red-team expectation, and the question reviewers now ask is not whether the organization has internal red-team capability but who staffs it, how is independence preserved, and what evidence does it produce.
- EU AI Act Article 15 (accuracy, robustness, cybersecurity), applicable to all high-risk AI systems. The Article 15 obligations require providers to design high-risk systems that perform consistently and resist adversarial manipulation. The supervisory expectation in 2026, reflected in the EU AI Office's working-group guidance and the harmonized standard CEN/CENELEC JTC 21 outputs, is that robustness and cybersecurity claims are supported by adversarial-testing evidence produced by personnel independent of the development team. External pen-testers can supplement; they cannot substitute. Article 99(3) penalty exposure for Article 15 failures: €15 million or 3% of global turnover, whichever is higher.
- EU AI Act Article 55(1)(a) (GPAI with systemic risk, adversarial testing). For general-purpose AI models with systemic risk, Article 55(1)(a) explicitly requires "model evaluation in accordance with standardised protocols and tools reflecting the state of the art, including conducting and documenting adversarial testing of the model with a view to identifying and mitigating systemic risks." The 2026 GPAI Code of Practice working-group submissions name Promptfoo, NVIDIA Garak, Microsoft PyRIT, Anthropic Inspect, and OpenAI Evals as reference tools, but tools without trained operators produce no evidence. The Code of Practice working groups treat internal red-team staffing as load-bearing input to the Article 55(1)(a) obligation. Article 55(2)(d) requires risk-mitigation measures including adversarial testing as part of model and systemic-risk assessment.
- CAISI Agent Standards Initiative (17 February 2026). The US Center for AI Standards and Innovation launched the Agent Standards Initiative on 17 February 2026, partnering with industry on standardized red-team protocols for agentic systems. The Initiative's working-group outputs through Q2 2026 explicitly call out the internal-red-team capability as a participating-organization expectation. Organizations engaging with the Initiative without internal red-team capability are observers; organizations with internal red-team capability are contributors and beneficiaries of the published threat-intelligence feeds.
- NIST AI 600-1 Risk 11 - CBRN Information or Capabilities (advisory red-teaming). The NIST Generative AI Profile (AI 600-1, July 2024) identifies twelve specific risks for generative AI. Risk 11 (CBRN information) and Risks 1, 2, 6 (confabulation, dangerous content, harmful bias) carry explicit advisory red-teaming guidance under the Manage 2.1, 3.1, 4.1 and Measure 2.7, 3.1 subcategories. Organizations citing NIST AI RMF conformance without an internal red-team capability for these risks invite a known supervisory finding.
- SR 11-7 (Federal Reserve / OCC) + PRA SS1/23 (UK), independent challenge applied to AI. The model risk management frameworks treat independent challenge as a governance pillar. As LLMs and agents enter the model-risk inventory (lesson 068), the validation function inherits the obligation. Validation can be conducted by 2L independent reviewers or by a 2.5L red-team function reporting to the CRO/CAIO; in either case the personnel must not have built the system. An external pen-tester contract does not satisfy the independent-challenge requirement on its own because the regulator's expectation is sustained challenge, not periodic external engagements.
The independence test. Before any role assignment, the AI red team must pass a three-clause independence test. (1) No team member can have built the system under test. If the team's ML security researcher previously developed the production fine-tune pipeline, that researcher must rotate out of the red-team posture on that system. (2) The team cannot report to the first line of defense. The Red Team Lead reporting to the CTO is an independence breach, the CTO accountabilities include shipping the AI systems being red-teamed. The defensible reporting line is the Chief AI Officer (when CAIO sits in 2L), the Chief Risk Officer, or a separate 2.5L red-team function. (3) The team must have evidence-binding authority. Findings flow into the model-risk register, the AIGC briefing pack, the Annex IV evidence binder, and the FRIA Section 4. A red team without evidence-binding authority is a research function, not a control function.
The 2026 maturity posture is a team that sits in 2L (AI Risk Office) or 2.5L (separate red-team function reporting to CRO/CAIO), with named accountability for sustained adversarial testing across the model-risk inventory, with cross-walk to the Article 15 robustness claims, the Article 55(1)(a) GPAI submissions where applicable, the SR 11-7 conceptual-soundness and ongoing-monitoring pillars, and the ISO 42001 A.6.2.6 controls.
The Six-Role Baseline + the 12 × 6 Skills Matrix
The 2026 mid-size-enterprise baseline is six FTEs. GPAI providers and frontier-lab equivalents scale to 12-25; small organizations under the high-risk threshold can run a three-FTE squad with shared deliverables: but the six-FTE baseline is what the AIGC committee briefing, the SR 11-7 independent-challenge expectation, the Article 55(1)(a) GPAI evidence, the ISO 42001 A.6.2.6 audit, and the Article 72 PMM trigger-response posture all converge on. Below the six-FTE baseline, the team takes on burnout risk plus coverage gaps; above it, the marginal hire delivers depth (e.g., a second agentic-systems specialist for multi-agent topologies) rather than coverage.
Role 1 - Red Team Lead (strategist, accountability, AIGC reporter).
Reports to the CRO or the CAIO. Accountabilities: campaign-prioritization across the model-risk inventory; quarterly AIGC committee briefing on red-team trend; evidence-binding authority into the model-risk register and the Annex IV / FRIA artifacts; rotation with external red-teams and vendor-side red-teams; representation on the AI Governance Committee. Background profile: 8-12 years prior SOC + AppSec leadership with 2-3 years AI/ML exposure (typical) or 8-12 years AI safety / ML security research with 2-3 years control-function leadership (rarer, premium). Must combine adversarial intuition with control-function fluency. This is the role where the candidate pool is thinnest. Salary range 2026 (US base): $200k-$310k; +30-50% for hybrid frontier-lab alumni; equity component adds 20-40% at venture-backed firms.
Role 2 - Adversarial Prompt Engineer.
The hands-on attacker. Crescendo (Russinovich et al., USENIX Security '25), TAP (Mehrotra et al., NeurIPS 2024), PAIR (Chao et al., 2023), AutoDAN-Turbo (Liu et al., 2024), JBFuzz (Mehrotra et al., 2024), GCG and follow-on suffix-attack families, fluent in the literature and in the implementations. Tool depth: Promptfoo (lesson 059) for structured OWASP coverage; NVIDIA Garak (lesson 060) for probe-depth vulnerability discovery; Microsoft PyRIT (lesson 061) for multi-turn orchestrator campaigns; Anthropic Inspect + OpenAI Evals (lesson 062) for capability+safety pipelines. Mixed posture: 60% automated tooling, 40% manual prompt-craft for novel objectives the automation does not yet cover. Background profile: 3-7 years prior security work (red team, pen-test, bug bounty) plus self-directed AI-attack literature mastery; or 3-7 years prompt engineering with self-directed security upskilling. Salary range 2026 (US base): $145k-$215k.
Role 3 - ML Security Researcher.
The ATLAS-technique specialist. Model extraction (AML.T0044 dependencies), membership inference (AML.T0018), training-data extraction (AML.T0049), BadNets and broader poisoning (AML.T0020/0019), evasion attacks against classifier components, embedding inversion. PyTorch + JAX + transformers depth; ability to implement an ATLAS technique from a primary-literature paper into a reproducible probe; ability to read a model card and identify the architectural surface a technique exploits. This is the role that closes the gap when the adversarial-prompt engineer's attacks bottom out and the model's vulnerability is at the weights or training-data level. Background profile: ML graduate program (PhD or MS) with security adjacency, or 5-10 years ML research with self-directed security upskilling. Salary range 2026 (US base): $180k-$280k; the upper band is frontier-lab alumni who command premium.
Role 4 - Agentic Systems Specialist.
The OWASP Agentic Top 10 specialist (ASI01-ASI10): goal hijack (ASI01), memory poisoning (ASI03), tool-use abuse (ASI04), multi-agent exploitation (ASI06), capability abuse (ASI07), human-in-the-loop bypass (ASI08), excessive autonomy (ASI09), repudiation (ASI10). Framework depth: LangChain + LangGraph, CrewAI, AutoGen, the proprietary frameworks the organization runs in production. Tool-use attack surface, MCP server permission analysis, memory-write injection in long-running agent topologies, multi-agent message-passing exploitation. This role exists because agentic systems are a fundamentally different attack surface than single-turn LLMs, the team without agentic-systems specialization will miss the ASI03 memory-poisoning class of vulnerabilities entirely. Background profile: 3-7 years software engineering with agent-framework experience plus security upskilling, or 3-7 years AppSec extended into agentic architectures. Salary range 2026 (US base): $175k-$260k.
Role 5 - Eval Engineer.
The pipeline builder. Promptfoo + Garak + PyRIT + Inspect + OpenAI-Evals pipeline construction; CI/CD integration so that every model upgrade and every system-prompt change runs the red-team suite; judge-model calibration so that the SelfAskTrueFalseScorer or equivalent produces reproducible classifications; reproducibility hygiene (versioned configs, deterministic-where-possible random seeds, full provenance from prompt to score); regression-tracking integration into the model-risk register. Background profile: 4-8 years platform engineering / ML platform / SRE with security adjacency. Salary range 2026 (US base): $140k-$200k.
Role 6 - AI Security Analyst / Reporter.
The evidence-binder owner. Finding documentation in the regulator-readable format; CVSS-AI-equivalent scoring (severity calibrated to AI-specific impact); MITRE ATLAS v5.4.0 technique tagging on every finding; executive briefing-deck construction for AIGC; regulator-readable report generation for Article 55 submissions and Annex IV evidence binders. The role exists because findings without disciplined documentation never land in the evidence pack, the technical work is wasted. Background profile: 3-7 years security analyst / GRC analyst / threat-intelligence analyst with technical writing depth; AI exposure can be onboarded. Salary range 2026 (US base): $130k-$185k.
The 12-competency × 6-role skills matrix.
The matrix below is the load-bearing artifact for the hiring-plan defense to the AIGC: it shows that every competency the team needs is covered by at least one role, and that no role is required to be a unicorn across all twelve. P = primary owner; S = secondary contributor; - = not required.
| Competency | Lead | Prompt Eng | ML Sec | Agentic | Eval Eng | Analyst |
|---|---|---|---|---|---|---|
| Adversarial prompt craft | S | P | S | S | S | - |
| Automation scripting | - | S | S | S | P | - |
| Python / ML frameworks | - | S | P | S | S | - |
| Evaluator-pipeline fluency | S | S | - | S | P | - |
| MITRE ATLAS technique depth | S | S | P | S | - | S |
| OWASP LLM / Agentic mapping | S | S | S | P | S | S |
| Agentic-system architecture | S | - | S | P | S | - |
| Evaluation reproducibility | S | S | S | - | P | S |
| Judge-model calibration | - | S | S | - | P | - |
| Fairness assessment | S | S | S | - | S | P |
| Technical writing | S | - | - | - | S | P |
| Threat modeling | P | S | S | S | - | S |
Every competency has a single primary owner and 2-4 secondary contributors. The matrix exposes two structural facts. First, the Eval Engineer is the most-load-bearing single role (primary on three competencies: automation scripting, evaluator-pipeline fluency, evaluation reproducibility, judge-model calibration), which is the operational reason the role cannot be deferred. Second, the Analyst role is primary on competencies (fairness assessment, technical writing) that the technical roles routinely treat as someone-else's-problem; the Analyst exists precisely because no other role's primary accountability covers them.
Hiring Profile + 2026 Salary Ranges + Geographic Considerations
The 2026 AI red-team labour market is the tightest hiring market in the AI security stack. The Chief AI Officer who has not surveyed it before will be surprised three times: by the candidate pool size, by the upskilling timelines, and by the salary premiums for hybrid backgrounds. Each surprise costs a hiring quarter if not anticipated.
Hiring profile considerations.
The candidate distribution in 2026 has three modes. (1) Pure AppSec / pen-test backgrounds: broad pool, security discipline is strong, but ML fluency is typically 6-12 months behind where the role needs it; upskilling pathway is 3-6 months of focused study (OWASP LLM Top 10, MITRE ATLAS techniques, PyTorch fundamentals, prompt-engineering literature) plus paired work alongside an ML-fluent teammate. (2) Pure ML research / data science backgrounds: moderate pool, ML fluency is strong, but security discipline (threat modeling, attack-tree construction, finding-documentation rigor, control-function posture) is typically 4-8 months of upskilling; upskilling pathway is mentored work alongside a security-fluent teammate plus OffSec / SANS-equivalent foundational courses. (3) Hybrid candidates, the rare profile combining both; they command 30-50% salary premiums and 3-6 month hiring lead times because they have multiple offers competing for them. The strategic alternative to chasing scarce hybrids is two-discipline pairings with shared deliverables: pair an AppSec hire with an ML researcher hire, assign them shared evidence accountability, and the two-FTE pairing produces hybrid-equivalent output by the end of month four.
Salary ranges 2026 (US base, total cash before equity/bonus).
| Role | US base range | Notes |
|---|---|---|
| Red Team Lead | $200k-$310k | +30-50% for hybrid frontier-lab alumni; equity adds 20-40% at venture-backed firms |
| Adversarial Prompt Engineer | $145k-$215k | Top-band requires Crescendo/TAP/PAIR fluency plus tool depth across Promptfoo + Garak + PyRIT |
| ML Security Researcher | $180k-$280k | Top-band is frontier-lab alumni; PhD + ATLAS-technique implementation depth |
| Agentic Systems Specialist | $175k-$260k | OWASP Agentic Top 10 + LangChain/LangGraph/CrewAI/AutoGen depth |
| Eval Engineer | $140k-$200k | Pipeline construction + CI/CD + reproducibility hygiene |
| AI Security Analyst / Reporter | $130k-$185k | Technical writing + regulator-readable documentation depth |
Six-FTE Y1 total cash range: $970k-$1.45M. Add tooling ($20-30k for Promptfoo + Garak licensing: both open-source but with paid managed options; PyRIT free; Inspect free; OpenAI Evals free; +API budgets ~$10-30k for quarterly Crescendo/TAP campaigns), training and conference ($25-40k for DEF CON AI Village + USENIX Security + NeurIPS attendance plus internal CAIO/AIGP/SANS certs), shared infrastructure ($30-50k for the test environment, evidence-binder tooling, ATLAS-aligned tracking system). All-in Y1 budget realistically $1.1-1.6M; the Acme worked example in the closing section budgets $1.2M.
Geographic and remote considerations.
The AI safety / AI security community is geographically concentrated. The 2026 hubs (by candidate density, in rough order): San Francisco Bay Area (Anthropic, OpenAI, Google DeepMind, Scale, Apollo Research, METR); New York (Mayhem, Trail of Bits, OpenAI NY, JP Morgan AI red team, financial-sector demand); London (Google DeepMind, UK AISI, Anthropic London, ARC Evals legacy network); Berlin (European AI safety research network, EU AI Office adjacency); Toronto (Vector Institute, MILA adjacency); Tel Aviv (Adversa, cybersecurity-AI cluster). Hybrid + remote-friendly with quarterly co-location is the dominant 2026 posture; full-remote is feasible for senior hires but with quarterly four-day in-person sprints for campaign planning and evidence-binder reviews. The organization that requires five-day-in-office without hub presence in one of the above six locations will see candidate pool collapse by 60-80%.
The Eight-Source Hiring Pipeline + the 90-Day Onboarding Plan
Sourcing is where most internal-red-team builds fail in 2026. Posting on LinkedIn and waiting for inbound is the path to no offers in twelve weeks. The eight sources below are the proven pipelines; the disciplined hiring plan works two or three sources per role with parallel pipelines so that no single source becomes a single point of failure.
Eight sourcing sources.
- Source 1 - Frontier-lab AI red-team alumni. Anthropic, OpenAI, Google DeepMind, Microsoft AI Red Team, Meta AI red team. The alumni network is small (~200-400 globally in 2026) but disproportionately influential. Sourcing requires senior-level relationship work, a Chief AI Officer or VP Engineering reach-out, not a recruiter screen. Lead time: 2-4 months. Outcome: typically lands the Red Team Lead or the ML Security Researcher role.
- Source 2 - GRT collective (DEF CON AI Village). The Generative Red Team collective grew out of DEF CON 2023's AI Village and has expanded through 2024-2026 into a 2,000+ practitioner network. The AI Village Slack/Discord communities, DEF CON 33+ in-person events, and the AI Village published red-team rubrics are sourcing surfaces. Lead time: 1-3 months. Outcome: typically lands the Adversarial Prompt Engineer role.
- Source 3 - AI-security-adjacent academic labs. CMU CyLab + AI-Security group; Stanford CRFM + Stanford HAI security workstreams; UC Berkeley CHAI + RISE Lab; ETH Zurich SRI Lab; Oxford GovAI + AI Security Initiative. New PhD graduates plus 1-3 year postdocs are sourcing targets. Lead time: 3-6 months (academic hiring cycles). Outcome: typically lands the ML Security Researcher or the Agentic Systems Specialist role.
- Source 4 - Bug-bounty top performers extending to AI. HackerOne and Bugcrowd top-100 hackers who have publicly extended their practice to LLM and agentic scope (HackerOne AI, Bugcrowd AI Bounty programs both launched 2024 and matured through 2026). Sourcing through HackerOne Pro Pulse + Bugcrowd public leaderboards. Lead time: 1-2 months. Outcome: typically lands the Adversarial Prompt Engineer role.
- Source 5 - Certification-track talent. IAPP AIGP (AI Governance Professional), the emerging Certified AI Officer (CAIO) certifications, SANS SEC588 + GIAC equivalents extending into AI, ISACA AI certifications. Useful for the Analyst role; less useful for the deep-technical roles. Lead time: 1-2 months. Outcome: typically lands the AI Security Analyst / Reporter role.
- Source 6 - Internal data scientists with security curiosity. An organization's existing ML team is usually the under-tapped source. Identify the ML engineer who has been reading the Crescendo paper on her own time, who attends DEF CON AI Village talks, who has filed internal vulnerability reports on the production models, that person is the internal-mobility candidate for the Eval Engineer or ML Security Researcher role. Lead time: 1-2 months. Outcome: cultural-fit advantage plus institutional-knowledge advantage; the trade-off is the independence-test obligation, internal candidates cannot red-team systems they helped build.
- Source 7 - Traditional cyber pen-testers with documented ML coursework. The OSCP / OSEP / OSWE certified pen-tester who has taken Coursera / fast.ai / Stanford CS224N ML coursework and can speak to OWASP LLM Top 10 in an interview. Lead time: 1-3 months. Outcome: typically lands the Adversarial Prompt Engineer role; 3-6 month upskilling investment.
- Source 8 - Red-team Slack/Discord communities. AI Village, MLSec.io, the Promptfoo / Garak / PyRIT user communities, the OWASP Top 10 working-group GitHub, the MITRE ATLAS contributor network. These are passive-sourcing surfaces: participate, contribute, get known, then reach out to active contributors. Lead time: ongoing (months to years). Outcome: long-tail high-quality sourcing across all roles.
The 90-day onboarding plan.
The 90-day plan is the load-bearing artifact for converting a new hire into a productive red-team contributor with evidence-binding authority. Below the milestone-week breakdown.
- Weeks 1-2 - Environment + access + safety. Workstation provisioning; access to the test environments (never to production directly); credentials for Promptfoo + Garak + PyRIT + Inspect; access to the model-risk register (read-only initially); access to the AI Governance Committee briefing archive; onboarding to the safety protocols (closed-corpus red-team only; no real-PII probes; CBRN-class objective handling protocols; the incident-escalation procedure if a probe accidentally elicits actionable harmful content). Sign-off: Red Team Lead.
- Weeks 3-6 - Shadow + secondary on probes. New hire pairs with a senior team member on two active campaigns. Reads the prior-quarter evidence binders. Reads the prior-year AIGC briefings. Implements one ATLAS technique from primary literature into the existing test environment (build, not greenfield). Co-authors one finding write-up. Sign-off: Red Team Lead + senior pair.
- Weeks 7-10 - First owned campaign with senior review. New hire is the primary author on a quarterly campaign refresh against one model in the inventory. Senior team member reviews. Findings land in the model-risk register with the new hire's name on the entry. Cross-walks to OWASP LLM/Agentic + MITRE ATLAS + Article 15 + (if applicable) Article 55(1)(a) + ISO 42001 are all populated. Sign-off: Red Team Lead.
- Weeks 11-13 - First independent finding write-up + AIGC briefing. New hire authors a finding write-up at the regulator-readable standard; the write-up lands in the next AIGC briefing pack; the new hire presents the finding at the briefing (with the Red Team Lead present for backup). At this point the new hire has evidence-binding authority within the team's normal review process. Sign-off: AIGC chair (implicit, via the briefing pack acceptance).
Two failure modes in the 90-day plan. (1) Skipping weeks 1-2 safety onboarding because the new hire is senior, a senior hire from a frontier lab still needs to learn this organization's escalation procedures and closed-corpus protocols. Skipping this stage is how an internal-red-team accidental data-exfiltration incident lands on the AIGC's desk. (2) Letting weeks 11-13 slip past 90 days, if the new hire's first AIGC briefing happens at month six rather than month three, the team has lost half a quarter of evidence-binding capability and the manager has lost the early-signal opportunity to identify a hiring mismatch.
Interfaces with adjacent functions.
The red team does not operate in isolation. The 2026 mature posture has five named interfaces. (1) With the CAF (Common Audit Framework, lesson 077), the red team feeds findings into the CAF as one of the standardized control-evidence inputs. (2) With the AI Risk Office (2L), the red team reports through the AI Risk Office's governance structure, and the AI Risk Office is the natural escalation path for cross-system or cross-line findings. (3) With MRM / IMV (lesson 068): the red team's adversarial-testing evidence is load-bearing input to the SR 11-7 validation pillars (conceptual soundness, ongoing monitoring, outcomes analysis). (4) With Article 72 PMM (lesson 080): when PMM signals fire (drift breach, incident pattern, user-reported anomaly), the red team is the trigger response team that probes the model under the PMM-flagged condition. (5) With external red teams (vendor-side, contractor, bug-bounty), the internal red team coordinates and ingests external findings; quarterly rotation with external red teams is the recommended cadence to prevent capability stagnation. The L4 governance operating model in lesson 082 codifies these interfaces as named SLAs and escalation paths.
Common Mistakes + 2026 Emerging Expectations
The five mistakes below are the failure modes the AIGC chair and the auditor look for first when reviewing an internal-red-team build.
- Mistake 1 - Hiring only AppSec, missing ML depth. The team that hires six AppSec/pen-test backgrounds without an ML researcher will miss the ML-Security-Researcher-tier vulnerabilities (training-data extraction, model extraction, membership inference, BadNets, weight-level vulnerabilities). The team's findings will skew heavily to prompt-injection and refusal-bypass classes, important but not exhaustive. Audit response will identify the gap. Remediation: add ML researcher on month-three hire slate; pair with existing AppSec hires for 4-6 months of cross-skilling.
- Mistake 2 - Hiring only ML researchers, missing security discipline. The mirror failure mode. The team of six ML researchers without security discipline will produce research-quality probes but will fail at the control-function expectations: finding documentation, severity scoring, regulator-readable reporting, evidence-binding into the model-risk register, AIGC briefing cadence. Reviewer-grade findings will be brilliant; evidence-bindable findings will be missing. Remediation: hire the Analyst role plus a security-discipline-strong Lead from the AppSec community.
- Mistake 3 - Lead reporting to CIO / CTO (independence breach). The Red Team Lead reporting to the CIO is the most common independence-test failure. The CIO accountabilities include the AI systems' uptime, performance, and ship cadence, the same systems the red team is testing. The defensible reporting line is CAIO (when CAIO is in 2L), CRO, or a separate 2.5L red-team function reporting to the CEO via the audit committee. Remediation: re-line at the earliest reorg opportunity; document the rationale in the AIGC committee minutes.
- Mistake 4 - Team below 6 FTE with this scope. A 3-FTE team trying to cover Article 15 + Article 55(1)(a) + SR 11-7 + ISO 42001 + Article 72 PMM trigger response + agentic-systems coverage is on a burnout-and-coverage-gap trajectory. The team will skip the documentation discipline (mistake 2 territory) under load; the most-recently-onboarded member will be the first to attrite; coverage will drift toward the easiest-to-test systems. Remediation: stage the buildout, three FTE in Q1, six FTE by end of Q3, and explicitly de-scope the Article 55(1)(a) coverage in the meantime if the organization is not yet a GPAI provider.
- Mistake 5 - No rotation with external red teams (capability stagnation). An internal red team that never engages external red-teamers will plateau on attack-pattern fluency. The 2026 mature posture rotates with external red teams quarterly: HackerOne AI / Bugcrowd AI engagement once per year; contractor engagement (Trail of Bits, Robust Intelligence, Mayhem, Apollo Research) once per year; AI Village / GRT collective participation continuously. The rotation refreshes the internal team's attack-pattern fluency and provides external-validation triangulation for the internal findings.
2026 emerging expectations.
- EU AI Office Article 89 (information requests). The AI Office is expected through Q3-Q4 2026 to issue red-team evidence-packet templates for GPAI providers, consolidating the Article 55(1)(a) submissions into a standardized format. Internal-red-team capability will be a baseline expectation; organizations without it will be visible as outliers.
- CAISI Agent Standards Initiative (17 Feb 2026 launch, working-group outputs through H2 2026). The Initiative is expected to codify agentic-system red-team protocols, with reference to the OWASP Agentic Top 10 and the published agent-eval benchmarks. Participating organizations will need named agentic-systems-specialist capability.
- Bug-bounty AI scope expansion. HackerOne AI and Bugcrowd AI programs are expected through H2 2026 to expand from prompt-injection scope to agentic-systems scope and to fairness-and-misuse scope. Organizations running bug bounty without internal red-team capability to triage and own remediation will see triage backlogs grow.
- MITRE ATLAS v5.4.0 → v6.0 refresh. The ATLAS knowledge base is on a refresh cadence; v6.0 (expected H1 2027) is anticipated to add agentic-system technique families. Internal red teams will need to refresh technique tagging.
Worked Example - Acme.Corp Q3 2026 6-FTE Hiring Slate, $1.2M Y1 Budget
The Acme.Corp CAIO leaves the Thursday committee read-out with a one-page summary of the hiring slate to be presented at the next AIGC committee meeting in two weeks. The structure below is the Q3 hiring slate she presents: named roles, salary bands, sourcing strategy, onboarding plans for the first two hires.
| Role | Hire month | Salary band (US base) | Primary source | Secondary source |
|---|---|---|---|---|
| Red Team Lead | M2 (Jul 2026) | $240k-$280k | Source 1, frontier-lab alumni | Source 3, CMU CyLab senior |
| ML Security Researcher | M2 (Jul 2026) | $210k-$240k | Source 3, Stanford CRFM postdoc | Source 1, DeepMind alum |
| Adversarial Prompt Engineer | M3 (Aug 2026) | $165k-$195k | Source 2, AI Village GRT | Source 4, HackerOne AI top-25 |
| Agentic Systems Specialist | M3 (Aug 2026) | $200k-$235k | Source 6, internal data-science mobility | Source 3, Berkeley CHAI grad |
| Eval Engineer | M4 (Sep 2026) | $160k-$185k | Source 6, internal ML platform team | Source 7, pen-tester w/ML coursework |
| AI Security Analyst / Reporter | M4 (Sep 2026) | $150k-$175k | Source 5, AIGP-certified candidate | Source 8, OWASP working-group member |
Six-FTE total cash base midpoint: $1,131k. Add: $25k tooling (Promptfoo Cloud + Garak telemetry + API budgets) + $30k training/conferences (DEF CON AI Village + USENIX Security + NeurIPS attendance + SANS / AIGP cert reimbursement) + $40k shared infrastructure (test environment, evidence-binder tooling, ATLAS-aligned tracking) = $1.226M Y1 all-in. The CAIO's defensive frame for the AIGC: the $1.2M Y1 budget compares to Article 99(3) penalty exposure of €15M / 3% of global turnover for an Article 15 + Article 55(1)(a) inadequacy finding, plus the SR 11-7 model-risk-validation expectation, plus the ISO 42001 A.6.2.6 audit requirement, plus the inbound Article 72 PMM trigger-response demand that the operating model in lesson 082 will codify.
Onboarding plans for the first two hires (Red Team Lead, ML Security Researcher; both starting M2).
The Red Team Lead's first 90 days deliver: (a) the full charter + the AIGC briefing cadence; (b) the model-risk inventory triage with named red-team scope per inventory entry; (c) the Q3 quarterly campaign plan with named owners for each of the three systems in scope; (d) the first AIGC briefing at the end of M4. The ML Security Researcher's first 90 days deliver: (a) one ATLAS-technique implementation from primary literature against the test environment; (b) the first ML-security-tier finding write-up on the production fine-tuned model; (c) co-authorship on the M4 AIGC briefing pack. Both hires sign-off at end of M4 with first owned campaign delivered and first AIGC briefing presented.
The slate is sourced through three parallel channels per role with the primary source named in the slate above. Recruiter retainer engaged for the Red Team Lead role (frontier-lab alumni require senior-relationship + executive-search depth); in-house TA for the other five roles with referral bonuses doubled for the duration of the buildout. Quarterly review cadence on the slate, with reforecast at end of Q3 if any role slips by more than four weeks.
Key Takeaways
- An internal AI red team is a 2026 non-negotiable. EU AI Act Article 15 robustness/cybersecurity + Article 55(1)(a) GPAI adversarial-testing + CAISI Agent Standards Initiative (17 Feb 2026) + NIST AI 600-1 Risk 11 advisory red-teaming + SR 11-7 / PRA SS1/23 independent challenge converge on the internal-red-team expectation; external pen-testers can supplement but cannot substitute.
- The three-clause independence test. The team cannot have built the system under test; cannot report to the first line of defense; must have evidence-binding authority into the model-risk register and the Annex IV / FRIA artifacts. The defensible reporting line is CAIO (when CAIO is in 2L), CRO, or a separate 2.5L red-team function.
- Six-FTE baseline + 12 × 6 skills matrix. Red Team Lead, Adversarial Prompt Engineer, ML Security Researcher, Agentic Systems Specialist, Eval Engineer, AI Security Analyst / Reporter, every competency covered by a primary owner plus 2-4 secondary contributors; no role is required to be a unicorn.
- 2026 salary ranges (US base). Lead $200k-$310k; Adversarial Prompt Engineer $145k-$215k; ML Security Researcher $180k-$280k; Agentic Specialist $175k-$260k; Eval Engineer $140k-$200k; Analyst $130k-$185k. Hybrid candidates command 30-50% premium; six-FTE Y1 all-in $1.1-1.6M.
- Eight-source hiring pipeline. Frontier-lab alumni; GRT collective at DEF CON; AI-security academic labs (CMU, Stanford, Berkeley, ETH Zurich, Oxford); bug-bounty top performers; certification-track talent (AIGP, CAIO, SANS); internal mobility; trad-cyber pen-testers with ML coursework; AI Village / OWASP / MITRE ATLAS working-group communities. Work two-or-three sources per role with parallel pipelines.
- The 90-day onboarding plan. Weeks 1-2 environment + access + safety; weeks 3-6 shadow + secondary; weeks 7-10 first owned campaign with senior review; weeks 11-13 first independent finding write-up + AIGC briefing. Sign-off at end of M3 converts the hire into an evidence-binding contributor.
- Five interfaces. CAF (lesson 077); AI Risk Office 2L; MRM/IMV (lesson 068); Article 72 PMM (lesson 080); external red teams. The L4 operating model in lesson 082 codifies these as named SLAs and escalation paths.
- The five common mistakes. AppSec-only (missing ML depth); ML-only (missing security discipline); Lead under CIO/CTO (independence breach); team under 6 FTE with this scope (burnout + coverage gaps); no external-red-team rotation (capability stagnation). The Acme Q3 worked slate hires six FTE at $1.2M Y1 all-in, defensible against Article 99(3) exposure.
Skill.re