The Eval-and-Red-Team Platform Consolidation: Promptfoo, OpenAI, and the Build-vs-Buy Inflection
On March 9, 2026, OpenAI acquired Promptfoo. The open-source CLI and library remain open source; the enterprise red-team product and Frontier integrations move inside OpenAI. The acquisition is the largest reshaping of the agent-evals landscape since the category emerged in 2023. Every architect building agents now factors three options into their 12-month commitment: stay on open-source Promptfoo, ride OpenAI's enterprise integration, or hedge with Braintrust / Confident AI / Maxim AI. This lesson walks through the inflection point, the decision criteria, and the two-page decision memo your team should write before locking in a year of eval-and-red-team spend.
Why This Acquisition Mattered
Promptfoo, founded in 2023, became the de facto open-source CLI for LLM evaluation. By late 2025, Promptfoo had:
- The most-downloaded open-source LLM-eval library (3M+ weekly npm + PyPI installs combined).
- A growing red-team product (Promptfoo Cloud) with 800+ enterprise customers including most Fortune 500 names that publicly disclosed agent deployments.
- A plugin ecosystem - the privilege-escalation, destructive-action, and indirect-injection plugins covered in earlier lessons in this chapter were Promptfoo plugins.
- Deep integrations with the major eval-platform competitors (LangSmith, Braintrust, Weights & Biases, Arize, Confident AI's DeepEval).
On March 9, 2026, OpenAI announced the acquisition. The terms publicly disclosed:
- The open-source CLI and library stay open source. Apache 2.0 license. No new restrictions on use. Continued development. This was the load-bearing commitment that prevented the immediate exodus of the open-source user base.
- The enterprise red-team product moves inside OpenAI. Promptfoo Cloud's red-team capabilities become part of OpenAI's safety-evaluation stack, integrated with the Frontier program and the Preparedness team's evaluations.
- Existing enterprise customers transition. Customers on Promptfoo Cloud have a 12-month migration window to either move to OpenAI's enterprise platform, stay on a vendor-supported version of the open-source CLI, or migrate to a competitor.
- The plugin ecosystem continues. Third-party plugins (including red-team plugins like Mindgard's Promptfoo integration) continue to work; OpenAI commits to plugin-API stability.
The acquisition immediately created three strategic options for every team running agent evals: stay on open source, move to OpenAI, or hedge to a competitor. This lesson is about how to choose.
The Three Options
Option 1: Stay on open-source Promptfoo
The CLI and library remain Apache 2.0. You self-host the eval orchestration. The plugin ecosystem still works. The acquisition doesn't reduce your current functionality.
Pros: Zero vendor lock-in. Continued open-source development (OpenAI has committed to this, and the community has forked-but-mainline behavior built in). Costs scale with your infrastructure, not per-seat licensing. You retain the option to switch vendors at any time without data migration.
Cons: No managed red-team service. You build the red-team infrastructure yourself (scheduling, result storage, alerting). For teams without an ML/security operations capability, this is non-trivial. Enterprise features that were in Promptfoo Cloud (SSO, RBAC, audit logs, SLA-backed result delivery) become DIY.
Best for: Teams with strong infra/security operations, multi-vendor model deployment, regulatory constraints that prevent sending data to external services, or teams that want maximum optionality.
Option 2: Move to OpenAI's enterprise integration
The hosted red-team product, now part of OpenAI, with Frontier program integration. OpenAI's safety-evaluation stack and Preparedness-team rigor become available to enterprise customers.
Pros: Managed service. The most rigorous safety evaluations available outside Anthropic's and Google's internal programs. Native integration with OpenAI models. Continuous updates as new attack classes emerge. Frontier-program-grade red-teaming on demand.
Cons: Vendor lock-in to OpenAI's ecosystem. If you also run Anthropic, Google, or open-source models, the evaluations are still possible but the integration is less first-party. Pricing has moved into OpenAI's enterprise tier (no public list price, but anchored to OpenAI's typical enterprise quotes). Data residency considerations for non-US customers.
Best for: Teams primarily or exclusively on OpenAI models, regulated industries that benefit from Frontier-program rigor, organizations that prefer managed-service security operations.
Option 3: Hedge with a competitor (Braintrust, Confident AI, Maxim AI)
The eval-platform competitor landscape didn't stand still during the acquisition. Three vendors moved fast to capture customers who didn't want OpenAI dependency.
Braintrust - the most-mature enterprise eval platform after Promptfoo. Strong on multi-model (anthropic, OpenAI, Google, open-source), strong on UI/observability, expanding red-team in 2026.
Confident AI - the company behind DeepEval (open-source). Took Promptfoo's exit as an opening; aggressive customer acquisition. Heavy investment in agentic-eval (multi-turn, tool-using, RAG-grounded).
Maxim AI - smaller, more specialized. Strong on observability-eval-redteam fusion. Picking off teams that want one integrated platform across observability and eval.
Pros: Vendor independence from any single model provider. Competitive feature roadmaps as vendors fight for share. Pricing under pressure - vendors are willing to discount aggressively to win Promptfoo refugees.
Cons: Smaller scale than OpenAI; less rigorous safety-evaluation depth. Migration cost from Promptfoo (data + scripts + plugins). Some of the Promptfoo plugins haven't yet been ported to these platforms.
Best for: Multi-model teams who want vendor-independent eval infrastructure, teams who value vendor competition and ongoing pricing pressure, teams on Anthropic/Google models primarily.
The Decision Criteria
The decision isn't one-dimensional. Five criteria, weighted by your team's context:
1. Model exposure
Where are your models actually running? If you're 100% OpenAI, the OpenAI integration becomes very attractive - first-party tooling, Frontier-program access, single vendor relationship. If you're multi-model (which most enterprise teams are by 2026), the OpenAI integration loses some of its edge; a vendor-neutral platform (Braintrust, Confident AI) or open-source DIY (Promptfoo CLI) starts to compete.
2. Red-team rigor required
What level of safety-evaluation rigor does your domain demand? A consumer-facing chatbot in a low-stakes domain can get by with the open-source Promptfoo plugins. A financial-services agent that touches customer money needs Frontier-grade evaluations. A medical-advisory agent should run continuous red-teaming with specialist-vendor depth (Mindgard, HiddenLayer, or OpenAI's enterprise tier). Match the rigor to the stakes.
3. Ops capability
Do you have the team to run self-hosted eval infrastructure? Self-hosted Promptfoo isn't expensive in dollars, but it's not free in operations time. Scheduling, result storage, alerting, on-call for eval failures, dashboard maintenance. If your team is 3 people, this is a real cost. If you're a 30-person AI platform team, this is a fraction of someone's job.
4. Data residency and regulatory
Can your eval data leave your environment? Regulated industries (finance, health, public sector) often can't send raw user prompts or production traces to a vendor's cloud. Self-hosted open-source is the default in these cases. Some vendors offer dedicated tenants, BYOC (bring your own cloud) deployments, or on-prem options - but these vary by vendor and tier.
5. Plugin and integration needs
Which red-team plugins does your stack require? If you've built workflows around specific Promptfoo plugins (privilege-escalation, destructive-action, indirect-injection, custom internal plugins), check whether the alternative supports the same plugin or has an equivalent. Migration cost can be days or weeks depending on integration depth.
The Two-Page Decision Memo
The artifact your team should produce is a two-page decision memo. The memo's job: force a structured decision before drift sets in. The template:
Page 1 - Context and decision
- Current state. Which eval platform are we using today? What is our annual spend? What is in our red-team test set? Who owns it? (3-4 sentences.)
- The Promptfoo acquisition implication for us. Specifically: what does it change for our deployment? (3-4 sentences.)
- Our model exposure. Percent OpenAI / Anthropic / Google / open-source. The next 12-month projection. (Table.)
- Our red-team rigor requirement. Domain stakes, regulatory context, current incident history. (4-6 sentences.)
- Decision. Which of the three options? Why this option over the other two? (5-8 sentences.)
Page 2 - Implementation and risk
- Migration path. If changing platforms, what's the sequence? Who owns each step? When does each step land?
- Cost projection. 12-month total cost (licenses + ops time + migration cost) for the chosen option vs. status quo.
- Reversibility plan. If the chosen option doesn't work out, what's the exit path? Cost of unwinding?
- Key risks and mitigations. 3-5 risks with concrete mitigations. (E.g., "OpenAI raises enterprise pricing during contract" → "Negotiate price-protection clause in initial contract; reserve right to move to open-source CLI if unilateral price increase exceeds X%.")
- Decision-revisit date. When will we re-evaluate? (Recommend: 6 months out, with light-touch monthly checkpoint.)
The memo's audience is everyone in the loop on the decision: engineering leadership, security, the architect who owns the agent stack, finance for cost validation. The two-page limit is enforced. If you can't fit the decision in two pages, you haven't actually decided.
The Build vs Buy Inflection
The Promptfoo acquisition crystallizes a build-vs-buy inflection that was already happening. Before March 9, 2026, the question was "use Promptfoo open-source CLI or pay for Promptfoo Cloud's managed service." Now the question reshapes:
- Buy from OpenAI = bet on OpenAI's roadmap, accept OpenAI lock-in, get the most rigorous safety evaluations.
- Buy from a competitor = bet on competitor's roadmap, accept different lock-in, get vendor-neutral but less rigorous evals.
- Build on open-source CLI = bet on your team's capability, accept ops overhead, get maximum optionality.
The right answer depends on where you are in the agent-deployment maturity curve. Early-stage teams should usually buy - the ops overhead of self-hosted is high relative to the value at early scale. Mature teams should evaluate buy vs build more seriously - at scale, the ops overhead is small relative to license costs, and the optionality of open-source is more valuable.
The inflection isn't "everyone should buy" or "everyone should build." It's "the choice landscape has changed and you should explicitly re-decide."
The Anthropic and Google Positioning
Two other frontier-model providers' responses to the Promptfoo acquisition matter for your decision:
Anthropic announced in April 2026 an expansion of its Trust & Safety evals as an enterprise offering accessible alongside Claude API usage. The pitch: if you're on Claude, you can use Anthropic's internal eval rigor without going to OpenAI. The product is less mature than OpenAI's Promptfoo-fused offering as of May 2026, but it's a credible alternative for Anthropic-centric teams.
Google integrated its responsible-AI evals more tightly with Vertex AI Studio in Q2 2026. For Google Cloud / Gemini customers, the eval and red-team capabilities became more first-party. The depth is less than OpenAI's, but the integration is tight for teams already on Google Cloud.
For multi-model teams, this means three frontier-provider eval offerings (OpenAI/Promptfoo, Anthropic, Google) plus vendor-neutral platforms (Braintrust, Confident AI, Maxim AI) plus the open-source CLI. The choice is denser than it was in 2025. The decision memo helps cut through the density.
Continuous Red-Teaming as the Real Discipline
Step back from the platform question. The deeper truth: red-teaming is a continuous discipline, not a one-time activity. Whichever platform you choose, the practice is what matters.
The practice in 2026:
- A fixed red-team test set that ships with the agent. Pulled from Garak, Pyrit, Promptfoo plugins, and your internal incidents. Runs in CI on every PR that touches agent prompts, tool registry, or model. Target: 100% pass rate.
- A continuous red-team set that runs daily or weekly against production traffic patterns. Catches novel attacks that show up in real traffic. Reviewed weekly by the team that owns the agent.
- An external red-team partnership - either a security-specialist vendor (Mindgard, HiddenLayer, Lakera) or a periodic engagement with a specialist firm. Brings adversarial creativity your internal team can't replicate.
- Incident-driven test expansion. When an incident occurs, the test that would have caught it is added to the fixed set. The set grows monotonically over the agent's lifetime.
- Quarterly red-team review. What's the pass rate trend? Which categories regressed? Are there new attack classes in the wild that we don't yet test? Re-baseline the priorities.
The platform is the tool. The practice is the work. The acquisition reshapes the tool landscape; the practice continues regardless.
What the Acquisition Means for Individual Builders
If you're a builder rather than an architect — building agents for your team, not setting platform strategy — the acquisition is mostly continuity. The Promptfoo CLI you use locally is unchanged. The plugins you've come to rely on still work. Your CI workflow doesn't change.
What does change:
- The hosted Promptfoo Cloud, if you were on it, is moving inside OpenAI. Your manager will need to decide whether to stay on the OpenAI-hosted version, move to a competitor, or move to self-hosted.
- The enterprise-tier features (SSO, RBAC, hosted result storage) that were behind Promptfoo Cloud are now OpenAI-tier. Pricing will change. Procurement will care.
- The plugin ecosystem evolution will reflect OpenAI's priorities. Plugins targeting OpenAI-specific failure modes will get more attention; plugins targeting other models will continue but with less first-party support.
For a builder, the practical advice: keep using Promptfoo for local development and CI. Watch the platform-strategy conversation for your team's decision. Be ready to migrate your eval scripts to whatever platform your team picks - the eval shape is portable across platforms, but the orchestration layer isn't.
The 12-Month Commitment Horizon
Why a 12-month horizon for this decision? Because three things are moving on roughly that timescale:
- OpenAI's enterprise integration matures. The Frontier-program integration, the Preparedness-team evals, the hosted red-team product - these will be substantially better in 12 months than they are in May 2026. A commitment now is a commitment to a less-mature product than the same commitment in 12 months would be.
- The competitor landscape consolidates. Of Braintrust, Confident AI, Maxim AI - and the platforms not listed - at least one will be acquired, merge, or pivot in the next 12 months. The vendor you pick today may not be the vendor 18 months from now.
- The open-source CLI evolves. OpenAI's stewardship of Promptfoo open-source will be tested over 12 months. Either OpenAI invests substantially in the open-source CLI (validating the bet) or the community forks (validating the alternative bet). The verdict will be clear by mid-2027.
A 12-month commitment with a 6-month checkpoint is the right cadence. Commit hard enough to actually use the platform; revisit before lock-in becomes structural.
The Deeper Pattern
The Promptfoo-OpenAI acquisition is the latest case study of a recurring pattern in the agent-tools landscape: open-source projects build category-defining products, get acquired by frontier-model providers, and the open-source layer becomes the lock-in escape hatch.
The same pattern played out with LangChain (which raised but didn't sell), with Hugging Face (which partnered but stayed independent), with Pinecone (which scaled to enterprise on the back of vector-search open-source). The pattern in 2026: the open-source layer is the strategic insurance. You build on it because you can switch providers. You use the managed service because it's faster than running infrastructure. You watch the M&A landscape because it tells you which managed services will reshape under you.
For agent builders, the discipline is to build on portable abstractions wherever possible. Your eval test cases should be portable. Your red-team probes should be portable. Your tool registry should be portable. The orchestration layer can be platform-specific; the work product should not be.
Key Takeaways
- OpenAI acquired Promptfoo on March 9, 2026. The open-source CLI/library stays open source (Apache 2.0); the enterprise red-team product and Frontier-program integration move inside OpenAI. Existing Promptfoo Cloud customers have a 12-month migration window.
- Three options for every team: stay on open-source Promptfoo (zero lock-in, DIY ops), move to OpenAI's enterprise integration (managed service, OpenAI lock-in, Frontier-grade evals), or hedge with Braintrust / Confident AI / Maxim AI (vendor-independent, less rigorous depth).
- Five decision criteria: model exposure, red-team rigor required, ops capability, data residency / regulatory, plugin / integration needs.
- The two-page decision memo: page 1 (context + decision), page 2 (implementation + risk). Forces explicit re-decision before drift sets in.
- Anthropic and Google responded with their own eval offerings (Anthropic Trust & Safety as enterprise, Google's responsible-AI evals in Vertex AI Studio). Multi-model teams now have three frontier-provider options plus vendor-neutral platforms plus open-source.
- The deeper practice is continuous red-teaming: fixed CI set + continuous traffic set + external partnership + incident-driven expansion + quarterly review. The platform is the tool; the practice is the work.
- For individual builders, the local Promptfoo workflow is mostly unchanged. The hosted-Promptfoo-Cloud decisions land with managers. Plugin ecosystem evolves toward OpenAI priorities.
- 12-month commitment horizon with 6-month checkpoint - OpenAI's integration matures, the competitor landscape consolidates, and the open-source stewardship is tested on roughly that timescale.
- The deeper pattern: build on portable abstractions (test cases, probes, tool registry). Orchestration layer can be platform-specific; the work product cannot.
Skill.re