Promptfoo logo

Promptfoo Review (2026)

The de-facto open-source CLI for LLM eval and red-teaming, driven by declarative YAML. MIT-licensed, ~23.5k stars, and acquired by OpenAI in March 2026 - still open source, now part of OpenAI Frontier.

Hands-on tested

Rating

4.0

Starting Price

$0

Free Plan

Yes

SDKs & Frameworks

5

Deployment

4

Best For

Security and CI teams who want config-driven LLM eval plus serious red-teaming, from an OSS tool with no seat cost

Last Updated:

10 Things You Should Know About Promptfoo

  1. 1 Acquired by OpenAI on 9 March 2026; both sides state it stays open source and folds into OpenAI "Frontier"
  2. 2 MIT-licensed with roughly 23.5k GitHub stars, the highest of the major eval tools
  3. 3 Ships 50+ red-team attack plugins with OWASP LLM Top 10, OWASP Agentic and NIST presets
  4. 4 Acts as its own OTLP receiver with a built-in trace viewer - no external Jaeger or Tempo
  5. 5 Raised ~$23.4M before acquisition - a16z $5M seed, then an $18.4M Series A led by Insight Partners in July 2025

Pros & Cons

Pros

  • MIT-licensed and free forever - all core eval and red-teaming included
  • The de-facto OSS red-teaming CLI, with the strongest OWASP mapping in the category
  • Declarative YAML means evals live in version control, not scattered scripts
  • Its own OTLP receiver ships a trace viewer, so no Jaeger or Tempo to run
  • Backed and now owned by OpenAI, with a public commitment to stay open source

Cons

  • Now an OpenAI company - long-term OSS governance is the open question for buyers
  • Enterprise and On-Premise are contact-sales only, no public pricing
  • Free red-teaming is capped at 10k probes per month
  • YAML config is less ergonomic than a Python test suite for programmatic teams
  • Some agent-SDK providers still do not export traces to the receiver (issue

Features

Declarative YAML eval configs, not SDK-first
Side-by-side model and prompt comparison
Red-teaming with 50+ attack plugins
Built-in OWASP LLM Top 10, NIST and OWASP Agentic presets
Acts as its own OTLP receiver with a built-in trace viewer
CI/CD integration for eval gates

What Promptfoo actually is

Promptfoo is an open-source command-line tool for two jobs - evaluating LLM apps and red-teaming them for security. You describe your tests in declarative YAML, point it at your prompts and models, and it runs comparisons, assertions and vulnerability scans from the CLI or in CI. It is MIT-licensed, it has roughly 23.5k GitHub stars - the most of the major eval tools - and OpenAI/vendor figures put it at 350k+ developers.

The headline you need first, though, is the ownership. OpenAI announced the acquisition of Promptfoo on 9 March 2026. Both OpenAI and Promptfoo state the project stays open source under its current MIT license, existing customers keep being supported, and the technology folds into OpenAI’s “Frontier” work on agentic security testing. That commitment is on the record. What it is worth to you is a judgment call - you are now betting on OpenAI’s stewardship of an OSS project, and long-term governance under a single large owner is a fair thing to weigh before you build a multi-year workflow on it.

Promptfoo was co-founded in 2024 by Ian Webster and Michael D’Angelo. It raised a $5M seed from a16z, then an $18.4M Series A led by Insight Partners in July 2025 - roughly $23.4M in total before the acquisition.

The distinctive part: YAML config plus red-teaming

Two things set Promptfoo apart from the pytest-style crowd.

It is config-driven, not SDK-first. Your evals are declarative YAML that lives in version control next to your prompts. You are not scattering assertions through Python test files - you are describing what “good” looks like in one place a reviewer can read in a pull request. Security and CI teams tend to love this shape; engineers who want programmatic Python test suites tend to prefer DeepEval’s approach. It is a genuine fork in the road, not a better-or-worse thing.

Red-teaming is the signature specialty. This is where Promptfoo pulls ahead of general eval tools. It ships 50+ attack plugins and built-in presets for the OWASP LLM Top 10, plus OWASP API Top 10, OWASP Agentic and NIST. It probes for harmful content across 20+ subcategories, PII leakage, prompt injection, jailbreaks, excessive agency, hallucination and competitor mentions, and you can write custom policy probes. If security testing is part of your eval story, this is the tool built for it.

There is a third detail worth flagging: Promptfoo acts as its own OTLP receiver with a built-in trace viewer. You send OpenTelemetry traces and Promptfoo displays them - no separate Jaeger or Tempo to run. One caveat from the tracker: issue #7333 reports some agent-SDK providers do not export traces to the receiver, so verify your stack works before you rely on it.

Pricing: free where it counts, opaque above it

TierPriceWhat you getCap
Community$0All core eval, every provider, red-teaming, self-host10k red-team probes/mo
EnterpriseContact salesDashboards, team controls, monitoring, SSO, compliance, SLACustom
On-PremiseContact salesEnterprise features inside your own network, air-gappedCustom

The free tier is the real product for most people. Community is MIT-licensed and free forever, with all core eval and red-teaming included - the only hard limit is 10k red-team probes per month. Gated behind Enterprise are the things bigger teams need: centralized dashboards, team management, continuous monitoring, custom plugins, compliance frameworks and a dedicated runner.

The honest gap: Enterprise and On-Premise have no public pricing - both route to “contact sales.” We could not find a figure, so we are not going to invent one. If you are budgeting, assume a sales conversation.

Promptfoo versus DeepEval

These are the two names people compare, and the split is clean. Promptfoo is YAML-config, red-teaming-led, and free with no seat cost. DeepEval is pytest-style, metrics-led, and pushes you toward the Confident AI cloud at $200/mo for real collaboration. If your team thinks in security tests and CI gates, Promptfoo fits the way you already work. If your team writes Python and wants research-backed scoring metrics like G-Eval, DeepEval is the more natural home.

One more difference that matters post-2026: Promptfoo is now an OpenAI company, while DeepEval’s Confident AI is an independent YC startup. If vendor independence is part of your calculus, that is a point on the board for DeepEval.

Should you use it?

Use Promptfoo if red-teaming or security testing is part of your eval work, you want your evals in version control as declarative config, or you want a genuinely capable CLI with no seat cost. For OWASP-mapped LLM security scanning specifically, nothing else in the open-source category is this complete.

Think twice if you want programmatic Python test suites - DeepEval is more ergonomic for that - or if long-term OSS governance under OpenAI is a dealbreaker for your organization. Neither is a reason not to try it, but both are worth naming before you commit.

Bottom line: it is the strongest open-source eval-plus-red-teaming tool available, the free tier is the real thing, and the OpenAI acquisition is a fact to weigh rather than a reason to walk away. Start with Community and see how far the probe cap gets you.


Pricing and features verified against promptfoo.dev on 23 July 2026. This category ships breaking changes monthly - we re-verify every 30 days.

Pricing Plans

Community

$0

  • MIT-licensed, free forever
  • All core eval and every model provider
  • Vulnerability scanning and red-teaming
  • 10k red-team probes per month
  • Local and self-hosted
Most Popular

Enterprise

Contact sales

  • Centralized dashboard and team controls
  • Continuous monitoring
  • Custom attack profiles and plugins
  • SSO, API access, compliance frameworks
  • SLA support and a dedicated runner

On-Premise

Contact sales

  • Enterprise features inside your own network
  • Air-gapped deployment
  • Custom pricing

SDKs & Frameworks

YAML config (declarative) Node / CLI OpenAI, Anthropic, Azure, Bedrock Gemini, Ollama and many more providers OpenTelemetry (OTLP) tracing

Deployment

Community edition - free, MIT, local or self-hosted Enterprise cloud (contact sales) On-Premise / air-gapped (contact sales) CI/CD pipelines

Eval Methods

Prompt and model comparison Assertions and LLM-as-judge Red-teaming with 50+ attack plugins OWASP LLM Top 10 presets PII, prompt injection, jailbreak probes

Our Verdict

The clearest pick if red-teaming and security testing are part of your eval story, and its YAML-in-git approach fits CI cleanly. The OpenAI acquisition in March 2026 is the thing to weigh - the commitment to stay MIT is on the record, but you are now betting on OpenAI's stewardship of the project. For the free tier alone it is still an easy yes.

Similar Tools

Frequently Asked Questions

Is Promptfoo still open source after the OpenAI acquisition?

Yes. OpenAI announced the acquisition on 9 March 2026, and both OpenAI and Promptfoo state the project remains open source under its current MIT license, with existing customers still supported. The technology folds into OpenAI's "Frontier" work on agentic security testing. The honest caveat - long-term governance of an OSS project under a single large owner is a real question, so weigh it if you are betting a multi-year workflow on it.

Can I self-host Promptfoo?

Yes, and the Community edition is built for it. It is MIT-licensed and runs entirely local or self-hosted, with all core eval and red-teaming included and a 10k red-team probe monthly cap. Centralized dashboards, team management, continuous monitoring and compliance frameworks are the Enterprise and On-Premise features, both contact-sales.

Does Promptfoo support OpenTelemetry?

Yes, and it goes further than most. Promptfoo supports OTLP tracing and acts as its own OTLP receiver with a built-in trace viewer, so you do not need to stand up Jaeger or Tempo to see traces. One known limitation - a GitHub issue (#7333) reports some agent-SDK providers do not export traces to the receiver.

How is Promptfoo different from DeepEval?

Approach. Promptfoo is config-driven - you write declarative YAML - which security and CI teams like because evals live in version control. DeepEval is SDK-first and pytest-style, which suits engineers who want programmatic Python test suites. Promptfoo also leads on red-teaming and OWASP mapping; DeepEval leads on research-backed scoring metrics.

How much does Promptfoo cost?

The Community edition is free forever under MIT, capped at 10k red-team probes per month. Enterprise and On-Premise are both contact-sales with no public price, adding centralized dashboards, SSO, compliance, SLA support and a dedicated runner. If your eval and red-teaming fit within the probe cap, you can run the real product at no cost.

Related Articles

alternatives

4 Promptfoo Alternatives for a Vendor-Neutral Eval Stack in 2026

Promptfoo is the de-facto open-source eval and red-teaming CLI, MIT-licensed with the most GitHub stars of the major eval tools - but OpenAI acquired it in March 2026. If you want a vendor-neutral eval framework, here are the alternatives matched to why teams look.

July 26, 2026

guide

Promptfoo Pricing in 2026 - What's Actually Free and When You Pay

Promptfoo is MIT-licensed and free forever, with one hard cap - 10k red-team probes a month. Here's how the pricing really works, the contact-sales gap above the free tier, and two eval tools with public pricing when you outgrow it.

July 26, 2026

comparison

Promptfoo vs Langfuse in 2026 - Which One You Actually Need

Promptfoo is a config-driven eval and red-teaming CLI. Langfuse is a self-hostable observability backend. They get compared constantly, but they solve different problems - here is which one fits your job, and when you want both.

July 26, 2026

comparison

DeepEval vs Promptfoo in 2026 - Pytest or YAML for LLM Evals

DeepEval is pytest-style, SDK-first, and metrics-led. Promptfoo is YAML-config, CLI-driven, and red-teaming-led. Both are free and open source. Here is which one fits your team, and where Braintrust beats both.

July 26, 2026

comparison

DeepEval vs Promptfoo vs Braintrust in 2026 - The Eval Tool Showdown

Three eval tools, three philosophies. DeepEval is pytest for LLM apps, Promptfoo is YAML-driven red-teaming, and Braintrust is the turnkey regression platform. Here is which one fits your team and where each one bites.

July 26, 2026

best-of

The Best LLM Eval Frameworks in 2026, Ranked for How You Actually Test

Four LLM eval frameworks judged on the fork in the road that decides your workflow - pytest-style SDK, declarative YAML, turnkey CI gates, or eval bolted onto observability. Plus the billing and ownership gotchas each one hides.

July 26, 2026

how-to

How to Red-Team an LLM in 2026 - A Step-by-Step Workflow

Red-teaming an LLM is not random prompt-poking - it is a repeatable pipeline of an attack taxonomy, an adversarial dataset and automated scans you rerun on every change. Here is the exact workflow, plus the two OSS tools that ship the attacks so you are not inventing jailbreaks by hand.

July 28, 2026

glossary

OWASP Top 10 for LLM Applications Explained (2026)

A plain-English walkthrough of all ten OWASP Top 10 risks for LLM applications - what each one actually means, a concrete example, and how eval, red-teaming and guardrail tools help you catch or mitigate it.

July 28, 2026

comparison

Braintrust vs DeepEval in 2026 - The Honest Eval Platform Comparison

Braintrust is the turnkey eval platform with CI quality gates that block bad merges. DeepEval is pytest for LLM apps, free and open source. Here is which one fits your team, and where each one bites.

July 26, 2026