Every LLM eval platform, running the same app
We instrumented one identical application across the major platforms and measured what each captures, what it misses, and what it really costs per million spans. Every other tool in the category is researched from primary sources and labelled as such. No sponsors. No ads.
Editor's picks
Best open source
Langfuse
5.0Free, MIT-licensed, and genuinely feature-complete when self-hosted - the open-source default.
Best value cloud
Opik
5.0The cheapest paid cloud in the category at $19/mo, with a permissive Apache-2.0 self-host option.
Best for evals
Braintrust
4.0The most turnkey for regression testing and dataset-driven evals - just watch the processed-data billing.
All 73 tools
Inspect AI
Free · Eval Frameworks
Langfuse
Free · Observability & Tracing Tested
LangWatch
Free · Agent Evaluation
LiteLLM
Free · LLM Gateways
Opik
Free · Observability & Tracing Tested
Pydantic Logfire
Free · Observability & Tracing
Agenta
Free · Prompt Management
AgentOps
Free · Agent Evaluation
Arize AX
Free · Observability & Tracing
Arize Phoenix
Free · Observability & Tracing Tested
Arthur
Free · Guardrails & Safety
Azure AI Foundry Evaluation
Model inference cost only · Agent Evaluation
Braintrust
Free · Eval Frameworks Tested
Cloudflare AI Gateway
Free · LLM Gateways
Confident AI
Free · Eval Frameworks
Databricks Agent Evaluation
Databricks consumption · Agent Evaluation
Confident AI (DeepEval)
Free · Eval Frameworks Tested
Evidently
Free · Eval Frameworks
Fiddler AI
Not published · Guardrails & Safety
Freeplay
Not published · Prompt Management
Galileo
Free · Observability & Tracing Tested
Giskard
Free · Eval Frameworks
Guardrails AI
Free · Guardrails & Safety
HiddenLayer
Not published · Guardrails & Safety
HoneyHive
Free · Observability & Tracing
Lakera
Free · Guardrails & Safety
Laminar
Free · Observability & Tracing Tested
Latitude
Free · Prompt Management
LM Evaluation Harness
Free · Eval Frameworks
MLflow
Free · Observability & Tracing
NVIDIA NeMo Guardrails
Free · Guardrails & Safety
Not Diamond
Free · LLM Gateways
OpenLIT
Free · Observability & Tracing
OpenRouter
Free · LLM Gateways
Patronus AI
Not published · Eval Frameworks
Portkey
Free · LLM Gateways Tested
Promptfoo
Free · Eval Frameworks Tested
PromptLayer
Free · Prompt Management
Ragas
Free · Eval Frameworks
SigNoz
Free · Observability & Tracing
TrueFoundry
Free · LLM Gateways
Vercel AI Gateway
Free · LLM Gateways
Vertex AI Gen AI Evaluation Service
Per token plus GCP compute · Agent Evaluation
W&B Weave
Free · Observability & Tracing
Maxim AI
Free · Agent Evaluation Tested
Aporia
Not published · Guardrails & Safety
Amazon Bedrock Evaluations
Per evaluation job · Agent Evaluation
CalypsoAI
Not published · Guardrails & Safety
Datadog LLM Observability
Free · Observability & Tracing
Kong AI Gateway
Free · LLM Gateways
LangSmith
Free · Observability & Tracing Tested
Langtail
Free · Prompt Management
Langtrace
Free · Observability & Tracing
Lunary
Free · Observability & Tracing
Openlayer
Not published · Eval Frameworks
Parea AI
Free · Prompt Management
Prompt Security
Not published · Guardrails & Safety
PromptHub
Free · Prompt Management
Traceloop
Free · Observability & Tracing
TruLens
Free · Eval Frameworks
UpTrain
Free · Eval Frameworks
Helicone
Free · Observability & Tracing Tested
LLM Guard
Free · Guardrails & Safety
Martian
Not offered · LLM Gateways
New Relic AI Monitoring
Free · Observability & Tracing
OpenAI Evals
Free · Eval Frameworks
Pezzo
Free · Prompt Management
Sentry
Free · Observability & Tracing
WhyLabs
Free · Observability & Tracing
Baserun
Unavailable · Prompt Management
Gentrace
Discontinued · Eval Frameworks
Literal AI
Discontinued · Observability & Tracing
Vellum
Not published · Prompt Management
By category
Eval Frameworks
15 toolsLibraries for scoring model and agent output.
Observability & Tracing
22 toolsTrace, log and monitor LLM applications in production.
Agent Evaluation
7 toolsTools built specifically for multi-step agent traces.
LLM Gateways
9 toolsRouting, caching and cost control proxies.
Prompt Management
10 toolsVersion, test and deploy prompts.
Guardrails & Safety
10 toolsRuntime validation, PII and injection defence.
Top 12 comparison
| Tool | Rating | Free | Price | Category | SDKs & Frameworks |
|---|---|---|---|---|---|
Inspect AI | 5.0 | Yes | $0 (open source) | Eval Frameworks | 4+ |
Langfuse | 5.0 | Yes | $29/mo | Observability & Tracing | 7+ |
LangWatch | 5.0 | Yes | Event-based, rates not published | Agent Evaluation | 4+ |
LiteLLM | 5.0 | Yes | $0 self-hosted | LLM Gateways | 3+ |
Opik | 5.0 | Yes | $19/mo | Observability & Tracing | 5+ |
Pydantic Logfire | 5.0 | Yes | $49/mo | Observability & Tracing | 7+ |
Agenta | 4.0 | Yes | $0 self-hosted | Prompt Management | 4+ |
AgentOps | 4.0 | Yes | $40/mo | Agent Evaluation | 3+ |
Arize AX | 4.0 | Yes | $50/mo | Observability & Tracing | 7+ |
Arize Phoenix | 4.0 | Yes | $0 | Observability & Tracing | 5+ |
Arthur | 4.0 | Yes | $60/mo | Guardrails & Safety | 3+ |
Azure AI Foundry Evaluation | 4.0 | - | Model inference cost only | Agent Evaluation | 4+ |
How we test
We cover this category at two depths, and we label every page so you always know which one you are reading.
Hands-on tested (12 platforms)
We instrument a single identical LLM application - the same traces, the same eval set, the same volume - across every platform, then compare what each one captures, how it scores, what it costs at realistic span volumes, and how much work self-hosting actually takes. Same app, same day, published methodology. These pages carry a Hands-on tested badge.
Researched (61 platforms)
The category is far larger than any team can instrument at once, and a tool nobody has written up honestly is still worth knowing about. These pages are built from primary sources - vendor pricing pages, documentation, licenses, funding records and acquisition filings - and we state plainly when we could not verify something rather than filling the gap. They carry a Researched badge and no performance numbers we did not measure ourselves.
Researched entries get promoted to tested over time. The distinction exists so that a directory listing is never mistaken for a benchmark.
What to look for
- Corporate status - is the vendor alive, acquired, or quietly in maintenance mode
- The real billing unit - spans, traces, seats, events or GB ingested change the bill by orders of magnitude
- License terms - MIT, Apache 2.0 and AGPL 3.0 have very different commercial consequences
- Self-hosting - whether it genuinely exists, and what is gated behind the cloud tier
- Pricing - what it actually costs at your volume, not the headline number
Categories we cover
- Observability & Tracing - Trace, log and monitor LLM applications in production.
- Eval Frameworks - Libraries for scoring model and agent output.
- Prompt Management - Version, test and deploy prompts.
- Guardrails & Safety - Runtime validation, PII and injection defence.
- LLM Gateways - Routing, caching and cost control proxies.
- Agent Evaluation - Tools built specifically for multi-step agent traces.
Our recommendation
There's no single best tool - it depends on team size, stack, and budget. Browse the guides for detailed recommendations by team type.
Popular guides
View all →LLM Evaluation Guide: Metrics, Methods and Workflow
comparison
AI Agent Observability with Langfuse: 2026 Guide
review
Evaluation of LLM Applications: A Practical 2026 Guide
guide
10 Observability Signals for Multi-Step LLM Systems
comparison
BLEU vs ROUGE vs BERTScore - Which to Use and Why All Three Fail on Chat
guide
Context Precision vs Recall Explained - Diagnosing RAG Retrieval in 2026
guide
The Faithfulness Metric Explained - How to Catch RAG Hallucination in 2026
guide
G-Eval Explained - How Chain-of-Thought LLM Scoring Works in 2026
guide
Latest articles
View all →LLM Evaluation Guide: Metrics, Methods and Workflow
AI Agent Observability with Langfuse: 2026 Guide
Evaluation of LLM Applications: A Practical 2026 Guide
10 Observability Signals for Multi-Step LLM Systems
BLEU vs ROUGE vs BERTScore - Which to Use and Why All Three Fail on Chat
Context Precision vs Recall Explained - Diagnosing RAG Retrieval in 2026
The Faithfulness Metric Explained - How to Catch RAG Hallucination in 2026
G-Eval Explained - How Chain-of-Thought LLM Scoring Works in 2026
Free Newsletter
Get the LLM Evals Newsletter
Platform comparisons, pricing changes and eval technique deep-dives. No spam.
Frequently Asked Questions
How do you test these platforms?
We instrument a single identical LLM application - the same traces, the same eval set, the same volume - across the platforms we test hands-on, then compare what each one captures, how it scores, what it costs at realistic span volumes, and how much work self-hosting actually takes. Same app, same day, published methodology. Those pages carry a Hands-on tested badge.
What is the difference between a tested and a researched page?
Tested means we instrumented the platform with our reference application and measured it ourselves. Researched means we verified the page against primary sources - vendor pricing pages, documentation, licenses, funding records and acquisition filings - but have not instrumented it yet. Researched pages never carry performance numbers we did not measure, and they say plainly when we could not verify a claim. Every page is labelled, so you always know which one you are reading. The category is far bigger than any team can instrument at once, and a tool nobody has written up honestly is still worth knowing about.
Are these comparisons sponsored?
No. No platform pays for placement or for a rating. Sponsored listings, where they exist, are clearly labelled as sponsored and never affect rankings or scores.
How often is this updated?
This category ships breaking changes monthly, so we re-verify pricing and feature claims every 30 days and re-run the full instrumentation test quarterly. Every page shows a Last Verified date.
Which platforms can genuinely be self-hosted?
Several vendors advertise self-hosting but gate meaningful features behind their cloud tier. We test the actual self-hosted build of every platform that claims to support it and document exactly what you lose.