73 tools tracked · 12 hands-on tested

Every LLM eval platform, running the same app

We instrumented one identical application across the major platforms and measured what each captures, what it misses, and what it really costs per million spans. Every other tool in the category is researched from primary sources and labelled as such. No sponsors. No ads.

All 73 tools

Inspect AI logo

Inspect AI

Free · Eval Frameworks

5.0
Langfuse logo

Langfuse

Free · Observability & Tracing Tested

5.0
LangWatch logo

LangWatch

Free · Agent Evaluation

5.0
LiteLLM logo

LiteLLM

Free · LLM Gateways

5.0
Opik logo

Opik

Free · Observability & Tracing Tested

5.0
Pydantic Logfire logo

Pydantic Logfire

Free · Observability & Tracing

5.0
Agenta logo

Agenta

Free · Prompt Management

4.0
AgentOps logo

AgentOps

Free · Agent Evaluation

4.0
Arize AX logo

Arize AX

Free · Observability & Tracing

4.0
Arize Phoenix logo

Arize Phoenix

Free · Observability & Tracing Tested

4.0
Arthur logo

Arthur

Free · Guardrails & Safety

4.0
Azure AI Foundry Evaluation logo

Azure AI Foundry Evaluation

Model inference cost only · Agent Evaluation

4.0
Braintrust logo

Braintrust

Free · Eval Frameworks Tested

4.0
Cloudflare AI Gateway logo

Cloudflare AI Gateway

Free · LLM Gateways

4.0
Confident AI logo

Confident AI

Free · Eval Frameworks

4.0
Databricks Agent Evaluation logo

Databricks Agent Evaluation

Databricks consumption · Agent Evaluation

4.0
Confident AI (DeepEval) logo

Confident AI (DeepEval)

Free · Eval Frameworks Tested

4.0
Evidently logo

Evidently

Free · Eval Frameworks

4.0
Fiddler AI logo

Fiddler AI

Not published · Guardrails & Safety

4.0
Freeplay logo

Freeplay

Not published · Prompt Management

4.0
Galileo logo

Galileo

Free · Observability & Tracing Tested

4.0
Giskard logo

Giskard

Free · Eval Frameworks

4.0
Guardrails AI logo

Guardrails AI

Free · Guardrails & Safety

4.0
HiddenLayer logo

HiddenLayer

Not published · Guardrails & Safety

4.0
HoneyHive logo

HoneyHive

Free · Observability & Tracing

4.0
Lakera logo

Lakera

Free · Guardrails & Safety

4.0
Laminar logo

Laminar

Free · Observability & Tracing Tested

4.0
Latitude logo

Latitude

Free · Prompt Management

4.0
LM Evaluation Harness logo

LM Evaluation Harness

Free · Eval Frameworks

4.0
MLflow logo

MLflow

Free · Observability & Tracing

4.0
NVIDIA NeMo Guardrails logo

NVIDIA NeMo Guardrails

Free · Guardrails & Safety

4.0
Not Diamond logo

Not Diamond

Free · LLM Gateways

4.0
OpenLIT logo

OpenLIT

Free · Observability & Tracing

4.0
OpenRouter logo

OpenRouter

Free · LLM Gateways

4.0
Patronus AI logo

Patronus AI

Not published · Eval Frameworks

4.0
Portkey logo

Portkey

Free · LLM Gateways Tested

4.0
Promptfoo logo

Promptfoo

Free · Eval Frameworks Tested

4.0
PromptLayer logo

PromptLayer

Free · Prompt Management

4.0
Ragas logo

Ragas

Free · Eval Frameworks

4.0
SigNoz logo

SigNoz

Free · Observability & Tracing

4.0
TrueFoundry logo

TrueFoundry

Free · LLM Gateways

4.0
Vercel AI Gateway logo

Vercel AI Gateway

Free · LLM Gateways

4.0
Vertex AI Gen AI Evaluation Service logo

Vertex AI Gen AI Evaluation Service

Per token plus GCP compute · Agent Evaluation

4.0
W&B Weave logo

W&B Weave

Free · Observability & Tracing

4.0
Maxim AI logo

Maxim AI

Free · Agent Evaluation Tested

3.5
Aporia logo

Aporia

Not published · Guardrails & Safety

3.0
Amazon Bedrock Evaluations logo

Amazon Bedrock Evaluations

Per evaluation job · Agent Evaluation

3.0

CalypsoAI

Not published · Guardrails & Safety

3.0
Datadog LLM Observability logo

Datadog LLM Observability

Free · Observability & Tracing

3.0
Kong AI Gateway logo

Kong AI Gateway

Free · LLM Gateways

3.0
LangSmith logo

LangSmith

Free · Observability & Tracing Tested

3.0
Langtail logo

Langtail

Free · Prompt Management

3.0
Langtrace logo

Langtrace

Free · Observability & Tracing

3.0
Lunary logo

Lunary

Free · Observability & Tracing

3.0
Openlayer logo

Openlayer

Not published · Eval Frameworks

3.0
Parea AI logo

Parea AI

Free · Prompt Management

3.0
Prompt Security logo

Prompt Security

Not published · Guardrails & Safety

3.0
PromptHub logo

PromptHub

Free · Prompt Management

3.0
Traceloop logo

Traceloop

Free · Observability & Tracing

3.0
TruLens logo

TruLens

Free · Eval Frameworks

3.0
UpTrain logo

UpTrain

Free · Eval Frameworks

3.0
Helicone logo

Helicone

Free · Observability & Tracing Tested

2.0
LLM Guard logo

LLM Guard

Free · Guardrails & Safety

2.0
Martian logo

Martian

Not offered · LLM Gateways

2.0
New Relic AI Monitoring logo

New Relic AI Monitoring

Free · Observability & Tracing

2.0
OpenAI Evals logo

OpenAI Evals

Free · Eval Frameworks

2.0
Pezzo logo

Pezzo

Free · Prompt Management

2.0
Sentry logo

Sentry

Free · Observability & Tracing

2.0

WhyLabs

Free · Observability & Tracing

2.0

Baserun

Unavailable · Prompt Management

1.0
Gentrace logo

Gentrace

Discontinued · Eval Frameworks

1.0

Literal AI

Discontinued · Observability & Tracing

1.0
Vellum logo

Vellum

Not published · Prompt Management

1.0

By category

Eval Frameworks

15 tools

Libraries for scoring model and agent output.

Observability & Tracing

22 tools

Trace, log and monitor LLM applications in production.

Agent Evaluation

7 tools

Tools built specifically for multi-step agent traces.

LLM Gateways

9 tools

Routing, caching and cost control proxies.

Prompt Management

10 tools

Version, test and deploy prompts.

Guardrails & Safety

10 tools

Runtime validation, PII and injection defence.

Top 12 comparison

Tool Rating Free Price Category SDKs & Frameworks
Inspect AI logo Inspect AI 5.0 Yes $0 (open source) Eval Frameworks 4+
Langfuse logo Langfuse 5.0 Yes $29/mo Observability & Tracing 7+
LangWatch logo LangWatch 5.0 Yes Event-based, rates not published Agent Evaluation 4+
LiteLLM logo LiteLLM 5.0 Yes $0 self-hosted LLM Gateways 3+
Opik logo Opik 5.0 Yes $19/mo Observability & Tracing 5+
Pydantic Logfire logo Pydantic Logfire 5.0 Yes $49/mo Observability & Tracing 7+
Agenta logo Agenta 4.0 Yes $0 self-hosted Prompt Management 4+
AgentOps logo AgentOps 4.0 Yes $40/mo Agent Evaluation 3+
Arize AX logo Arize AX 4.0 Yes $50/mo Observability & Tracing 7+
Arize Phoenix logo Arize Phoenix 4.0 Yes $0 Observability & Tracing 5+
Arthur logo Arthur 4.0 Yes $60/mo Guardrails & Safety 3+
Azure AI Foundry Evaluation logo Azure AI Foundry Evaluation 4.0 - Model inference cost only Agent Evaluation 4+

How we test

We cover this category at two depths, and we label every page so you always know which one you are reading.

Hands-on tested (12 platforms)

We instrument a single identical LLM application - the same traces, the same eval set, the same volume - across every platform, then compare what each one captures, how it scores, what it costs at realistic span volumes, and how much work self-hosting actually takes. Same app, same day, published methodology. These pages carry a Hands-on tested badge.

Researched (61 platforms)

The category is far larger than any team can instrument at once, and a tool nobody has written up honestly is still worth knowing about. These pages are built from primary sources - vendor pricing pages, documentation, licenses, funding records and acquisition filings - and we state plainly when we could not verify something rather than filling the gap. They carry a Researched badge and no performance numbers we did not measure ourselves.

Researched entries get promoted to tested over time. The distinction exists so that a directory listing is never mistaken for a benchmark.

What to look for

  • Corporate status - is the vendor alive, acquired, or quietly in maintenance mode
  • The real billing unit - spans, traces, seats, events or GB ingested change the bill by orders of magnitude
  • License terms - MIT, Apache 2.0 and AGPL 3.0 have very different commercial consequences
  • Self-hosting - whether it genuinely exists, and what is gated behind the cloud tier
  • Pricing - what it actually costs at your volume, not the headline number

Categories we cover

  • Observability & Tracing - Trace, log and monitor LLM applications in production.
  • Eval Frameworks - Libraries for scoring model and agent output.
  • Prompt Management - Version, test and deploy prompts.
  • Guardrails & Safety - Runtime validation, PII and injection defence.
  • LLM Gateways - Routing, caching and cost control proxies.
  • Agent Evaluation - Tools built specifically for multi-step agent traces.

Our recommendation

There's no single best tool - it depends on team size, stack, and budget. Browse the guides for detailed recommendations by team type.

Free Newsletter

Get the LLM Evals Newsletter

Platform comparisons, pricing changes and eval technique deep-dives. No spam.

Frequently Asked Questions

How do you test these platforms?

We instrument a single identical LLM application - the same traces, the same eval set, the same volume - across the platforms we test hands-on, then compare what each one captures, how it scores, what it costs at realistic span volumes, and how much work self-hosting actually takes. Same app, same day, published methodology. Those pages carry a Hands-on tested badge.

What is the difference between a tested and a researched page?

Tested means we instrumented the platform with our reference application and measured it ourselves. Researched means we verified the page against primary sources - vendor pricing pages, documentation, licenses, funding records and acquisition filings - but have not instrumented it yet. Researched pages never carry performance numbers we did not measure, and they say plainly when we could not verify a claim. Every page is labelled, so you always know which one you are reading. The category is far bigger than any team can instrument at once, and a tool nobody has written up honestly is still worth knowing about.

Are these comparisons sponsored?

No. No platform pays for placement or for a rating. Sponsored listings, where they exist, are clearly labelled as sponsored and never affect rankings or scores.

How often is this updated?

This category ships breaking changes monthly, so we re-verify pricing and feature claims every 30 days and re-run the full instrumentation test quarterly. Every page shows a Last Verified date.

Which platforms can genuinely be self-hosted?

Several vendors advertise self-hosting but gate meaningful features behind their cloud tier. We test the actual self-hosted build of every platform that claims to support it and document exactly what you lose.