best-of

The Best LLM Observability Tools in 2026, Ranked and Road-Tested

Eight LLM observability platforms judged on the four things that actually decide the bill and the migration - self-host reality, OpenTelemetry support, pricing at scale, and eval depth. One clear winner, one you should not start on.

Published:

I have wired a lot of these tools into production stacks, and the pattern is always the same. The demo looks great. Then one of four things bites - the self-host turns out to be a hobbled core, the OpenTelemetry support is a marketing checkbox, the bill explodes at real trace volume, or the eval tooling is thin when you finally need to score outputs. So that is exactly how I judged these eight.

Four axes, weighted by what actually causes migrations:

  • Self-host reality. Not “is there a repo” but “do you keep the real features when you run it yourself, and under what license.”
  • OpenTelemetry support. Whether it is native architecture or a bolted-on receiver, because that decides how locked-in you are.
  • Pricing at scale. The number at 1M traces or 100k spans, not the friendly entry tier.
  • Eval capability. Whether it scores output quality or just shows you what happened.

Here is the ranked list, one clear winner, and one tool I am telling you not to start on.

The short version

ToolBest forSelf-hostOTelStarting price
LangfuseThe all-round open-source defaultFree, MIT, near-completeOTLP backend (HTTP)Free / $29/mo
OpikCheapest managed cloud, permissive OSSFree, Apache-2.0, full featuresOne ingestion pathFree / $19/mo
Arize PhoenixFast OSS tracing and RAG evalFree, but ELv2 serverNativeFree (OSS)
LangSmithLangChain and LangGraph appsEnterprise onlyOTLP receiver$39/seat/mo
GalileoEnterprise eval intelligence + guardrailsEnterprise onlyYes (OTLP)Free / $100/mo
LaminarAgent and browser-agent tracingFree, full stackNativeFree / $30/mo
PortkeyGateway-first teams who also want tracesGateway only, not observabilityIngests + enrichesFree / $49/mo
HeliconeNobody new - maintenance mode, plan an exitFree, Apache-2.0, but frozenPartialFree / $79/mo

1. Langfuse - the winner for most teams

Langfuse wins because it is the rare open-source tool where self-hosting gets you the real product. Only three features are enterprise-gated in the self-host build - tracing, evals, prompt management, human annotation and RBAC are all free under MIT. That is the differentiator, and it is why “what’s the self-hostable alternative” almost always resolves to Langfuse.

The economics seal it. At 1M events a month it runs about $101/mo managed, against LangSmith’s roughly $2,514/mo for comparable volume - the widely-cited ~25x gap. Self-host and the per-trace cost disappears. It is framework-agnostic and runs as an OpenTelemetry backend on an OTLP endpoint, mapping the GenAI semantic conventions.

The honest catch is operational, not commercial. Langfuse v3 needs Postgres plus ClickHouse, Redis and S3-compatible storage - four services, and the migration to that architecture is where self-hosters get stuck. gRPC OTLP is not supported yet, HTTP only. And it is now a ClickHouse subsidiary after the January 2026 acquisition, which I would file away for a multi-year bet. If you can run the stack, nothing else gives you this much for free. If you cannot, the $29/mo Core cloud tier sidesteps it and still beats LangSmith badly.

2. Opik - the cheapest managed cloud, no license asterisk

If the appeal of Langfuse is the openness but you do not want to run four services, Opik is the move. It is Comet’s open-source observability and eval platform, and it does something no one else here quite matches - the OSS build is Apache-2.0 with the full feature set self-hosted, unlimited spans, members and retention, no gates. That is the most permissive license in the set, more so than Phoenix’s Elastic License.

On the cloud, Pro is $19/mo for 100k spans, the cheapest paid tier of the major platforms, with $5 per additional 100k. It is also the fastest-growing project of its peers at roughly 20.8k GitHub stars. The eval side is real too - LLM-as-judge, code metrics, online evaluation, guardrails and an Agent Optimizer.

The gotcha is per-seat pricing at scale. The $19 headline is the small-team configuration, and the recurring complaint is that seat costs climb as the team grows. A few users report UI slowdown on very large projects. OTel is one ingestion path here, not the native architecture. None of that undercuts the core value - just model the seat cost if you are a big team.

3. Arize Phoenix - fastest to try, best RAG eval

Arize Phoenix is the OSS pick that starts fastest - a working trace UI on your laptop in under a minute. It is genuinely OpenTelemetry-native, built on OTel and Arize’s own OpenInference conventions, so it is framework-agnostic in a way LangSmith is not. And its 50+ pre-built eval metrics include what reviewers call the best RAG evaluation in the category - serious retrieval and answer scoring without writing your own judge prompts.

The catch is the license, and it is a real one. The main server repo is Elastic License 2.0 - source-available, not OSI-approved open source. Only the client and eval subpackages are Apache-2.0. ELv2 forbids offering Phoenix as a hosted service to third parties. For internal use it behaves like open source and the features are not gated. But if your plan is to resell it as a service, read the license first. There are also reports of ingest lag before traces appear. Arize the company is well-funded - a $70M Series C in February 2025 - so it is not going anywhere.

4. LangSmith - the best experience for LangChain, at a price

LangSmith is the most turnkey observability you can point at a LangChain app. Add a callback and every chain, tool call and agent step shows up traced with zero extra work. Nothing touches it for depth on LangChain and LangGraph code, because nobody else ships LangChain. The eval side is strong too, and Align Evals - calibrating an LLM judge against human scores - is genuinely useful.

Two things keep it at fourth. The bill runs roughly $2,514/mo at 1M base traces on one seat, about 25x Langfuse, with base traces at $2.50 per 1,000 and extended at $5.00. And you cannot self-host your way out - LangSmith is fully closed source and self-hosting is Enterprise-only. It accepts OpenTelemetry as a receiver, so you are not forced onto the LangChain SDK, but the zero-config magic only shows up when you use it. Backed by a $1.25B company (roughly $260M raised), so survival is not the question. Value-for-money is. If you live in LangChain and the bill does not scare you, stay. Otherwise the alternatives are cheaper.

5. Galileo - enterprise eval intelligence and guardrails

Galileo is the best-funded platform in the space at roughly $68M raised, and the most research-forward. Its bet is the proprietary Luna and Luna-2 eval models - small models fine-tuned for eval tasks like hallucination and prompt-injection detection - cheap and fast enough to run on every request, which is what turns offline evals into real-time production guardrails. The free tier is generous at 5,000 traces with unlimited users.

The reservations are commercial. Above the $100/mo Pro tier (50,000 traces, billed yearly) everything is contact-sales, and self-host is Enterprise-only with no open-source version. It supports OpenTelemetry across CrewAI, LangGraph and the OpenAI Agents SDK. One warning that matters more here than anywhere - there is a separate, unrelated Google-acquired design tool with the same name, so verify you are reading about galileo.ai before you trust any price or review. This is a platform you buy through a rep, not a card.

6. Laminar - built for agents, self-host the whole thing

Laminar is the most purpose-built option for AI agents, especially browser agents. It is OpenTelemetry-native, written in Rust for low overhead, and it is the only platform here you can self-host in full - the whole stack is open source, not just a gateway or a hobbled core. Browser Use, one of the most popular open-source browser agents, documents Laminar as its observability integration, and OTel co-creator Ben Sigelman is an angel investor.

The gotcha is the billing. You pay on two axes - data by the GB and “Signals,” and Signals are metered by the tokens spent reading your traces, not the tokens your agent spends. That depends on Laminar’s own trace-compression claims, which makes a monthly forecast genuinely hard - the least predictable pricing in the category. Self-hosting removes the usage bill entirely. It is also the youngest and smallest here (2024, YC S24, $3M seed), so weigh maturity against the fact that full self-host gives you an exit. Cloud starts at $30/mo.

7. Portkey - a gateway first, observability second

Portkey is a different animal. It is an LLM gateway that routes to 1,600+ models with fallbacks, caching, budgets and 50+ guardrails, and it happens to include observability. If your problem is “we call five providers, costs are exploding, and we need fallbacks and budgets in one place,” it is aimed squarely at you. Its OpenTelemetry support is strong - it ingests OTel and enriches traces with cost and token metrics per the GenAI conventions.

But know the split before you plan a deployment. The open-source gateway is Apache-2.0 and self-hosts free - but that gets you routing and a basic dashboard, not real observability. Logs, traces, analytics and retention live on the managed Production tier at $49/mo for 100k logs. The meter caps logs, not requests, so past the cap your traffic keeps flowing but your observability quietly goes dark. Plenty of teams run Portkey as the gateway and feed a dedicated observability backend over OTel. That pairing is the right mental model.

8. Helicone - do not start here

I am including Helicone only to tell you not to build on it. It was a clean, open-source, proxy-based observability tool. Then Mintlify acquired it on 3 March 2026 and put it in maintenance mode - security patches, bug fixes and new-model support only, no roadmap, and Mintlify is actively helping customers migrate off.

For a new buyer that single fact overrides everything else. The Apache-2.0 self-host still works, but self-hosting a frozen codebase whose own creator is pointing people to the exit is a weak bet. The proxy model was also a structural single point of failure - if Helicone is down, your LLM calls fail even when the provider is healthy, and every proxied call adds latency (Helicone cites ~10ms). If you are already on it, use the runway to plan your migration. If you are evaluating fresh, skip to Langfuse or Opik.

So which one?

  • You want the best all-round open-source platform and can run the self-host stack - Langfuse. It is the default for good reason, and roughly 25x cheaper than LangSmith at scale.
  • You want managed hosting for the least money, with a clean license - Opik at $19/mo.
  • You want to try something in the next 60 seconds, or you need the best RAG eval - Arize Phoenix, as long as you are not reselling it.
  • You live in LangChain and LangGraph - LangSmith, if the bill does not scare you.
  • You are an enterprise that got burned by production hallucinations - Galileo, and confirm you have the right company.
  • You are building agents, especially browser agents - Laminar, self-hosted to sidestep the Signals meter.
  • You need a multi-provider gateway more than a tracer - Portkey, and expect to pay for the real observability.
  • You are on Helicone today - plan your exit. Do not start there fresh.

Every price and date on this page was read from each vendor’s own materials on 23 July 2026 and links to our full tool reviews. This category ships breaking changes monthly, so we re-verify every 30 days. The one gap that decides most migrations - the ~25x between LangSmith and Langfuse at 1M traces - has held for a while. If cost is why you are here, it is not close.

Frequently Asked Questions

What is the best LLM observability tool in 2026?

For most teams it is Langfuse. It is MIT-licensed, self-hosts free with almost every feature intact, and runs roughly $101/mo at 1M events versus LangSmith's ~$2,514/mo for comparable volume. If you want managed hosting with no ops work, Opik is the cheapest cloud in the category at $19/mo for 100k spans. Pick Langfuse if you can run the self-host stack, Opik if you would rather not.

Which LLM observability tools support OpenTelemetry?

All eight here support OTel in some form, but the depth varies. Arize Phoenix and Laminar are OpenTelemetry-native, built on it from the ground up. Langfuse runs as an OTLP backend (HTTP only, no gRPC yet). Portkey ingests OTel and enriches traces with cost and token metrics. LangSmith accepts OTLP as a receiver. Opik treats OTel as one ingestion path among 60+ integrations. If strict OTel-native architecture is the requirement, Phoenix or Laminar are the cleanest fits.

What is the cheapest LLM observability tool?

Self-hosting Langfuse, Opik, Arize Phoenix, Laminar or the Helicone code is free under their open licenses - you pay only for infrastructure. For managed cloud with no ops work, Opik's Pro tier at $19/mo for 100k spans is the cheapest of the major platforms, ahead of Langfuse Core at $29/mo and far ahead of LangSmith, whose trace bill reaches roughly $2,514/mo at 1M traces.

Should I use Helicone in 2026?

Not for a new project. Mintlify acquired Helicone on 3 March 2026 and put it in maintenance mode - only security patches, bug fixes and new-model support ship, there is no roadmap, and Mintlify is actively helping customers migrate off. The proxy-based product was clean, but building fresh on a frozen tool whose own owner is pointing people to the exit is a dead end. Start on Langfuse or Opik instead.

Explore More

Free Newsletter

Get the LLM Evals Newsletter

Platform comparisons, pricing changes and eval technique deep-dives. No spam.

Related Articles