comparison

AI Observability for LLMs and Agents: A 2026 Guide

AI observability for LLMs and agents explained: what to trace, OTel GenAI conventions, real failure modes from open issues, and what each tool actually costs.

Published:

AI Observability for LLMs and Agents in 2026

Published 30 September 2026. Pricing, repository statistics and GitHub issue statuses each carry their own “checked on” date, because all three change monthly.

Methodology and disclosure: this guide is built from public documentation, published pricing pages, repository metadata, open issue trackers and public practitioner commentary. Nothing here was benchmarked, deployed or measured first-hand. Where a claim comes from a vendor, it is labelled as a vendor claim. It is never restated in our own voice.

AI observability for LLMs and agents, defined

AI observability for LLMs and agents is the practice of capturing structured telemetry from every step an LLM or agent takes: traces, spans, sessions and evaluation scores. The point is to reconstruct behaviour after the fact, score it, and change it. The steps that matter are prompt construction, retrieval, tool calls, agent handoffs and final output.

That definition covers the mechanism. The reason the category exists is sharper. A run can finish with HTTP 200, zero exceptions and full latency SLO compliance, and still return an answer a domain expert would reject outright. Error-rate dashboards cannot see that kind of failure at all.

Three terms get conflated constantly. They are not the same thing, and they sit at different points in one pipeline.

TermQuestion it answersPrimary artefactCatches a fluent, on-time, wrong answer?
MonitoringDid a number move?Metrics, alertsNo
ObservabilityWhy did behaviour change?Traces, spans, sessions, payloadsOnly if a human reads the trace
EvaluationWas the output actually good?Scores against rubrics or referencesYes, by design

LangChain’s observability guide and Sentry’s FAQ both draw a version of this distinction. Neither shows how the three connect as one pipeline, which is the part that matters. Instrumentation emits spans. Evaluation attaches scores to those spans. Monitoring alerts on aggregates of those scores. The score is a span attribute, not a separate system.

Agent observability is the subset where the unit of analysis is a multi-step decision path with memory and tool use, rather than a single request and response. An LLM call has one input and one output. An agent turn has a plan, a retrieval, three tool calls, two retries and a reflection step. Any one of those can be correct on its own while the path as a whole goes nowhere.

Reference architecture for AI observability for LLMs and agents: five layers stacked vertically, from instrumentation, to data protection (masking, redaction, sampling), to the telemetry pipeline (OTLP collector), to analysis (traces, evals, datasets, annotation queues), to operations (APM, SIEM, incident management).

Figure 1. The masking layer sits before export. That is the detail most published diagrams get wrong. No external data; structure synthesised from the layer model described in the Oracle A-Team article.

Why APM breaks on agents (and where it still wins)

Four properties of agent systems break instrumentation designed for request and response services. Each one is a reason AI observability for LLMs and agents became its own product category rather than an APM feature flag.

Non-determinism. The same input produces a different execution path on a different day. Trace diffing against a “normal” shape stops working, because there is no normal shape.

Silent success. The failure has no exception attached to it. Nothing in the HTTP layer knows.

Unbounded fan-out. An agent decides at runtime how many tool calls to make. Three expected, fifty executed, no code change involved.

Cost as a first-class failure mode. In a conventional service, a slow endpoint costs latency. In an agent, a looping endpoint costs money per iteration, and the bill is the alert.

Oracle’s A-Team write-up describes the shape of a real incident well. An agent retried tools roughly 20 times, looped while loading skills, retrieved the wrong document, and still returned a confident answer (ateam-oracle.com). Every step there has a span attribute that would have caught it. Retry count on the tool span catches the 20 retries. Step count per trace catches the skill-loading loop. Retrieval source and document IDs on the retrieval span catch the wrong document. A groundedness score catches the confident answer. All four are cheap to capture. None of them appear in a default APM install.

Now the honest negative case. If you are testing one prompt against a handful of examples, a notebook and a CSV beat a platform rollout. The threshold is when the system gains retrieval, tool calls, retries and memory, and more than one person is changing prompts.

Where APM still wins is real. Infra-layer incident triage, cross-service correlation, on-call routing and paging all live in Datadog APM, Datadog RUM or equivalents. None of that moves to an LLM observability tool. The realistic end state is hybrid. Four of the ten pages currently ranking for this query imply replacement. That is a vendor’s preference, not an architecture.

The telemetry model: spans, traces, sessions and scores

The data model behind AI observability for LLMs and agents nests span inside trace inside session. It is worth working through one concrete agent turn.

A user asks a support agent to find their last invoice and email it. The trace is that one turn. Inside it:

  1. User input span. Raw message, session ID, tenant ID, deployment version.
  2. Planner LLM span. System prompt (versioned), model ID, temperature, input and output tokens, the plan text.
  3. Retrieval span. Query embedding, index name, top-k document IDs, retrieval scores.
  4. Tool-call span 1. lookup_invoice, arguments, latency, status, retry count.
  5. Tool-call span 2. send_email, arguments, status.
  6. Reflection LLM span. Did the plan complete, tokens consumed.
  7. Output span. Final text, guardrail verdict, evaluator scores.

The session is the whole conversation across turns. Sessions matter separately because individual traces can each look healthy while the conversation degrades. The agent answers turn one correctly. It answers turn two correctly. By turn five it has forgotten the invoice number the user gave in turn one. No span-level score detects that. Session-level evals ask whether the agent carried earlier context forward, whether the user reached their stated goal, and how many turns it took. You cannot rebuild any of those by averaging span scores.

The most consequential instrumentation decision is the fork between payload-bearing spans and metadata-only spans. Payload-bearing spans carry prompt and completion text, retrieved chunks and tool arguments. Metadata-only spans carry IDs, counts, latencies and verdicts. Payload spans are what make a trace debuggable. They also make your trace store a second copy of your most sensitive data, and 50 to 200 times larger. Decide this per span type, not globally.

For readers coming from classic ops, the AI-specific signals map onto MELT cleanly:

MELT signalAI-specific content
MetricsToken counts (input, output, cached), cost per run, step count, TTFT, evaluator score aggregates
EventsTool failures, guardrail trips, human handoffs, prompt-version deploys
LogsPrompt/completion pairs as structured records, retrieved chunk text
TracesAgent step sequences, tool-call nesting, multi-agent handoffs

Span attributes worth capturing (and the ones that will bite you)

AttributeWhy it earns its place
trace_id, session_idJoin key for everything; sessions enable multi-turn evals
user_id / tenant_idPer-customer quality and cost breakdown
prompt_versionAnswers “did quality change after the prompt edit”
deployment_versionAnswers “did quality change after the release”
model_id, model_versionSilent provider upgrades shift behaviour
temperature, decoding paramsExplains variance you would otherwise blame on the prompt
input_tokens, output_tokens, cached_tokensCost attribution; cached tokens change the arithmetic materially
tool_name, tool_args_hashTool-selection accuracy without storing raw arguments
retrieval_source, doc_ids, scoresDistinguishes bad retrieval from bad generation
guardrail_verdictLets you separate blocked from wrong
evaluator_scoresMakes triage queries possible at all
retry_count, step_indexLoop detection

Two of these get skipped and then regretted: prompt version and deployment version. Without both, you cannot answer “did quality change after the release”, which is the question that justifies buying the tool at all. Add them in the first commit. Retrofit them later and every trace from before the retrofit is uncomparable.

On indiscriminate parameter capture, OpenLIT’s own documentation warns that database parameter capture can expose passwords, API keys and personal information, as summarised in Oracle A-Team’s tooling comparison (ateam-oracle.com). Auto-instrumentation that captures “everything” is capturing your connection strings too.

There is a cardinality trap that catches teams late. Org and tenant identifiers get attached as free-form trace metadata, and then the dashboard cannot break down by them. Langfuse issue #12614 requests exactly that capability, using metadata fields such as organizationId as a breakdown dimension in dashboard widgets. It has 56 reactions and was open when we checked. Design your dimensions before you instrument, and promote anything you will group by from metadata to a first-class attribute.

OpenTelemetry for GenAI: what is actually standardised in 2026

Every competing page repeats that OTel is the standard. Here is the layering, plainly stated.

OpenTelemetry is two things at once. It is a transport, meaning OTLP, the wire protocol and the SDKs. It is also a semantic-convention registry of agreed attribute names, including the gen_ai.* namespace. OpenInference is Arize’s AI-specific attribute layer that rides on top of OTel spans. Vendor SDKs may emit OTel conventions, OpenInference conventions, both, or a proprietary schema with an OTLP exporter bolted on.

Portability buys you exactly one thing. OTLP export means your spans can reach a second backend without re-instrumenting the application. That is real and worth having.

What it does not buy you is everything the platform layer adds. Evaluator definitions, annotation queues, dataset links, experiment records and prompt-version registries are almost entirely vendor-proprietary. Lock-in does not live in the traces. It lives in the workflow artefacts built on top of them.

Check the status before you rely on it. Large parts of the GenAI convention set are marked experimental, which means attribute names can change between releases. Read the live OpenTelemetry semantic conventions registry, then state the status per attribute group with the date you read it. A Medium guide on this topic claims OTel “has become foundational” with no status detail at all (medium.com). That is not a checkable statement and should not be repeated.

The Arize Phoenix team’s 2024 retrospective states that OpenTelemetry became the LLM observability standard during that year, alongside agents becoming the norm and evaluation moving from niche to product requirement (Hacker News post). That is a vendor-team perspective from a company whose product is built on OTel, not a neutral finding.

The only honest portability test takes an afternoon. Emit one agent turn. Export it to two backends at once. Diff which attributes survive in each UI. Whatever does not appear in backend two is the part you would rebuild on migration.

Four instrumentation paths, and how to choose

Every deployment of AI observability for LLMs and agents lands on one of these four, or on a deliberate mix of them.

SDK and framework instrumentation. You install a tracing SDK or a framework callback handler. LangSmith with LangChain and LangGraph, Langfuse with LlamaIndex, TruLens with a Python app. You get the deepest application context available: the planner’s reasoning structure, in-process tool execution, intermediate state. The cost is discipline, and async edge cases break it. See the failure catalogue below.

Gateway and proxy. LiteLLM, Helicone and Portkey sit on the request path. Adoption is fast and app changes are near zero, which is why this path wins in organisations with many teams and no central instrumentation mandate. The limit is structural. A gateway only sees what crosses the model API boundary. In-process agent state, local tool execution and an agent-to-agent handoff that never hits a model API are all invisible to it.

APM-native. Datadog Agent Observability, Sentry AI Observability and Pydantic Logfire give you correlation with services, infrastructure and user sessions for free, because your infra telemetry is already there. Eval, annotation and dataset workflows generally are not free. In some products they do not exist.

Auto-instrumentation via OTel distros. AWS Distro for OpenTelemetry feeding Amazon CloudWatch GenAI Observability, with Amazon Bedrock AgentCore as the reference example, is the lowest-code path. It is also the least controlled. You get what the distro decides to capture, which is the wrong default when payloads contain regulated data. OpenLLMetry and OpenLIT sit in the same family with more configurability.

Decision tree for choosing an instrumentation path, four questions deep so it renders on mobile. Q1: do tool calls execute in-process? Q2: is prompt and completion text regulated data? Q3: do non-engineers need to review outputs? Q4: is your org already standardised on one APM? Leaves map to SDK-platform, gateway, APM-native, or OTel auto-instrumentation.

Figure 2. No external data; synthesised from the four paths above.

Read the leaves this way. In-process tools plus non-engineer reviewers points at an SDK-platform. Regulated payloads plus no in-process tools points at a gateway with client-side masking. Existing APM standardisation plus engineering-only review points at APM-native.

Where LLM observability actually breaks: a failure catalogue from open issue trackers

Not one of the ten pages ranking for this query cites a single GitHub issue, bug report or public incident. Every “observability is hard” claim on those pages is abstract.

Framing first, because this section is easy to misread. These are open issues in widely used, actively maintained projects. A high reaction count indicates a popular project with engaged users as much as it indicates a serious defect. The point is not that Langfuse or the Datadog Agent are bad software. The point is that instrumentation is itself software, and it fails. The failures cluster into four groups: context propagation, model-format drift, query and UI limits, and runtime compatibility. A fifth group sits below the LLM layer entirely, in collector delivery.

Status caveat. Whether the context-propagation issues below still reproduce against current SDK releases needs hands-on reproduction, and this article did not perform any. Every issue below is described as open with the stated reaction count when we checked on 30 September 2026, not as present-tense breakage in your stack. Reproducing #3961 and #8780 on your own versions is a 30-minute job. Do it before quoting them to your team.

Context propagation

FastAPI StreamingResponse resets the tracing context after the first yielded chunk, producing several fragmented traces for what was one request. Langfuse issue #3961, 49 reactions. The consequence is worse than a missing trace. Your p95 “latency per trace” silently measures chunks rather than requests, so the dashboard looks excellent while the user waits. Any streaming chat UI built on FastAPI is in scope.

Async LangChain callbacks raise repeated “Failed to detach context” errors under ainvoke. Langfuse issue #8780, 62 reactions. This hits the exact async pattern most production LangGraph deployments use, which is what makes it expensive rather than merely annoying.

Model-format drift

Experiments and evaluators fail to parse Anthropic “thinking” blocks from reasoning models such as claude-haiku-4-5, so both break the moment extended reasoning is enabled. Langfuse issue #11109, 62 reactions. The lesson runs past any one tool. Eval harnesses are coupled to provider response schemas, and a reasoning-model rollout is a schema change. Treat “we enabled thinking mode” as a change that needs an eval-pipeline regression test.

Query and UI limits

Filtering traces by numeric evaluator score returns wrong or incomplete result sets. Langfuse issue #11874, 50 reactions. “Show me every trace my judge scored below 1” is the core triage query of this entire discipline. If it cannot be trusted, the triage queue is built on sand. Verify the count against a raw export before you trust a filtered view.

Images uploaded via LangfuseMedia reach the S3 bucket but render as a button rather than inline in the UI. Issue #4555, 41 reactions. Minor, until you are reviewing 200 vision-agent traces. At that volume one extra click per trace is the difference between a review session and an abandoned one.

Runtime compatibility

A Pydantic v1 dependency blocks Python 3.14 compatibility. Issue #9618, 179 reactions, the highest-signal issue in the set. The Python SDK omits the py.typed marker, breaking PEP 561 type-checking for consumers: issue #2169, 51 reactions. Jest users hit dynamic-import failures when importing the JS SDK in tests: issue #5704, 57 reactions. All three are build-pipeline problems rather than runtime ones, which means they surface on the day you upgrade, not the day you install.

Collector and agent layer

These sit below the LLM layer, in the transport itself. They are the ones that lose the data you most need.

Trace delivery is not guaranteed. “Failed to send, dropping traces to intake after 3 retries (Errno 111 connection refused)”: DataDog/datadog-agent issue #17310, 215 reactions.

The agent does not flush buffered data on SIGTERM: issue #3940, 144 reactions. Read that alongside a Kubernetes scale-down event. The traces from the run that caused the incident are precisely the ones you lose, because the pod that produced them is the pod that just got terminated.

The agent does not start on read-only filesystems: issue #15127, 199 reactions. A hostname-resolution failure in v7.40.0 caused a Kubernetes crash loop: issue #14152, 112 reactions. Both argue for the same practice. Pin your collector version, and test collector upgrades as deliberately as you test application upgrades.

Failure-mode catalogue table with columns for symptom, affected stack (framework, runtime or provider), project, issue number, reaction count, open or closed status, date checked, and known workaround, covering the eleven Langfuse and datadog-agent issues cited above.

Figure 3. Re-open each URL on the day of writing, record current status, reaction count and the maintainer’s last comment, and put the check date in the table caption.

The diligence method, reusable

Do not take this list as the definitive set. Take the method. It costs ten minutes per tool you are evaluating.

  1. Search the tool’s issue tracker for your web framework (FastAPI, Next.js, Django), sorted by reactions.
  2. Search for your async pattern (ainvoke, asyncio, streaming).
  3. Search for your model provider (anthropic, bedrock, vertex, thinking).

Three searches. It is the highest-yield diligence available before a procurement decision, and anyone with a browser can run it.

The evaluation layer: offline, online and human

AI observability for LLMs and agents without evaluation gives you a searchable archive of outputs nobody has judged.

Offline evals run versioned golden datasets against a candidate prompt or model before deploy. They catch regressions on known-hard cases and they gate CI. They cannot catch anything your dataset does not contain, which is most of what production will throw at you.

Online evals score a sample of live traffic. They catch distribution shift, novel user behaviour and the slow degradation that follows a silent provider model update. They cannot prevent a bad deploy, because by the time they fire, the users already saw it.

Most teams need both. The offline suite is the gate. The online sample is the smoke detector.

LLM-as-a-judge is appropriate for subjective properties: tone, helpfulness, whether an explanation is pitched at the right level, whether a summary preserved the point. A deterministic code evaluator is strictly better wherever the property is checkable: JSON schema validity, required-field presence, tool-argument correctness, citation IDs actually appearing in the retrieved set. Code evaluators cost nothing per run, never drift and never hallucinate a score. Use the judge only where code cannot reach.

Judge drift is the failure teams do not plan for. Your judge model gets silently upgraded by the provider. Its scoring distribution shifts. Your quality baseline moves without any change to your application. Pin the judge model version explicitly, store it as a span attribute, and keep a fixed calibration set you re-score whenever the judge changes.

Metric families and what each actually measures

MetricWhat it measures
Hallucination rate (intrinsic)Output contradicts the provided source material
Hallucination rate (extrinsic)Output asserts facts absent from source and unverifiable
Groundedness / faithfulnessEvery claim traceable to retrieved context
Toxicity scoreHarmful, abusive or unsafe language in the output
Task completion rateThe user’s stated goal was achieved end to end
Tool-usage correctnessRight tool selected, arguments well-formed, result used
Step-by-step reasoning accuracyIntermediate steps are individually valid, not just the answer

For RAG specifically, the RAG Triad of context relevance, answer relevance and groundedness remains the compact default. It decomposes failure usefully. Bad context relevance is a retrieval problem. Bad groundedness with good context is a generation problem.

Three frameworks worth knowing apart:

  • RAGAS. Reference-free RAG scoring, so you can evaluate without golden answers.
  • DeepEval. Pytest-style unit testing for LLM outputs, which fits teams that want evals in CI rather than in a UI.
  • TruLens. Python-first instrumentation plus feedback functions, strong when you want evaluation logic living in application code.

For toxicity, the tooling options include the OpenAI Moderation API, Detoxify, the roberta_toxicity_classifier, Azure AI Content Safety, and the Vectara hallucination evaluation model for groundedness. JetBrains’ guide states that Google’s Perspective API will no longer be in service after 2026 (blog.jetbrains.com), a claim the original leaves unsourced. If you depend on Perspective, the five options above are where to look.

The loop most teams never close

Production trace, annotated as a failure in an annotation queue, added to a versioned dataset, running as a regression test before the next deploy.

That is the whole value chain. It is where tools differentiate far more than on trace-viewer aesthetics. A platform without a one-click path from trace to dataset leaves the hardest, most repetitive work manual, and manual work in this loop does not happen. Check the path explicitly in a trial. Can a non-engineer reviewer promote a bad trace into the eval set without asking an engineer?

Four-stage closed loop: production trace, annotation queue verdict, versioned dataset entry, regression test in CI, with an arrow returning to the next production trace and a callout on the trace-to-dataset step that most platforms leave manual.

Figure 4. No external data; the four stages are the ones described in this section.

Agent-specific metrics beyond latency and tokens

MetricDefinitionComputed from
Task completion rateRuns achieving the stated goal ÷ total runsSession-level evaluator score, or explicit success event
Tool-call success rateNon-error tool spans ÷ total tool spansstatus on tool spans
Tool-selection accuracyCorrect tool chosen ÷ tool calls, judged against a rubrictool_name vs expected, code or judge evaluator
Step count per task (and p95)Count of agent-step spans per tracestep_index max per trace_id
Retry rateRetried tool spans ÷ total tool spansretry_count
Human-handoff rateTraces ending in escalation ÷ total tracesHandoff event on the trace
Cost per completed taskTotal token cost ÷ completed tasksToken attributes ÷ completion evaluator
Time to first tokenFirst streamed chunk timestamp − request timestampSpan event on the LLM span

Three of these deserve argument rather than definition, and together they are most of what AI observability for LLMs and agents adds on top of an APM dashboard.

Cost per completed task beats cost per request. A cheap run that fails, prompting the user to retry three times, costs more in total than one expensive run that succeeds first time. Request-level dashboards show the cheap failing run as the better outcome. They are measuring the wrong denominator. Divide by completed tasks and the ranking flips to match reality.

Step count p95 is your loop detector. Mean step count barely moves when 1% of runs go pathological. The p95 moves immediately. Tie that back to the Oracle A-Team scenario of an agent retrying tools roughly 20 times. A p95 step count alert would have fired on the first day the behaviour appeared, before anyone read a single trace.

Human-handoff rate is a capability-gap signal, not a failure signal. Treating it as failure pushes teams to suppress handoffs, which is the wrong incentive. Cluster the handoff reasons instead. The largest cluster names the tool or knowledge source you should build next.

What observability actually costs: a worked model from published pricing

No competing page prices a workload end to end. LangChain’s comparison page lists rivals’ list prices and never runs a workload through any of them, so a reader finishes it still unable to answer “what will this cost me”.

The reference workload

Hold this constant across every option:

AssumptionValue
Agent runs per month100,000
Average spans per run12
Total spans per month1,200,000
Engineering seats4
Non-engineer reviewer seats2
Retention90 days
Average payload text per span3.5 KB
Total payload volume/month~4.2 GB
Online eval sampling1 run in 10 (10,000 judged runs)
Judge prompt size~2,000 input / 200 output tokens

Edit those numbers and every figure below moves proportionally. That is the point of publishing them.

The three pricing models

All list prices below are as published on LangChain’s comparison page (langchain.com), which makes them competitor-authored figures. Re-verify each one on the vendor’s own pricing page and date-stamp the check before you make a purchase decision.

OptionPricing modelListed priceWorkload cost/month
LangSmith PlusPer seat$39/seat/mo6 seats = $234, plus usage overages
LangSmith DeveloperPer seat$0Free tier, not viable at 1.2M spans
Datadog Agent Observability ProPer spanFrom $160/mo annually for the first 100k LLM spans1.2M spans is 12 times the first tier, so roughly $1,920/mo if pricing is linear
Langfuse CoreTiered platform$29/moCheck the span allowance against 1.2M
Langfuse ProTiered platform$199/moLikely tier for this workload
Langfuse EnterpriseTiered platform$2,499/moSSO, residency, support
Braintrust ProTiered platform$249/mo
Helicone Pro / TeamTiered platform$79 / $799/moGateway architecture
Portkey ProductionTiered platform$49/moGateway architecture
Arize AX ProTiered platform$50/mo

Two unit rates fall out of that table. The $234 seat line works out at $2.34 per 1,000 agent runs. The $1,920 span line works out at $0.0016 per span. Those are the numbers to compare against your own gross margin per run.

The Datadog line assumes linear scaling past the first 100,000 spans, which is exactly the kind of assumption that is wrong in practice. Volume tiers usually step down. Treat that cell as an upper bound and get a quote.

The line item everyone omits: the judge itself

The evals are LLM calls. Nobody prices them.

Sample one run in ten out of 100,000 and you make 10,000 judge calls per month. At roughly 2,000 input and 200 output tokens each, that is 20 million input tokens and 2 million output tokens per month.

Plug in your judge model’s published per-million-token rates:

Monthly judge cost = (20 × input price per 1M tokens) + (2 × output price per 1M tokens)

Work it for a mid-tier judge at, say, $3 per million input and $15 per million output: (20 × $3) + (2 × $15) = $90/month, or $0.009 per judged run. At a frontier-model judge priced at $15 and $75, the same workload costs (20 × $15) + (2 × $75) = $450/month, or $0.045 per judged run. Both rate pairs are illustrative. Substitute your own judge model’s published rates before you trust the total.

Compare those figures against a $199/month platform tier. At frontier-judge pricing, the evaluation tokens cost more than twice the observability subscription. That is the finding, and it is why the sampling rate is a budget decision rather than a quality-team decision. Run more than one evaluator per trace, say groundedness plus toxicity plus tool correctness as three judge calls, and multiply accordingly.

The other omitted line item: self-hosting is not free

Self-hosting Langfuse removes the subscription and adds an operating cost with four components:

ComponentWhat drives it
ComputeApplication containers plus workers; scales with ingest rate
Object storage4.2 GB/month of payloads at 90-day retention, so about 12.6 GB steady state, plus media
Analytical storageClickHouse-class columnar store for span queries; the dominant infra line at volume
Engineer-hoursUpgrades, backups, access control, retention policy enforcement, on-call

The engineer-hours line gets estimated at zero and is never zero. Use the release cadence as evidence. langfuse/langfuse shipped v4.48.0 on 2026-09-30, with a last push the same day (github.com/langfuse/langfuse, checked 2026-09-30). A project shipping that frequently is a project you will be upgrading frequently. Put a stated number of hours per month and a stated internal rate in the model. At 4 hours and $150/hour, the self-host path carries $600/month of labour before a single dollar of infrastructure.

Laminar’s open-source stack illustrates what self-hosting means architecturally: RabbitMQ, Postgres, ClickHouse and Qdrant, in Rust (Hacker News). Four stateful services to operate. That is the honest shape of “free”.

A price anchor from the low end

Oodle.ai publicly pitched agent-trace storage at $10 per million agent traces, citing a columnar storage engine built to avoid sampling (Hacker News). That is a vendor’s own claim, made in a promotional context, and it prices storage rather than a workflow platform. Use it only as an order-of-magnitude marker for what raw trace retention can cost when retention is all you are buying.

Cost model table pricing one synthetic workload of 100,000 runs a month, 1.2 million spans, 6 users, 90-day retention and one-in-ten online eval sampling across per-seat, per-span, tiered-cloud and self-hosted options, with separate rows for judge-model token spend and self-host engineer-hours.

Figure 5. Every cell carries a date stamp, and the self-host row states its hourly rate assumption.

Sensitivity: which variable actually moves the total

Sensitivity bar chart of monthly total cost under five configurations of the same workload: full payload retention at 90 days, payload retention at 14 days, online eval sampling at one run in ten, online eval sampling at 3%, and rule-based tail retention with a 5% baseline sample.

Figure 6. Derived entirely from the cost table. The formula is published inline above so readers can rerun it with their own volumes.

Two variables dominate every other choice: payload retention and eval sampling rate.

Drop full-payload retention from 90 days to 14 and your payload storage footprint falls by roughly 85%, while every metadata field and every score stays for the full 90 days. Drop online eval sampling from one run in ten to 3% and judge token spend falls by 70%, from $450/month to $135/month at frontier-judge prices, on 3,000 judged runs instead of 10,000.

Either move usually saves more than switching vendors does. Both together nearly always do. Run the retention and sampling exercise before you run the procurement one.

The tool landscape: what each one is actually for

The disclosure that matters most here: four of the ten pages currently ranking for this query are published by vendors who rank their own product first. LangChain, Datadog, Arize and Sentry. Their comparative claims are marketing. This publication does not sell any tool named below.

One concrete example of why that matters. LangChain’s page asserts that SmithDB is “12-15x faster on core LangSmith workloads”, with P50 trace-tree loads around 92ms and P50 single-run loads around 71ms (langchain.com). That is a self-measured, self-published benchmark of a product about itself. No methodology, no hardware specification, no competitor baseline, no definition of “core workloads”. It may well be accurate. It is not evidence, and it should never be repeated as a neutral fact.

Feature checkboxes age badly, so the table below organises the AI observability for LLMs and agents market by architecture class instead.

ClassToolsWorkflow it is built aroundWhat you still build yourselfA checkable fact
SDK-platformLangSmith, Braintrust, W&B WeaveTrace → dataset → experiment → deploy gate, tightly coupled to the SDKInfra correlation; anything outside the framework’s call graphLangSmith Plus listed at $39/seat/mo on LangChain’s own comparison page
Open-source self-hostableLangfuse, Arize Phoenix, OpenLIT, OpenLLMetryFull platform you run yourself, with OTel/OpenInference ingestUpgrades, backups, access control, retention enforcementlangfuse/langfuse: 35,230 stars, 968 open issues, licence listed as NOASSERTION (checked 2026-09-30)
GatewayHelicone, Portkey, LiteLLMDrop-in proxy: caching, routing, cost control, request logsIn-process agent state; local tool spans; handoffsHelicone Pro $79 / Team $799 as listed by LangChain; Portkey Production $49/mo
APM-nativeDatadog Agent Observability, Sentry AI Observability, Pydantic LogfireCorrelating LLM spans with services, infra and user sessionsDatasets, annotation queues, experiment trackingDataDog/datadog-agent: 3,754 stars, 777 open issues, Apache-2.0, release 7.83.3 (checked 2026-09-30)
LibraryTruLens, DeepEval, RAGASEvaluation logic as code, running in CIStorage, UI, retention, reviewer workflowDeepEval runs as pytest tests; no server required

That NOASSERTION licence label on langfuse/langfuse deserves attention rather than a footnote. It means GitHub’s licence detector did not resolve the LICENSE file to a clean SPDX identifier. Read the actual LICENSE file before you assume OSI-approved terms, particularly if you are planning a commercial self-hosted deployment.

Ownership changes are roadmap risk. LangChain’s guide raises Langfuse’s acquisition by ClickHouse and Helicone’s acquisition by Mintlify as buyer-diligence questions. That framing comes from a direct competitor to both, so treat it as such, and then check it anyway from the repositories themselves. Commit frequency. Issue response latency. Whether the last three releases shipped on the historical cadence. Acquisitions cut both ways. ClickHouse owning a product whose analytical layer is ClickHouse is a plausible strengthening rather than an obvious risk.

The open-source tail moves monthly, which is the real caveat on any table like this one. Laminar shipped a Rust-based open-source observability stack (HN), and self-hosted agent-topology tools such as AgentLens have surfaced since (HN). Arize’s changelog dates “Jev-as-a-Judge” structured high-volume evaluators to 26 September 2026 (arize.com/changelog), a useful marker that eval tooling ships new primitives on a monthly cycle. Other names worth knowing in the surrounding ecosystem: Lunary, Alyx, IBM BeeAI, Microsoft AutoGen, Mastra, Pydantic AI, Strands Agents, Deep Agents, Llama Agents, the OpenAI Agents SDK and the Vercel AI SDK. Most of those now emit OTel spans natively or through a community instrumentation package. Oracle’s stack covers the same ground with Autonomous AI Database Select AI Agent, OCI Log Analytics, LoganAI and DBMS_OBSERVABILITY.

Annotated agent trace showing one turn expanded into its seven spans, with callouts marking the attributes that make it debuggable: prompt version, deployment version, tool arguments, step count and evaluator score.

Figure 7. Use a public product screenshot, or an open-source Phoenix or Langfuse demo instance that permits reuse, with attribution. A schematic mock also works. Do not present it as a production system.

PII, masking, residency and retention

Prompt and completion payloads are the highest-value observability data you have. They are also, frequently, regulated personal data. Your trace store quietly becomes a second copy of your most sensitive dataset, held under weaker access controls than the primary one, and usually nobody in compliance has been told.

Client-side masking redacts before telemetry leaves the application process. You can never over-collect, because the sensitive bytes never crossed the wire. The cost is that every application team must implement it correctly, and a missed field is a silent leak.

Ingestion-side masking applies one central policy at the collector or platform boundary. Policy stays consistent and auditable. The raw data crossed the network and touched the vendor’s ingest path first, which for some regulatory regimes is already the violation. Use ingestion-side masking as a backstop for client-side masking, not as a replacement for it.

The architectural alternative gets less attention than it deserves: store references rather than payloads. Row IDs, content hashes, retrieval document IDs, guardrail policy decisions, short summarised observations. It is cheaper and safer. It is also meaningfully harder to debug, because reconstructing what the model actually saw means joining back to a system that may have mutated the row since. Name that tradeoff in your design doc rather than discovering it during an incident.

Conditional tracing is the ability to skip tracing entirely for requests matching a rule. Check for it by name during evaluation. It is not sampling. It is not masking. It is the ability to say “requests from this tenant, or carrying this classification flag, produce no trace at all”.

Residency is a genuine differentiator rather than a checkbox. Langfuse publishes distinct EU, US and JP cloud base URLs, documented in the Oracle A-Team walkthrough (ateam-oracle.com). Verify which region your SDK’s default base URL points at. The default is rarely the one your DPA specifies.

Coding-agent telemetry is its own hazard and deserves a separate decision. Tracing plugins for OpenAI Codex and Claude Code upload transcript data wholesale: prompts, assistant messages, reasoning summaries, tool-call inputs and outputs, model metadata and token usage. The A-Team guide explicitly advises against enabling it for sessions containing data you are not comfortable storing in the vendor’s platform (same source). A coding agent’s tool outputs include file contents. File contents include your .env handling, your customer fixtures and your internal source. Enable this per project. Never org-wide by default.

Sampling and retention strategy at volume

Naive head sampling is the wrong default for agents, and the reason is a one-liner. You sample away the 0.5% of runs that are pathological, and those are the only runs worth keeping. A uniform 5% head sample keeps 5% of your looping, hallucinating, retry-storming runs and 5% of your boring successful ones, at exactly the moment you need all of the former and none of the latter.

Use rule-based tail retention instead. Keep 100% of runs that:

  • errored or timed out,
  • exceeded a step-count threshold,
  • triggered a guardrail,
  • scored below an evaluator threshold,
  • received negative user feedback.

Then sample the healthy majority at 1% to 5% for baseline trend data. The pathological tail is small by definition, so this costs far less than it sounds.

Tier your payload retention separately from your metadata retention. Full payloads for 7 to 14 days, which covers essentially every debugging session anyone actually runs. Metadata and scores for 90 to 365 days, which is what trend analysis and regression comparison need. This is usually the single biggest lever on total observability spend, and it requires no vendor change.

One interaction catches teams out. Sampling and evaluation interfere. Sample first, evaluate what survived, and your online quality metric describes a biased subset. Rule-based retention makes that bias severe, because the retained set is deliberately enriched with failures. Evaluate on a separate, genuinely random sample. Retain on rules. Keep the two pipelines distinct and label which sample each dashboard number came from.

Tie this back to the delivery hazard from the failure catalogue. DataDog agent issue #3940, no flush on SIGTERM, 144 reactions, means buffered telemetry can vanish precisely when a pod is scaled down mid-incident. A retention policy that assumes complete delivery is a retention policy that is wrong on the worst day of the quarter. Design for gaps. Alert on ingest-rate drops, not only on the metrics computed from ingested data.

Three surfaces the guides forget: voice, MCP servers and coding agents

Voice agents. The relevant span chain is STT, then LLM, then TTS. The metric that matters is time to first token per call, not aggregate request latency. A 400ms TTS delay is inaudible in a p95 latency chart and unmistakable to a caller. Practitioners running LiveKit voice agents describe building their own tooling precisely because they had no per-call TTFT, no per-call cost and no latency visibility across the STT, LLM and TTS chain (Hacker News, Whispey team). Instrument each of the three legs as its own span with its own first-byte timestamp. Attribute cost per leg separately, because STT and TTS are priced by audio duration while the LLM is priced by tokens.

MCP servers. Agent tools increasingly live behind the Model Context Protocol, which moves tool latency, tool error rate and per-client usage outside your application process. Your in-process SDK sees “tool call took 4 seconds”. It does not see which of the MCP server’s six downstream dependencies was slow, or that a different client is saturating the same server. Sentry documents MCP server monitoring as a first-class case (sentry.io). Only one of the ten pages ranking for this query mentions MCP at all, which is remarkable given how quickly it became the default tool-distribution mechanism. Instrument the MCP server itself, propagate trace context across the MCP boundary, and tag spans with the calling client ID.

Coding agents. Teams now trace Claude Code, Codex and OpenCode sessions for token spend and tool activity. Sentry’s cookbook and the Oracle A-Team Langfuse-for-Codex walkthrough both document working setups, and the JetBrains PyCharm AI Agents Debugger sits in the same space for step-through inspection. Pair any of them with the PII warning above before you enable anything org-wide.

Agent memory is the under-instrumented surface. Memory writes and memory injections are decisions that change downstream behaviour, and almost nobody captures them as spans. When an agent behaves differently today than yesterday on the same input, the cause is frequently a memory record written three sessions ago that nobody can see. The Threadbase founder frames debugging AI memory as the unsolved part of the problem (Hacker News). That is a vendor’s problem statement, and an accurate description of a gap. Emit a span for every memory write and every memory read injected into a prompt, with the record ID and the reason it was selected.

Lock-in: what you can take with you when you switch

Four of the ten pages ranking for this query are written by tools you might one day switch away from. None of them explains what you can export.

ArtefactPortable?Mechanism
Spans and tracesYesOTLP export, dual-write to a second backend
DatasetsUsuallyJSONL or CSV export via API
Evaluator scoresYes, as dataAPI export; the numbers move, the meaning needs re-deriving
Evaluator definitionsRarelyJudge prompts sometimes exportable; scoring config usually not
Annotation queues and reviewer feedbackOften trappedVaries wildly; check the API docs, not the marketing page
Prompt registry and version historyUsually trappedVersion lineage rarely has an export endpoint
Dashboards and monitorsNeverRebuild by hand

Portability matrix with rows for spans, evaluator definitions, evaluator scores, datasets, annotation and feedback, prompt versions, and dashboards or monitors, and per-vendor columns carrying a Yes, Partial or No verdict, the documented export format (JSONL, CSV or API), and a link to that vendor's export documentation page.

Figure 8. Every cell needs a link to the vendor’s own export docs, not to a marketing page.

Run the migration test during the trial, not after the contract. Export one dataset and one annotation set from the tool you are evaluating, then try to import both into a second tool. Time it. That elapsed number, multiplied by the volume you will have accumulated in two years, is your exit cost. It is the only version of that figure anyone will ever give you.

Sentry documents exporting production agent conversations via CLI or MCP into Braintrust, Langfuse, promptfoo or Phoenix as eval datasets (sentry.io). That is evidence trace-to-eval portability is achievable when a vendor chooses to support it.

The structural defence is simple. Emit OTLP to a vendor-neutral collector you control, and fan out from the collector to your backends. Switching a backend then becomes a collector configuration change rather than an application re-instrumentation project. It costs you one hop of latency and one component to operate. Worth it.

A 30-day rollout that does not stall

Week 1: instrument one critical path end to end. Not the whole application. One path, fully. Required attributes only: trace ID, session ID, user or tenant ID, model and prompt version, token counts, tool name, deployment version. Masking on from the first commit, never retrofitted, because retrofitted masking leaves every trace you already collected unmasked.

Week 2: build the triage queue. Write the auto-retain rules. Errors, step count above threshold, guardrail trips, negative user feedback. Then name an owner. LangChain’s guide is right on this point: trace visibility without a named reviewer produces nothing (langchain.com). A queue nobody owns is an archive.

Week 3: turn the first 30 triaged failures into a golden dataset and wire one offline eval into CI. Define the pass bar before you look at the scores, or you will define it around the scores you got.

Week 4: add online eval on a 3% to 10% sample, set exactly two alerts (evaluator score drop, cost-per-completed-task spike), and run the cost model from this article against your real span volume. You now have actual spans per run rather than an assumed 12, which usually changes the answer. That week is where AI observability for LLMs and agents stops being a purchase and starts being a practice.

Four anti-patterns reliably stall a rollout:

  • Instrumenting everything at once, producing 40 GB of traces nobody reads.
  • Alerting on raw latency before you have any quality signal.
  • Buying a platform before anyone owns the triage queue.
  • Enabling payload capture on regulated data “temporarily”.

What would change this recommendation

This page is maintained against four moving parts. Ownership changes: further acquisitions in the open-source tier alter roadmap risk, and should be re-checked from repository activity rather than press releases. Licence changes: the NOASSERTION status on langfuse/langfuse is exactly the kind of thing that either resolves or hardens. OTel convention stabilisation: if gen_ai.* attributes reach stable, the portability argument strengthens materially and the “lock-in lives above the traces” framing needs revisiting. Pricing: every figure in the cost model is a list price subject to change, and per-span tiers in particular tend to restructure.

Second-hand market statistics circulate on this topic, including a KPMG survey on agent adoption and a Gartner forecast on agentic AI in enterprise software. Both are cited by IBM without links to the primary reports, so neither is used in this article’s argument. The same applies to Datadog’s marketed “40% lower token usage per task” and to Sentry’s Anthropic-attributed productivity and incident-resolution figures. Vendor marketing, no methodology published, not load-bearing here.

Frequently Asked Questions

What is AI observability for LLMs and agents?

AI observability for LLMs and agents means capturing structured telemetry from every step an LLM or agent takes: traces, spans, sessions and evaluation scores. Behaviour can then be reconstructed, scored and improved after the fact. Error rates alone are insufficient, because agent runs routinely fail silently while returning HTTP 200.

What is the difference between LLM monitoring, observability and evaluation?

Monitoring detects that a metric changed: latency, cost, error rate. Observability explains why, by exposing traces, prompts, retrieval results and tool calls. Evaluation judges whether the output was correct or useful against a rubric or reference. Only evaluation catches a fluent, on-time, confidently wrong answer.

What are the best open-source AI agent observability tools?

Langfuse (self-hostable full platform with datasets and evals), Arize Phoenix (OTel and OpenInference tracing plus evaluation), OpenLIT (OTel-native broad auto-instrumentation), OpenLLMetry (lightweight OTel instrumentation) and TruLens (Python feedback functions). Licences vary. langfuse/langfuse is listed as NOASSERTION with 35,230 stars, so read the LICENSE file directly.

Is OpenTelemetry enough for LLM observability on its own?

Necessary but not sufficient. OpenTelemetry standardises transport and span attributes, which gives you real backend portability through OTLP. It does not give you datasets, evaluators, annotation queues or prompt versioning, and those are where vendor lock-in actually lives. Check the current gen_ai.* convention stability status and date-stamp what you find.

How much does LLM observability cost?

Three models. Per seat, with LangSmith Plus listed at $39/seat/month. Per span, with Datadog Agent Observability from $160/month for the first 100,000 LLM spans. Tiered platform, with Langfuse at $29, $199 and $2,499. The omitted line item is judge tokens: sample one run in ten out of 100,000 and judge spend of $90 to $450 a month frequently exceeds the subscription.

How is AI agent observability different from LLM observability?

LLM observability’s unit is a single call: prompt in, completion out. Agent observability’s unit is a multi-step decision path with memory, tool calls, retries and handoffs between agents. So you need step count, tool-selection accuracy and session-level evaluation, none of which exist at the single-call level.

Can I use Datadog, Sentry or Azure for AI agent observability instead of a dedicated tool?

Yes, if your priority is correlating LLM behaviour with services, infrastructure and user sessions, and your on-call process already lives there. No, if your core workflow is datasets, experiments, annotation queues and regression evals, which are thinner or absent in APM-native products. Hybrid is the common real outcome.

How do I observe an AI agent without leaking PII into the trace store?

Four controls. Client-side masking before telemetry leaves the app. Conditional tracing that skips sensitive requests entirely. Storing references, hashes and summaries instead of full payloads. Tiered retention that drops payloads after 7 to 14 days. Coding-agent tracing plugins are a separate hazard: they upload prompts, reasoning summaries and tool outputs wholesale.

What metrics should I put on an AI agent dashboard first?

Six, in this order: task completion rate, cost per completed task, step count p95, tool-call success rate, evaluator score trend, and time to first token for streaming interfaces. Cost per completed task beats cost per request, because a cheap failed run the user retries three times costs more than one expensive success.

Will my traces transfer if I switch observability vendors?

Spans usually transfer via OTLP, and datasets often export as JSONL. Evaluator definitions, annotation history, prompt-version lineage and dashboards usually do not. Run an export-and-import test during the trial rather than after signing, and treat the elapsed time as your measured exit cost.


Related reading: LLM evaluation metrics explained · Langfuse vs LangSmith comparison · build golden datasets · gen_ai attribute field guide · LLM-as-a-judge costs · multi-agent debugging guide · RAG evaluation with RAGAS · MCP server monitoring · voice agent TTFT tracing · telemetry masking patterns · LLM cost control tactics

The most useful thing you can do after reading this is the ten-minute diligence pass. Search your candidate tool’s issue tracker for your web framework, your async pattern and your model provider, sorted by reactions. Then run the retention-and-sampling arithmetic from the cost model against your own span volume. AI observability for LLMs and agents is a data-engineering decision about what you capture, mask, sample and retain, and those two exercises answer more of it than any vendor comparison table will.

Frequently Asked Questions

What is AI observability for LLMs and agents?

AI observability for LLMs and agents means capturing structured telemetry from every step an LLM or agent takes: traces, spans, sessions and evaluation scores. Behaviour can then be reconstructed, scored and improved after the fact. Error rates alone are insufficient, because agent runs routinely fail silently while returning HTTP 200.

What is the difference between LLM monitoring, observability and evaluation?

Monitoring detects that a metric changed: latency, cost, error rate. Observability explains why, by exposing traces, prompts, retrieval results and tool calls. Evaluation judges whether the output was correct or useful against a rubric or reference. Only evaluation catches a fluent, on-time, confidently wrong answer.

What are the best open-source AI agent observability tools?

Langfuse (self-hostable full platform with datasets and evals), Arize Phoenix (OTel and OpenInference tracing plus evaluation), OpenLIT (OTel-native broad auto-instrumentation), OpenLLMetry (lightweight OTel instrumentation) and TruLens (Python feedback functions). Licences vary. langfuse/langfuse is listed as NOASSERTION with 35,230 stars, so read the LICENSE file directly.

Is OpenTelemetry enough for LLM observability on its own?

Necessary but not sufficient. OpenTelemetry standardises transport and span attributes, which gives you real backend portability through OTLP. It does not give you datasets, evaluators, annotation queues or prompt versioning, and those are where vendor lock-in actually lives. Check the current `gen_ai.*` convention stability status and date-stamp what you find.

How much does LLM observability cost?

Three models. Per seat, with LangSmith Plus listed at $39/seat/month. Per span, with Datadog Agent Observability from $160/month for the first 100,000 LLM spans. Tiered platform, with Langfuse at $29, $199 and $2,499. The omitted line item is judge tokens: sample one run in ten out of 100,000 and judge spend of $90 to $450 a month frequently exceeds the subscription.

How is AI agent observability different from LLM observability?

LLM observability's unit is a single call: prompt in, completion out. Agent observability's unit is a multi-step decision path with memory, tool calls, retries and handoffs between agents. So you need step count, tool-selection accuracy and session-level evaluation, none of which exist at the single-call level.

Can I use Datadog, Sentry or Azure for AI agent observability instead of a dedicated tool?

Yes, if your priority is correlating LLM behaviour with services, infrastructure and user sessions, and your on-call process already lives there. No, if your core workflow is datasets, experiments, annotation queues and regression evals, which are thinner or absent in APM-native products. Hybrid is the common real outcome.

How do I observe an AI agent without leaking PII into the trace store?

Four controls. Client-side masking before telemetry leaves the app. Conditional tracing that skips sensitive requests entirely. Storing references, hashes and summaries instead of full payloads. Tiered retention that drops payloads after 7 to 14 days. Coding-agent tracing plugins are a separate hazard: they upload prompts, reasoning summaries and tool outputs wholesale.

Explore More

Free Newsletter

Get the LLM Evals Newsletter

Platform comparisons, pricing changes and eval technique deep-dives. No spam.

Related Articles