comparison

LLM Tracing for AI Agents: A Practical Guide

LLM tracing and AI tracing for agents: the span data model, which OpenTelemetry GenAI attributes are stable, four ways to instrument, and what it costs.

Published:

LLM tracing records every operation inside a single request as a timestamped span: the model call, the retrieval query, the tool invocation, the guardrail check, the sub-agent handoff. Spans share one trace ID and nest through parent span IDs, producing a tree that reproduces the execution path in the order it actually ran. LLM tracing AI tracing for agents is the version of that practice tuned for systems that decide their own control flow, where the thing you need to inspect is not a stack trace but the arguments a model invented for a tool call.

Last reviewed: 30 September 2026. Sourcing note: every capability claim below is attributed to the vendor’s own documentation, every failure mode to a public issue tracker or changelog, and every number to whoever published it. Vendor performance figures are labelled as vendor claims throughout, including the ones that flatter tools I otherwise recommend. Nothing here was measured hands-on: no benchmarks were run, no latency compared, no reliability tested. Where an answer requires that, the article says so and tells you what to measure.

What LLM tracing is, in one paragraph

A span is one unit of work with a start time, an end time, a status, a parent, and a payload. The payload is the structural difference from ordinary logging. An APM span carries durations and status codes; an LLM span carries the actual prompt, the actual retrieved chunks with their similarity scores, the actual JSON the model produced for a function call. That semantic payload is what makes the record diagnostic instead of merely chronological.

For a single-turn RAG call, a trace is a convenience. You could reconstruct most of it from three log lines. For an agent that makes twelve tool calls and two sub-agent handoffs before answering, the trace is the only artefact that tells you which of the twelve decisions was wrong, and LLM tracing AI tracing for agents exists because that difference is categorical rather than one of degree.

One more contrast the incumbent definitions miss: a trace is recorded as execution happens, in order, by instrumentation inside the process. It is not reconstructed afterwards by correlating timestamps across log files, which is the technique it replaces and the reason it survives async fan-out that log correlation cannot.

This guide covers the data model, the OpenTelemetry GenAI attribute situation as of the September 2026 semconv registry, four instrumentation approaches with code, the failure modes visible in public issue trackers, a cost model built from published price lists, and a capability matrix across ten tools.

Trace, span, thread, log, metric: five words teams use interchangeably

LevelWhat it isThe question it answersWhere it lives
LogA discrete timestamped eventDid this happen?Log backend, usually unstructured text
MetricA pre-aggregated number over timeIs something wrong right now?Time-series store
SpanOne unit of work with inputs and outputsWhich step failed?Trace store
TraceThe connected span tree for one requestWhy did this request go the way it did?Trace store, keyed by trace ID
Thread / sessionThe ordered set of traces in one conversationDid the conversation succeed?Backend-specific grouping, not OTel-standardised

Ladder diagram of the five telemetry levels from log up to thread Figure 1. The five levels stacked from log at the base to thread at the top, each rung labelled with the question it answers and the store it lives in. The top rung is shaded differently to mark it as backend-specific rather than specified by OpenTelemetry.

The thread row is where agent failures actually live, and it is the level almost nobody instruments on day one. An agent can emit ten individually defensible replies and still fail the user: it forgets the constraint given in turn two, re-asks a question already answered, or never resolves the goal. Per-request LLM tracing AI tracing for agents will score all ten of those turns as fine.

Thread grouping is a backend concern rather than a specification concern. The OpenTelemetry GenAI conventions standardise span attributes and do not define conversation identity, so session and thread IDs are named differently by each platform. Check the current registry at https://opentelemetry.io/docs/specs/semconv/gen-ai/ before claiming either coverage or absence, because this is one of the areas that moves.

The practical instruction: stamp a stable thread ID on every trace from day one, even when nothing reads it yet. Retro-fitting conversation identity onto historical traces is usually impossible, because the join key was never written.

Metrics still beat traces at two things. Cost-per-day trends and p95 latency by model are aggregate questions, and answering them by scanning a trace explorer is slow and, on usage-billed platforms, expensive.

Why agents break the assumptions traditional tracing was built on

Failure signal inversion comes first. APM alerting is built around exceptions, timeouts and 5xx responses. The canonical agent failure is a clean HTTP 200 carrying an answer that is wrong, ungrounded, or against policy. Nothing throws. No error budget moves. This is why LLM tracing AI tracing for agents cannot be reduced to APM with new field names: the signal APM keys on is absent by construction.

Non-determinism is second. The same input can take a different path on two consecutive runs, picking a different tool or making four retrieval calls instead of one, which makes trace shape itself a diagnostic signal and makes two traces of the “same” request non-comparable in the way two APM traces are.

Third is error propagation distance. A malformed tool argument generated at step three corrupts every step after it, so the span where the symptom appears is almost never the span where the cause is. The debugging move that works: start at the wrong output and walk upward through parents until you find the first span whose output was already wrong. That span, not the last one, is the bug.

The cost axis changes too. APM cost dimensions are CPU, memory and network. Agent cost is tokens and per-call model spend, attributable per span, which is how you discover the sub-agent that silently retries four times on every third request.

Then payload size. An APM span is tens of bytes of metadata. An LLM span can carry a 30 KB prompt plus twenty retrieved chunks, and that single fact is the root of most storage, truncation and UI-performance problems described later in this guide.

Side-by-side comparison of an APM trace and an agent trace Figure 2. Left, a microservice trace: uniform shape across runs, failure marked by a red 5xx span, payloads of tens of bytes. Right, the same request handled by an agent: variable shape between two runs of identical input, every span green, the defect sitting inside a tool span’s generated arguments, payloads measured in kilobytes.

The practitioner framing is worth quoting rather than paraphrasing. The author of Agd, posting on Hacker News, describes the status quo this way: “Every agent framework gives you logs (each its own flavour of logs). Unstructured text. Maybe some spans if you’re lucky. When your agent breaks something, you get to grep through a wall of output in some proprietary system” (news.ycombinator.com/item?id=47224560).

Anatomy of an agent trace: the six span types and what each must capture

Take one concrete request: a support agent receives “my payment method expired and I was double charged”. It classifies the ticket, retrieves policy documents, calls a billing API, hands off to a refund sub-agent, passes a guardrail check, and answers. Six span types, each with a minimum payload. Treat the list as a specification you can check your own instrumentation against, because this is the part of LLM tracing AI tracing for agents that determines whether anything downstream works.

Root / CHAIN span. User input, final output, thread ID, app version, prompt version, model version, environment, user or session ID. The version metadata earns its place the first time someone asks whether the regression started at prompt v14: that question is only answerable if v14 was stamped on the span when it ran.

LLM span. Model name and version, the full prompt or an explicitly redacted form of it, the completion, input and output token counts, sampling parameters, finish reason, and for streamed calls, time-to-first-token recorded separately from total duration.

Retrieval span. Query text, embedding model, top-k, the returned chunk IDs with their scores, and either the chunk text or a stable reference to it. Without scores you cannot tell “the retriever returned nothing relevant” apart from “the model ignored relevant context”, and those two bugs have nothing in common.

Tool span. Tool name, the exact arguments the model generated, the raw result, status, retry count, and whether the call had a side effect. The generated arguments are the single most diagnostic field in an agent trace. A refund issued for the wrong amount is visible there and nowhere else.

Agent / handoff span. Which agent ran, what instruction it received from its parent, its step count, and its termination reason: goal met, max steps exhausted, or error. Max-step exhaustion renders as success in most trace UIs unless termination reason is recorded explicitly, which makes it the quietest agent failure there is.

Custom spans. Guardrail verdicts, cache hit or miss, schema validation, post-processing. Cache spans matter for a structural reason: a cached response produces a trace with no LLM span, which looks identical to broken instrumentation.

A minimal tool span, written out so the shape is unambiguous:

{
  "name": "billing.get_charges",
  "span_kind": "TOOL",
  "trace_id": "8f2c1ab4...",
  "parent_span_id": "d41a9e02...",
  "start_time_unix_nano": 1759190400000000000,
  "end_time_unix_nano": 1759190400412000000,
  "status": "OK",
  "attributes": {
    "gen_ai.tool.name": "billing.get_charges",
    "tool.arguments": "{\"customer_id\":\"c_8812\",\"window_days\":30}",
    "tool.result.size_bytes": 2144,
    "tool.retry_count": 1,
    "tool.side_effect": false,
    "app.prompt_version": "v14",
    "session.thread_id": "th_5590c1"
  }
}

Annotated span tree for the double-charge support request Figure 3. The root/CHAIN span at the top, then the classify tool span, a retrieval span showing three chunk IDs with their similarity scores, the billing API tool span with its generated arguments, the refund sub-agent handoff span with two children of its own, the guardrail span and the response span. Each callout lists the attributes that span type must carry, with names taken from the current OTel GenAI semconv registry.

OpenTelemetry GenAI conventions: what is actually stable, and what has moved

The portability argument is simple and it is the only real one: spans that comply with a shared convention can be exported to more than one backend and survive a vendor switch, which is the sole hedge against paying for instrumentation twice. Every ranking page says the industry is standardising on OpenTelemetry. None of them says which attributes you can safely build a dashboard on.

AttributeHoldsStatus as publishedRename historyNotes
gen_ai.operation.nameOperation type (chat, embeddings, invoke_agent)Development / experimentalAdded after the original draft setEnumerated values have grown release to release
gen_ai.request.modelModel requested by the callerDevelopment, widely implementedStable name since early releasesSafest single attribute to key queries on
gen_ai.response.modelModel the provider actually servedDevelopment—Differs from request when a provider aliases versions
gen_ai.usage.input_tokensPrompt tokensDevelopmentRenamed from gen_ai.usage.prompt_tokensOld name still emitted by pinned instrumentation
gen_ai.usage.output_tokensCompletion tokensDevelopmentRenamed from gen_ai.usage.completion_tokensCost dashboards break silently on this pair
gen_ai.tool.nameTool invokedDevelopmentAdded with the tool/agent groupPair with your own arguments attribute
gen_ai.agent.*Agent id, name, descriptionDevelopment, least settled groupAdded later than the model groupHandoff semantics remain thin
Provider identityWhich provider served the callDevelopmentgen_ai.system superseded by gen_ai.provider.nameCheck both when querying historical traces
Prompt / completion contentMessage contentContested; moved between span events and attributesSee release notes for the events-to-attributes moveDetermines whether prompts survive a backend that drops events

Both columns come from the registry at https://opentelemetry.io/docs/specs/semconv/gen-ai/ and the release diffs at https://github.com/open-telemetry/semantic-conventions/releases. Read the diff for the release tag attached to each rename and note it beside your own instrumentation, because Confident AI’s guide raises the churn caveat without naming a single version and you will need one.

Map of old to new GenAI attribute names Figure 4. Old name on the left, current name on the right, one row per rename in the table above, with an arrow for each. A shaded band marks the pair that silently breaks cost dashboards. Blank tag fields sit beside each arrow so a team can fill in the release they upgraded on.

The operational consequence is three instructions. Pin your semconv version in the dependency file rather than floating it. Read the release diff before bumping. Expect every saved query and dashboard keyed on an attribute name to break on the bump, and budget an hour for it.

The content-location question deserves its own answer because it is load-bearing. If message content lives in span events and your backend samples or drops events, your prompts are gone and your durations remain, which is the worst possible outcome for LLM tracing AI tracing for agents. Resolve it from the spec, not from a vendor’s quickstart.

What the spec does not cover: conversation and thread grouping, evaluation scores attached to spans, and most multi-agent handoff semantics. Those parts of your data model are portable in name only. Acceldata’s page says “no proprietary format, no lock-in” on the strength of OTLP ingest (acceldata.io/ai-observability/agent-llm-tracing). OTLP ingest says nothing about whether your eval scores, curated datasets and human annotations can leave, and that is where lock-in actually lives.

Keep three layers distinct: the semantic conventions are the naming standard, OpenInference and OpenLLMetry are instrumentation libraries that emit spans, and OTLP is the wire protocol. Sitting underneath all three is W3C Trace Context, the traceparent header, which is what keeps one trace intact when your agent calls your own service rather than a model provider. Related reading: OpenTelemetry GenAI attribute reference.

Four ways to instrument, with code for each

Auto-instrumentation. One line patches the provider client.

import mlflow
mlflow.openai.autolog()          # https://mlflow.org/docs/latest/genai/tracing/

Coverage of every model call in an afternoon. Retrieval, tool dispatch and orchestration stay invisible, which for an agent is most of the interesting surface.

Decorator or manual SDK spans.

from mlflow import trace

@trace(span_type="TOOL")
def get_charges(customer_id: str, window_days: int) -> dict:
    ...
from langfuse import observe

@observe(name="refund_subagent")
def run_refund_agent(instruction: str) -> str:
    ...

“Langfuse observe decorator” is a related search, which means readers arrive already looking for this pattern. What decorators cannot see: span boundaries inside one long function, and anything running in a thread or task their context does not reach.

Framework integrations. LangChain and LangGraph callbacks, LlamaIndex, DSPy, Pydantic AI, CrewAI, AutoGen, the OpenAI Agents SDK, the Vercel AI SDK, AWS Strands, and Google’s Agent Development Kit:

from vertexai.preview.reasoning_engines import AdkApp
app = AdkApp(agent=root_agent, enable_tracing=True)   # spans land in Google Cloud Trace

Spans match the framework’s own abstractions, which is pleasant until you need a boundary the framework does not model. Coverage stops at the framework edge, so your pre- and post-processing is dark. For programmatic analysis, pull the Cloud Trace spans into a DataFrame and group by span name and status; the pattern is credited to the public Agent Engine walkthrough at medium.com/@martinhadid, whose pinned gemini-2.0-flash now dates it.

Raw OpenTelemetry plus a collector. Emit standard spans, route OTLP to one backend or two. You get multi-homing and reuse of an existing collector fleet, and you own the semconv-version problem yourself.

ApproachSetup costWhat it seesWhat stays darkProne to
Auto-instrumentationMinutesEvery provider call, tokens, latencyRetrieval, tools, orchestration, your own logicDuplicate spans when combined with a framework integration
Decorator / manualHours to daysExactly the boundaries you chooseAnything crossing a context the decorator does not followOrphan spans across async and thread boundaries
Framework integrationUnder an hourFramework-shaped spans, handoffs includedCode outside the framework; granularity is fixedVersion coupling; span names change on framework upgrades
Raw OTel + collectorDaysWhatever you emit, exportable anywhereNothing, if you do the workAttribute drift on semconv bumps; collector config errors

The decision rule, stated rather than hedged: auto-instrument for model-call coverage, add manual spans at every point where a model output becomes an action (tool arguments, SQL, API payloads, handoff instructions), and use the framework integration only if you are not already emitting OTel spans yourself. Enabling both usually produces duplicated or double-nested span trees, and that confusion costs more than the coverage gains.

On friction, one practitioner’s view deserves attribution rather than dismissal: the developer behind tracium wrote that he looked at “LangSmith and LangFuse” and “they felt like they are overcomplicated to integrate and that there should be a simpler way” (news.ycombinator.com/item?id=47247631). That is one opinion, not a benchmark. The part he is most likely describing is project and key configuration plus deciding where span boundaries belong, not the install, and no tool can decide the second one for you.

Decision flowchart for choosing an instrumentation approach Figure 5. “Which instrumentation approach should I start with?” branching on whether a supported framework is in use, whether an OTel collector already exists, whether payloads contain regulated data, and whether the app is multi-turn. Each terminal node names an approach from the table above and the one capability it leaves dark.

What actually breaks: a failure-mode catalogue for agent tracing

Seven of the eight pages currently ranking for this topic are published by a vendor selling the tool being described, and not one documents a single failure mode. Here are the six that matter. Each row gives the mechanism and the search that surfaces the public reports; no issue number is invented here, and when you run the search yourself, record the issue URL, its number, its open or closed state and the date you read it.

Failure modeSymptom in the trace UICauseWhere the evidence livesWorkaround
Broken parent-child links across async boundariesSpans appear as orphan root traces, not a treeContext does not propagate implicitly into a new task, thread or processSearch github.com/langfuse/langfuse issues for “nested observation” and “async”; Arize-ai/openinference for context propagationExtract the context at the boundary and inject it explicitly in the child
Traces lost at process exitRun completes, no trace ever appearsBatching exporter has not flushed when the interpreter exitsSearch github.com/mlflow/mlflow issues labelled area/tracing for “flush”Call the SDK’s explicit flush in a finally block or an atexit hook
Silent payload truncationPrompt or chunk array ends mid-string; no warningPer-span or per-attribute size limits in SDK, collector or backendEach platform’s documented limits pageStore large payloads by reference; log a truncation marker attribute
Misleading streamed-span durationsLatency looks impossibly fast or uniformly slowSpan closes on first token, or closes on stream end with no first-token recordFramework streaming issue trackersRecord time-to-first-token and total duration as separate fields
Duplicate or double-nested tracesTwo trees per request, or a tree inside itselfAuto-instrumentation and a framework integration both activeReproducible from any two-integration setupPick one emitter per layer; assert one root span per request in CI
Self-hosted storage and attachment problemsMedia fails to upload; database grows without boundObject-store configuration, retention jobs not runningSelf-hosting issue labels in the relevant reposVerify the object-store path before production; monitor table growth

Diagram of a trace tree broken at an async boundary Figure 6. Above, the intended tree: root span, tool span, and a child span created inside an async task. Below, what the UI shows when context does not cross the boundary: two unrelated root traces, the second one stripped of its parent and of every attribute the root carried. An inset shows the explicit extract-and-inject that repairs it.

Mode two rewards a closer look, because it is the same mechanism as a headline feature. Langfuse documents that its SDKs send data asynchronously, queueing events locally and flushing in batches so application response time is not affected (langfuse.com/docs/observability/overview). That design is exactly why tracing adds no request latency and exactly why a short-lived script, a serverless handler or a notebook cell loses spans when it exits before the queue drains. One property, two consequences.

The honest limit of this exercise: an issue tracker shows what people reported, not how often it happens. It cannot rank platforms by reliability, and it cannot tell you how any SDK behaves under sustained load. Those questions require hands-on testing this page has not done. If you need them answered, run a trial that replays a realistic span volume through two candidates and measures dropped-span rate and ingest lag yourself.

Correction route: if you maintain a tool named above and a row is wrong, send the issue link and the current state to the contact address on this site. Corrections are logged inline with the date they were applied.

Sampling, cardinality and retention: keeping trace volume survivable

Coverage decisions come before sampling decisions. Instrument the steps most likely to fail silently first: retrieval, every point where a model output becomes a tool argument, and the final response. Expand to intermediate reasoning and sub-agent calls once those three are clean.

Naive head sampling is worse for agents than for microservices. At 5%, you almost never hold the trace for the rare failure you were hired to explain, because the failures that matter are by definition uncommon. The pattern that works instead: always capture a lightweight trace with metadata and structure, and retain the full payload only for traces that error, trip a low-confidence or guardrail signal, fail an eval, or get flagged by a reviewer. Failure-biased retention is the single highest-leverage policy in LLM tracing AI tracing for agents.

Sample at the trace level, never the span level. Independently sampled spans produce trees with holes, and a tree with holes is worse than no tree because it looks complete.

Retention tiering, by the question each tier answers: full payloads hot for days answer “what exactly went wrong in this incident”; metadata and eval scores warm for months answer “when did this regress”; daily aggregates kept indefinitely answer “is quality trending down”. Regulated environments are the exception where full capture is mandatory and sampling is not a lever at all. The levers there are redaction and self-hosting.

MLflow documents sampling-ratio and async-logging controls, which is the shape of control to look for in any SDK (mlflow.org/docs/latest/genai/tracing/). Check whether the control is documented for the production package specifically, since MLflow ships mlflow-tracing separately from the full library. See also: self-hosting storage and retention burden.

Prompts are PII: redaction, masking and self-hosting

The whole value of an LLM span is its payload, and the payload is user text that may contain names, account numbers, medical detail or card data. Redact everything and LLM tracing AI tracing for agents degrades into APM with a larger bill.

Graduated options, and what each costs you diagnostically:

OptionKeepsLosesAdded risk
Field-level maskingStructure, shape, lengthsThe specific fact that caused the failureLow
Entity-level redaction via classifierMost diagnostic valueEntities the classifier misses stay in the payloadInline latency, plus a new component that can fail
HashingJoin-ability across spans and tracesAll contentLow; re-identification risk if the space is small
Tokenised reference to your own bucketEverything, under your controlConvenience; the UI needs a resolverTwo systems to keep in sync

The rule is: redact, do not drop. Spans with durations and no payload cannot answer the question you added tracing to answer.

Verify one thing in any vendor’s documentation before trusting it: whether redaction runs client-side in the SDK before transmission, or server-side after ingestion. Only the first keeps raw prompts off the vendor’s network, and only the first satisfies most data-residency reviews. The answer differs per tool, and it lives in the docs rather than on the marketing page.

Self-hosting is a data-residency and payload-sensitivity decision, not a price decision. It also moves every failure mode in the catalogue above from the vendor’s on-call rotation to yours, particularly the object-store and retention-job rows. One more argument for capturing tool spans in full: the author of Telos argues that agents are routinely given “shell access and API keys, relying on system prompts or Docker for security”, and that prompt-injected agents exfiltrate data using signed binaries like curl or base64 that look ordinary to the operating system (news.ycombinator.com/item?id=47246950). The tool-call span is often the only place that behaviour is legible. Related: PII redaction for LLM applications.

What LLM tracing costs: a unit-based model from published pricing

No page in the current top ten states a price. Here is a model you can re-run.

Assumptions box. 50,000 agent requests per month. 14 spans per request: one root, two retrieval, six LLM, four tool, one guardrail. Average LLM span payload 6 KB; average request payload across all spans roughly 50 KB. That is 700,000 spans per month, 8.4 million per year, and about 2.5 GB of raw payload per month (30 GB per year before compression). Change any of the four inputs and the arithmetic below follows directly.

Billing units are the trap. “Units”, “events”, “spans”, “traces” and “ingested GB” are not the same thing, so headline prices are not comparable until you normalise. At 14 spans per request, a platform billing per trace charges you for 50,000 and a platform billing per span charges you for 700,000: a 14x difference in the same workload before any discount.

Progress Telerik Agent Engineering publishes concrete tiers, which makes it the only usable anchor on the page: Free at 10,000 units with 7-day retention, Starter at 200,000 units with 30-day retention and $8 per additional 100,000 units, Pro at 1,000,000 units with 60-day retention, and Enterprise from $3,000 per month with infinite retention (telerik.com/ai-engineering/agent-tracing-observability, read September 2026). Their docs define what a unit is, and everything below turns on that definition, so read it first. If one unit is one span, the reference workload overruns Starter by 500,000 units, or five overage blocks at $8, so $40 per month in overage on top of whatever Starter’s base costs, and fits inside Pro with 300,000 units of headroom. At 10x volume (7 million spans per month), Pro plus 6 million overage units is 60 blocks, $480 per month in overage alone, or $5,760 a year before the base.

For Langfuse, LangSmith, Arize and Datadog LLM Observability, pull the rate and the billing unit from each pricing page and record the date read; they are not modelled here because no rate is quoted in a source I can cite. Datadog’s Agent Observability page routes pricing entirely to FAQ links rather than stating a rate (datadoghq.com/products/ai/agent-observability/), which is itself the finding: when a vendor’s real rate is not publishable in a table, say so instead of guessing.

Self-hosted, the licence line is zero and the other lines are not. 30 GB of raw payload a year compresses well in a columnar store, so storage is the cheapest part of the bill. The real costs are a Postgres instance, a ClickHouse or equivalent analytics store, an object store for attachments, and engineer hours per upgrade cycle that only your own timesheets can price.

The line item everyone forgets is evaluation. Running an LLM judge over a 5% sample of the reference workload means 35,000 judged spans a month; at roughly 2,500 tokens per judge call that is about 87.5 million tokens a month, billed by your model provider, not your tracing vendor. It scales with sample rate, not trace count. Datadog’s own FAQ carries a dedicated “What do evals cost?” question, which tells you how often this surprises buyers.

Retention is the steepest driver. A 7-day or 30-day window means the trace for last quarter’s incident is simply gone, and extended retention is where every price list bends upward.

Vendor efficiency claims are not inputs to this model. MLflow’s “95% smaller footprint” for the mlflow-tracing package and Datadog’s “40% lower token usage per task” are vendor self-reports with no published methodology. Telerik’s “85% faster root cause analysis” and “3x faster time to resolution” carry no sample size, method or date. Acceldata’s “40% reduction in pipeline downtime”, “30% faster time-to-model” and “99.9% SLA adherence” describe data pipelines rather than tracing and are reused as tracing proof points. Cite all of them as claims; model with none of them.

Grouped bar chart of monthly tracing cost at 700,000 and 7 million spans Figure 7. Monthly cost for the 700,000-span reference workload beside the same workload at 10x, one group per vendor that publishes a rate, plus self-hosted infrastructure as a separate bar. Each vendor’s billing unit is labelled under its bar and the spans-to-unit normalisation arithmetic runs inline above it. Vendors that publish no rate are absent rather than estimated.

More on the token side of the ledger: LLM cost tracking per request.

Tracing tools for agents, compared

This is a documented-capability matrix, not a quality ranking. A cell reads Yes only where the vendor documents the capability. Unconfirmed means the documentation was not read for that cell here, which is not the same as the capability being missing. Not documented is reserved for cases where a vendor says nothing about a capability at all.

ToolLicenceSelf-hostOTel GenAI ingestOTel export outThread groupingEvals on prod tracesDataset curationAnnotation queue
MLflow TracingApache 2.0, Linux FoundationYesUnconfirmedUnconfirmedUnconfirmedYesYesUnconfirmed
LangfuseOpen coreYes, documentedYesUnconfirmedYes, sessionsYesYesYes
LangSmithProprietaryUnconfirmedUnconfirmedUnconfirmedYes, threadsYesYesYes
Arize PhoenixOpen sourceYesYes, via OpenInferenceUnconfirmedUnconfirmedYesYesUnconfirmed
Confident AI / DeepEvalDeepEval open source; platform proprietaryUnconfirmedUnconfirmedUnconfirmedYes, threadsYesYesYes, annotation-triggered workflows
Datadog LLM ObservabilityProprietaryNoUnconfirmedVia Datadog export APIs, unconfirmedUnconfirmedYesUnconfirmedYes
Google Cloud Trace + Vertex Agent EngineProprietaryNoYes, OTel-nativeYes, via Cloud Trace APINot documented as threadsSeparate product surfaceUnconfirmedUnconfirmed
Progress Telerik Agent EngineeringProprietaryUnconfirmedUnconfirmedUnconfirmedUnconfirmedUnconfirmedUnconfirmedUnconfirmed
AcceldataProprietaryUnconfirmedYes, OTLPUnconfirmedUnconfirmedUnconfirmedUnconfirmedUnconfirmed
Monte CarloProprietaryNoUnconfirmedUnconfirmedUnconfirmedUnconfirmedUnconfirmedUnconfirmed

Two distinctions predict more than any feature cell. First, whether the product was built around tracing (MLflow, Langfuse, LangSmith, Phoenix, Confident AI) or whether tracing is a new surface on an existing observability platform (Datadog, Acceldata, Monte Carlo). That tells you which internal stakeholders you will be negotiating with and whose budget line it lands on.

Second, the export column. Many tools accept OTLP inbound; far fewer let you stream traces, eval scores and curated datasets back out. That asymmetry is the working definition of lock-in in this category, and it is the question to ask on the first sales call.

The open-source instrumentation layers deserve naming separately, because they are what you actually install: OpenInference from Arize and OpenLLMetry from Traceloop both emit spans for provider SDKs and frameworks, and either can point at any backend that speaks OTLP, including Jaeger, Grafana Tempo, Zipkin, Honeycomb, New Relic or Elastic. Other products sell into the same category and are absent from the matrix only because their documentation was not read for this page: Weights & Biases Weave, Braintrust, Helicone, Literal AI, LangWatch, Portkey, Laminar and HoneyHive.

Newer entrants, clearly early-stage and described in their authors’ own words: Auditi, “an open-source LLM tracing and evaluation platform”, whose author frames the gap as “Tracing tells you what happened. But I wanted to know how well it happened” (46974783); Agd, a content-addressed DAG for tracking agent actions (47224560); tracium, a one-line agent observability tool (47247631). No ranking implied. What is useful is the problem each author says existing tools left open.

One adjacent category gets shopped for by mistake constantly. Agent simulation and regression testing generates traffic; tracing observes production. The Cekura founders describe running voice agent simulation for 1.5 years to “simulate real user conversations, stress-test prompts and LLM behavior, and catch regressions” (47232903). Complements, not substitutes, and buyers frequently arrive wanting one while asking for the other. See simulation vs production tracing and the fuller LLM observability tools compared.

No cell in that table says anything about reliability under load. Nobody publishes that. Decide it with a trial.

Screenshot of a nested agent span tree in a trace explorer Figure 8. A nested agent span tree in the MLflow trace explorer, captured from the publicly hosted demo instance at demo.mlflow.org. The data is the demo’s own sample data, not a production system belonging to this publisher.

From traces to evals to datasets: the loop that pays for instrumentation

A trace store needs consumers or it becomes a write-only archive with a monthly invoice. There are exactly three worth building: debugging individual failures, running evaluations against production traces, and curating regression datasets from real traffic.

Evaluation works on traces precisely because spans carry semantic payloads. Faithfulness scores against the chunks in the retrieval span. Tool correctness scores against the arguments in the tool span. Answer relevancy scores against the root span’s input and output. Each of those metrics is computable only if the corresponding span captured its payload, which is why the span specification earlier in this guide is the load-bearing section rather than a reference appendix. See LLM evaluation metrics explained and LLM-as-a-judge on production traces.

Dataset curation is a pipeline, not a feature. Traces that fail an eval or get flagged in review become test cases, so the next prompt change is tested against failures that actually happened. Copy inputs, retrieval context and expected behaviour into the dataset; reference large attachments rather than duplicating them. Building regression datasets from traces covers the mechanics.

Human review is the third loop, and it is the part of the category that moved most recently, so read each changelog entry against its own date rather than trusting a summary. Confident AI’s changelog records annotation-triggered workflows that fire when a reviewer annotates, plus free-text search across traces and spans where filtering was previously condition-by-condition (confident-ai.com/docs/changelog). Langfuse’s changelog records running evaluators over historical observations and backfilling scores when a rule is attached after the fact, alerts on evaluator results, and an assistant that runs code over thousands of observations in a sandbox (langfuse.com/changelog). MLflow 3.16.0 shipped natural-language-defined custom trace views, a redesigned trace explorer as the default, and first-class span links for connecting related spans across traces (mlflow.org/releases). Span links matter specifically for agents: they are how a trace connects to the trace of a job it spawned, which is the one relationship a parent-child tree cannot express.

The honest limit: tracing tells you what happened, not whether the data feeding the request was fresh or correct. Stale embeddings and a broken upstream job produce perfectly well-formed traces. That is the legitimate core of Monte Carlo’s argument, and it holds regardless of the fact that the page making it is selling something.

A 90-minute first instrumentation, in order

  1. Turn on auto-instrumentation for your model provider and confirm one trace renders end to end. Nothing else is worth doing until a trace appears.
  2. Add an explicit flush at process exit. Short-lived runs lose spans until you do, and you will misdiagnose step 1 without it.
  3. Stamp app version, prompt version, model version and environment on the root span. Regressions become attributable from this moment forward, and never retroactively.
  4. Add a stable thread or session ID, for the same retro-fit-impossible reason.
  5. Add manual spans wherever a model output becomes an action: tool arguments, generated SQL, API payloads, handoff instructions.
  6. Add retrieval spans with chunk IDs and scores, which is what separates a retriever bug from a grounding bug.
  7. Set the redaction policy before pointing production traffic at the backend. Reversing this order means raw prompts already left your network.
  8. Pin your semconv version and note the tag in the repo.
  9. Attach one eval to one span type. One is enough to prove the loop closes.
  10. Create the review queue that will actually read the traces.

Steps 3 and 4 are the two teams skip and regret, because both are impossible to backfill.

The pre-production gate is a single question: given a user complaint and a timestamp, can someone reach the responsible span in under two minutes? If not, the instrumentation is unfinished no matter how many spans it emits.

Do this before the first incident, not during it. The value of LLM tracing AI tracing for agents is comparative, and a trace store with no history has nothing to compare against.

Five mistakes that make a tracing setup not worth its cost

Metadata without payloads. Durations and status codes only, which converts an expensive tracing bill into a slow APM.

No version metadata. The regression is real, the trace is there, and nothing connects it to the change that caused it.

Request-level tracing on a multi-turn product. Every turn passes, the conversation fails, and the measurement unit was wrong from the start.

Tracing with no consumer. No eval reads it, no alert fires on it, no human reviews it. Structured logging with a bigger invoice.

Retention shorter than your incident-review latency. If incidents get reviewed monthly and retention is 30 days, the evidence expires before anyone opens the ticket.

Unbounded attributes. A user ID or a full prompt placed in a span name or an indexed tag rather than a plain attribute is how trace search gets slow and bills get strange. Keep high-cardinality content in attributes, and keep span names enumerable.

LLM tracing AI tracing for agents: what to decide this week

Four decisions carry the rest. Pick the instrumentation approach from the matrix and commit to one emitter per layer. Pin a semconv version and write the tag down. Choose failure-biased retention over flat head sampling before volume forces the question. Decide redaction client-side or server-side before production traffic points anywhere. Everything else in LLM tracing AI tracing for agents can be changed later; those four are the ones that make you re-instrument.

Frequently Asked Questions

What is LLM tracing?

LLM tracing records every operation in one request as a timestamped span: model calls, retrieval queries, tool invocations, guardrail checks, sub-agent handoffs. Spans share a trace ID and nest into a tree that reproduces the execution path. The differentiator from logging is that spans carry semantic payloads, not just durations and status codes.

What is the difference between a trace, a span and a thread?

A span is one operation and answers “which step failed”. A trace is the span tree for one request and answers “what happened on this request”. A thread is the ordered set of traces in one conversation and answers “did the conversation succeed”. Threads are a backend concept, not an OpenTelemetry-standardised one.

How is AI tracing for agents different from tracing a normal application?

APM hunts exceptions and 5xx responses; agent tracing hunts clean 200s carrying wrong answers, because nothing throws. Three structural differences follow: spans must carry prompts and retrieved context, execution paths are non-deterministic so trace shape varies between runs, and cost is measured in tokens per span rather than CPU.

Do I need OpenTelemetry for LLM tracing?

No, but it is the safest default, because compliant spans stay portable across backends and survive a vendor switch. The caveat competitors omit: the GenAI conventions are still marked in development, attribute names have changed between semconv releases, and they standardise neither thread grouping nor eval scores. Pin your version.

What should an LLM span capture?

Operation type, start and end time, parent span ID, and the payload for that span type: prompt, completion, token counts and finish reason for LLM spans; query and scored chunks for retrieval spans; generated arguments and raw result for tool spans. Prompt, model and app version metadata is what makes a regression attributable.

Does adding tracing slow my agent down?

Not in steady state. SDKs queue spans locally and flush asynchronously in batches, so the request path is not blocked; Langfuse documents exactly this (langfuse.com/docs/observability/overview). The tradeoff nobody states: that same batching is why short-lived processes lose spans without an explicit flush, and inline redaction does add latency.

Do I need to trace every request?

No at high volume, yes in regulated environments. Better than flat sampling: capture lightweight traces on everything, then retain full payloads only for requests that error, trip a guardrail or low-confidence signal, fail an eval, or get flagged in review. Flat head sampling systematically discards the rare failures you need most.

What is the best open source LLM tracing tool?

There is no single winner; pick on criteria. Licence and self-host path, whether it exports OTel as well as ingesting it, thread grouping, evals on production traces, and dataset curation without manual export. MLflow Tracing, Langfuse, Arize Phoenix and DeepEval cover most needs, with OpenInference or OpenLLMetry as the instrumentation layer.

How much does LLM tracing cost?

For 50,000 requests a month at 14 spans each (700,000 spans), Telerik’s published tiers put the workload inside Pro’s 1,000,000 units, or roughly $40 a month in overage on Starter at $8 per 100,000 units. Three drivers dominate: billing unit, retention length, and eval inference billed separately.

How do I trace a multi-agent system with handoffs?

Give each agent invocation its own agent span containing its children. Propagate trace context explicitly across async, thread and process boundaries, because implicit propagation stops at those edges. Record the instruction passed on handoff plus the termination reason, carry one thread ID across the conversation, and use span links for spawned jobs.

Is tracing the same thing as LLM observability?

No. Tracing is the capture layer: what the application did on one request. Observability is what gets built on top, including dashboards, alerting, evaluation on traces, drift detection, dataset curation and human review. The consequence: tracing with no consumer is structured logging with a bigger invoice. See what is LLM observability.

Can I keep prompts out of my tracing vendor’s systems?

Yes, but verify how. Client-side redaction in the SDK before transmission keeps raw prompts off the vendor’s network; server-side redaction after ingestion does not. Fallbacks are hashing for join-ability, tokenised references to payloads in your own bucket, or self-hosting. The rule holds throughout LLM tracing AI tracing for agents: redact, do not drop.

Frequently Asked Questions

What is LLM tracing?

LLM tracing records every operation in one request as a timestamped span: model calls, retrieval queries, tool invocations, guardrail checks, sub-agent handoffs. Spans share a trace ID and nest into a tree that reproduces the execution path. The differentiator from logging is that spans carry semantic payloads, not just durations and status codes.

What is the difference between a trace, a span and a thread?

A span is one operation and answers "which step failed". A trace is the span tree for one request and answers "what happened on this request". A thread is the ordered set of traces in one conversation and answers "did the conversation succeed". Threads are a backend concept, not an OpenTelemetry-standardised one.

How is AI tracing for agents different from tracing a normal application?

APM hunts exceptions and 5xx responses; agent tracing hunts clean 200s carrying wrong answers, because nothing throws. Three structural differences follow: spans must carry prompts and retrieved context, execution paths are non-deterministic so trace shape varies between runs, and cost is measured in tokens per span rather than CPU.

Do I need OpenTelemetry for LLM tracing?

No, but it is the safest default, because compliant spans stay portable across backends and survive a vendor switch. The caveat competitors omit: the GenAI conventions are still marked in development, attribute names have changed between semconv releases, and they standardise neither thread grouping nor eval scores. Pin your version.

What should an LLM span capture?

Operation type, start and end time, parent span ID, and the payload for that span type: prompt, completion, token counts and finish reason for LLM spans; query and scored chunks for retrieval spans; generated arguments and raw result for tool spans. Prompt, model and app version metadata is what makes a regression attributable.

Does adding tracing slow my agent down?

Not in steady state. SDKs queue spans locally and flush asynchronously in batches, so the request path is not blocked; Langfuse documents exactly this ([langfuse.com/docs/observability/overview](https://langfuse.com/docs/observability/overview)). The tradeoff nobody states: that same batching is why short-lived processes lose spans without an explicit flush, and inline redaction does add latency.

Do I need to trace every request?

No at high volume, yes in regulated environments. Better than flat sampling: capture lightweight traces on everything, then retain full payloads only for requests that error, trip a guardrail or low-confidence signal, fail an eval, or get flagged in review. Flat head sampling systematically discards the rare failures you need most.

What is the best open source LLM tracing tool?

There is no single winner; pick on criteria. Licence and self-host path, whether it exports OTel as well as ingesting it, thread grouping, evals on production traces, and dataset curation without manual export. MLflow Tracing, Langfuse, Arize Phoenix and DeepEval cover most needs, with OpenInference or OpenLLMetry as the instrumentation layer.

Explore More

Free Newsletter

Get the LLM Evals Newsletter

Platform comparisons, pricing changes and eval technique deep-dives. No spam.

Related Articles