comparison

LangSmith Alternatives (2026): 12 Tools, Honestly Compared

Every LangSmith alternative guide is written by a vendor. This one isn't. 12 tools compared on real pricing, self-host cost, licences and open GitHub issues.

Published:


title: “What Is the Alternative to LangSmith? 12 Tools, Honestly Compared” description: “Compare what is the alternative to LangSmith across 12 tools by blocker — price, self-hosting, retention, CI gating and the AWS-native path. See which one clears yours.” slug: what-is-the-alternative-to-langsmith

What Is the Alternative to LangSmith? 12 Tools, Honestly Compared

The short answer to what is the alternative to LangSmith

What is the alternative to LangSmith depends on which LangSmith constraint is actually blocking you. Langfuse is the most widely adopted open-source replacement and the one most teams land on. MLflow is the Apache-2.0 choice if you already run a data platform and want no per-trace fee. Braintrust is strongest if eval results must gate CI. Arize Phoenix or Helicone will do if you only need tracing. CloudWatch with AgentCore Observability is the answer if your inference already runs on Amazon Bedrock.

What is the alternative to LangSmith is really five questions wearing one costume, so start from the blocker.

Your blockerWhat to look atWhy
$39/seat/month is taxing who can look at tracesLangfuse, Helicone, BraintrustUsage-priced or unlimited-user tiers; reviewers cost nothing to add
Prompt payloads cannot leave your infrastructureSelf-hosted Langfuse, MLflow, Arize PhoenixFree-tier self-host, no vendor DPA needed
Retention window is too short for your audit trailMLflow, self-hosted LangfusePoint them at your own S3, GCS or Azure Blob Storage
LangChain coupling worries youTraceloop/OpenLLMetry, Langfuse, PhoenixOpenTelemetry-native ingestion, portable spans
You just want to see traces on your laptopLocal Phoenix or Langfuse in Docker, JaegerNo account, no cloud round-trip

One thing to declare before you read further. This page sells none of these tools, and every other result currently on page one for this query is published by a vendor that appears in its own list.

Decision flowchart from blocker to tool Five yes/no branches (data must stay in our infra? Bedrock-native? bill is seats or volume? need merge-blocking gates? need runtime guardrails?) terminating in named tools, with a “stay on LangSmith” terminal node.

Read this first: who wrote the other results, and who they crowned

The search results you just clicked through are a vendor battlefield. Each page is a comparison written by one of the things being compared, and in every case the author wins.

PagePublisherCommercial interestRanked #1
langfuse.comLangfuseSells Langfuse CloudLangfuse
braintrust.devBraintrustSells Braintrust Pro/EnterpriseBraintrust (“Winner” in 4 of 5 rows of its own head-to-head)
mlflow.orgMLflow (Databricks-originated)Databricks managed MLflowMLflow
orq.aiOrq.aiSells Orq.aiOrq.ai, first of six
lunary.aiLunarySells Lunary Team planLunary, first of five
zenml.ioZenMLSells ZenML orchestrationZenML as the layer above both

The mechanism matters more than the motive. A vendor’s feature matrix is a selection of axes on which that vendor wins. Braintrust’s matrix includes CI/CD gating, a genuine Braintrust strength, and omits licence terms and self-host cost. Langfuse’s matrix includes licence and seat fees, and omits CI/CD gating entirely. Neither page is lying. Both chose the columns.

So here is the rule. Treat any claim a vendor makes about a competitor’s roadmap, storage engine or pricing ceiling as a lead to verify, not a fact. Langfuse’s page, for instance, makes granular dated claims about LangSmith’s storage engine, its retention cap, the metering unit on its tuned evaluators and which accounts get bulk export. Those are claims about LangChain made by LangChain’s direct competitor, and none of them are repeated as established fact here. Check LangChain’s own pricing page and changelog before you budget on any of them.

Two of the ranking pages are simply out of date. The dev.to comparison carries its author’s own warning that the data is accurate as of December 2024, and prices Langfuse at $100/user/month for self-hosted LLMOps and $60/user/month for cloud, figures that no longer match Langfuse’s published tiers. ZenML’s November 2025 table lists Langfuse at “17.9k+” GitHub stars against 35,158 on the repo today, and attributes “670+” GitHub stars to “LangSmith”, which cannot be the platform, because the platform is closed source.

Vendor-bias audit table For each of the ten ranking results: publisher, commercial interest, tool ranked #1, and the comparison axes the page chose to include or omit.

Five distinct complaints behind leaving LangSmith

These get lumped together as “cost”. They are five problems with five solutions.

Seat-based pricing taxes visibility. LangSmith Plus is published at $39/seat/month, so a ten-person team is $390/month before a single trace is billed (verify on LangChain’s pricing page; this figure is also repeated by competitors, which is not a source). The structural problem is who gets excluded. The people you most want inside an observability tool are PMs, domain reviewers and support leads, and those are exactly the seats hardest to justify to finance. You end up with a quality tool that only engineers can open.

Per-trace volume pricing is fine in a pilot and painful in production. Agent traffic does not grow linearly with users. One user turn in a deep-agent system can fan out into dozens of spans across sub-agents, retries and tool calls, so the trace count outruns the user count.

Retention behaves as a second meter. The included window is short, and keeping traces longer is an upgrade. That interacts badly with online evaluators, because scoring a trace tends to promote it into the tier you are paying extended retention on. Confirm the current windows on LangChain’s pricing docs and screenshot them for your finance case. These numbers move.

Self-hosting is an Enterprise motion, not a Docker Compose file. Developer and Plus users have no self-host option, which means trace payloads containing raw user prompts go to a vendor cloud. That is a compliance blocker, qualitatively different from a cost blocker. No discount fixes it.

Framework gravity cuts both ways. LangSmith’s zero-config path is one environment variable for LangChain and LangGraph, a real engineering advantage. The further your stack drifts from LangChain, the more of that advantage evaporates while the pricing stays exactly where it was.

Now the part the listicles skip. LangSmith is genuinely good at several things, and a page that cannot say so is not worth trusting on the criticism either. One-env-var tracing is the fastest instrumentation path in the category. LangSmith Studio’s visual agent building has no real open-source equivalent. Annotation queues are mature, native alerting ships destinations including PagerDuty and Dynatrace, and production insights clustering groups similar failures without you writing the grouping logic. MLflow’s own comparison page concedes most of this. The listicles mostly do not.

How we scored these tools, and the four axes vendors leave out

Six standard axes, one line each:

  1. Instrumentation surface. OTel-native, SDK-only, or proxy.
  2. Evaluation depth. Offline datasets, online judges, human annotation, CI gating.
  3. Prompt management. Versioning, deploy without a code change, playground.
  4. Pricing unit. Seat, trace, span, GB, or event.
  5. Deployment modes. SaaS, BYOC, self-host, air-gapped.
  6. Compliance artefacts. SOC 2 Type II, ISO 27001, HIPAA, and whether a BAA gets signed.

Four axes the vendor pages skip:

Licence reality at the file level. Open-core repos commonly ship an enterprise-licensed subdirectory alongside an MIT core. GitHub’s own metadata for langfuse/langfuse currently resolves the licence as NOASSERTION rather than MIT, which is exactly the thing to open the LICENSE file about before telling your legal team “it’s MIT”.

Total self-host cost, including every component you must operate and keep upgraded.

Failure-mode load. What is broken today according to the public issue tracker, with URLs and reaction counts, not “users report”.

Exit cost. Whether you can take your traces and your instrumentation somewhere else.

The pricing-unit trap decides more comparisons than any feature does. A tool that bills per trace and a tool that bills per span cross over at some spans-per-trace figure. A 5-span RAG chain and a 60-span deep-agent run will rank the same two vendors in opposite orders. Compute your average spans per trace before you compare anything. It is one query against whichever tool you run today.

Twelve answers to what is the alternative to LangSmith, by job

Every entry uses the same template: identity, best for, pricing unit and published entry price, deployment modes, licence, the strongest reason to pick it, the strongest reason not to, and where to verify. Every tool gets at least one concrete, sourced limitation.

Prices are list prices as of 29 September 2026 and should be re-checked against each vendor’s own page before you commit budget.

Langfuse, the default open-source answer

OTel-based ingestion, an observations-first data model, prompt management, LLM-as-a-judge and code evaluators, datasets and experiments, annotation queues. Usage-unit pricing with no seat fee: a free Hobby tier plus Core and Pro tiers published on langfuse.com/pricing, and a free self-host path on top.

Repo health is verifiable and strong: 35,158 GitHub stars, v4.46.0 released 2026-09-25, last push three days later.

Then the part no ranking page includes. The same repo carries 947 open issues, and these are the highest-signal ones.

IssueSymptomSignalAffects
#9618Python 3.14 incompatibility, attributed by the reporter to Pydantic v1 usage179Anyone planning a 3.14 upgrade
#11109Experiments and Evaluators fail to parse Anthropic “thinking” blocks from reasoning models such as claude-haiku-4-562Eval pipelines on reasoning models
#8780Repeated “Failed to detach context” from the LangChain CallbackHandler under async ainvoke62LangGraph agents
#5704Importing langfuse in Jest throws “A dynamic import callback was invoked without —experimental-vm-modules”57TypeScript test suites
#12614Dashboard widgets cannot break down by metadata fields such as organizationId56Multi-tenant reporting
#2169Python SDK ships no py.typed marker, so it is not PEP 561-compliant51mypy users downstream
#11874Filtering traces by numeric score does not return the expected traces50Triaging low-scoring judge output
#3961FastAPI StreamingResponse resets tracing context after the first chunk, splitting one request into multiple traces49Streaming APIs
#4555Multi-modal images uploaded via LangfuseMedia land in S3 but render as a button, not inline41Vision app debugging

Read that list as a map, not a verdict. Issue age is not severity, several may have shipped fixes since, and every one should be checked for current state before it changes your decision. The useful signal is the pattern. Async context propagation and multi-tenant reporting are the two areas to smoke-test in a pilot, because that is where the open reports cluster.

Self-hosting caveat, with its source labelled. Braintrust’s page claims production self-hosting of Langfuse needs PostgreSQL, ClickHouse, Redis and Kubernetes. That is a competitor’s characterisation, so confirm the component list against Langfuse’s own self-hosting docs. Whatever the exact list turns out to be, it is the cost line that “free and open source” hides.

Failure-mode catalogue table Open GitHub issue, reaction count, one-line symptom, affected workload type, and a “verify current status” link column, re-checked on the publish date.

MLflow if you already run a data platform

Checkable structural advantages, not claims: Apache 2.0 under the Linux Foundation with no open-core split, a server plus database plus object storage architecture so you pick Postgres, MySQL or SQLite and S3, GCS, Azure Blob Storage or HDFS, and autolog() one-line instrumentation across a wide integration surface. Managed variants exist on Databricks, Amazon SageMaker and Azure ML.

The real argument for MLflow is boring and strong. Traces land where the rest of your data lives, so you can join trace records to product analytics without an export step.

The differentiator worth understanding is automated prompt optimisation with GEPA and memAlign, plus judge alignment against human feedback and multi-turn conversation evaluation with simulation. MLflow’s page asserts LangSmith lacks these; treat that as a vendor claim and check LangChain’s docs. Judge alignment itself deserves explaining, since no other page on this SERP does it. You take labelled human scores, measure how well your LLM judge agrees with them, and re-tune the judge until it does. Skip that step and your judge produces confident numbers of unknown correlation to anyone’s actual opinion.

Label the marketing. “30 million monthly downloads” and “60%+ of the Fortune 500” are MLflow’s own figures about itself, published without a methodology link. So is the “$2K to over $200K/year” LangSmith cost range, which carries no source and should not be repeated as fact anywhere.

Limitations. MLflow’s heritage is ML experiment tracking, so the UI carries runs, experiments and a model registry as first-class concepts that a pure LLM team has to learn. The GenAI surface also moves fast: MLflow 3.16.0 replaced the default trace explorer outright and added custom trace views generated from plain-English prompts, per mlflow.org/releases. Pin versions if UI stability matters to your runbooks.

Braintrust, closed source with the strongest eval-to-CI story

What Braintrust does differently is a workflow, not a feature. A GitHub Action runs evals on every PR and blocks merges below a threshold. A production trace converts in one click into a permanent regression case. The Playground lets a PM compare outputs without an engineer in the loop, and Loop, its assistant, generates scorers and datasets from production data. Braintrust’s changelog notes that Python workflow evaluations from SDK v0.39.0 can route tasks and scorers through provider batch APIs for discounted eval runs, in public preview.

Published tiers: free at 1 GB processed data, 10K scores, unlimited users; Pro at $249/month with 5 GB and unlimited users; Enterprise custom (verify at braintrust.dev/pricing).

The unit change is the thing nobody on this SERP explains. Braintrust bills processed data volume, not traces and not seats. Verbose tool-call payloads, base64 images and long RAG contexts drive your bill; a chatty agent with tiny payloads is cheap. That is a genuinely different cost curve from every other tool here.

Honest cons, including ones Braintrust lists itself. It is closed source, and self-hosting is Enterprise-only. If the reason you are leaving LangSmith is that prompt payloads cannot leave your infrastructure, Braintrust reproduces the blocker at the same tier gate. The Notion/Stripe/Vercel/Dropbox/Zapier customer list is Braintrust’s own claim on its own page.

The lighter-weight and single-job tools: Phoenix, Helicone, Lunary, Agenta, Traceloop

Arize Phoenix (Arize AI) is open source, OTel and OpenInference-native, and strongest when you want local or self-hosted trace inspection with eval libraries attached rather than a full LLMOps suite. The December 2024 dev.to hands-on comparison records weak LangChain.js and TypeScript-framework integration and an ML-general UI that felt indirect for agent debugging. That evidence is dated by its own author, so TypeScript-first teams should verify current JS support before committing.

Helicone is open source and architecturally distinct. It can run as a proxy, so you swap a base URL instead of adding a library. Fastest possible path to seeing your calls. Best for cost and latency logging and multi-provider spend visibility. The tradeoff is structural: a proxy sees requests, not your application’s internal spans, so agent-step attribution is weaker than SDK instrumentation gives you. The proxy also becomes a component in your request path, which is an availability question your SRE will ask.

Lunary is self-hostable and chatbot-shaped, with cloud published from $20/user/month on the Team plan including 50,000 events. The dev.to comparison records four limits: no built-in evaluators, no way to add traces to datasets directly, exports gated to a Team licence, and automatic PII masking available only on Enterprise with no manual option. December 2024 data, re-check it.

Agenta has an MIT-licensed functional core, is self-hostable, and packages prompt playground, prompt versioning, evals and observability together with code-free prompt deploys. Free tier caps at 2 seats and 5,000 traces/month; Pro from $49/month; Business from $399/month with RBAC and SOC 2 (figures per Braintrust’s roundup, verify on agenta.ai/pricing). No native CI gating.

Traceloop / OpenLLMetry is the OTel-purist answer. It is an instrumentation SDK, so your spans are portable to Traceloop, Grafana Tempo, Datadog, Jaeger or Langfuse without touching application code. This also matters if your stack is not LangChain at all: OTel-native ingestion is how LlamaIndex, CrewAI, the OpenAI Agents SDK and the Vercel AI SDK reach any of these backends. Pick it when you want the instrumentation decision to be reversible even though the backend decision is not.

PromptLayer covers prompt versioning and a registry with a non-engineer-friendly editor. Lunary’s comparison notes that its pricing is not transparent on its own site and that it is not open source. Both are easy to confirm by visiting the page.

Weights & Biases Weave brings W&B’s experiment-tracking depth to LLM tracing and evaluation. Powerful and broad, priced in a unit that is awkward to map onto LLM usage, and heavier than an LLM-only team needs.

The enterprise and guardrail lane: Galileo, Fiddler AI, HoneyHive, Orq.ai

Galileo does production-scale evaluators and runtime guardrails, with small purpose-built scoring models (Luna-2) so you are not paying frontier-model prices per judgement, plus agentic metrics like flow adherence and task completion. Free tier 5,000 traces/month, paid from $100/month (verify on galileo.ai/pricing). It is not a prompt-iteration or version-management tool. If prompt iteration is your daily workflow, Galileo is not a LangSmith replacement.

Fiddler AI unifies classical-ML and LLM monitoring with explainability, bias detection, embedding drift, and guardrails for hallucination, toxicity, PII leakage and prompt injection. Right answer for finance, healthcare and defence teams who must govern both model classes under one story. Wrong answer if you want a playground, because it has none.

HoneyHive leans on user-tracking and engagement analytics, startup-friendly pricing, customisable dashboards and feedback capture. Pick it when your quality question is “what are real users doing” rather than “did this prompt regress”.

Orq.ai is an end-to-end generative AI collaboration platform launched February 2024: a gateway across 130+ models, playgrounds and experiments, deployments with guardrails and fallbacks, observability with drift detection, SOC 2 certification and GDPR and EU AI Act alignment claims, all from Orq.ai’s own page. That page concedes the maturity tradeoff, a newer vendor with a thinner community and a smaller third-party integration ecosystem.

Frame this lane honestly. These are platform purchases with procurement cycles, security reviews and annual contracts. If you are a three-person team whose LangSmith bill just doubled, nothing in this section is your answer.

Running it locally, the option nobody on page one covers

A Show HN post from this month put the frustration better than any vendor page: “Built this because LangSmith needs a cloud account just to see my own traces” (HN 48063206). The author’s tool, opensmith, installs with pip, serves traces at localhost, stores to SQLite, needs no account and no config, works fully offline, and the post reported 800+ downloads in two days. Treat that as evidence of demand, not as a recommendation. One maintainer, days old at the time of posting, no compliance artefacts.

Three local-first tiers you can act on today:

  1. Docker Compose a real UI. Run Arize Phoenix or Langfuse locally. You get a genuine trace explorer, datasets and evals, with no vendor account and no egress. This is the tier I would put in front of a regulated workload, once it is deployed properly rather than on a laptop.
  2. OTel to a local collector. Emit spans, read them in Jaeger or Grafana Tempo. No eval features at all, but if all you need is a span tree with timings, this is fifteen minutes of config on infrastructure your platform team already understands.
  3. Single-binary SQLite projects like the one above. Zero infrastructure, zero compliance story. Fine for a solo developer on a plane. Not fine for anything with a customer’s data in it.

A related Ask HN post is worth citing for the why rather than the what. Its author described his tool as “NOT like langsmith/langfuse”, because existing tools “tell which agent failed, but my tool tells why” (HN 49701608). That is the real gap. Trace viewers answer where a run broke. Root-cause attribution across agent steps is still unsolved across the category, which is why this list keeps growing rather than consolidating.

What about AWS? The Bedrock-native path

If your inference runs on Amazon Bedrock, the AWS-native path is Bedrock and AgentCore Observability emitting to Amazon CloudWatch, with AWS X-Ray for distributed traces. You get logs, metrics and traces inside the account and IAM model your team already audits, with no new vendor DPA to negotiate. This surface has moved quickly, so verify current AgentCore capabilities in AWS’s own documentation.

Be explicit about what you do not get. No dataset-driven experiments. No prompt versioning with a playground. No LLM-as-a-judge scoring wired to traces. No annotation queue. CloudWatch is an observability backend, not an evaluation platform, and that trade is precisely why this question keeps getting asked without a good answer anywhere on page one.

Two practical architectures:

(a) Split the job. CloudWatch and X-Ray carry ops telemetry: latency, errors, throughput, cost attribution by account. A self-hosted eval-capable tool inside the same VPC carries datasets, judges and annotation. Two systems, one account boundary.

(b) One Apache-2.0 tool. AWS offers managed MLflow alongside SageMaker, so you get traces plus evaluation under Apache 2.0 without leaving the account boundary or signing anything new.

The compliance argument is what makes this section matter. For a team whose blocker is “prompt payloads must not leave our account”, an in-account solution clears procurement in a way no SaaS does, including the clouds run by the open-source vendors.

The cost model: same workload, every option, including self-host infrastructure

Two fixed workloads, published so you can dispute or re-derive the numbers:

  • Workload A, small team: 50,000 traces/month, 6 spans per trace (300,000 spans), 3 users, 30-day retention.
  • Workload B, production: 500,000 traces/month, 10 spans per trace (5,000,000 spans), 1 score per trace, 8 users, 12-month retention.

Every SaaS line should be priced from the vendor’s own pricing page with an as-of date on each cell, because these pages change monthly. LangSmith, Langfuse Cloud, Braintrust, Agenta, Galileo and Lunary all publish enough to do this. Where a tier is “contact us”, the cell reads not published rather than being estimated.

Monthly cost per option for Workload A and Workload B Two grouped bar charts, with self-hosted options split into infrastructure cost versus engineer-hour cost so the “free” options are not drawn at zero. Assumptions block published alongside.

The line no competitor prices is self-host total cost. Build it from public cloud list prices and include every component: managed Postgres, a ClickHouse-capable node or ClickHouse Cloud, Redis, object storage for payloads and blob exports, a load balancer, and a realistic monthly engineer-hours figure for upgrades and on-call. Then state the crossover honestly. Below some monthly volume, self-hosting costs more than the SaaS tier it replaces, because your engineer’s time is more expensive than the bill. “Free and open source” is a licence fact, not an invoice.

The spans-per-trace crossover is the mechanism behind most of the disagreement between vendor pages. Tools that bill per trace get relatively cheaper as traces get deeper. Tools that bill per span or per unit get more expensive on exactly the same traffic. Find your own crossover with one line of arithmetic: multiply your monthly trace count by your average spans per trace, price both models, and see which side of the line you sit on. Langfuse’s own page concedes that it bills every span while LangSmith bills the trace, and that the gap narrows as traces deepen. A vendor conceding a weakness is the most reusable sentence on their page.

Braintrust is a third curve entirely. Neither traces nor spans: gigabytes of processed data. Payload size drives the bill, so a RAG app stuffing 8,000-token contexts into every span costs vastly more than an agent making the same number of calls with short payloads. Check your average payload size, not your call count.

Spans-per-trace crossover chart Monthly cost versus average spans per trace for a per-trace-billed tool, a per-span-billed tool and a per-GB-billed tool at fixed trace volume, with crossover points marked and the KB-per-span assumption stated.

What breaks this model: annual commitments, startup discounts (dev.to notes 50% startup discounts on both LangSmith and Langfuse), volume tiers, and negotiated enterprise floors. This is a list-price model. Nobody at scale pays list.

For a worked example of the disagreement, Langfuse’s page puts 500k traces/month with 5 users at $621 on Langfuse Pro against $3,895 on LangSmith Plus, roughly 6x cheaper. That is a vendor’s arithmetic about a competitor, with the vendor choosing the workload shape. Re-run it with your own spans-per-trace figure before quoting it to anyone.

Licences, self-hosting and compliance: check the LICENSE file, not the landing page

Start with the conflation that ruins forum answers. langchain-ai/langchain is MIT-licensed, with 147,215 GitHub stars, 577 open issues and langchain-core==1.6.5. That is the framework. LangSmith the platform is a separate, closed-source product. The framework being MIT tells you nothing about the platform’s licence, and this confusion appears constantly in Reddit and Hacker News answers.

The nuance that separates this page from every vendor page: GitHub’s licence metadata for langfuse/langfuse currently resolves to NOASSERTION, not MIT. The usual cause is an open-core repo mixing an MIT core with an enterprise-licensed subdirectory, typically under ee/. Before your legal review, open the repo’s LICENSE file and any licence file inside that subdirectory, and write down exactly what each covers: which features you may run, modify and self-host for free, and which require a commercial agreement. MLflow under Apache 2.0 via the Linux Foundation has no equivalent split, which is a real and underrated advantage.

The applicable test for open-core versus fully-open is one question. Which features live behind an enterprise licence flag? The dev.to comparison gives the concrete example: as of its December 2024 testing, the free self-hosted Langfuse image lacked the playground and the LLM-as-judge evaluators. Re-verify that today, because it may well have changed. The general point survives regardless. The feature you are switching for may be the feature that is gated.

Compliance checklist, and the verb is “ask for the document”.

ArtefactWhat to requestTrap
SOC 2 Type IIThe report under NDAA badge on a website is not a report
ISO 27001The certificate with scope statementScope may exclude the product you are buying
GDPRSigned DPA, sub-processor list, data residencySub-processors change; ask for change notification
HIPAAA signed BAA, naming plan and region”HIPAA-compliant per docs” ≠ “we will sign”

Langfuse states it signs a BAA from the Pro plan in a dedicated HIPAA region. That is a vendor statement. Get it in writing, with the region named.

Air-gap reality check. If you have no egress, every feature depending on vendor-managed inference is off the table regardless of licence. Before you commit to an air-gapped deployment, enumerate for your chosen tool which evaluators, assistants and MCP server integrations call out to whose models, and confirm each can be pointed at an in-network endpoint.

Migrating off LangSmith without re-instrumenting twice

Frame the choice architecturally, because the migration is only painful if you let two separable decisions merge. Your instrumentation layer and your backend are independent. Instrument with OpenTelemetry using the GenAI semantic conventions and the backend becomes a config change to an OTLP endpoint. Instrument with a vendor SDK and every future move is a code migration.

Before touching code, inventory what actually lives in LangSmith today.

AssetMigrate?Note
TracesUsually notExport only what audit requires
DatasetsYes, firstThis is your regression suite
Prompts and versionsYesVersion history is the part teams lose
Annotation queues and human labelsYesExpensive to recreate; nobody re-labels
Evaluator definitionsRewriteScorer APIs differ between tools
DashboardsRebuildRarely portable
Alert rules and destinationsRebuildRe-point PagerDuty/Dynatrace hooks

Check the export gotcha first. Bulk export is frequently an enterprise-tier feature, and retention windows mean older data may already be deleted. Your plan’s export capability and your current retention window together determine whether this is a copy or a cutover.

The parallel-run pattern is what nobody on this SERP describes. Dual-emit to both backends for two to four weeks using an OTel collector fan-out. Compare three things across the two systems: trace counts, computed cost figures, and a sample of eval scores. This is how you discover that span attribution or token-cost calculation differs between tools while you still have the old one running.

OpenTelemetry portability diagram App → OTel SDK → collector → fan-out to two backends simultaneously, with the three comparison checkpoints (trace counts, cost figures, sampled eval scores) labelled.

Test two hazards deliberately during the parallel run, both sourced from the issue tracker rather than imagined: async context propagation under LangGraph’s ainvoke (Langfuse #8780) and streaming responses splitting one request into several traces (Langfuse #3961). Both are exactly the shape an agent app has. A two-week dual-emit surfaces both immediately.

Rollback plan, in two sentences. Keep the old exporter configured behind a feature flag until a full billing cycle and at least one production incident have passed under the new tool. Delete it after that, not before.

Evaluation depth, the axis most comparison tables get wrong

“LLM-as-a-judge: ✅” is true of nearly everything on this list and therefore tells you nothing. Use a ladder instead.

RungCapabilityWho clearly does it
1Capture human feedback and scoresNearly all
2Offline datasets with versioning and baseline comparisonLangSmith, Langfuse, MLflow, Braintrust, Agenta
3Online judges scoring live trafficLangSmith, Langfuse, Braintrust, Galileo
4Judge calibration against human labels, plus judge versioningMLflow (claimed), Galileo’s Luna-2 on the cost side
5Results gating deployment automaticallyBraintrust

Evaluation depth ladder Five rungs from human feedback capture to automatic deployment gating, with each tool plotted at its highest verifiable rung.

Rung 4 is the one that matters and almost nobody has. An uncalibrated LLM judge produces confident scores with unknown correlation to your reviewers’ opinions, which means your regression suite may be measuring nothing at all while displaying a reassuring number. Calibration involves labelled examples, a measured agreement statistic between judge and humans, and re-alignment whenever the judge model changes underneath you. That last part is the ongoing cost people forget.

Rung 5 is Braintrust’s claim to the category, and the counter-consideration is worth stating. A merge-blocking gate on a noisy judge is an outage generator. Rung 4 must precede rung 5. Gate on a judge you have calibrated, or gate on nothing.

Judges are LLM calls, so they carry token cost, latency and version risk. A concrete, sourced failure: Langfuse #11109 reports Experiments and Evaluators unable to parse Anthropic “thinking” blocks from reasoning models, breaking evaluator runs outright. Reasoning-model output formats break eval pipelines in ways no feature matrix predicts.

The direction of travel is cheaper judging. Langfuse’s changelog lists “Jev as a judge”, using a typed decision model for calibrated scores at lower cost and latency than a full LLM judge, alongside running evaluators over historical observations to backfill scores (langfuse.com/changelog). Galileo’s Luna-2 attacks the same problem with small purpose-built scoring models. Judging is being decoupled from frontier-model pricing, which changes the economics of rung 3.

Master comparison table

ToolLicenceSelf-host on free tierPricing unitEntry price (as of 29 Sep 2026)InstrumentationEval rungPrompt mgmtCI gatingBest for
LangSmithClosed sourceNo (Enterprise)Seat + trace + retention$39/seat/mo PlusSDK + OTel3Yes (LangChain Hub)PartialLangGraph-native small teams
LangfuseOpen core, GitHub reads NOASSERTIONYesUnit/span, no seat feeFree Hobby tierOTel-native3YesVia APIMost teams leaving on price
MLflowApache 2.0 (Linux Foundation)YesFree; infra only$0 licenceSDK + autolog4 (claimed)YesVia APIData-platform teams wanting Apache 2.0
BraintrustClosed sourceNo (Enterprise)GB processed dataFree 1 GB / Pro $249/moSDK5YesNative GitHub ActionMerge-blocking eval gates
Arize PhoenixOpen sourceYesFree; infra only$0 licenceOTel/OpenInference2LimitedNoLocal trace inspection with evals
HeliconeOpen sourceYesRequestsFree tier publishedProxy + SDK1LimitedNoFastest path to seeing calls
LunaryOpen source, self-hostableYesSeat + events$20/user/mo Team, 50k eventsSDK1YesNoChatbot-shaped products
AgentaMIT coreYesSeat + tracesFree 2 seats / 5k traces; Pro $49/moSDK2Yes, code-free deploysNoPrompt-first teams wanting open source
GalileoClosed sourceNoTracesFree 5k traces/mo; $100/moSDK4NoNoRuntime guardrails at scale
Fiddler AIClosed sourceEnterpriseNot publishedNot publishedSDK3NoNoRegulated ML + LLM governance
HoneyHiveClosed sourceNot publishedNot publishedNot publishedSDK2YesNoUser engagement analytics
Orq.aiClosed sourceNoNot publishedNot publishedSDK3YesNoCollaborative gen-AI platform, SOC 2
Traceloop / OpenLLMetryOpen source SDKYes (SDK)Backend-dependent$0 SDKOTel-native1NoNoPortable instrumentation layer
CloudWatch / AgentCoreAWS serviceIn-accountAWS meteringAWS ratesOTel/X-Ray1NoNoBedrock-native, in-account audit

Rules for this table: every price cell carries an as-of date, any cell that could not be verified from a primary source reads “not published” rather than a guess, and there is no “Winner” column, because the winner depends on the blocker.

Master comparison matrix 14 rows by 12 columns including exact licence from the LICENSE file, self-host on free tier, pricing unit, entry price with capture date, eval ladder rung and compliance artefacts, all captured on one date.

Decision guide: pick one in five minutes

Prompt payloads cannot leave our infrastructure → self-hosted Langfuse or MLflow. Both run free in your VPC; MLflow if Apache 2.0 with no enterprise-gated core matters to legal.

We are Bedrock-native and audit by AWS account → CloudWatch with AgentCore Observability and X-Ray for telemetry, plus an in-VPC eval tool for datasets and judges.

Our bill is seats, not volume → any usage-priced tool, and run the spans-per-trace calculation before you choose between Langfuse and a per-trace competitor.

We ship daily and fear regressions → Braintrust for rung-5 gating, on the condition that you calibrate the judge first.

We need runtime guardrails in a regulated industry → Galileo for agentic metrics and cheap scoring, Fiddler AI if you must govern classical ML in the same story.

We just want to see traces today → Helicone’s proxy, or a local Phoenix container. Both are same-afternoon jobs.

Stay on LangSmith if you are all-in on LangGraph, your team is small enough that seat fees are rounding error, and the retention window covers your actual debugging horizon. Switching costs engineer weeks that the licence saving will not repay. Say that out loud before anyone opens a migration ticket.

What could not be verified without running a pilot

Three questions here genuinely need hands-on work, and nothing on this page should be read as answering them: actual ingestion latency and UI responsiveness at high span volume; whether the async-context and streaming trace-splitting issues still reproduce on current Langfuse versions; and whether today’s free self-hosted image ships the playground and LLM-as-judge evaluators. Each is a half-day of work in your own stack, and each is worth more than any comparison table, including this one.

Frequently Asked Questions

What is the alternative to LangSmith?

Langfuse is the most widely adopted open-source alternative, MLflow is the Apache-2.0 choice with no per-trace fee, Braintrust is strongest when eval results must gate CI, and CloudWatch with AgentCore is the pick if you are Bedrock-native. The right one depends on which constraint blocks you: pricing, self-hosting, retention or framework coupling.

What are some alternatives to LangSmith in AWS?

On AWS the native path is Bedrock and AgentCore Observability emitting to CloudWatch, with X-Ray for distributed traces, keeping prompt payloads inside your own account and IAM boundary. Managed MLflow on SageMaker covers traces plus evaluation under Apache 2.0, and any self-hostable tool runs in your VPC. CloudWatch alone gives you no datasets, no prompt versioning, no playground and no LLM-as-a-judge, so most teams pair it with an in-VPC eval tool.

Is there a free, fully open-source LangSmith alternative?

Yes, with a precise qualification. MLflow is Apache 2.0 under the Linux Foundation with no enterprise-gated core. Langfuse is free to self-host but is open-core, and GitHub currently resolves its repo licence as NOASSERTION, so read the LICENSE file and any enterprise subdirectory. Arize Phoenix, Helicone and Agenta’s functional core are also open source. A free licence is not a free bill. You still pay for the database, object storage and the engineer who runs the upgrades.

Is LangSmith free?

There is a free Developer tier, published as one seat with a limited monthly trace allowance and a short retention window, and it is genuinely enough for solo prototyping. Paid starts at the per-seat Plus tier, so cost arrives with your second teammate rather than your second thousand traces. Self-hosting is not available on the free or Plus tiers, so free-tier trace data goes to LangChain’s cloud. Check LangChain’s pricing page for current numbers rather than trusting any comparison page, this one included.

Can I run LangSmith locally?

Not on the free or Plus tiers. Self-hosting and BYOC are Enterprise-contract options requiring Kubernetes, so a laptop-local LangSmith is not a supported path.

Three real local options exist instead, in descending order of maturity. Run Langfuse or Arize Phoenix in Docker locally for a full UI with no vendor account. Emit OTel spans to a local collector and view them in Jaeger or Grafana Tempo if a span tree is all you need. Or use an early single-binary SQLite project like the pip-installable localhost tool posted to Show HN, which is single-maintainer, early-stage, and not appropriate for regulated workloads.

LangSmith vs Langfuse, which should I choose?

Pick Langfuse if you need self-hosting on any tier, want usage-based pricing with no seat fee so adding reviewers is free, or want instrumentation that survives a change of agent framework. Pick LangSmith if you are committed to LangChain and LangGraph. Its one-env-var trace path and Studio’s visual agent building have no open-source equal, and per-seat pricing is noise for a team of four. The caveat no vendor page includes: verify Langfuse’s async-context (#8780) and score-filter (#11874) issues against your own stack first.

What do developers on Reddit and Hacker News actually recommend instead?

Langfuse dominates public discussion as the open-source default. Two recurring complaints target the whole category rather than any one tool: that trace viewers tell you which agent step failed without telling you why (HN 49701608), and that seeing your own traces should not require a cloud account (HN 48063206).

Two warnings about forum answers. They routinely confuse LangChain’s MIT licence with LangSmith’s closed-source platform, and most listicles a searcher lands on are vendor-published.

How hard is it to migrate off LangSmith?

Two separable jobs. Re-pointing instrumentation is hours if you already emit OpenTelemetry, or a code change if you use the vendor SDK. The harder job is assets: datasets, prompt versions, annotation labels, evaluator definitions, dashboards and alert rules each migrate differently. Bulk export is frequently enterprise-tier, and retention windows may already have deleted older traces, so check export capability and retention as step one. Dual-emit to both backends for two to four weeks via an OTel collector, compare trace counts, cost figures and sampled eval scores, then cut over with the old exporter behind a flag for one billing cycle.

Is self-hosting actually cheaper than LangSmith?

It depends on volume, and there is a crossover. Above some monthly trace count self-hosting wins decisively, because you pay infrastructure rates instead of per-trace rates. Below it, the managed tier is cheaper than your engineer’s time. The real cost lines are managed Postgres, ClickHouse-capable compute or ClickHouse Cloud, Redis, object storage for payloads, a load balancer, plus upgrades and on-call. Every open-source recommendation on page one of Google prices all of those at zero.



Every price, licence and feature claim here carries an as-of date of 29 September 2026 and should be re-verified against the vendor’s own page before you spend money. Corrections from vendors and readers are welcome and will be logged with dates. The alternative to LangSmith worth buying is whichever one clears the blocker you started with, priced at your own spans-per-trace figure.

Frequently Asked Questions

What is the alternative to LangSmith?

Langfuse is the most widely adopted open-source alternative, MLflow is the Apache-2.0 choice with no per-trace fee, Braintrust is strongest when eval results must gate CI, and CloudWatch with AgentCore is the pick if you are Bedrock-native. The right one depends on which constraint blocks you: pricing, self-hosting, retention or framework coupling.

What are some alternatives to LangSmith in AWS?

On AWS the native path is Bedrock and AgentCore Observability emitting to CloudWatch, with X-Ray for distributed traces, keeping prompt payloads inside your own account and IAM boundary. Managed MLflow on SageMaker covers traces plus evaluation under Apache 2.0, and any self-hostable tool runs in your VPC. CloudWatch alone gives you no datasets, no prompt versioning, no playground and no LLM-as-a-judge, so most teams pair it with an in-VPC eval tool.

Is there a free, fully open-source LangSmith alternative?

Yes, with a precise qualification. MLflow is Apache 2.0 under the Linux Foundation with no enterprise-gated core. Langfuse is free to self-host but is open-core, and GitHub currently resolves its repo licence as NOASSERTION, so read the LICENSE file and any enterprise subdirectory. Arize Phoenix, Helicone and Agenta's functional core are also open source. A free licence is not a free bill. You still pay for the database, object storage and the engineer who runs the upgrades.

Is LangSmith free?

There is a free Developer tier, published as one seat with a limited monthly trace allowance and a short retention window, and it is genuinely enough for solo prototyping. Paid starts at the per-seat Plus tier, so cost arrives with your second teammate rather than your second thousand traces. Self-hosting is not available on the free or Plus tiers, so free-tier trace data goes to LangChain's cloud. Check LangChain's pricing page for current numbers rather than trusting any comparison page, this one included.

Can I run LangSmith locally?

Not on the free or Plus tiers. Self-hosting and BYOC are Enterprise-contract options requiring Kubernetes, so a laptop-local LangSmith is not a supported path. Three real local options exist instead, in descending order of maturity. Run Langfuse or Arize Phoenix in Docker locally for a full UI with no vendor account. Emit OTel spans to a local collector and view them in Jaeger or Grafana Tempo if a span tree is all you need. Or use an early single-binary SQLite project like the pip-installable localhost tool posted to [Show HN](https://news.ycombinator.com/item?id=48063206), which is single-maintainer, early-stage, and not appropriate for regulated workloads.

LangSmith vs Langfuse, which should I choose?

Pick Langfuse if you need self-hosting on any tier, want usage-based pricing with no seat fee so adding reviewers is free, or want instrumentation that survives a change of agent framework. Pick LangSmith if you are committed to LangChain and LangGraph. Its one-env-var trace path and Studio's visual agent building have no open-source equal, and per-seat pricing is noise for a team of four. The caveat no vendor page includes: verify Langfuse's async-context ([#8780](https://github.com/langfuse/langfuse/issues/8780)) and score-filter ([#11874](https://github.com/langfuse/langfuse/issues/11874)) issues against your own stack first.

What do developers on Reddit and Hacker News actually recommend instead?

Langfuse dominates public discussion as the open-source default. Two recurring complaints target the whole category rather than any one tool: that trace viewers tell you which agent step failed without telling you why ([HN 49701608](https://news.ycombinator.com/item?id=49701608)), and that seeing your own traces should not require a cloud account ([HN 48063206](https://news.ycombinator.com/item?id=48063206)). Two warnings about forum answers. They routinely confuse LangChain's MIT licence with LangSmith's closed-source platform, and most listicles a searcher lands on are vendor-published.

How hard is it to migrate off LangSmith?

Two separable jobs. Re-pointing instrumentation is hours if you already emit OpenTelemetry, or a code change if you use the vendor SDK. The harder job is assets: datasets, prompt versions, annotation labels, evaluator definitions, dashboards and alert rules each migrate differently. Bulk export is frequently enterprise-tier, and retention windows may already have deleted older traces, so check export capability and retention as step one. Dual-emit to both backends for two to four weeks via an OTel collector, compare trace counts, cost figures and sampled eval scores, then cut over with the old exporter behind a flag for one billing cycle.

Explore More

Free Newsletter

Get the LLM Evals Newsletter

Platform comparisons, pricing changes and eval technique deep-dives. No spam.

Related Articles