MLflow logo

MLflow Review (2026)

The open-source ML platform that grew a serious GenAI half. MLflow 3 adds OpenTelemetry-compatible tracing, LLM judges and review apps - free and self-hostable, with Databricks selling the managed version. The best zero-cost option if you can run infrastructure.

Researched

Rating

4.0

Starting Price

$0 (open source)

Free Plan

Yes

SDKs & Frameworks

8

Deployment

4

Best For

Teams that already run MLflow for classical ML, Databricks customers, and anyone who wants a genuinely free and complete self-hosted platform and has the operational capacity to run it.

Last Updated:

10 Things You Should Know About MLflow

  1. 1 MLflow Tracing is open source and OpenTelemetry-compatible with no vendor lock-in
  2. 2 A single line of code captures prompts, retrievals, tool calls, responses, latency and token counts
  3. 3 Databricks now recommends storing traces in Unity Catalog for new and production workloads
  4. 4 Unity Catalog traces land in OpenTelemetry Delta tables with no storage cap and SQL queryability
  5. 5 Open-source telemetry collection was introduced in MLflow 3.2.0 and is disabled on Databricks by default
  6. 6 In MLflow 3 the Agent Evaluation SDK methods were integrated with Databricks-managed MLflow

Pros & Cons

Pros

  • Genuinely free and Apache 2.0 - the open-source build is a complete platform, not a crippled teaser
  • The single most widely deployed ML tooling in existence, so institutional knowledge and hiring are easy
  • OpenTelemetry-compatible tracing means your trace data is portable rather than trapped
  • Review apps closing the loop from expert feedback to aligned automated judges is a strong workflow few competitors match
  • Same platform covers classical ML experiment tracking and GenAI, which matters for teams doing both
  • Building eval datasets straight from production traces is the right shape for continuous improvement

Cons

  • You are operating it - database, artifact store, scaling, upgrades and availability are yours
  • The UI is functional rather than polished, and noticeably dated next to Braintrust or Logfire
  • The best-governed experience (Unity Catalog trace storage, no storage cap, SQL queryability) is Databricks-only
  • Historically an experiment-tracking tool, so some GenAI workflows still feel grafted on
  • Databricks steering the roadmap is a soft dependency even though the license is permissive

Features

One-line tracing capturing prompts, retrievals, tool calls, responses, latency and token counts
OpenTelemetry-compatible with no vendor lock-in
Built-in and custom LLM judges, alignable against expert judgment
Review apps to collect expert feedback and build labelled evaluation datasets
Build evaluation datasets directly from production traces
Interactive timeline views and side-by-side run comparison
Shared lineage with classical ML experiment tracking and the model registry

What changed with MLflow 3

MLflow spent years as the default open-source experiment tracker for classical machine learning. If your mental model is “the thing that logs training runs,” it is out of date.

MLflow 3 is a genuine GenAI platform. It unifies tracking, evaluation and observability across the development and production lifecycle - real-time trace logging, built-in and custom scorers, human feedback collection, and version tracking.

MLflow Tracing gives end-to-end observability into GenAI applications including complex agent systems. A single line of code captures prompts, retrievals, tool calls, responses, latency and token counts. It automatically instruments popular GenAI libraries or ingests traces directly, and offers interactive timeline views and side-by-side comparison.

Critically, tracing is OpenTelemetry-compatible with no vendor lock-in. Your trace data is standard telemetry, not a proprietary format you would have to re-instrument to escape.

The evaluation workflow is the underrated part

Most comparisons stop at “MLflow has evals.” The workflow around them is what makes it interesting.

You get built-in and custom LLM judges and scorers, so you can define what quality means for your use case rather than accepting generic metrics. Fine, everyone has that.

The differentiated piece is review apps. Domain experts review outputs and provide feedback. That feedback builds labelled evaluation datasets. Those datasets are then used to align the automated judges against expert judgment.

That loop - expert opinion becoming labelled data becoming a calibrated automatic judge - is the genuinely hard problem in LLM evaluation, and it is the difference between an LLM-as-judge score you trust and one you do not. Few competitors handle it this coherently.

You can also build evaluation datasets directly from production traces, which is the correct shape for continuous improvement: the failures your users actually hit become the regression suite.

Free, and genuinely complete

This needs saying plainly because the category is full of misleading self-hosting claims.

MLflow’s open-source build is a complete platform, not a gated teaser. Apache 2.0, full tracing, full evaluation, prompt registry, run comparison. Several commercial vendors advertise self-hosting and then withhold the useful half behind their cloud tier. MLflow does not.

What Databricks sells on top is operational and governance value - managed hosting, production-scale infrastructure, lakehouse integration, Unity Catalog governance. That is a legitimate commercial line to draw, and it means the free version is genuinely usable in production rather than a trial.

The honest cost of free is operational. You own the tracking server, the backing database, the artifact store, scaling, upgrades, backups and availability. At small scale that is a container and a Postgres instance. At high trace volume it is a system someone has to own, and trace data grows quickly. Budget engineering time, not licence fees - and be honest about whether that ownership will survive the person who set it up moving teams.

Where open source and Databricks diverge

This is the detail most comparisons miss, and it matters because the shared name hides it.

Databricks now recommends storing traces in Unity Catalog for new and production workloads. Traces land in OpenTelemetry Delta tables, which gives you three real things:

  • No storage cap
  • SQL queryability over raw trace data
  • Unity Catalog-governed access control

That is arguably the best trace storage architecture in the entire category. It is also a Databricks capability. Self-hosted open-source MLflow does not get it.

So “MLflow” describes two meaningfully different experiences depending on where you run it, and the gap is in governance and scale rather than core features. If you are evaluating on the strength of the Unity Catalog story, be clear that you are evaluating Databricks.

Two smaller version notes worth knowing: open-source telemetry collection was introduced in MLflow 3.2.0 and is disabled on Databricks by default, and in MLflow 3 the Agent Evaluation SDK methods were integrated with Databricks-managed MLflow, which matters if you are migrating from the older SDK.

Where it is weak

The UI. It is functional rather than pleasant, and it looks dated next to Braintrust or Logfire. MLflow was built for ML engineers reading experiment tables, and application developers used to modern observability tooling notice the difference immediately.

The lineage shows. Some GenAI workflows still feel grafted onto an experiment-tracking substrate rather than designed for LLM applications from scratch. This is improving but it is visible.

Databricks stewardship is a soft dependency. The license is permissive and you can genuinely run independently, but roadmap priorities reflect a commercial sponsor. Compare with Langfuse, where the open-source project and the paid cloud are the same feature set.

Should you use it?

Use MLflow if you already run it for classical ML, you are a Databricks customer, or you want a complete free self-hosted platform and have the operational capacity to run it.

Don’t use it if you have no infrastructure capacity, UI quality matters a lot to your team’s adoption, or you only do LLM application work and would prefer a purpose-built tool.

Bottom line: the strongest genuinely-free option in the category, and unusually honest about what the open-source build includes. The evaluation loop - expert review to labelled dataset to aligned judge - is better than its reputation suggests. Pay for it in operational time rather than licence fees, and understand that the best-governed version of it lives on Databricks.


Features and architecture verified against MLflow and Databricks documentation on 31 July 2026. Managed pricing is bundled into Databricks consumption and is not published as a standalone rate. This is a researched directory entry - we have not yet instrumented this platform with our reference application.

Pricing Plans

Open source

$0

  • Full tracing, evaluation and prompt registry
  • Apache 2.0 licensed
  • Self-host anywhere
  • No feature gating against the managed version's core
Most Popular

Managed MLflow (Databricks)

Bundled with Databricks

  • Fully managed hosting and production scaling
  • Unity Catalog governance
  • Trace storage in OTel Delta tables with no storage cap
  • Priced as part of Databricks consumption

Enterprise (Databricks)

Custom

  • Lakehouse integration
  • Governed access via Unity Catalog
  • Contact sales

SDKs & Frameworks

Python SDK TypeScript SDK OpenAI, Anthropic and other provider auto-instrumentation LangChain LlamaIndex DSPy CrewAI OpenTelemetry ingest

Deployment

Self-hosted, open source under Apache 2.0 Databricks Managed MLflow Any cloud or on-prem Unity Catalog trace storage (Databricks)

Eval Methods

Built-in LLM judges and scorers Custom scorers Review apps for expert human feedback Datasets built from production traces Version comparison across runs

Our Verdict

MLflow is the strongest zero-cost option in the category, with the caveat that free software is not free to operate. MLflow 3 turned what was an experiment-tracking tool into a real GenAI platform - OpenTelemetry-compatible tracing from a single line of code, built-in and custom LLM judges, review apps that collect expert feedback and align automated judges against it, and evaluation datasets built directly from production traces. It is Apache 2.0 and the open-source build is complete rather than a gated teaser, which is more than can be said for several commercial competitors advertising self-hosting. Two honest caveats. The UI is functional rather than pleasant, and it shows its lineage as a tool built for ML engineers rather than application developers. And the genuinely best-governed experience - Unity Catalog trace storage in OTel Delta tables, no storage cap, SQL queryable - is available on Databricks, which is where the commercial gravity sits. If you already run MLflow or Databricks, this is close to automatic.

Similar Tools

Frequently Asked Questions

Is the open-source version actually complete?

Yes, and this is worth stating clearly because the category is full of vendors who advertise self-hosting and then gate the useful half behind cloud. MLflow's open-source build under Apache 2.0 gives you tracing, evaluation with built-in and custom judges, the prompt registry and run comparison. What Databricks sells on top is managed hosting, production-level scaling, lakehouse integration and Unity Catalog governance - operational and governance value rather than core features withheld from the free build. For most teams the open-source version is genuinely sufficient.

What is the Unity Catalog trace storage change?

Databricks now recommends storing traces in Unity Catalog for new and production workloads rather than the older backend. Traces land in OpenTelemetry Delta tables, which brings three concrete benefits - no storage cap, SQL queryability over your raw trace data, and Unity Catalog-governed access control. This is a genuinely good architecture and arguably the best trace storage story in the category. The catch is that it is a Databricks capability. Self-hosted open-source MLflow does not get it, so the governance and scale story diverges between the two deployments in a way the shared name obscures.

How good are the evals?

Better than most people expect, because MLflow's reputation is still anchored to experiment tracking. You get built-in LLM judges and custom scorers, and you can define quality dimensions per use case. The genuinely strong part is the workflow around them - review apps let domain experts label outputs, those labels build evaluation datasets, and the datasets align the automated judges against expert judgment. Closing that loop is the hard part of LLM evaluation and few competitors handle it this coherently. You can also build eval datasets directly from production traces, which is the correct shape for continuous improvement.

What does self-hosting actually cost me?

Not money, but real operational time. You own the tracking server, the backing database, the artifact store, scaling, upgrades, backups and availability. At small scale this is a container and a Postgres instance and it is genuinely easy. At production scale with high trace volume it is a system somebody has to own, and trace data grows fast. Budget engineering time rather than licence fees, and be honest about whether you have that capacity - the most common self-hosting failure is not technical difficulty but nobody owning it after the person who set it up moves on.

Is MLflow tied to Databricks?

Legally no, practically somewhat. The license is Apache 2.0, the project is genuinely open source, and you can run it entirely independently on any infrastructure with no Databricks relationship. But Databricks stewards the project and sells the managed version, which means roadmap priorities reflect commercial interests - the Unity Catalog trace storage recommendation is a reasonable example, since it is the better architecture and also only available on their platform. This is not a criticism so much as a thing to understand. Compare it with Langfuse, where the open-source project and the commercial cloud are the same feature set.

How does it compare with Langfuse?

They are the two serious free self-hosted options and the choice mostly comes down to what else you do. Langfuse is LLM-native, has a more modern UI, is MIT licensed and is purpose-built for this category. MLflow is broader - it covers classical ML experiment tracking and the model registry alongside GenAI, so a team fine-tuning models and shipping LLM applications gets one platform instead of two. If you only do LLM application work, Langfuse is the more pleasant tool. If you do both, MLflow's shared lineage is a real advantage.