LLM Tool Cost Calculator

Published prices in this category are not comparable, because every vendor meters something different. Enter your workload once and see what 30 tools would actually charge for it.

We normalise 11 different billing meters - event, span, record, use-case, gb-month, markup-pct, request, gb-ingested, job, seat, prompt - into one annual figure. Tools that publish no pricing are marked rather than guessed at.

ToolMeterPer monthPer yearBasis
Inspect AI-Free$0Free, self-hosted · MIT, UK AISI. No commercial tier.
LangWatcheventFree$0Within free allowance · Apache-2.0 core. Per-event rates not published.
LiteLLM-Free$0Free, self-hosted · MIT, zero markup. Cost is infrastructure.
Agenta-Free$0Free, self-hosted · MIT self-host; cloud pricing not verified.
Arize Phoenix-Free$0Free, self-hosted · Self-host free (ELv2). Online monitoring requires AX.
Azure AI Foundry Evaluation-Free$0No usage metering · Model inference and tool calls only; no runtime fee.
Databricks Agent Evaluation-Free$0No usage metering · Inside Databricks consumption.
Confident AI (DeepEval)-Free$0Free, self-hosted · Apache 2.0.
Evidently-Free$0Free, self-hosted · Apache 2.0; cloud pricing not published.
Fiddler AI-Free$0No usage metering · No public pricing. VPC from Enterprise.
Freeplay-Free$0No usage metering · No public pricing.
GalileospanFree$0Within free allowance · Pricing not published.
Giskard-Free$0Free, self-hosted · Apache 2.0. Hub pricing not published.
Guardrails AI-Free$0Free, self-hosted · Apache 2.0.
HiddenLayer-Free$0No usage metering · No public pricing.
HoneyHiveeventFree$0Within free allowance · Free tier only; paid pricing not published.
Lakera-Free$0No usage metering · No public pricing; Check Point procurement.
LaminarspanFree$0Within free allowance · Open source.
Latitude-Free$0Free, self-hosted · MIT self-host; cloud pricing not verified.
LM Evaluation Harness-Free$0Free, self-hosted · Free, EleutherAI.
MLflow-Free$0Free, self-hosted · Apache 2.0. Managed via Databricks consumption.
NVIDIA NeMo Guardrails-Free$0Free, self-hosted · Apache 2.0. Cost is extra model calls per rail.
Not Diamond-Free$0No usage metering · Pricing sources conflict.
OpenLIT-Free$0Free, self-hosted · Apache 2.0, self-hosted.
Patronus AI-Free$0No usage metering · No public pricing.
Portkey-Free$0Free, self-hosted · Apache-2.0 gateway; observability is paid.
Promptfoo-Free$0Free, self-hosted · MIT. Now an OpenAI company.
Ragas-Free$0Free, self-hosted · Free. Cost is judge model calls.
Vertex AI Gen AI Evaluation Service-Free$0No usage metering · Per token plus GCP compute.
W&B Weavegb-ingestedFree$0Within free allowance · Per-GB ingested plus seats. Rates not published.
Maxim AIspanFree$0Within free allowance · Contact sales.
Aporia-Free$0No usage metering · Inside Coralogix platform.
CalypsoAI-Free$0No usage metering · Absorbed into F5 platform.
Langtrace-Free$0Free, self-hosted · Cloud currently free; AGPL-3.0 server.
Openlayer-Free$0No usage metering · No public pricing.
Prompt Security-Free$0No usage metering · Absorbed into SentinelOne Singularity.
TraceloopspanFree$0Within free allowance · OpenLLMetry Apache-2.0. Now ServiceNow.
TruLens-Free$0Free, self-hosted · MIT, Snowflake-maintained.
UpTrain-Free$0Free, self-hosted · Apache 2.0. Managed pricing unclear.
LLM Guard-Free$0Free, self-hosted · MIT but archived 9 July 2026.
Martian-Free$0No usage metering · Router no longer offered.
OpenAI Evals-Free$0No usage metering · Hosted platform shuts down 30 Nov 2026.
Pezzo-Free$0Free, self-hosted · Unmaintained since mid-2025.
WhyLabs-Free$0Free, self-hosted · Company shut down; open-sourced.
Baserun-Free$0No usage metering · Acquired by LlamaIndex.
Gentrace-Free$0No usage metering · Shut down; MIT source released.
Literal AI-Free$0No usage metering · Discontinued.
Vellum-Free$0No usage metering · No longer a developer platform.
Cloudflare AI Gatewayrecord$5.00$60800,000 records/mo, 100,000 free · $5 Workers Paid = 1M logs. No token markup.
PromptHubrequest$9.00$108100,000 requests/mo, 2,000 free · Pro. Free tier makes prompts public.
Opikspan$19$228800,000 spans/mo · Cheapest paid cloud; Apache-2.0 self-host.
Sentryevent$26$312800,000 events/mo, 5,000 free · Team tier. Not an LLM observability tool.
Langfuseevent$29$348800,000 events/mo, 50,000 free · Core tier; self-host free under MIT.
Lunaryevent$30$360800,000 events/mo, 30,000 free · Free tier is 1,000/DAY not monthly. Apache-2.0.
AgentOpsevent$40$480800,000 events/mo, 5,000 free · Event = each LLM/tool call, not each run.
Pydantic Logfirerecord$49$588800,000 records/mo, 10,000,000 free · Team tier. Records = spans + logs + metrics.
SigNozgb-ingested$49$5885 GB/mo · Cloud from $49; self-host free (ClickHouse).
Arthuruse-case$60$7200 use-cases/mo, 4 free · Premium: 100 use cases, unlimited seats.
Heliconerequest$79$948100,000 requests/mo, 10,000 free · Pro. Overage rate not published. Maintenance mode.
Langtailprompt$99$1,1880 prompts/mo, 2 free · Pro caps at 20 prompts, not requests.
Vercel AI Gateway-$100$1,200Flat platform fee · No token markup. $5/mo free credits.
LangSmithseat$195$2,3405 seats · Plus tier, per seat.
Confident AIgb-month$204$2,4465 GB/mo, 1 free · Starter, unlimited seats, 5 GB-months included.
Braintrustspan$249$2,988800,000 spans/mo · Pro tier. Billed on processed data.
Parea AIevent$250$3,000800,000 events/mo, 3,000 free · Team: 100k logs, 3 seats then $50 each.
OpenRoutermarkup-pct$275$3,3005.5% of $5,000 provider spend · 5.5% of provider spend, no volume discounts.
PromptLayerrequest$342$4,098100,000 requests/mo, 2,500 free · Pro. $0.003 per transaction overage.
New Relic AI Monitoringgb-ingested$495$5,9405 GB/mo, 100 free · Plus CCU meter for AI features.
TrueFoundryrequest$499$5,988100,000 requests/mo, 50,000 free · Pro tier.
Kong AI Gateway-$500$6,000Flat platform fee · Konnect ~$500-2500/mo. OSS gateway free.
Arize AXspan$800$9,600800,000 spans/mo, 25,000 free · AX Pro, 50k spans. Seats not metered.
Datadog LLM Observabilityspan$1,216$14,592800,000 spans/mo, 40,000 free · Pro tier, 100k spans. New pricing from 1 May 2026.

The estimate almost everyone gets wrong

Teams model cost from request volume. Most tools in this category do not bill requests. They bill spans, and a single agentic request produces 20 to 50 of them. That one substitution is routinely an order-of-magnitude error, and it is why free tiers that look generous evaporate in days.

Before you commit to anything here, instrument one representative request and count the spans it actually emits. That number, not your traffic, is what determines your bill.

Frequently Asked Questions

Why can I not just compare the published prices?

Because vendors meter different things. Datadog bills per LLM span, AgentOps bills per event, W&B Weave bills per GB ingested, Pydantic Logfire bills records, Confident AI bills GB-months, Langtail bills the number of prompts you have, OpenRouter takes a percentage of provider spend, and LangSmith bills seats. A published price of $50 a month means something completely different in each case. This calculator normalises them to one workload so the numbers are actually comparable.

Why does spans per request matter so much?

Because it is the difference between a manageable bill and an unmanageable one, and it is the single most common estimating mistake in this category. A simple completion produces roughly 3 to 5 spans. An agentic workflow with tool calls, retrieval and reasoning steps commonly produces 20 to 50. Any tool that meters spans therefore charges an agent workload up to ten times what it charges a completion workload at identical request counts. Datadog's 40,000-span free tier is a few thousand simple requests, or a day or two of agent traffic.

Why does trace size affect the cost?

Because several vendors bill volume rather than events. W&B Weave charges per GB ingested and Confident AI charges GB-months, which means two applications making identical numbers of calls can have very different bills. A RAG system logging a long system prompt, ten retrieved chunks and a lengthy completion produces a far larger payload than a classification endpoint returning one word. If you bill by volume, verbosity is the cost driver.

How accurate are these figures?

Treat them as a modelling tool rather than a quote. They are computed from published rates we verified against vendor pricing pages, but real bills depend on negotiated terms, annual commitments, regional pricing and tier boundaries we cannot see. Several vendors publish no pricing at all and are marked as such rather than guessed at. Use this to find the shape of the answer and to spot the tools that are an order of magnitude wrong for your workload, then confirm with the vendor.

Why do some tools show as free?

Either they are open source and self-hostable at no licence cost, or your workload fits inside their free allowance. Free on licence is not free to operate - self-hosting Langfuse, SigNoz or LiteLLM means running the infrastructure, which is real money and real engineering time even though no vendor invoices you. The calculator shows licence and subscription cost, not total cost of ownership.

What is not included in these numbers?

Judge model inference, which is frequently the largest hidden cost. Evaluation frameworks like Ragas and guardrails like NeMo make additional model calls for every check they run, and those land on your model provider invoice rather than on the tool vendor's bill. A thousand-case eval suite across four metrics is several thousand judge calls per run. Self-hosting infrastructure is also excluded, as is engineering time.