LLM Tool Cost Calculator
Published prices in this category are not comparable, because every vendor meters something different. Enter your workload once and see what 30 tools would actually charge for it.
We normalise 11 different billing meters - event, span, record, use-case, gb-month, markup-pct, request, gb-ingested, job, seat, prompt - into one annual figure. Tools that publish no pricing are marked rather than guessed at.
| Tool | Meter | Per month | Per year | Basis |
|---|---|---|---|---|
| Inspect AI | - | Free | $0 | Free, self-hosted · MIT, UK AISI. No commercial tier. |
| LangWatch | event | Free | $0 | Within free allowance · Apache-2.0 core. Per-event rates not published. |
| LiteLLM | - | Free | $0 | Free, self-hosted · MIT, zero markup. Cost is infrastructure. |
| Agenta | - | Free | $0 | Free, self-hosted · MIT self-host; cloud pricing not verified. |
| Arize Phoenix | - | Free | $0 | Free, self-hosted · Self-host free (ELv2). Online monitoring requires AX. |
| Azure AI Foundry Evaluation | - | Free | $0 | No usage metering · Model inference and tool calls only; no runtime fee. |
| Databricks Agent Evaluation | - | Free | $0 | No usage metering · Inside Databricks consumption. |
| Confident AI (DeepEval) | - | Free | $0 | Free, self-hosted · Apache 2.0. |
| Evidently | - | Free | $0 | Free, self-hosted · Apache 2.0; cloud pricing not published. |
| Fiddler AI | - | Free | $0 | No usage metering · No public pricing. VPC from Enterprise. |
| Freeplay | - | Free | $0 | No usage metering · No public pricing. |
| Galileo | span | Free | $0 | Within free allowance · Pricing not published. |
| Giskard | - | Free | $0 | Free, self-hosted · Apache 2.0. Hub pricing not published. |
| Guardrails AI | - | Free | $0 | Free, self-hosted · Apache 2.0. |
| HiddenLayer | - | Free | $0 | No usage metering · No public pricing. |
| HoneyHive | event | Free | $0 | Within free allowance · Free tier only; paid pricing not published. |
| Lakera | - | Free | $0 | No usage metering · No public pricing; Check Point procurement. |
| Laminar | span | Free | $0 | Within free allowance · Open source. |
| Latitude | - | Free | $0 | Free, self-hosted · MIT self-host; cloud pricing not verified. |
| LM Evaluation Harness | - | Free | $0 | Free, self-hosted · Free, EleutherAI. |
| MLflow | - | Free | $0 | Free, self-hosted · Apache 2.0. Managed via Databricks consumption. |
| NVIDIA NeMo Guardrails | - | Free | $0 | Free, self-hosted · Apache 2.0. Cost is extra model calls per rail. |
| Not Diamond | - | Free | $0 | No usage metering · Pricing sources conflict. |
| OpenLIT | - | Free | $0 | Free, self-hosted · Apache 2.0, self-hosted. |
| Patronus AI | - | Free | $0 | No usage metering · No public pricing. |
| Portkey | - | Free | $0 | Free, self-hosted · Apache-2.0 gateway; observability is paid. |
| Promptfoo | - | Free | $0 | Free, self-hosted · MIT. Now an OpenAI company. |
| Ragas | - | Free | $0 | Free, self-hosted · Free. Cost is judge model calls. |
| Vertex AI Gen AI Evaluation Service | - | Free | $0 | No usage metering · Per token plus GCP compute. |
| W&B Weave | gb-ingested | Free | $0 | Within free allowance · Per-GB ingested plus seats. Rates not published. |
| Maxim AI | span | Free | $0 | Within free allowance · Contact sales. |
| Aporia | - | Free | $0 | No usage metering · Inside Coralogix platform. |
| CalypsoAI | - | Free | $0 | No usage metering · Absorbed into F5 platform. |
| Langtrace | - | Free | $0 | Free, self-hosted · Cloud currently free; AGPL-3.0 server. |
| Openlayer | - | Free | $0 | No usage metering · No public pricing. |
| Prompt Security | - | Free | $0 | No usage metering · Absorbed into SentinelOne Singularity. |
| Traceloop | span | Free | $0 | Within free allowance · OpenLLMetry Apache-2.0. Now ServiceNow. |
| TruLens | - | Free | $0 | Free, self-hosted · MIT, Snowflake-maintained. |
| UpTrain | - | Free | $0 | Free, self-hosted · Apache 2.0. Managed pricing unclear. |
| LLM Guard | - | Free | $0 | Free, self-hosted · MIT but archived 9 July 2026. |
| Martian | - | Free | $0 | No usage metering · Router no longer offered. |
| OpenAI Evals | - | Free | $0 | No usage metering · Hosted platform shuts down 30 Nov 2026. |
| Pezzo | - | Free | $0 | Free, self-hosted · Unmaintained since mid-2025. |
| WhyLabs | - | Free | $0 | Free, self-hosted · Company shut down; open-sourced. |
| Baserun | - | Free | $0 | No usage metering · Acquired by LlamaIndex. |
| Gentrace | - | Free | $0 | No usage metering · Shut down; MIT source released. |
| Literal AI | - | Free | $0 | No usage metering · Discontinued. |
| Vellum | - | Free | $0 | No usage metering · No longer a developer platform. |
| Cloudflare AI Gateway | record | $5.00 | $60 | 800,000 records/mo, 100,000 free · $5 Workers Paid = 1M logs. No token markup. |
| PromptHub | request | $9.00 | $108 | 100,000 requests/mo, 2,000 free · Pro. Free tier makes prompts public. |
| Opik | span | $19 | $228 | 800,000 spans/mo · Cheapest paid cloud; Apache-2.0 self-host. |
| Sentry | event | $26 | $312 | 800,000 events/mo, 5,000 free · Team tier. Not an LLM observability tool. |
| Langfuse | event | $29 | $348 | 800,000 events/mo, 50,000 free · Core tier; self-host free under MIT. |
| Lunary | event | $30 | $360 | 800,000 events/mo, 30,000 free · Free tier is 1,000/DAY not monthly. Apache-2.0. |
| AgentOps | event | $40 | $480 | 800,000 events/mo, 5,000 free · Event = each LLM/tool call, not each run. |
| Pydantic Logfire | record | $49 | $588 | 800,000 records/mo, 10,000,000 free · Team tier. Records = spans + logs + metrics. |
| SigNoz | gb-ingested | $49 | $588 | 5 GB/mo · Cloud from $49; self-host free (ClickHouse). |
| Arthur | use-case | $60 | $720 | 0 use-cases/mo, 4 free · Premium: 100 use cases, unlimited seats. |
| Helicone | request | $79 | $948 | 100,000 requests/mo, 10,000 free · Pro. Overage rate not published. Maintenance mode. |
| Langtail | prompt | $99 | $1,188 | 0 prompts/mo, 2 free · Pro caps at 20 prompts, not requests. |
| Vercel AI Gateway | - | $100 | $1,200 | Flat platform fee · No token markup. $5/mo free credits. |
| LangSmith | seat | $195 | $2,340 | 5 seats · Plus tier, per seat. |
| Confident AI | gb-month | $204 | $2,446 | 5 GB/mo, 1 free · Starter, unlimited seats, 5 GB-months included. |
| Braintrust | span | $249 | $2,988 | 800,000 spans/mo · Pro tier. Billed on processed data. |
| Parea AI | event | $250 | $3,000 | 800,000 events/mo, 3,000 free · Team: 100k logs, 3 seats then $50 each. |
| OpenRouter | markup-pct | $275 | $3,300 | 5.5% of $5,000 provider spend · 5.5% of provider spend, no volume discounts. |
| PromptLayer | request | $342 | $4,098 | 100,000 requests/mo, 2,500 free · Pro. $0.003 per transaction overage. |
| New Relic AI Monitoring | gb-ingested | $495 | $5,940 | 5 GB/mo, 100 free · Plus CCU meter for AI features. |
| TrueFoundry | request | $499 | $5,988 | 100,000 requests/mo, 50,000 free · Pro tier. |
| Kong AI Gateway | - | $500 | $6,000 | Flat platform fee · Konnect ~$500-2500/mo. OSS gateway free. |
| Arize AX | span | $800 | $9,600 | 800,000 spans/mo, 25,000 free · AX Pro, 50k spans. Seats not metered. |
| Datadog LLM Observability | span | $1,216 | $14,592 | 800,000 spans/mo, 40,000 free · Pro tier, 100k spans. New pricing from 1 May 2026. |
The estimate almost everyone gets wrong
Teams model cost from request volume. Most tools in this category do not bill requests. They bill spans, and a single agentic request produces 20 to 50 of them. That one substitution is routinely an order-of-magnitude error, and it is why free tiers that look generous evaporate in days.
Before you commit to anything here, instrument one representative request and count the spans it actually emits. That number, not your traffic, is what determines your bill.
Frequently Asked Questions
Why can I not just compare the published prices?
Because vendors meter different things. Datadog bills per LLM span, AgentOps bills per event, W&B Weave bills per GB ingested, Pydantic Logfire bills records, Confident AI bills GB-months, Langtail bills the number of prompts you have, OpenRouter takes a percentage of provider spend, and LangSmith bills seats. A published price of $50 a month means something completely different in each case. This calculator normalises them to one workload so the numbers are actually comparable.
Why does spans per request matter so much?
Because it is the difference between a manageable bill and an unmanageable one, and it is the single most common estimating mistake in this category. A simple completion produces roughly 3 to 5 spans. An agentic workflow with tool calls, retrieval and reasoning steps commonly produces 20 to 50. Any tool that meters spans therefore charges an agent workload up to ten times what it charges a completion workload at identical request counts. Datadog's 40,000-span free tier is a few thousand simple requests, or a day or two of agent traffic.
Why does trace size affect the cost?
Because several vendors bill volume rather than events. W&B Weave charges per GB ingested and Confident AI charges GB-months, which means two applications making identical numbers of calls can have very different bills. A RAG system logging a long system prompt, ten retrieved chunks and a lengthy completion produces a far larger payload than a classification endpoint returning one word. If you bill by volume, verbosity is the cost driver.
How accurate are these figures?
Treat them as a modelling tool rather than a quote. They are computed from published rates we verified against vendor pricing pages, but real bills depend on negotiated terms, annual commitments, regional pricing and tier boundaries we cannot see. Several vendors publish no pricing at all and are marked as such rather than guessed at. Use this to find the shape of the answer and to spot the tools that are an order of magnitude wrong for your workload, then confirm with the vendor.
Why do some tools show as free?
Either they are open source and self-hostable at no licence cost, or your workload fits inside their free allowance. Free on licence is not free to operate - self-hosting Langfuse, SigNoz or LiteLLM means running the infrastructure, which is real money and real engineering time even though no vendor invoices you. The calculator shows licence and subscription cost, not total cost of ownership.
What is not included in these numbers?
Judge model inference, which is frequently the largest hidden cost. Evaluation frameworks like Ragas and guardrails like NeMo make additional model calls for every check they run, and those land on your model provider invoice rather than on the tool vendor's bill. A thousand-case eval suite across four metrics is several thousand judge calls per run. Self-hosting infrastructure is also excluded, as is engineering time.