Portkey Pricing Explained (2026) - What You Actually Pay
Portkey's meter caps logs, not requests, so your traffic keeps flowing while your observability quietly goes dark past the limit. Here is how the $49/mo Production tier really works, a worked bill, and cheaper picks for real tracing.
Published:
Portkey’s pricing looks simple - free, then $49/mo - but there is one meter detail that changes how you plan a deployment, and one structural gotcha about what “self-host” gets you. Let me decode both, walk a worked bill, then point at cheaper options if what you actually want is tracing depth.
Start with what Portkey is, because it decides whether the price is even relevant to you. Portkey is an LLM gateway, not an eval tool. It routes to 1,600+ models with fallbacks, caching, budgets and 50+ guardrails, and observability is a bundled bonus. If your problem is “we call five providers and costs are exploding,” this is aimed at you. If your problem is “score whether our agent’s answers are good,” that is a different tool.
The pricing model, decoded
Four tiers, and the free ones are more generous than they look.
| Tier | Price | Logs/mo | Log retention | Overage |
|---|---|---|---|---|
| Open Source | $0 | self-hosted | - | Apache 2.0 |
| Developer | $0 | 10,000 | 3 days | none |
| Production | $49/mo | 100,000 | 30 days | $9/100k requests |
| Enterprise | Custom | 10M+ | configurable | - |
The quota unit is logs, not requests, and that distinction is the whole game. The open-source gateway self-hosts free and does the hard networking - routing, retries, fallbacks, guardrails. The managed Developer tier is free forever for 10,000 logs a month with a 3-day retention window.
The trap is what happens at the cap. On Production you get 100,000 logs a month. Past that, your requests keep routing - Portkey never drops your traffic - but logs stop recording unless you pay the overage, which is $9 per additional 100,000 requests. That is friendlier than a hard request cap in one sense, because your app never breaks. But it means your observability silently goes dark at volume while everything looks healthy. There is no annual discount published.
Estimate your bill
Work three cases.
Small app, kicking the tires. Under 10,000 logs a month is $0 on the Developer tier. Fine for prototyping, but 3-day log retention is tight - you cannot look back at last week’s incident.
Production app at moderate volume. 100,000 logs a month is the flat $49/mo Production tier with 30-day retention. Clean and predictable, as long as you stay under the cap.
High-volume app. Say you route 500,000 requests a month and want them all logged. That is 100,000 included plus 400,000 overage, and the overage bills per 100k requests at $9, so 4 x $9 = $36 on top of the $49 base. Roughly $85/mo. Not bad - but the number you have to watch is whether logging keeps up, because past 100k the meter is on you, not Portkey.
The self-host gotcha
This is the part that catches people who assume open-source means free observability. The Apache 2.0 gateway gives you routing and a basic dashboard. Real observability lives on the paid tier.
Self-host Portkey and you get the pipe - routing, retries, fallbacks, load balancing, guardrails. What you do not get is meaningful logging, traces, analytics or retention. For that you move to managed Production or stand up your own logging stack. It is honest open-core, and the free part is genuinely useful, but it inverts the usual self-host promise. With Langfuse, self-host gets you the whole product. With Portkey, self-host gets you the pipe, not the dashboard.
One more number to treat as a claim. Portkey markets sub-1ms added latency and a 122kb footprint, but a benchmark run by Kong, a competing gateway vendor, reported Portkey at about 65% higher latency than Kong’s own gateway. Both parties are interested, so test on your own path before committing to a latency budget.
Cheaper alternatives for real observability
If what you actually want is tracing and eval depth rather than a gateway, two options serve that better.
Langfuse is the open-source observability default, and here self-host means the whole product. It is MIT-licensed and free to run, or $29/mo managed on its Core tier, and at 1M events a month it lands around $101/mo. It also ingests OpenTelemetry, which makes the common pairing easy - run Portkey as your gateway and feed Langfuse over OTel for the tracing. The full decode is in Langfuse pricing.
Helicone is the other proxy-based option, and I mention it only to warn you off. Mintlify acquired it on 3 March 2026 and put it in maintenance mode - security patches and bug fixes only, no roadmap, and Mintlify is actively helping customers migrate off. Its Apache-2.0 self-host still works, but building fresh on a frozen tool whose own owner points people to the exit is a dead end. Do not start there.
So which one?
- You call many providers and want one place to route, cache, budget and guardrail - Portkey, and expect to pay for the Production tier once you need real logs.
- You want tracing and eval depth you can self-host free - Langfuse, paired with Portkey over OpenTelemetry if you also need the gateway.
- You were about to pick Helicone for the proxy model - do not start there. It is frozen.
For the wider field, see Portkey alternatives and Portkey vs Langfuse. Every price here was read from portkey.ai on 26 July 2026. This category ships breaking changes monthly, so we re-verify every 30 days.
Frequently Asked Questions
How much does Portkey cost?
The self-hosted open-source gateway and the managed Developer tier are both free. Production is $49/mo for 100,000 logs, then $9 per additional 100,000 requests. Enterprise is custom for 10M+ logs plus SSO, VPC deployment and compliance reports. The detail that bites - the meter caps logs, not requests, so past the cap your calls still route, you just stop recording logs beyond the limit.
Can I self-host Portkey for free observability?
No. The gateway is open-source under Apache 2.0 and self-hosts free, giving you routing, retries, fallbacks, load balancing and guardrails plus a basic dashboard. But meaningful observability - full logs, traces, analytics and retention - is not in the open-source build. For that you move to the managed Production tier or build your own logging stack. Self-hosting Portkey does not get you self-hosted observability.
What does the Portkey log cap actually mean?
Production includes 100,000 logs a month. Past that, your requests keep routing normally - Portkey never breaks your traffic - but logging stops unless you pay the overage of $9 per additional 100,000 requests. The upside is your app never fails at the cap. The downside is your observability quietly goes dark at volume unless you are watching the meter.
Is Portkey an observability tool or a gateway?
A gateway first. Portkey routes to 1,600+ models with fallbacks, caching, budgets and 50+ guardrails, and it happens to include observability. If your core need is scoring or tracing depth, a purpose-built observability tool like Langfuse serves you better. Many teams run Portkey as the gateway and feed a dedicated observability backend over OpenTelemetry.
Explore More
Tool Reviews
Related Articles
- The 2026 LLM Observability Consolidation Map - Who Got Bought, Who Stayed Free
- The Best LLM Monitoring Tools in 2026, Ranked for Production Cost and Reliability
- The Best LLM Observability for OpenAI Apps in 2026, by Use Case
- How to Reduce LLM Costs in 2026 - 6 Levers That Actually Move the Bill
- 3 OpenRouter Alternatives for Teams That Outgrew the Hosted Router (2026)
Free Newsletter
Get the LLM Evals Newsletter
Platform comparisons, pricing changes and eval technique deep-dives. No spam.
Related Articles
Evaluation of LLM Applications: A Practical 2026 Guide
A vendor-neutral guide to evaluation of LLM applications: metric selection, dataset sizing math, judge calibration, cost models and a tool comparison.
August 9, 2026
guideBLEU vs ROUGE vs BERTScore - Which to Use and Why All Three Fail on Chat
BLEU counts precision, ROUGE counts recall, BERTScore compares embeddings. Here is how each one actually computes a score, a worked example on the same sentence, and why none of them can grade an open-ended LLM answer.
July 28, 2026
guideContext Precision vs Recall Explained - Diagnosing RAG Retrieval in 2026
Context precision punishes noise, context recall punishes gaps. Here is how each retrieval metric is computed, a worked example, and how the two scores together tell you whether your retriever is over-fetching or missing documents.
July 28, 2026
Portkey Review
Langfuse Review