Arthur Review (2026)
Open-sourced its real-time evaluation engine and, unusually for this segment, publishes actual prices. Free tier has unlimited seats; Premium is $60/mo. The meter is use cases rather than volume, which is worth understanding.
Rating
Starting Price
$60/mo
Free Plan
Yes
SDKs & Frameworks
3
Deployment
4
Best For
Teams that want open-source guardrails running inside their own stack with the option of an inexpensive managed tier, and who value published pricing and unlimited seats.
Last Updated:
10 Things You Should Know About Arthur
- 1 The Arthur Engine is open source and described as always free, running locally inside your own stack
- 2 Arthur maintains two open-source tools under permissive licences - Arthur Bench for evaluation and Arthur Engine for monitoring and guardrails
- 3 The Free tier includes core performance metrics, cloud data connectors, up to 4 use cases and unlimited seats
- 4 Premium is $60 per month with customisable metrics and dashboards, custom alerting and webhooks, and up to 100 use cases
- 5 Enterprise offers dedicated and managed VPC options
- 6 Arthur Shield, an enterprise LLM firewall, is a separate product with custom pricing
- 7 The Engine collects OpenInference OpenTelemetry traces from any agentic workflow
- 8 Agent Discovery and Governance launched in December 2025
- 9 Engine releases shipped in February and April 2026
Pros & Cons
Pros
- ✓ Publishes real prices, which almost nothing else in this segment does - Lakera, Fiddler, HiddenLayer, CalypsoAI and Prompt Security are all sales-gated
- ✓ The Engine is open source under a permissive licence and runs locally inside your stack, so prompts never leave your environment
- ✓ Unlimited seats on the free tier, so giving people visibility into quality costs nothing
- ✓ Premium at $60/mo is inexpensive for what it covers
- ✓ Ingests OpenInference OTEL traces from any agentic workflow, so it is standards-based rather than requiring proprietary instrumentation
- ✓ Actively developed, with substantive Engine releases in February and April 2026
Cons
- ✕ Metering by use cases rather than volume is unusual, and what counts as a use case is not obvious from the pricing page
- ✕ Arthur Shield, the enterprise LLM firewall, is a separate product with custom pricing, so the published prices do not cover the full guardrails story
- ✕ The open-source Engine and the commercial platform overlap confusingly, and it takes effort to work out which capability sits where
- ✕ Smaller presence than the acquired competitors now backed by Palo Alto, Cisco and Check Point
- ✕ Hallucination and prompt injection detection quality is not independently verified
Features
Transparent pricing, which is remarkable here
Start with the thing that distinguishes Arthur before any feature comparison.
| Tier | Price | Included |
|---|---|---|
| Evals Engine | $0, open source | PII, sensitive data, custom LLM and regex rules; self-hosted |
| Free | $0 | Core metrics, cloud connectors, 4 use cases, unlimited seats |
| Premium | $60/mo | Custom metrics and dashboards, alerting and webhooks, 100 use cases |
| Enterprise | Custom | Dedicated and managed VPC |
Now compare with the rest of this segment: Lakera, Fiddler, HiddenLayer, CalypsoAI and Prompt Security publish nothing. Every one requires a sales conversation before you can learn what anything costs.
That makes the commercial AI guardrails market effectively opaque to anyone doing self-directed comparison. Arthur is the exception, and it is a real advantage independent of product quality - it means Arthur gets evaluated by teams that would never enter a sales cycle for a guardrail.
Unlimited seats on the free tier deserves specific credit. Charging per seat for a quality and safety tool means fewer people look at the data, which is the opposite of what you want. Arthur removes that friction entirely at the bottom tier.
The meter is use cases, which needs a question
The one thing to establish before budgeting: Arthur meters use cases, not requests, traces or seats. Four on Free, a hundred on Premium.
That is an unusual unit, in the same family as Langtail charging by prompt count, and it shares the weakness that the unit does not track volume or cost incurred.
A hundred use cases is generous enough that most teams will never approach it. But “use case” could plausibly mean a monitored model, an application, or a configured evaluation - and those are very different ceilings. Ask directly rather than assuming.
Three products, and the boundaries are muddy
Worth untangling, because the overlap is genuinely confusing:
- Arthur Engine - open source, always free, self-hosted. Real-time evaluation and guardrails covering PII, sensitive data, hallucination, prompt injection and toxicity, with custom LLM and regex rules.
- The commercial platform (Free / Premium / Enterprise) - managed monitoring, dashboards, alerting, integrations layered on top.
- Arthur Shield - a separate enterprise LLM firewall with its own custom pricing.
So the published prices cover the monitoring platform, not Shield, and the open-source Engine overlaps parts of both. If you are quoting a number to anyone, work out which product you actually need first, because “Arthur is $60 a month” is only true for one of three things.
Arthur also maintains Arthur Bench, a separate open-source LLM evaluation tool.
It runs in your own stack
The Engine was released to run locally inside your own stack, rather than requiring data be sent to a third-party platform, and that architecture appears unchanged.
For a guardrail this matters more than usual: the prompts you most want to inspect are frequently the ones you least want to transmit. A managed detection API means every sensitive prompt is shipped to a vendor precisely because it is sensitive.
That puts Arthur alongside Fiddler’s in-VPC models and the now-archived LLM Guard, and against managed APIs like Lakera where inspection happens on the vendor’s infrastructure. With LLM Guard gone, the set of credible self-hosted options has shrunk, and Arthur is one of the remaining few.
It ingests OpenInference OpenTelemetry traces from any agentic workflow, so instrumentation is standards-based rather than proprietary - the same portability argument that makes OpenLLMetry and Langtrace attractive.
Actively maintained
Confirmed with recent substance, which is worth checking in a segment where LLM Guard was archived in July 2026:
- 28 April 2026 - unified Evaluators interface, bulk evaluation testing across trace IDs, automated 24-hour compliance checks, configurable trace retention, an Engine chatbot assistant, Apple Silicon SentenceTransformer support
- 19 February 2026 - flexible analytics across models, agents and datasets, an Agent Span Count metric, redesigned Trace Viewer, dark mode
- December 2025 - Agent Discovery and Governance launched for managing agentic AI in production
Arthur or Guardrails AI?
They overlap substantially, and the deciding factors are mostly not technical.
Both are permissively licensed, both run in your environment, both cover PII, toxicity and injection detection.
Guardrails AI has the larger validator library - 50+ pre-built checks in its Hub - more community adoption, and custom validators as plain Python classes.
Arthur brings OTEL trace ingestion, continuous evaluation against live production traces, and a cheap managed tier with published pricing if you want to stop self-hosting.
If you want the widest validator selection, Guardrails AI. If you want open-source guardrails with a transparent, inexpensive managed option behind them, Arthur.
Should you use it?
Use Arthur if you want open-source guardrails running in your own stack, you value published pricing and unlimited seats, or you want continuous evaluation over OTEL traces from agentic workflows.
Don’t use it if you need the enterprise firewall specifically - Shield is separately priced and not covered by the transparent tiers - or you need independently verified detection quality.
Bottom line: the most transparent vendor in an opaque segment, with a genuinely open-source engine that keeps your prompts in your environment. Clarify what a use case means and which of the three products you are buying, and the value is easy to see.
Pricing tiers, open-source status, release history and architecture verified against the vendor’s pricing page, repository and newsroom on 3 August 2026. Detection quality claims are vendor-stated and have not been independently tested. Arthur Shield pricing is not published. This is a researched directory entry - we have not yet instrumented this platform with our reference application.
Pricing Plans
Evals Engine (Open Source)
$0
- Always free and open source
- PII, sensitive data, custom LLM and regex rules built in
- Self-serve deployment
- Runs locally inside your own stack
Free
$0
- Core performance metrics
- Cloud data connector integrations
- Monitoring for up to 4 use cases
- Unlimited seats
Premium
$60/mo
- Customisable performance metrics and dashboards
- Custom alerting and webhook integrations
- Monitoring for up to 100 use cases
Enterprise
Custom
- Dedicated and managed VPC options
- Arthur Shield sold separately with custom pricing
SDKs & Frameworks
Deployment
Eval Methods
Our Verdict
Arthur is the most transparent vendor in a segment that has almost none. It publishes actual prices - Free at zero with unlimited seats and four use cases, Premium at $60 a month for a hundred - while Lakera, Fiddler, HiddenLayer, CalypsoAI and Prompt Security all require a sales conversation before you can learn anything about cost. It also open-sourced the Arthur Engine, a real-time evaluation and guardrails service covering PII and sensitive data leakage, hallucination, prompt injection and toxicity, which runs locally inside your own stack rather than shipping prompts to a third party. That combination of published pricing, permissive licensing and in-environment execution is genuinely rare here. Two things to understand. The meter is use cases rather than requests or volume, and what constitutes a use case is not obvious from the pricing page, so establish that before budgeting. And Arthur Shield, the enterprise LLM firewall, is a separate product with custom pricing, so the transparent prices do not cover the whole guardrails story.
Similar Tools
Fiddler AI
Regulated enterprises that need guardrails running inside their own environment, particularly those already running predictive ML alongside LLM systems, and who can work with enterprise procurement.
Guardrails AI
Teams that want composable, permissively licensed output validation living next to their application code, especially those who prefer plain Python over a policy DSL.
HiddenLayer
Enterprises and government buyers who need model supply chain security, AI asset discovery and airgapped operation, rather than application-layer output filtering.
Lakera
Teams that want managed, low-latency prompt injection defence with enterprise support and are comfortable with a proprietary API and an enterprise sales process.
Frequently Asked Questions
What counts as a use case?
This is the thing to establish before you budget, and the pricing page does not make it obvious. Arthur meters use cases rather than requests, traces or seats - four on Free and a hundred on Premium. That is an unusual meter, in the same family as Langtail charging by number of prompts, and it shares the same weakness that the unit does not track volume or cost incurred. A hundred use cases is generous enough that most teams will not hit it, but a use case could plausibly mean a monitored model, an application, or a configured evaluation, and those are very different ceilings. Ask directly rather than assuming.
Why does publishing prices matter so much here?
Because virtually nothing else in this segment does. Lakera, Fiddler, HiddenLayer, CalypsoAI and Prompt Security all require a sales conversation before you can learn what anything costs, which means the commercial AI guardrails market is effectively opaque to anyone doing self-directed comparison. Arthur publishing Free at zero with unlimited seats and Premium at $60 a month lets you make a decision in an afternoon. That is a real advantage independent of the product, because it means Arthur gets evaluated by teams that will never enter a sales cycle for a guardrail.
What is the difference between the Engine, Shield and the platform?
This is genuinely confusing and worth untangling. The Arthur Engine is the open-source, always-free service you self-host, providing real-time evaluation and guardrails including PII, sensitive data, hallucination, prompt injection and toxicity detection with custom LLM and regex rules. The commercial platform - Free, Premium and Enterprise - adds managed monitoring, dashboards, alerting and integrations on top. Arthur Shield is a separate enterprise LLM firewall with its own custom pricing. So the published prices cover the monitoring platform, not Shield, and the open-source Engine overlaps with parts of both. Work out which product you actually need before quoting anyone a price.
Does the open-source Engine send my data anywhere?
No, and this is one of its better properties. The Engine was released to run locally inside your own stack, unlike solutions that require sending data to a third-party platform, and that architecture appears unchanged. For a guardrail this matters more than usual, because the prompts you most want to inspect are frequently the ones you least want to transmit. It puts Arthur in the same architectural category as Fiddler's in-VPC models and the now-archived LLM Guard, and against managed APIs like Lakera where inspection happens on the vendor's infrastructure.
Is it actively maintained?
Yes, with recent substance. The Engine shipped an update on 28 April 2026 adding a unified Evaluators interface, bulk evaluation testing across trace IDs, automated 24-hour compliance checks, configurable trace retention policies, an Engine chatbot assistant and Apple Silicon SentenceTransformer support. A February 2026 release added flexible analytics across models, agents and datasets, an Agent Span Count metric, a redesigned Trace Viewer and dark mode. Arthur also launched Agent Discovery and Governance in December 2025 for managing agentic AI in production. That is a healthy cadence, and worth confirming given that LLM Guard in this same segment was archived in July 2026.
How does it compare with Guardrails AI?
They overlap substantially and the deciding factors are mostly non-technical. Both are permissively licensed, both run in your own environment, both cover PII, toxicity and injection detection. Guardrails AI has a larger validator library through its Hub with 50+ pre-built checks, more community adoption, and custom validators as plain Python classes. Arthur brings OTEL trace ingestion, continuous evaluation on live production traces, and a managed tier with published pricing if you want one. If you want the widest validator selection, Guardrails AI. If you want open-source guardrails with a cheap, transparent managed option behind them, Arthur.