Amazon Bedrock Evaluations Review (2026)
AWS-native model and agent evaluation, billed per evaluation job. The recurring criticism is fragmentation - CloudWatch, Prompt Flows and Evaluations look like one product on paper and behave like three in practice.
Rating
Starting Price
Per evaluation job
Free Plan
No
SDKs & Frameworks
2
Deployment
5
Best For
Teams already committed to Bedrock who want evaluation inside their existing AWS account and IAM boundary, and who value procurement simplicity over best-in-class agent tooling.
Last Updated:
10 Things You Should Know About Amazon Bedrock Evaluations
- 1 Bedrock offers Model Evaluation jobs billed per job, with custom model hosting billed per inference
- 2 AgentCore launched in preview in July 2025 and reached general availability on 13 October 2025
- 3 AgentCore added Policy and Evaluations in preview in December
- 4 Bedrock uses on-demand per-token billing with no minimum commitment
- 5 Provisioned throughput is available at a fixed hourly rate with model-specific throughput units
- 6 A recurring criticism is fragmentation across CloudWatch metrics, Bedrock Prompt Flows and Bedrock Evaluations
Pros & Cons
Pros
- ✓ Everything stays inside your AWS account and IAM boundary, which resolves data residency and procurement questions in one step
- ✓ No new vendor relationship, contract or security review if you are already on AWS
- ✓ Billed per evaluation job rather than as a subscription, so occasional use is genuinely cheap
- ✓ AgentCore reached GA in October 2025, so agent capability is no longer preview-only
- ✓ Human evaluation workflows are built in, which several dedicated tools lack
Cons
- ✕ Fragmented across CloudWatch, Bedrock Prompt Flows and Bedrock Evaluations - three surfaces that read as one product and behave as three
- ✕ AgentCore Policy and Evaluations were still preview features as of December, so agent evaluation is newer than the GA date suggests
- ✕ Weaker than dedicated tools on agent trajectory analysis and simulation
- ✕ Ties your evaluation workflow to AWS, which is awkward if you evaluate models across clouds
- ✕ Per-job pricing is hard to forecast without knowing your job cadence
Features
The argument is the boundary, not the features
Bedrock Evaluations will not beat a dedicated evaluation platform on capability. That is not why you would choose it.
Everything stays inside your AWS account and IAM perimeter. No new vendor relationship. No new contract. No new security review. No new data processing agreement. No prompts crossing a boundary you have already cleared.
For a regulated enterprise that spent months getting AWS approved, adding a specialist evaluation vendor can be a multi-month procurement project for a tool that costs less than the review costs. In that situation, the AWS-native option wins on grounds that have nothing to do with its feature list.
If that does not describe you, dedicated tools are better at evaluation and this page is straightforward: use one of them.
Fragmentation is the real complaint
The recurring criticism is fair and worth understanding before you commit.
CloudWatch metrics for Bedrock, Bedrock Prompt Flows, and Bedrock Evaluations are described as three surfaces that feel like one on paper but three in practice.
That is the characteristic experience of AWS-native tooling. Each component works. Assembling them into a coherent workflow is your job - trace in one place, evaluate in another, monitor in a third, correlate manually.
Compare a dedicated platform like LangWatch or Braintrust, where the trace, the dataset and the evaluation are the same object with the same identity. On AWS you get capable primitives and the integration work.
Whether that trade is acceptable depends on whether you have platform engineers who will own the stitching. If nobody owns it, you get three tools nobody uses together.
The agent story is younger than it looks
Worth being precise about the timeline:
| Date | Milestone |
|---|---|
| July 2025 | AgentCore launches in preview |
| 13 October 2025 | AgentCore reaches GA |
| December | Policy and Evaluations added in preview |
So AgentCore is generally available, but the agent evaluation capability specifically arrived in preview afterwards.
Preview features on AWS carry no SLA and can change. If agent evaluation is the reason you are here, confirm what is GA today rather than reading the AgentCore GA announcement as covering it.
Against LangWatch’s simulation architecture - Agent Under Test, User Simulator, Judge, running under pytest - Bedrock’s agent trajectory analysis is considerably thinner.
Pricing, and what to actually model
Model Evaluation jobs are billed per job, with model inference charged separately and no subscription. Bedrock overall uses on-demand per-token billing with no minimum commitment, with provisioned throughput available at a fixed hourly rate.
Per-job billing is excellent for occasional use - a team evaluating weekly pays very little.
The forecasting difficulty is that cost tracks job cadence rather than traffic, and cadence is a decision rather than a measurement. Move to evaluating on every pull request and both the job count and the underlying inference rise sharply.
Model your intended cadence, not your current one. That is the same advice this site gives about Ragas, NeMo Guardrails and LangWatch simulations, and it is the most consistently underestimated cost in this entire category.
The single-provider constraint
The tooling is built around Bedrock-hosted models and the AWS estate.
That is fine if Bedrock is where your models live. It is awkward if you are comparing across providers - and comparing models is one of the main reasons teams evaluate at all.
This is structurally the same limitation OpenAI Evals had, and the reason provider-agnostic tools like promptfoo, DeepEval and Ragas are more flexible. If your evaluation question is which provider should we use, do not run it inside one provider’s platform.
Comparing the three clouds
One practical recommendation worth repeating, because the three platforms use different units and comparison is harder than it should be:
Price three real workloads before signing - a routine production call, a long-running agent workflow, and a high-volume reserved-capacity path.
The cheapest platform on the marketing page is frequently not the cheapest after AWS commitments, Azure EA terms or GCP reserved-throughput planning. Bedrock is per-token on demand plus hourly provisioned throughput; Azure AI Foundry prices OpenAI models identically to OpenAI direct with no separate runtime fee; Vertex is per-token plus GCP compute.
Should you use it?
Use Bedrock Evaluations if you are already committed to Bedrock and want evaluation inside your existing AWS account and IAM boundary, and procurement simplicity outweighs tooling quality.
Don’t use it if you evaluate across providers, you need strong agent trajectory analysis, or you have no platform engineering capacity to stitch the surfaces together.
Bottom line: competent AWS-native primitives with an integration burden you inherit, and a genuine advantage in never leaving your account boundary. Confirm what is GA before planning around agent evaluation.
Pricing model, AgentCore timeline and platform criticisms verified against AWS documentation and third-party comparison analyses on 3 August 2026. Preview feature status changes frequently and should be reconfirmed. This is a researched directory entry - we have not yet instrumented this platform with our reference application.
Pricing Plans
Model Evaluation
Billed per job
- Charged per evaluation job run
- Model inference billed separately
- No subscription
AgentCore
Usage-based
- GA since 13 October 2025
- Policy and Evaluations added in preview in December
- Priced within the wider Bedrock usage model
SDKs & Frameworks
Deployment
Eval Methods
Billing Unit
Our Verdict
Bedrock Evaluations is the sensible choice for a specific buyer - the team already on Bedrock - and unremarkable for anyone else. The argument is not capability, it is boundary. Everything stays inside your AWS account and IAM perimeter, there is no new vendor, no new contract and no new security review, and evaluation is billed per job rather than as a subscription so occasional use costs very little. For a regulated enterprise that has already cleared AWS, that combination resolves in one step what a specialist tool turns into a procurement project. The recurring criticism is fragmentation, and it is fair. CloudWatch metrics for Bedrock, Bedrock Prompt Flows and Bedrock Evaluations are described as three surfaces that feel like one on paper and three in practice, and that is the characteristic experience of AWS-native tooling - each piece works, and stitching them into a coherent workflow is your job. On agents specifically, AgentCore reached GA in October 2025 but Policy and Evaluations arrived in preview in December, so the agent evaluation story is newer than the headline suggests.
Similar Tools
LangWatch
Teams building multi-turn or multi-agent systems who need evaluation that models conversations rather than scoring single outputs, and who want it running in CI.
AgentOps
Teams debugging multi-agent systems who want session replay and broad framework coverage with minimal instrumentation effort, and who will move to the paid tier quickly.
Azure AI Foundry Evaluation
Enterprises already standardised on Azure that need governance, audit trails and red-teaming around AI usage as much as they need the models themselves.
Databricks Agent Evaluation
Existing Databricks customers building agents over their own governed data, where inheriting Unity Catalog permissions and MLflow lineage is worth more than best-in-class conversation simulation.
Frequently Asked Questions
What is the fragmentation problem?
That the pieces are separate products wearing a shared name. CloudWatch metrics for Bedrock, Bedrock Prompt Flows and Bedrock Evaluations are described as three surfaces that feel like one on paper but three in practice, and that matches the general experience of AWS-native tooling. Each component works. Assembling them into a coherent workflow - trace here, evaluate there, monitor somewhere else, correlate manually - is your responsibility. A dedicated platform like LangWatch or Braintrust gives you one product where the trace, the dataset and the evaluation are the same object. On AWS you get capable primitives and the integration work.
Is the agent evaluation actually ready?
Newer than the dates imply, so check current status before planning around it. AgentCore launched in preview in July 2025 and reached general availability on 13 October 2025, which sounds settled. But Policy and Evaluations were added in preview in December, meaning the agent evaluation capability specifically is considerably younger than AgentCore's GA milestone. Preview features on AWS carry no SLA and can change. If agent evaluation is the reason you are looking at this, confirm what is GA today rather than relying on the AgentCore announcement.
When does staying inside AWS beat a specialist tool?
When procurement and data boundaries dominate the decision. Everything running inside your AWS account and IAM perimeter means no new vendor relationship, no new contract, no new security review, no new data processing agreement and no prompts leaving a boundary you have already cleared. For a regulated enterprise that has spent months getting AWS approved, adding a specialist eval vendor can be a multi-month project for a tool that costs less than the review. If that describes you, Bedrock Evaluations wins on grounds that have nothing to do with its features. If it does not, dedicated tools are better at evaluation.
How does per-job billing work out?
Well for occasional use, less predictably at scale. Model evaluation jobs are billed per job, with model inference charged separately, and there is no subscription - so a team running evaluations weekly pays very little. The forecasting difficulty is that your cost tracks job cadence rather than traffic, and job cadence is a decision rather than a measurement. If you move to running evaluations on every pull request, the number of jobs rises sharply and the underlying inference rises with it. Model your intended cadence rather than your current one, which is the same advice we give about every judge-based evaluation.
Can I use it to evaluate models outside Bedrock?
Not naturally. The tooling is built around Bedrock-hosted models and the AWS estate, which is fine if Bedrock is where your models live and awkward if you are comparing across providers. Since comparing models is one of the main reasons teams evaluate at all, that constraint matters - it is the same structural limitation OpenAI Evals had with its single-provider coupling, and the reason provider-agnostic tools like promptfoo, DeepEval and Ragas are more flexible. If your evaluation question is which provider to use, do not run it inside one provider's platform.
How should I compare the three cloud platforms on cost?
By pricing real workloads rather than reading rate cards, because the three use different units and comparison is harder than it should be. Bedrock is on-demand per-token with no minimum commitment, plus provisioned throughput at a fixed hourly rate. Azure AI Foundry prices OpenAI models identically to OpenAI direct and charges no separate runtime fee. Vertex AI is per-token plus GCP compute. One practical recommendation worth following is to price three real workloads before signing - a routine production call, a long-running agent workflow, and a high-volume reserved-capacity path - because the cheapest platform on the marketing page is frequently not the cheapest after AWS commitments, Azure EA terms or GCP reserved-throughput planning.