Azure AI Foundry Evaluation logo

Azure AI Foundry Evaluation Review (2026)

Microsoft's evaluation, red-teaming and observability layer, positioned as a production lifecycle tool rather than a model API. Charges no separate runtime fee - but Foundry Memory was still in preview as of Q1 2026.

Researched

Rating

4.0

Starting Price

Model inference cost only

Free Plan

No

SDKs & Frameworks

4

Deployment

4

Best For

Enterprises already standardised on Azure that need governance, audit trails and red-teaming around AI usage as much as they need the models themselves.

Last Updated:

10 Things You Should Know About Azure AI Foundry Evaluation

  1. 1 Microsoft frames Foundry as a full production lifecycle tool covering evaluation, red-teaming and observability alongside inference
  2. 2 It prices OpenAI models identically to OpenAI's direct API while adding Azure's enterprise SLA
  3. 3 Foundry charges no separate runtime fee, billing only for underlying model inference and tool calls
  4. 4 Pay-as-you-go is the default, with Provisioned Throughput Units available for guaranteed capacity
  5. 5 Foundry Memory was still in Public Preview as of Q1 2026, not generally available
  6. 6 Total costs can span model inference, orchestration, retrieval, evaluations, observability, storage and connected Azure services
  7. 7 Foundry Agent Service reached an enterprise GA milestone between October 2025 and Q1 2026

Pros & Cons

Pros

  • Built-in red-teaming is unusual for a cloud platform's evaluation layer and addresses a real gap
  • No separate runtime fee for the agent service - you pay model inference and tool calls, which is cleaner than a per-agent-hour charge
  • OpenAI models priced identically to OpenAI's direct API while adding Azure's enterprise SLA, so the SLA is effectively free
  • Governance and audit trails are first-class, which matches the enterprise buyer this is aimed at
  • Everything stays inside your Azure tenant and Entra ID boundary

Cons

  • Foundry Memory was still Public Preview as of Q1 2026, so SLAs and pricing were not finalised for teams evaluating then
  • Total cost can span inference, orchestration, retrieval, evaluations, observability, storage and connected Azure services, which makes forecasting genuinely hard
  • Tied to the Azure estate, so cross-cloud model comparison is awkward
  • Agent evaluation is less specialised than a dedicated tool like LangWatch
  • Enterprise Agreement terms mean list pricing may not reflect what you actually pay, in either direction

Features

Evaluation, red-teaming and observability alongside inference
Governance and audit trails around AI usage
Foundry Agent Service with no separate runtime fee
OpenAI models at parity pricing with OpenAI direct, plus Azure SLA
Provisioned Throughput Units for guaranteed capacity

Microsoft’s own framing is the useful signal

Microsoft presents Foundry less as a model API and more as a full production lifecycle tool, covering evaluation, red-teaming and observability alongside inference.

That framing tells you who it is for, and it is worth taking at face value. The buyer is a large enterprise that needs governance and audit trails around AI usage as much as it needs the models themselves.

If that is not you, Foundry will feel heavy. If it is you, the alternatives will feel like toys with a compliance gap.

Red-teaming is the differentiated part

Most cloud evaluation offerings score outputs and stop. Foundry includes red-teaming alongside evaluation and observability.

That matters because adversarial testing before deployment catches a different class of problem from measuring quality after it. Scoring tells you the system performs well on cases you thought of; red-teaming tells you what happens when someone actively tries to break it.

Pairing the two is what Giskard does in open source and what CalypsoAI did before F5 acquired it. Very little else in this category attempts both, and no other cloud platform’s evaluation layer does.

For a regulated enterprise that must demonstrate it tested for misuse rather than only for accuracy, having red-teaming inside the same governed platform is a compliance artefact, not just a feature.

The pricing has one quietly good property

No separate runtime fee. You pay for model inference and tool calls, not for the agent service itself.

That compares well with alternatives charging a per-agent or per-hour runtime fee on top of inference.

Combined with OpenAI models priced identically to OpenAI’s direct API, the effect is that Azure’s enterprise SLA costs nothing extra. Same token price as going direct, plus an SLA, governance and Entra ID integration.

For an enterprise that would otherwise procure OpenAI separately - with its own contract, security review and payment relationship - that is straightforwardly good.

Provisioned Throughput Units are available for guaranteed capacity on latency-sensitive workloads.

Why the total is hard to predict

The offsetting problem: the bill is assembled from many services rather than one meter.

Depending on workload architecture, costs can span model inference, orchestration, retrieval, evaluations, observability, storage and connected Azure services.

So a headline token price tells you very little about what you will actually pay. This is the general condition of cloud platform pricing rather than a Foundry-specific failing, but it means any single figure is close to meaningless.

Price a real workload end to end, including the storage and observability services it pulls in.

Preview status, and a note on the whole segment

Foundry Memory was still Public Preview as of Q1 2026, not GA - so SLAs and pricing were not finalised for teams evaluating at that point. Preview on Azure carries no service level commitment.

Worth zooming out on this, because it applies to the category rather than just Microsoft. All three cloud agent offerings - AgentCore, Foundry Agent Service and Vertex AI Agent Engine - reached enterprise GA milestones between October 2025 and Q1 2026.

This entire segment is recently matured rather than long settled. Anything you read about it that is more than a few months old may describe preview behaviour, and the components within each platform reached GA at different times. Confirm current status for the specific pieces you depend on rather than trusting a platform-level GA announcement.

Where a specialist beats it

Foundry is broader and less specialised.

It covers the production lifecycle with governance, audit trails, red-teaming and observability, which dedicated tools generally do not.

What it does not match is depth on agent trajectory analysis. LangWatch’s simulation architecture - Agent Under Test, User Simulator Agent, Judge Agent, running under pytest in CI - is a fundamentally better shape for evaluating multi-turn agents, because it evaluates the path rather than the destination.

The split is clean: governance and tenancy → Foundry. Understanding why your agent misbehaves across a conversation → a specialist.

Comparing the three clouds

Same advice as on the Bedrock page, because it is the single most useful thing to say about cloud evaluation pricing:

Price three real workloads before signing - a routine production call, a long-running agent workflow, and a high-volume reserved-capacity path.

The three platforms use different units, and the cheapest on the marketing page is frequently not cheapest after Azure EA terms, AWS commitments or GCP reserved-throughput planning. Enterprise Agreement terms in particular mean list pricing may not reflect what you pay, in either direction.

Should you use it?

Use Azure AI Foundry Evaluation if you are standardised on Azure, governance and audit trails matter as much as model access, and built-in red-teaming is valuable to you.

Don’t use it if you evaluate across clouds, or agent trajectory analysis is your primary need.

Bottom line: the most enterprise-shaped of the cloud evaluation offerings, with red-teaming as a genuine differentiator and an SLA that costs nothing over going direct to OpenAI. Confirm GA status on the components you need, and price a real workload rather than a token rate.


Positioning, pricing structure and preview status verified against Microsoft documentation and third-party platform comparisons on 3 August 2026. Preview status changes frequently and should be reconfirmed. Enterprise Agreement terms mean published pricing may not reflect negotiated rates. This is a researched directory entry - we have not yet instrumented this platform with our reference application.

Pricing Plans

Pay-as-you-go

Model inference and tool calls

  • No separate runtime fee for the agent service
  • OpenAI models priced identically to OpenAI's direct API
  • Adds Azure's enterprise SLA
Most Popular

Provisioned Throughput Units

Reserved capacity

  • Guaranteed capacity for latency-sensitive workloads

Enterprise

Custom (EA terms)

  • Enterprise Agreement pricing
  • Costs may span inference, orchestration, retrieval, evaluations, observability and storage

SDKs & Frameworks

Python SDK C# and .NET JavaScript REST API

Deployment

Azure AI Foundry Azure OpenAI models Wider Azure estate and Entra ID Foundry Agent Service

Eval Methods

Automated evaluation Red-teaming Observability Human review workflows

Positioning

Production lifecycle tool rather than a model API

Our Verdict

Azure AI Foundry is the most convincingly enterprise-shaped of the three cloud evaluation offerings, and Microsoft's own framing is the clearest signal - it presents Foundry less as a model API than as a full production lifecycle tool covering evaluation, red-teaming and observability alongside inference. That matches its buyer, which is a large organisation needing governance and audit trails around AI usage as much as it needs the models. Built-in red-teaming is genuinely unusual for a cloud platform's evaluation layer, and the pricing has one quietly good property - there is no separate runtime fee for the agent service, so you pay for model inference and tool calls rather than a per-agent-hour charge, and OpenAI models cost the same as going to OpenAI directly while gaining Azure's SLA. Two cautions. Foundry Memory was still Public Preview as of Q1 2026, so SLAs and pricing were unsettled for anyone evaluating then. And total cost can span inference, orchestration, retrieval, evaluations, observability, storage and connected Azure services, which makes a single number close to meaningless.

Similar Tools

Frequently Asked Questions

What does no separate runtime fee actually mean?

That you are not charged for the agent service itself, only for the model inference and tool calls it makes. That is cleaner than it sounds and compares well with alternatives that charge a per-agent or per-hour runtime fee on top of inference. Combined with OpenAI models being priced identically to OpenAI's direct API, the practical effect is that Azure's enterprise SLA comes at no premium over going direct - you pay the same token price and get an SLA, governance and Entra ID integration. For an enterprise that would otherwise need to procure OpenAI separately, that is a straightforwardly good deal.

Why is cost forecasting hard here?

Because the bill is assembled from many services rather than one meter. Depending on workload architecture, costs can span model inference, orchestration, retrieval, evaluations, observability, storage and connected Azure services - so a headline token price tells you very little about what you will actually pay. This is the general condition of cloud platform pricing rather than a criticism of Foundry specifically, but it means any single figure is close to meaningless. Price a real workload end to end, including the storage and observability services it pulls in, rather than modelling from token rates.

How significant is the preview status?

It mattered for anyone evaluating in early 2026 and should be rechecked now. Foundry Memory was still Public Preview as of Q1 2026, not GA, which means SLAs and pricing were not finalised. Preview on Azure carries no service level commitment and can change without the guarantees enterprises normally require. More broadly, all three cloud agent offerings - AgentCore, Foundry Agent Service and Vertex AI Agent Engine - reached enterprise GA milestones between October 2025 and Q1 2026, so this whole segment is recently matured rather than long settled. Confirm current GA status for the specific components you depend on.

Is the red-teaming worth having?

It is unusual enough to be a differentiator. Most cloud evaluation offerings score outputs and stop; Microsoft includes red-teaming alongside evaluation and observability. Adversarial testing before deployment catches a different class of problem from measuring quality after it, and pairing the two in one platform is what Giskard does in open source and CalypsoAI did before F5 acquired it. For a regulated enterprise that has to demonstrate it tested for misuse rather than only for accuracy, having red-teaming inside the same governed platform as everything else is a meaningful compliance artefact rather than just a feature.

How does it compare with a dedicated agent eval tool?

Foundry is broader and less specialised. It covers the production lifecycle with governance, audit trails, red-teaming and observability, which a dedicated tool generally does not. What it does not match is depth on agent trajectory analysis - LangWatch's simulation architecture, with an Agent Under Test, a User Simulator and a Judge running under pytest, is a fundamentally better shape for evaluating multi-turn agents. If your priority is governance and staying inside your Azure tenant, Foundry. If your priority is understanding why your agent behaves badly across a conversation, a specialist tool.

Should I compare it directly with Bedrock and Vertex on price?

Not from rate cards, because the three use different units and the comparison is harder than it should be. Bedrock is on-demand per-token with provisioned throughput at an hourly rate. Foundry prices OpenAI models at parity with OpenAI direct and adds no runtime fee. Vertex is per-token plus GCP compute. A practical recommendation worth following is to price three real workloads before signing - a routine production call, a long-running agent workflow, and a high-volume reserved-capacity path - since the cheapest on the marketing page is frequently not cheapest after Azure EA terms, AWS commitments or GCP reserved-throughput planning.