Learn LLM Evaluation
A free, practical course on evaluating LLM applications - from your first eval set to production tracing, LLM-as-judge scoring and regression testing.
9 chapters
Why LLM Evaluation Matters
Why shipping an LLM feature without evaluation is a gamble, and what evals actually buy you.
11 min read
Offline vs Online Evaluation
The two modes of LLM evaluation, what each catches, and why you need both in a real system.
12 min read
Building Evaluation Datasets
How to build the datasets your LLM evals actually depend on, from first cases to a living golden set.
13 min read
Using LLM-as-a-Judge
How to use a strong language model to grade the output of another model, and how to keep that judge honest.
12 min read
Core Evaluation Metrics
The handful of metrics that actually tell you whether an LLM application is working, and how each one is computed.
13 min read
Evaluating RAG Systems
How to evaluate retrieval-augmented generation by grading the retriever and the generator as separate stages.
13 min read
Evaluating AI Agents
How to evaluate multi-step AI agents when the output is a trajectory of tool calls and decisions, not a single answer.
13 min read
Observability and Tracing in Production
How tracing turns opaque LLM and agent behavior in production into structured spans you can search, score, and debug.
13 min read
Choosing Your Eval Stack
A decision framework for picking LLM eval and observability tools, plus how everything in this course fits into one workflow.
12 min read
Free Newsletter
Get the LLM Evals Newsletter
Platform comparisons, pricing changes and eval technique deep-dives. No spam.