how-to

LLM Evaluation Techniques: A Practitioner's Guide

A practitioner's guide to LLM evaluation techniques: benchmarks, overlap and semantic metrics, verifiers, LLM judges, RAG and agent evals — plus dataset sizing, judge reliability and per-run cost.

Published:

LLM Evaluation Techniques That Produce Numbers You Can Act On

LLM evaluation techniques divide into five families: deterministic reference matching, statistical overlap (BLEU, ROUGE, METEOR), model-based scoring that is not a judge (BERTScore, NLI entailment), LLM-as-a-judge, and executable verifiers. Two questions settle which family you pick. Do you have reference answers? Can a program check the result rather than the wording? Everything else here is detail on those two choices, plus the four things the rest of the field skips: how many samples you need before a score means anything, how to prove your judge agrees with humans, what a judge suite costs per run, and which open-source framework implements which technique.

Pick your technique in 60 seconds

The table maps your situation to the LLM evaluation techniques that fit it.

Your situationTechnique family to start withFirst metrics to implement
Reference answers exist, one right answerDeterministic matching plus classification metricsExact match after normalisation, accuracy, per-class F1
References exist, many valid phrasingsSemantic similarity plus a reference-guided judgeBERTScore or cosine similarity as a gate, judge for the pass/fail call
No references, open-ended outputReference-free LLM-as-a-judge with a written rubricBinary criterion scores, faithfulness against provided context
Code, SQL or mathVerifiersUnit-test pass rate, boxed-answer extraction, result-set diff
A retrieval or recommendation stepRanking metricsPrecision@K, Recall@K, nDCG@K, Hit Rate@K

Every one of these LLM evaluation techniques needs the same three parts. A dataset. A scorer. An aggregation rule with a threshold. Teams build the first two and skip the third, which is why so many suites produce a number nobody can act on. An aggregation rule says proportion of samples passing, names the slices it covers, and states what value fails the build.

A score is worth what the dataset under it and the judge above it are worth. Both get their own sections below. Both are where eval programmes quietly fail.

Decision tree for choosing among the five families of LLM evaluation techniques

Visual spec: Decision-tree SVG spanning all five technique families. Branches: do reference answers exist? → is there exactly one correct answer? → is the output executable or programmatically checkable? → is the task open-ended? Terminal nodes name the family and its starting metrics, using the exact metric names from this article. Ship with a text-equivalent nested list beneath for screen readers and crawlers.

What ‘LLM evaluation’ means, and the two distinctions that change everything

Two axes explain most of the confusion in this field. They are independent of each other, and both cut across every one of the LLM evaluation techniques below.

Model versus system. Model evaluation scores the model alone: no retriever, no tools, no guardrails. It catches capability regressions after fine-tuning. It answers “is this checkpoint worse at arithmetic than the last one”. It cannot see a chunking bug, a stale index, or a prompt template that drops the user’s constraint. System evaluation scores the model inside its harness, so it sees all of those. It cannot tell you which component broke unless you instrument the components separately. IBM and LaunchDarkly both draw this line. The operational version: model evals gate a model swap, system evals gate a release.

Offline versus online. Offline evaluation runs a fixed candidate against a curated dataset before exposure, usually with ground truth available. Online evaluation scores live responses continuously. Ground truth almost never exists there, so you substitute reference-free scorers and user signals. All four combinations are real: an offline model eval (a benchmark run), an offline system eval (your CI suite), an online model eval (a canary comparing two model versions on sampled traffic), an online system eval (production quality monitoring).

Evaluation is not benchmarking. A benchmark compares models on someone else’s task with someone else’s data. An eval tests your system on your task with your data. A model can lead MMLU and still fail your support flow, because the failure was that your retriever returned the wrong chunk.

The reason there are so many LLM evaluation techniques rather than one is stochastic generation. Ask for an explanation of quantum mechanics for an eight-year-old. Three excellent answers will share almost no surface form. Classification-style thinking assumes one correct label. Generation breaks that assumption, and every family below is a different attempt to score correctness without assuming a single string.

The five families of LLM evaluation techniques

Family 1: deterministic and reference matching. Exact match, fuzzy match, word or item containment, JSON key-value match, regex assertions, JSON-schema validation. Near-zero cost, fully interpretable, brittle to phrasing. The right answer for structured output.

Family 2: statistical overlap. BLEU, ROUGE, ROUGE-L, METEOR, Levenshtein distance, and perplexity as a model-level descriptor. Cheap, deterministic, semantics-blind.

Family 3: model-based scoring that is not an LLM judge. BERTScore, MoverScore, COMET, Sentence Mover Similarity, BLEURT, NLI and entailment scoring, and fine-tuned classifiers for toxicity or bias. One small model call per item, no prompt to maintain, weak interpretability.

Family 4: LLM-as-a-judge and its prompting frameworks. Direct scoring, pairwise comparison, reference-guided grading, G-Eval, GPTScore, DAG-style decomposition, juries of models, fine-tuned judges such as Prometheus.

Family 5: verifiers and executable checks. Unit-test pass rate, boxed-answer extraction, tool-call trace assertions, SQL result-set diffs, SWE-bench-style patch-and-run. Zero judge variance where the task admits it.

FamilyNeeds ground truth?Compute cost per 1k samplesSemantic sensitivityWhat a failing score tells you
Deterministic matchingYes, exactNegligible: string operations onlyNone: “four” ≠ “4”Exactly which field mismatched, so debugging is immediate
Statistical overlapYesNegligible: n-gram countingLow: rewards shared surface formAlmost nothing, only that wording moved
Model-based scoringUsually, except NLI against contextLow: one embedding or classifier call per itemModerate to highA number without a reason, so triage is manual
LLM-as-a-judgeNo, rubric replaces itHigh: one or more generative calls per criterion per itemHighA written reason you can audit, if you force a reason field
VerifiersYes, as a check not a stringLow compute, engineering-heavy to buildNot applicable, checks the resultThe failing test name, the most actionable output available

No family wins. Production suites mix families, which is why the five-metric discipline later in this guide matters.

Competitor taxonomies disagree in vocabulary, not substance. Evidently splits LLM evaluation techniques into reference-based and reference-free. Microsoft Learn splits the same way and adds a timeline of metric development. Sebastian Raschka splits on benchmark-based versus judgment-based. Map them. Reference-based covers families 1 to 3 plus reference-guided judging. Reference-free covers rubric judging, SelfCheckGPT and text descriptors. Benchmark-based is family 1 or 5 applied to public datasets. Judgment-based is family 4. Same object, three camera angles.

Matrix comparing the five families of LLM evaluation techniques on ground truth, cost, sensitivity and debuggability

Visual spec: Technique family matrix as a graphic, five rows by five columns, mirroring the table above, with each qualitative rating carrying a one-clause justification inside the cell so no rating reads as bare opinion.

Deterministic and reference-matching techniques

These are the most under-taught LLM evaluation techniques and often the right answer. Exact match after normalisation. Fuzzy match with a similarity floor. Word or item containment for list outputs. JSON key-value match for extraction. Regex assertions for format. JSON-schema validation for structure. Unit-test pass rate for code.

Where each is genuinely correct: function-calling arguments, where the argument either equals the expected value or does not; extraction into a fixed schema, where a missing key is a bug whatever the prose quality; classification labels; SQL that must parse before anything else matters; citation IDs that must exist in the retrieved set.

The classic failure is a perfect answer scoring zero because the model wrote “four” and the reference said “4”. Fix it with a normalisation checklist applied identically to candidate and reference: case folding, whitespace collapse, punctuation stripping, numeral form (spelled versus digit), unit form (kg versus kilograms), and article removal for short answers. Write the normaliser once, test it, and version it with the dataset. An un-versioned normaliser shifts scores silently.

Text-stat descriptors belong here too: output length, sentiment, reading level, banned phrases, refusal detection. Use them as pass/fail gates, never as quality scores. “No response exceeds 400 tokens” is enforceable. An average reading level is trivia.

Overlap metrics, and why BLEU, ROUGE, METEOR and Levenshtein mislead

Of all the LLM evaluation techniques, these are the cheapest and the most often misread.

BLEU is modified n-gram precision between candidate and reference, averaged geometrically across n-gram orders and multiplied by a brevity penalty so short outputs cannot game precision. It was built for machine translation with multiple references. Two variants matter and only the technical references cover them. SacreBLEU standardises tokenisation so two papers’ numbers are comparable. NIST reweights n-grams by information value, giving rare n-grams more credit than function words. A BLEU score quoted without its tokenisation scheme is not a number. It is a rumour.

ROUGE inverts the emphasis. ROUGE-N is n-gram recall against the reference, which suits summarisation, where coverage matters more than precision. ROUGE-L uses the longest common subsequence, so it rewards in-order overlap without requiring contiguity. Report both precision and recall forms, not F-measure alone. If ROUGE-L recall rose and precision fell, the model got more verbose, and the F-score hides that.

METEOR combines unigram precision and recall with stemming and WordNet synonym matching, then penalises word-order scrambling. That synonym step is why METEOR tracks human judgement better than BLEU on single-reference tasks.

Levenshtein distance and its derived ratios (simple, partial, token-sort, token-set) count edit operations. Edit distance earns its place in OCR correction, spelling normalisation, and character-level formatting tasks where the target genuinely is a string. Elsewhere it is noise.

The field openly disagrees here. Confident AI calls statistical scorers “non-essential”. The NVIDIA developer blog and Microsoft Learn both present F1 and ROUGE as first-class metrics. The resolution: overlap metrics are regression tripwires, not quality measures. They are deterministic, they cost nothing, and they run on every commit. So they are excellent at answering “did today’s prompt change rewrite every output”. They should never be the number a stakeholder sees, because a stakeholder reads 0.41 ROUGE-L as a grade.

The trap shows best in the Apollo 11 example from the arXiv practitioner guide. A candidate that names the wrong astronaut but reuses the reference’s phrasing can outscore a factually correct paraphrase that shares few n-grams. One sentence of overlap arithmetic beats three paragraphs of caveats. A metric that ranks a false statement above a true one is not measuring quality.

Perplexity deserves its own note. It measures how surprised a model is by held-out text. That makes it useful for tracking pretraining or domain adaptation, and useless for scoring an answer’s helpfulness. Low perplexity on fluent nonsense is common.

Semantic similarity and entailment techniques

This middle tier of LLM evaluation techniques costs far less than a judge and breaks far less than n-grams.

Embedding-based scoring. Cosine similarity between candidate and reference embeddings is the blunt version, and it depends entirely on which embedding model you chose. BERTScore improves on it. It matches tokens greedily between candidate and reference using contextual embeddings, then reports precision, recall and F1, so you can see whether the candidate omitted reference content or added its own. MoverScore treats the comparison as an Earth Mover’s Distance problem over embedded tokens. Sentence Mover Similarity does the same at sentence granularity. COMET is translation-specific and takes source, reference and hypothesis together, which is why it beats BLEU on human correlation for MT. BLEURT is a learned metric fine-tuned on human ratings.

NLI and entailment scoring is the version that matters for hallucination detection. Classify each output sentence against the retrieved context as entailed, contradicted or neutral. SummaC aggregates sentence-level NLI scores for summary consistency. FactCC trains a classifier on synthetically corrupted summaries. DAE (Dependency Arc Entailment) checks entailment at the level of individual dependency arcs, which localises the error. All three degrade on long inputs. NLI models were trained on sentence pairs, and a very long context is not a sentence pair. Chunk the context and score per chunk.

Factuality via question generation. QAFactEval, QuestEval and SRLScore generate questions from the candidate and check whether the context answers them the same way. Reference-free summarisation quality has its own cluster. SUPERT scores against a pseudo-reference built from the source. BLANC measures how much a summary helps a masked-language model reconstruct the source. ROUGE-C compares the summary against the source document instead of a human summary.

Microsoft Learn’s metric reference states plainly that contextualised-embedding metrics, BERTScore and MoverScore and Sentence Mover Similarity among them, suffer from poor correlation with human evaluators, limited interpretability, inherent bias, and poor adaptability across tasks. The practical consequence: a single similarity number tells you nothing about what broke. You cannot open a ticket against it.

Which is why these belong in a gate, not a report. Flag every item below a 0.85 similarity floor for judge review and let the rest pass untouched. That turns a weak metric into a cheap triage filter, and it removes most of your judge spend on regression suites where most outputs are unchanged.

Classification and ranking metrics, the half of eval that is just ML

A large share of an LLM application is ordinary supervised learning in a generative costume, and the corresponding LLM evaluation techniques are the ones ML engineers already know.

Classification hides in intent detection, agent routing, ticket triage, content moderation, and every guardrail decision. These have labels. Use accuracy, precision, recall, F1, and always a per-class breakdown.

Accuracy is the wrong default. For an unsafe-content gate, recall answers “did we catch everything harmful” and precision answers “how much harmless traffic did we block”. Which one you optimise is a product decision with a cost on each side, and the choice belongs in writing next to the threshold. Take a moderation gate at 94 percent accuracy. If the misses concentrate in one harm category, that gate is catastrophic, and only a per-class table exposes it.

For the retrieval step, ranking metrics:

MetricRank-aware?What it answers
Precision@KNoOf the K chunks retrieved, what fraction are relevant
Recall@KNoOf all relevant chunks in the corpus, what fraction appear in the top K
Hit Rate@KNoIn what fraction of queries does at least one relevant chunk appear in the top K
MRR@KYesHow high up the first relevant chunk sits, averaged as a reciprocal rank
nDCG@KYesGraded relevance discounted by position, so a relevant chunk at rank 1 beats the same chunk at rank 8

Rank-agnostic metrics treat the top K as a set. Rank-aware metrics penalise burying the right chunk at position 9. Choose K by how many chunks you actually place in the prompt. If the generator sees five chunks, Recall@5 predicts whether the answer is possible, and nDCG@5 predicts whether a reranker is earning its latency. Evaluating at K=20 when you pass 5 flatters the retriever.

Per-class and per-segment slicing is the highest-yield habit in this guide. Aggregate scores hide the one intent class, one language, or one customer tier that regressed. Tag every eval row at creation time so slicing is a groupby, not a re-labelling project.

LLM-as-a-judge, the technique everyone uses and few implement well

LLM-as-a-judge is the most used of the LLM evaluation techniques, because it is the only family that scores open-ended output without references. It is also the family where implementation quality varies most.

Three scoring modes.

ModePrompt shapeBest forMain weakness
Direct scoringOne output, one rubric, emit a score or pass/failContinuous monitoring, absolute thresholds, online scoringScale drift between runs and models
Pairwise (H2H)Two outputs, pick a winner, with an explicit tie optionPrompt and model selectionPosition bias, no absolute level
Reference-guidedOutput plus a golden answer plus a rubricRegression suitesNeeds a maintained golden set

Omit the tie option in pairwise comparison and you force a coin flip on identical answers, which inflates apparent differences. Include it.

Low-cardinality rubrics beat Likert scales. The G-Eval paper’s own observation is the concrete reason. On a 1-to-5 scale the token probability mass concentrates on “3”, so the score distribution is narrow and unstable whatever the output quality. Two responses to that finding: define binary criteria, or weight the emitted digits by their output-token probabilities to get a continuous score. Binary is available to everyone. Probability weighting needs logprob access your provider may not expose.

G-Eval mechanics, concretely. Give the model a criterion definition. Have it generate chain-of-thought evaluation steps for that criterion. Fill a form-style prompt with the task, the steps and the output under test. Emit a 1-to-5 score. Optionally normalise by output-token probabilities. Confident AI reports that G-Eval reaches higher Spearman and Kendall-Tau correlation with human judgement than traditional non-LLM scorers, and notes that the original experiments used only GPT-3.5 and GPT-4. Read the correlation figures as a property of those two judges on those datasets, not a general guarantee.

Decomposition. DAG-style judging turns one vague criterion into a decision tree of closed questions, each answerable yes or no, so the verdict has an audit trail. FineSurE checks summaries fact by fact. QAG generates close-ended questions from the claims in an output, answers each against the context, then scores the proportion answered affirmatively. Decomposition is what converts “score coherence 1 to 5” into something a human reviewer can dispute at a specific step. Three related prompting patterns are worth naming. Reason-then-Score (RTS) has the model write its reasoning before the score. Multiple Choice Question Scoring (MCQ) converts grading into option selection. GEMBA is the translation-quality judge prompt family.

Self-consistency. SelfCheckGPT samples N responses to the same prompt at non-zero temperature, with nucleus sampling to make the variation meaningful, and measures mutual agreement between them. Disagreement between samples is evidence of fabrication, and it needs no reference at all. Juries of heterogeneous models are the cheaper cousin. Three mid-tier judges from different families (one OpenAI model, one Anthropic model, one open-weights model such as Llama or Qwen) disagreeing is more informative than one frontier judge agreeing with itself.

Fine-tuned judges. Prometheus is an open judge model trained for rubric-based grading. Reward models such as the Nemotron-4 reward and instruct variants score preference directly. The tradeoff is hosting cost and latency against a frontier judge’s per-token price. It flips in favour of fine-tuned judges at high call volume, which the cost section turns into arithmetic.

A naming clean-up no competitor makes clearly: G-Eval, GPTScore and QAG are prompting techniques, not metrics. Evidently makes this point, and it prevents real confusion, because “we use G-Eval” says nothing about what criterion was scored. DeepEval’s metric names such as Turn Faithfulness and Plan Adherence are one vendor’s vocabulary, not industry standards. Use them if you use DeepEval, and translate when you read anyone else’s docs.

Then the finding that should temper framework enthusiasm. Evidently’s guide cites work showing that simply asking an LLM to think step by step can outperform G-Eval (Chiang et al., 2023, cited author-year only on that page, with no URL or title, so trace the paper before quoting it). A/B your judge prompt against plain chain-of-thought on a human-labelled sample before adopting any named framework. The framework may win. It may also cost you three extra calls per item for nothing.

For prompt patterns and bias mitigations in depth, see our judge prompt patterns guide.

Verifiers, the technique with no judge and no ambiguity

A verifier is a deterministic program that checks an answer’s result rather than its text. Extract the boxed final answer and compare it to the key. Run generated code against unit tests. Execute the SQL and diff the result set against the expected rows. Apply the patch and run the repository’s test suite, which is the SWE-bench pattern. Among all LLM evaluation techniques, verifiers have the lowest variance, because there is no scorer to calibrate.

The mechanics break at extraction, which is the part Raschka’s walkthrough makes concrete in code. The model wraps its answer in prose, emits two candidate letters, restates the options before choosing, or answers correctly in a language your regex does not cover. Three extraction rules, three different scores:

  • First match credits the model’s initial instinct and punishes self-correction.
  • Last match credits self-correction and can be gamed by a model that enumerates options.
  • Delimiter-bounded extraction (\boxed{}, <answer> tags, a required final line) is the only fair rule. It requires the prompt to demand the delimiter and the score to record parse failures separately from wrong answers.

Report parse failures as their own bucket. A double-digit parse-failure rate looks like a capability gap and is actually a prompt-formatting gap.

Multiple choice has the same problem one level deeper. Scoring by log-probability over the option tokens and scoring by matching the generated letter give different numbers for the same model on the same benchmark. The first measures preference among the given options. The second measures instruction following plus preference. A reported benchmark figure without its scoring variant and harness version is not reproducible.

Verifiers extend well beyond code and math: tool-call trace assertions (did the agent call refund with this order ID), JSON-schema conformance, citation-existence checks against the retrieved set, and idempotency checks that confirm an agent did not send the same email twice.

The design goal follows from the arithmetic. Every criterion you move from a judge to a verifier removes a source of variance and a recurring per-run cost. Keep a running list of judge criteria. Each quarter, ask which of them a program could now check.

Academic benchmarks: what they measure and what they hide

Benchmarks are the LLM evaluation techniques with the best branding and the narrowest use. Group them by capability:

CapabilityBenchmarksHow scored
Core knowledgeMMLU, MMLU-Pro, HellaSwag, WinoGrande, GLUE, SuperGLUEAutomatic against a key
Reasoning and QAARC Challenge, GPQA, TruthfulQA, TriviaQA, BIG-Bench Hard, SQuADAutomatic, with TruthfulQA partly judge-scored
MathGSM8K, MGSMAutomatic via answer extraction
CodeHumanEval, CodeXGLUE, SWE-benchVerifier: tests executed
Instruction followingIFEval, MT-Bench, MT-Bench-101IFEval heuristic-verifiable; MT-Bench LLM-judged
MultilingualXNLI, MGSMAutomatic against a key
Long contextLongGenBench, ZeroSCROLLSAutomatic, task-dependent
Multimodal and domainMMMU, FinanceBenchAutomatic plus human or judge review
Reward modelsRewardBenchAutomatic against preference pairs
Open-ended preferenceChatbot Arena / LM Arena, AlpacaEval 2.0Human pairwise votes; AlpacaEval 2.0 is LLM-judged with length control

Sizes worth knowing. MMLU spans 57 subjects and roughly 16,000 multiple-choice questions, and a uniform random guesser scores 25 percent on four-option items, which is the floor any reported MMLU number must clear (Raschka). WinoGrande contains 44,000 commonsense reasoning problems, IFEval uses 500 heuristically verifiable prompts, and TruthfulQA spans 38 topics (NVIDIA). GSM8K measures multi-step mathematical reasoning over grade-school word problems. Aggregators such as Stanford’s HELM and the Hugging Face Open LLM Leaderboard republish these runs, and they inherit every caveat below.

How a benchmark is scored determines what its number can be compared against. Automatic-against-a-key scores are comparable across labs if the harness and prompt format match. Human-preference scores from Chatbot Arena are comparable only within the same arena population and period. LLM-judged scores such as AlpacaEval 2.0 and MT-Bench inherit every bias in the judge section below, which is why AlpacaEval 2.0 added length control.

Four documented limitations, each with a mechanism:

  1. Data leakage and contamination. Benchmark items appear in pretraining corpora scraped from the web, so the model recalls rather than reasons. The mechanism is simple and the effect is unbounded.
  2. Overfitting to the benchmark. Once a benchmark drives publication, training data selection drifts toward it. The score rises without the underlying capability moving.
  3. Language and cultural bias. Items written in English and rooted in Western context measure English-language, Western-context performance. Scores for other languages typically trail, so an English-only benchmark table is not evidence about your Spanish-speaking users.
  4. Null models. Trivial constant responses score respectably on some benchmarks, which means the benchmark partly measures format compliance. Run a null baseline on any benchmark you adopt internally, and subtract what a fixed string earns.

Saturation is the practical problem. NVIDIA’s own post concedes that academic benchmarks “become saturated quickly”, and it carries no per-benchmark dates or version numbers. The operational consequence: a benchmark table older than about 12 months is a historical document. Before citing any leaderboard, check the harness version, the prompt format, the scoring variant, and the date.

Use benchmarks for two jobs. As a shortlist filter when choosing which two or three models to evaluate on your own data. As a knowledge-regression check after fine-tuning, to confirm you did not trade general capability for a narrow gain. Never as evidence that your system works.

Evaluating RAG systems component by component

Most people searching for LLM evaluation techniques are evaluating a RAG pipeline, often built on LangChain or LlamaIndex, and the highest-value move is to score the retriever and the generator separately.

Retriever criteria. Contextual precision asks whether relevant chunks rank above irrelevant ones. Contextual recall asks whether the retrieved context contains what the correct answer needs. Contextual relevancy asks what proportion of retrieved sentences are actually relevant, which catches over-retrieval inflating your token bill.

Generator criteria. Faithfulness (groundedness) asks whether every claim in the output is entailed by the retrieved context. Answer relevancy asks whether the output addresses the question without padding.

Once you have those five, diagnosis becomes a lookup rather than an argument:

Observed patternLikely broken componentNext test to run
Low faithfulness, good contextual recallGenerator or prompt: the model is ignoring or embellishing the contextRe-run with a stricter grounding instruction and a citation requirement; if faithfulness recovers, it was the prompt
Good faithfulness, wrong answerRetrieval recall: the context never contained the answerRecall@K against a labelled query set; raise K or fix chunking
Good recall, poor contextual precisionRanking: relevant chunks exist but sit below noisenDCG@K before and after the reranker; if the delta is small, the reranker is not earning its latency
Good precision and recall, low answer relevancyPrompt or generator verbosityLength-controlled judge pass, tighten the output spec
High scores offline, complaints in productionDataset representativenessCorrelate offline scores against in-product satisfaction on the same period
Faithful and relevant, but staleRetrieval freshnessAge distribution of retrieved documents versus document update times

RAG failure diagnosis grid linking each symptom to the LLM evaluation techniques that confirm it

Visual spec: RAG failure diagnosis grid rendered as a graphic from the table above, with the metric definitions footnoted to their sources and the mapping logic stated in prose beneath, so a reader can audit each row rather than trust it.

Chunking, embedding and reranking are separately evaluable stages. Score Precision@K, Recall@K and nDCG@K for the embedding model, then run the identical metrics after the reranker. The reranker’s contribution becomes a measured delta rather than an assumption. Add temporal freshness of retrieved documents as its own descriptor. A pipeline that retrieves a superseded policy page is faithful and operationally wrong.

You can build retrieval ground truth without hand-labelling the whole corpus. Generate relevance labels for query-document pairs with an LLM, and generate question-answer pairs from your own documents so the source chunk is known by construction. Both make your eval dependent on the labeller’s biases, and both need a human-audited sample. Label 50 pairs by hand, compare against the LLM labels, and compute agreement before you trust the other 5,000. If that agreement is weak, your retrieval metrics are measuring the labeller.

One deterministic guardrail catches a large class of hallucination for free: citation-existence verification. Require the generator to cite chunk IDs, then assert that every cited ID appears in the retrieved set. A fabricated citation becomes a hard failure found by a set-membership check, not a judge. Details in our RAG retriever diagnosis guide.

Evaluating agents and multi-turn conversations

Agent work stretches these LLM evaluation techniques furthest, and published guidance here is thin. The structural idea is span-level versus trace-level scoring. A span is one step: a tool call, a retrieval, a single model turn. A trace is the whole run. Score both. Trace-level task completion tells you that the agent failed. Span-level assertions tell you where.

Criteria worth implementing, with the scoring family each belongs to:

CriterionHow to scoreFamily
Task completionJudge over the full trace against a completion definitionLLM-as-a-judge
Tool correctnessExact match against an expected tool list per taskDeterministic verifier
Argument correctnessKey-value match on the tool arguments, with normalisationDeterministic verifier
Plan qualityJudge the proposed plan before execution against a rubricLLM-as-a-judge
Plan adherenceCompare executed steps against the stated planDeterministic on step names, judge on intent
Step efficiencySteps, tokens and cost per successful taskArithmetic from the trace

Tool correctness and argument correctness are the two criteria most often handed to a judge and most cheaply handled by a verifier. Move them.

Multi-turn conversation adds two things single-turn evaluation cannot see. State carried across turns, where turn 4 depends on a constraint the user gave in turn 1. And error compounding: the arXiv practitioner guide notes that errors compound with each agent execution, which is why per-turn scoring beats whole-conversation scoring for debugging. One overall score for a conversation gives you nothing. A conversation where turn 3 dropped the user’s stated budget gives you a fix.

Compute RAG criteria per assistant turn, conditioned on prior turns. Faithfulness of turn n given the context retrieved for turn n. Relevancy of turn n given the full prior exchange. Aggregate as the proportion of passing turns, never as a mean. A mean across 20 turns hides the one turn that quoted the wrong refund window.

Cost and latency belong in the eval output, not in an ops dashboard. An agent completing 95 percent of tasks in 40 tool calls and 90 seconds may be unshippable at your price point, and you want that visible in the same table as the quality score. Report tokens per successful task, because tokens per task rewards an agent that fails cheaply.

Be honest about the ceiling. Multi-agent coherence and negotiation quality have no accepted automated measure. Databricks names this as a future direction in its agent evaluation post rather than claiming a solution. That is the correct posture. If you run a multi-agent system, your options today are human trace inspection on a sampled basis and outcome-level task completion. Our agent trace evaluation guide covers span instrumentation in detail.

The eval dataset: sizing, sourcing, and decontamination

No LLM evaluation techniques survive a bad dataset. Every technique above inherits the quality of the rows beneath it, and this is the part of the SERP where only the arXiv paper does real work.

Three sourcing routes.

RouteStrengthHonest weakness
Public benchmarksImmediate, comparable, freePossibly contaminated, rarely use-case specific
Human-annotated golden setHighest fidelity, defensibleSlow, expensive, needs annotator-agreement management
Synthetic silver setScales to thousands of rows fastInherits the generator’s biases, needs a human audit

The recommended path is a small golden seed, synthetic expansion around it, and a promotion pipeline that moves audited silver rows into gold over time. The golden set stays small and trusted. The silver set stays large and provisional.

Synthetic generation techniques worth naming: distillation from a frontier model, Evol-Instruct-style iterative complexity increase where each round rewrites a prompt to be harder, Constitutional-AI-style self-critique to produce both a flawed and a corrected response, persona prompting for voice diversity, and temperature and top-p variation to avoid 500 rows that read like one row.

Five properties to design for, each with a measurable test.

  1. Defined scope. Test: every row maps to a named capability tag, and every capability in your product spec has rows.
  2. Demonstrative of production usage. Test: correlation between offline scores and in-product satisfaction over the same period. High offline scores beside low satisfaction proves the dataset is unrepresentative, and that single check is worth more than adding 500 rows.
  3. Diverse. Test: cluster embeddings of your inputs, or run locality-sensitive hashing for near-duplicate detection, and report cluster count and duplicate rate.
  4. Decontaminated. Test: the contamination procedures below.
  5. Dynamic. Test: rows added per month from production failures, and the date of the oldest row.

The sample-size arithmetic the rest of the SERP skips. For a proportion metric such as pass rate, the required sample size at 95 percent confidence is:

n = z² × p × (1 − p) / e²

with z = 1.96, p the expected metric level, and e the margin of error you will tolerate. For a metric expected around 80 percent with a 5 percent margin: 1.96² × 0.8 × 0.2 / 0.05² = 3.8416 × 0.16 / 0.0025 = 246 samples. That is the figure the arXiv guide derives. Here is the rest of the table, computed from the same formula, so you can size your own set:

Expected metric±10% margin±5% margin±3% margin±2% margin
0.60933691,0252,305
0.80622466831,537
0.951973203457

Two things fall out. Tightening the margin from 5 percent to 2 percent multiplies the requirement by 6.25, because n scales with 1/e². And high-scoring metrics need fewer samples, because p(1−p) shrinks near the extremes. A safety gate expected to pass 99 percent of the time can be validated on far fewer rows than a helpfulness score hovering near 60 percent.

The corollary is the sentence to put in front of your team. A 40-row eval set cannot distinguish a 3-point regression from noise. At 40 rows and a pass rate near four in five, the margin of error is roughly ±12 points. Every “the new prompt improved us from 82 to 85” claim made on 40 rows is a claim about sampling error.

Required eval-set size against margin of error for LLM evaluation techniques that report a pass rate

Visual spec: Line chart of required eval-set size against margin of error, one curve each for expected metric values 0.6, 0.8 and 0.95, margins from 10% down to 2% at 95% confidence, with the 246-sample point annotated. Publish the formula and parameters beside the chart so every point is reproducible from arxiv.org/html/2506.13023v1.

Contamination detection, procedurally. Continuation testing: feed the first half of an eval item and see whether the model completes the rest verbatim, including formatting quirks a reasoner would not reproduce. Logprob and perplexity inspection: suspiciously low perplexity on a benchmark item, relative to comparable unseen text, suggests memorisation. Exact, substring and hash comparison against the training corpus where you have access, which usually means only models you trained.

The hard rule: the only reliable defence is a private eval set collected after the model’s training cutoff and never published. Keep a held-out slice you never put in a blog post, a prompt, or a support ticket.

Metadata that makes a dataset debuggable. Capability and segment tags for slicing. Language and locale. Difficulty. Source (production log, synthetic, hand-written). Grounding context for the judge, which must never come from the system under test, because then you are asking the model to grade itself against its own claim. Expected-information fields listing the facts a correct answer must contain. Those need auditing too, since a wrong expected-information field produces a confidently wrong score forever. See our eval dataset build guide for the annotation workflow.

Evaluate the judge before you trust the judge

Every guide says LLM judges are biased. Almost none gives a procedure. Judge-based LLM evaluation techniques are only as good as this check.

  1. Human-label a stratified sample. 100 to 300 outputs, stratified across the segments you care about, labelled against the exact rubric the judge will use. Two annotators on an overlapping subset, so you can compute inter-annotator agreement first. If humans cannot agree with each other, the rubric is broken and no judge will save it.
  2. Run the judge on the identical sample. Same items, same order, same rubric text.
  3. Compute agreement with Cohen’s kappa (two raters) or Krippendorff’s alpha (more raters, or ordinal scales), not raw accuracy. Raw agreement of 85 percent on a task where 80 percent of items pass is barely above chance.
  4. Inspect the confusion matrix for direction. A systematically lenient judge is a different problem from a noisy one, and only the matrix distinguishes them.
  5. Revise the rubric, not the model, and repeat. Treat the human labels as the judge’s regression suite, versioned alongside the eval set.

Cohen’s kappa is (po − pe) / (1 − pe), where po is observed agreement and pe is agreement expected by chance from the marginals. A worked template with illustrative values:

Sample IDHuman labelJudge labelAgreementError direction
001passpassyes—
002passpassyes—
003failpassnojudge lenient
004passpassyes—
005failfailyes—
006passpassyes—
007failpassnojudge lenient
008passpassyes—
009passpassyes—
010failfailyes—

These ten rows are an illustrative template, not measured results. Observed agreement po = 8/10 = 0.80. Humans passed 6 of 10, the judge passed 8. Chance agreement pe = (0.6 × 0.8) + (0.4 × 0.2) = 0.48 + 0.08 = 0.56. Kappa = (0.80 − 0.56) / (1 − 0.56) = 0.24 / 0.44 = 0.55. Moderate agreement, and both errors point the same way: the judge passes items humans failed. That is a rubric that has not defined the failing condition sharply enough. Tighten the failure definition. Do not lower the score threshold.

Judge calibration worksheet template for LLM evaluation techniques that use a model as grader

Visual spec: Judge-calibration worksheet as a copyable template, columns for sample ID, human label, judge label, agreement and error direction, with the kappa formula and the ten-row illustrative block above. Label it explicitly as a template with illustrative numbers.

Documented biases and their named mitigations.

BiasMechanismMitigation
Position biasIn pairwise comparison, the option shown first wins more oftenBalanced Position Calibration: run both orderings and average, and report the disagreement rate between orderings as a health metric
Verbosity biasLonger answers read as more thoroughLength control, or compare against a length-matched baseline as AlpacaEval 2.0 does
Self-enhancement biasA model prefers its own outputsNever let the generating model grade itself; use a jury from different model families
Poor numeric calibrationScores cluster on the middle of a 1-to-5 scaleBinary or few-valued rubrics, or probability-weighted scoring
Weak math and reasoningThe judge cannot check arithmetic it could not performMove numeric criteria to verifiers

Microsoft Learn lists positional bias, verbosity bias, self-enhancement bias, weak mathematical and reasoning ability, and difficulty assigning numerical scores as reported limitations of LLM evaluators. It names three mitigations: Multiple Evidence Calibration (score with multiple pieces of evidence and aggregate), Balanced Position Calibration, and Human In The Loop Calibration. It gives one sentence on each. The table above is the operational version.

Statistical significance between two candidates. Only the arXiv guide covers this, and it is the difference between a decision and a vibe. For paired binary judgements (pass/fail on the same items under two configurations), use McNemar’s test, which looks only at the discordant pairs. For continuous scores, a paired two-tailed t-test. For Likert-scale judgements, the Wilcoxon signed-rank test, because Likert data are ordinal and not normally distributed. Both candidates run on the same rows, so paired tests are always available and always more powerful than unpaired ones.

A 2-point win on 100 examples is usually nothing. With 100 examples at a pass rate near four in five, the margin of error is about ±8 points. Ship on the significance test, not the delta.

Non-determinism handling. Fix temperature and seed where the provider exposes them. Where it does not, run the evaluation three times and report a variance or confidence interval beside the point estimate. A single run’s number without a spread is unfalsifiable. Treat prompt sensitivity as a measurable property. Paraphrase your eval inputs and see whether the score moves. If it moves more than your release threshold, your threshold is measuring phrasing.

Rubric-writing rules that raise agreement. One criterion per judge call. Define the failing condition explicitly, not only the passing one, because “is the answer helpful” has no failure boundary. Include two or three anchor examples at the boundary rather than obvious cases. Force structured output with a reason field emitted before the score, so the reason constrains the score rather than rationalising it. Name the evidence the judge must quote from the context.

Reliability figures for specific frontier judges circulate widely in blog posts without primary links. Treat them as leads. This guide cites none of them as findings, and neither should your procurement document.

What a judge-based eval suite costs per run

No competitor in the top 10 prices evaluation, even though cost is the first objection raised when anyone proposes judge-based LLM evaluation techniques. The model is arithmetic:

cost per judged sample = (input tokens × input rate) + (output tokens × output rate)

Input tokens include the rubric, the question, the output under test, and, for RAG, the entire retrieved context. Output tokens include the reason field and the score.

Token assumptions used below, stated so you can change them: rubric 150 tokens, question 60, output under test 250, formatting overhead 40. Total 500 input tokens per non-RAG judge call, with 120 output tokens for a reason plus a score. The RAG variant adds 4,000 tokens of retrieved context, giving 4,500 input tokens.

Rate assumptions, illustrative only. A frontier judge at $3.00 per million input tokens and $15.00 per million output tokens. A small judge at $0.15 per million input and $0.60 per million output. Replace each pair with the figure copied from your provider’s own pricing page, whether that is OpenAI, Anthropic, Google or a host you run yourself, and record the access date beside it. The arithmetic below is reproducible from whatever rates you substitute.

The suite. 300 eval samples × 5 criteria × 2 orderings for position debiasing × 3 runs for variance = 9,000 judge calls per suite execution.

Suite variantJudge callsFrontier judgeSmall judge
300 × 5 criteria, single run, no retrieved context1,500$4.95$0.22
× 2 orderings3,000$9.90$0.44
× 2 orderings × 3 runs9,000$29.70$1.32
Full 9,000 calls, RAG variant (4,500 input tokens)9,000$137.70$6.72

Per-call arithmetic for the first row: 500 × $3/1M = $0.0015 input, 120 × $15/1M = $0.0018 output, $0.0033 per call, × 1,500 = $4.95. The RAG variant: 4,500 × $3/1M = $0.0135 plus $0.0018 = $0.0153 per call, × 9,000 = $137.70. Every figure in the table comes from those two lines and the stated assumptions.

Cadence turns those figures into a budget line:

CadenceFrontier, RAG variantSmall judge, RAG variant
Nightly (30 runs/month)$4,131$202
Every PR, 20 PRs/day (600 runs/month)$82,620$4,032

Three levers dominate the bill.

  • Retrieved-context length in the judge prompt. Going from 500 to 4,500 input tokens multiplied the per-run cost by 4.6× here. It is the least-noticed lever, because nobody counts the context they paste into a rubric. Score faithfulness against only the chunks the generator actually cited, not the full retrieved set, and this line drops by most of its value.
  • Number of criteria. Cutting from 5 criteria to 3 removes 40 percent, linearly. Ask which criteria have ever changed a release decision.
  • Repetition for debiasing and variance. Dropping from 3 runs to 1 removes 67 percent, at the cost of your confidence interval. Compromise: 3 runs on a 100-row subset to establish variance, 1 run on the full set.

The alternatives on the same basis. Deterministic verifiers and overlap metrics are effectively free compute, bounded by CI minutes. Embedding-based scoring is one cheap embedding call per item, orders of magnitude below a generative call. Human labelling is priced per annotator-hour: 300 items at 2 minutes each is 10 annotator-hours per pass. Compare that against the frontier figures above at your own loaded hourly rate before assuming humans are the expensive option. At per-PR cadence they are not even close.

The tiered pattern that falls out of these numbers.

TierCadenceContentsRough cost profile
1Every commitDeterministic checks, schema validation, overlap tripwires, unit-test verifiersCI minutes only
2Every PREmbedding triage plus a small-judge pass on the pinned subsetSingle-digit dollars per run
3Nightly and pre-releaseFull frontier-judge suite with orderings and repeatsTens to low hundreds of dollars per run
4WeeklyHuman review of a fixed stratified sampleAnnotator-hours

Cost model for judge-based LLM evaluation techniques across four suite variants and two cadences

Visual spec: Cost-model table as a graphic, rows for the four suite variants and the two cadences, columns for a frontier judge and a small judge, with a footnote block listing token-count assumptions per judge call and each rate’s source URL and access date. Label it explicitly as arithmetic from published rates, not a measured invoice.

Method note: every rate in this section is an assumption, to be replaced with a figure quoted from the provider’s public pricing page with an access date. The model is arithmetic you can audit and re-run when prices change. Token accounting for production traffic as well as judges is covered in our inference cost modelling guide.

From scores to gates: evaluation in CI and in production

Scores that gate nothing are a hobby. This is where LLM evaluation techniques become engineering.

Offline gating in CI. Per-commit: deterministic checks, JSON-schema validation, and a regression run on a small pinned set. Those finish in seconds and cost nothing. Per-PR: the full offline suite with thresholds. Express thresholds as a proportion of samples passing, never as a mean score, because a mean lets one catastrophic output hide behind nine good ones. Then add the practice most teams miss. Fail the build on a drop relative to the last green run, not only on an absolute floor. Absolute floors get lowered when they are inconvenient. A relative-drop gate catches the slow erosion a lowered floor is designed to hide. See our CI regression testing guide for the pipeline layout.

Online evaluation without ground truth. Sampled judge scoring on live traffic, at a sampling rate your cost model can carry. Reference-free descriptors: output length distribution, refusal rate, sentiment floor, guardrail trip rate, citation-existence failures. User-signal proxies: thumbs, retries, edits, escalation to a human, session abandonment. None of these is truth. Together they move before your users tell you anything.

The five-stage production loop.

  1. Detect. A score or error rate crosses a threshold on sampled live traffic. Prerequisite: online scoring already running, with alerting at the slice level rather than the aggregate.
  2. Attribute. Every response carries the identity of the prompt version, model version and configuration that produced it. Prerequisite: response-level provenance.
  3. Bound. The change was rolled out progressively, so exposure is a known fraction. Prerequisite: progressive delivery.
  4. Act. Revert to the last known-good configuration without a redeploy. Prerequisite: runtime configuration, which is what products such as LaunchDarkly AI Configs sell, and what a config service plus a cache gives you in-house.
  5. Verify. Scores return to baseline, and the failing case joins the offline set so it cannot regress silently again.

Production evaluation loop showing where online LLM evaluation techniques attach to detect, attribute, bound, act and verify

Visual spec: Production evaluation loop diagram, detect → attribute → bound → act → verify, with the prerequisite named under each stage. Label it as a generalisation of a documented pattern and cite launchdarkly.com/blog/llm-evaluation as prior articulation rather than presenting it as original.

The prerequisite most teams lack is stage 2. If you cannot say which prompt version and model version produced a given response, online evaluation produces alarms nobody can action, and the on-call engineer’s only move is to guess. Provenance is a logging change. Make it before you buy anything.

Silent model updates are their own failure class. A provider changing the model behind a stable alias moves your scores with no change on your side, and your git history contains no explanation. Mitigations: pin model versions explicitly, keep a continuously running canary eval on a fixed 50-row set so a shift shows up as a step change with a timestamp, and record the resolved model version in every response’s provenance.

Compliance now sits inside this loop. EU AI Act transparency and documentation provisions and California SB 53 make retained evaluation evidence an obligation rather than hygiene: versioned datasets, recorded scores, audit trails of who changed which threshold. LaunchDarkly frames this well. Read the statutes themselves rather than any blog’s summary, because the obligations attach differently depending on your role in the supply chain and your risk classification. The engineering requirement is unglamorous. Your eval artefacts need retention, immutability and an owner.

The commercial version of the same point is the Air Canada tribunal decision, in which the airline was held to its chatbot’s misstatement of bereavement fare policy. Read the ruling itself rather than the blogs that cite it. The generalisable lesson for evaluation design: a policy-stating surface needs a grounding check per claim, and “the chatbot said it, not us” is not a defence.

Which framework implements which technique

Frameworks differ in which LLM evaluation techniques they implement out of the box, and the tooling question splits into three categories. Benchmark harnesses answer “which model is better at standardised tasks”. Application eval frameworks answer “is my system good enough on my data”. Tracing and observability platforms answer “what is happening in production and how do I score it”. A team that adopts a benchmark harness to test its support bot will spend a month building dataset adapters for a job an application framework does in an afternoon.

Method note for the table below. Populate each cell from that project’s own documentation or repository, with the URL and access date recorded per row, and mark “not documented” where the docs are silent rather than guessing. No vendor blog posts as sources for another vendor’s row. The version below is a starting scaffold, and this category ships weekly, so recheck every cell against the linked docs on the day of publication.

ToolLicenceCategoryBuilt-in judge metricsRAG metricsAgent / trace metricsCustom metricsCI integrationSelf-hostableHosted account for dashboard
DeepEvalApache-2.0App evalYes, extensiveYesYesYesPytest-nativeYesDashboard is Confident AI
EvidentlyApache-2.0App eval + monitoringYesYesNot documentedYesYesYesOptional cloud
RAGASApache-2.0App evalYesYes, its focusPartialYesYesYesNo
promptfooMITApp eval + red teamingYesPartialPartialYesYes, config-drivenYesOptional
OpenAI EvalsMITApp eval / benchmarkYesNot documentedNot documentedYesScriptableYesNo
EleutherAI lm-evaluation-harnessMITBenchmark harnessNoNoNoTask YAMLScriptableYesNo
Inspect AIMITApp eval + safety evalYesPartialYes, solver tracesYesScriptableYesNo
MLflow evaluateApache-2.0Tracking + app evalYesPartialTracingYesYesYesNo
TruLensMITApp eval + tracingYes, feedback functionsYesYesYesScriptableYesNo
Arize PhoenixOpen sourceTracing + evalYesYesYesYesPartialYesOptional
NVIDIA NeMo EvaluatorProprietaryBenchmark harnessYesYesNot documentedYesEnterpriseEnterprise deployYes
Databricks Agent EvaluationProprietaryApp eval, agent focusYesYesYesYesWorkspace-nativePlatform-boundYes
IBM FM-evalProprietaryApp evalYesNot documentedNot documentedNot documentedEnterprisePlatform-boundYes
Azure AI Foundry evaluationProprietaryApp eval + monitoringYesYesPartialYesYesPlatform-boundYes

LangSmith, Braintrust, Langfuse, Comet’s Opik and Weights & Biases Weave belong in the tracing-plus-eval column alongside Phoenix. Their value is dataset management, trace capture and human annotation queues rather than novel scorers. Stanford’s HELM is a benchmark suite, so read it next to lm-evaluation-harness rather than next to DeepEval. Giskard sits closer to promptfoo, with testing and red-teaming as its centre. If you already run on a cloud, check what ships there first: Amazon Bedrock evaluations, Google Vertex AI evaluation and Azure AI Foundry all cover the common judge and RAG criteria, and none of them needs a new vendor relationship.

Vendor self-claims, labelled. Evidently’s marketing page states that its open-source library has over 25 million downloads. That is self-reported, with no link to an independent counter, and PyPI’s own download statistics are the place to check it. DeepEval’s repository describes the library as letting anyone implement state-of-the-art LLM metrics in five lines of code. That is a maintainer’s framing of ergonomics, not a finding about metric quality, and the metrics themselves are the ones described earlier in this guide. Neither figure is repeated here as fact, and neither belongs in a procurement document without its source.

What has actually shipped recently, from changelogs rather than blog posts. Confident AI’s changelog records adding free-text search over traces and spans, where the product previously supported filtering only, plus annotation-triggered workflows. Read that as a signal about the whole category. Full-text trace search and human-annotation-in-the-loop are current gaps being filled, not solved problems. If your evaluation plan assumes you can grep your traces, check that your platform can.

For neutral learning material, practitioner-built reference implementations are worth reading before you commit. A community reference implementation for evaluating AI applications was discussed on Hacker News, and the thread’s central argument is that testing AI applications differs fundamentally from testing classic software, because the correctness oracle is statistical rather than exact. A standalone version of the capability table lives at our open-source eval frameworks comparison.

Failure modes: where evaluation setups break in practice

Most broken setups fail at the dataset or the judge, not at the choice of LLM evaluation techniques. This catalogue is built from public discussion and primary documents, each linked, because the abstract “challenges” section every competitor writes teaches nothing.

Hallucination detection is two different jobs, and teams conflate them. A Hacker News thread asking how one detects hallucinations lays out the split cleanly. Verifying correctness against world knowledge requires search or a knowledge base. Verifying groundedness against the provided context requires only the context. Only the second is testable offline in a RAG system, because the first depends on the state of the world at query time. Build the groundedness check, which is an NLI or QAG pass over the retrieved chunks. Treat world-knowledge verification as a separate, online, search-backed system, and do not expect one metric named “hallucination” to cover both. Our hallucination detection methods guide separates the two pipelines.

Ad-hoc personal evaluation is still the norm, including among sophisticated engineers. An Ask HN thread on how practitioners personally evaluate new models is full of people describing vibes-based assessment and day-to-day IDE usage as their real method, while explicitly distrusting public benchmarks. Read it before you write a best-practice document. It explains why formal LLM evaluation techniques get skipped: the informal method is fast and grounded in the evaluator’s actual work. The productive response is not to shame the practice but to formalise it. Twenty real failures from your own logs is the bridge between vibes and statistics, which is why the staged rollout below starts there rather than at a benchmark.

Framework-adoption friction shows up in the vendors’ own launch threads. Confident AI’s Launch HN positions DeepEval as “Pytest for LLMs”. The mental model is genuinely useful. Eval cases are test cases, assertions are metrics, CI runs them. The positioning is the vendor’s own, and the comment thread is where adopters raise the question that matters for a build-versus-buy decision: how much of the value sits in the open-source package versus the hosted dashboard. Read the thread, not the tagline.

Customer-logo claims are marketing until independently confirmed. Relari’s Show HN states that its LLM evaluation stack is used in production by AI teams at Vanta and PwC. That may well be accurate. It is still a vendor statement about its own traction, published by the vendor, and it is exactly the class of claim that should not appear in your evaluation of the tool. Ask for a reference call instead.

Your evaluation pipeline is an attack surface and a data-governance problem. Eval datasets contain production data. Eval pipelines hold credentials for every model provider you use. Eval runners execute model-generated code whenever you use verifiers. So access-control your eval datasets like production data, sandbox every verifier that runs generated code, and scope eval API keys separately from production keys. None of that is exotic, and none of it is usually done.

Adversarial inputs survive evaluation because eval sets contain none. Research on novel prompt-injection threats to application-integrated LLMs establishes the input classes that matter, including injection delivered through retrieved content rather than user input. If your eval set is 300 well-formed customer questions, it will report a healthy score while your system is one poisoned document away from leaking its system prompt. Every eval set needs a dedicated adversarial slice, scored separately, with its own threshold and its own pass-rate-per-category table.

An integrity note for whoever maintains this page. When compiling failure modes, verify that every issue tracker cited actually belongs to an LLM-evaluation project. Repository issue data bundled with research for this article (type-challenges, BLEUnlock) belongs to unrelated projects and must never be cited here.

Safety, robustness and red teaming as evaluation techniques

Safety testing is a set of LLM evaluation techniques, not a values statement, and it is scored the way everything else is.

Red teaming as a systematic suite. Build an adversarial dataset with named categories: jailbreak patterns, prompt injection delivered through retrieved content, PII extraction attempts, harmful-instruction elicitation, and refusal-bypass framing such as fictional or hypothetical wrappers. Track pass rate per category across releases. An aggregate safety score that hides a weak injection-via-retrieval category is worse than no score. Rotate in new attack patterns as they are published, and keep a held-out set so your prompt engineering cannot overfit to the attacks you already know. promptfoo and Inspect AI both ship red-teaming support, which makes this a configuration job rather than a build.

Toxicity and bias with named public datasets. RealToxicityPrompts provides prompts that elicit toxic continuations. ToxiGen provides implicitly toxic statements targeting demographic groups, which catch the failures keyword filters miss. Score with a fine-tuned classifier rather than a judge wherever a classifier exists. It is cheaper per item, deterministic between runs, and versionable, and those three properties matter more for a gate than marginal accuracy. Keep the classifier version in your provenance record, because a classifier upgrade shifts your historical trend line.

Robustness testing measures deltas, not levels. Perturb your eval inputs and report the score change: paraphrase invariance, typo and casing perturbation, input truncation, long-context stress by padding irrelevant material around the relevant chunk, and non-English translation of the same items. The delta is the robustness measure. A system that loses a quarter of its score under paraphrase was never as good as its headline number, and the paraphrase delta is also the cheapest available proxy for the prompt-sensitivity problem in the judge section.

Language and cultural coverage is a concrete gap, not a nicety. Benchmark results for English typically exceed results for other languages on the same tasks, and Western-weighted training data shows up in model defaults for names, units, dates and etiquette. XNLI and MGSM exist precisely because English-only measurement misleads. If you serve non-English users, an English-only eval set evaluates one segment of your product, not your product. Stratify by language, set per-language thresholds, and expect the gap to be largest on the languages with the least representation in pretraining data. Our red teaming checklist lists the adversarial categories in full.

A staged rollout: what to build in week one, month one, quarter one

Pick LLM evaluation techniques in this order.

Week one. Collect 20 to 50 real failures from production logs, support tickets or internal complaints. For each, write the pass/fail criterion in plain language. What would a correct response have contained, and what did this one do wrong. Run them by hand against your current system and record the results in a spreadsheet. This is your first eval set. It took a day, and it will outperform any borrowed benchmark, because every row is a failure someone actually cared about. Do not automate anything yet.

Month one, part one: pick five metrics and no more. One task-specific custom criterion written from your product spec. One or two architecture-driven criteria (faithfulness for RAG, tool correctness for agents). One safety gate. One cost-or-latency metric. More than five produces dashboards nobody reads and thresholds nobody enforces, because each threshold needs an owner who will block a release on it, and no team has 12 of those people.

Month one, part two: size and calibrate. Grow the set to the sample size your target margin of error requires, using the table earlier in this guide: 246 rows for a ±5 point margin on a metric expected near four in five, 683 rows for ±3 points. Add stratification tags at the same time, because retrofitting tags onto 600 rows is miserable. Then calibrate your judge against 100 human-labelled rows and compute Cohen’s kappa before any number from that judge gates anything.

Quarter one. Wire deterministic checks and verifiers into CI on every commit. Put the full offline suite behind a pre-release gate with pass-rate thresholds and a relative-drop rule. Start sampled online scoring, with response-level provenance shipped first. Then establish the discipline that closes the loop. Every production failure becomes an offline test case in the same week it is found, with the incident ID in the row’s metadata. That single habit is what makes an eval suite compound instead of decay.

Anti-patterns to name and kill. Reporting means instead of pass rates. An eval set small enough for the team to memorise, which turns evaluation into prompt-fitting. Letting the generating model grade itself. Chasing a public leaderboard position no customer will ever see. Treating the eval suite as a one-off project rather than a versioned asset with an owner, a changelog and a review cadence.

Frequently Asked Questions

What are the best LLM evaluation tools?

There is no single best tool, because the category splits three ways. Benchmark harnesses such as lm-evaluation-harness, Stanford HELM and NVIDIA NeMo Evaluator compare models on standardised tasks. Application eval frameworks such as DeepEval, RAGAS, promptfoo, Evidently and Inspect AI test your system on your data. Tracing platforms such as MLflow, Arize Phoenix, LangSmith, Langfuse and Braintrust score production traffic.

Choose on four verifiable criteria: licence, how easily you can author a custom criterion, whether it runs in CI without a hosted service, and whether the dashboard requires an account. The capability table above lists all fourteen against those columns. Any “best tool” claim sourced from a vendor’s own page is marketing, including the two labelled in this guide.

What are the different techniques used in LLM testing?

The LLM evaluation techniques in use today fall into five families: deterministic reference matching (exact match, schema validation), statistical overlap (BLEU, ROUGE, METEOR, Levenshtein), model-based scoring (BERTScore, NLI entailment, fine-tuned classifiers), LLM-as-a-judge (direct scoring, pairwise, G-Eval, juries), and verifiers (unit tests, boxed-answer extraction, trace assertions).

Red teaming and robustness perturbation layer across all five rather than sitting beside them. Both are ways of choosing inputs, and any family can score the outputs. Two facts about your task determine the family: whether ground truth exists, and whether a program can check the output.

What are the 5 basic types of evaluation?

That phrasing comes from programme evaluation, where the five types are formative, summative, process, outcome and impact evaluation. Each maps onto LLM work. Formative is offline iteration on a dev set. Summative is the pre-release gate. Process is component and span-level scoring. Outcome is task completion and online quality. Impact is business and user-satisfaction metrics.

The mapping is worth keeping, because it exposes the stage most teams skip. Almost everyone does formative and summative evaluation. Few correlate offline scores against in-product satisfaction, which is the impact layer, and without it you cannot tell whether your eval set represents your users at all.

What are LLM evaluation frameworks?

A framework is three things together: a dataset specification, a set of scorers, and an execution and reporting harness that runs them and stores results. That distinguishes it from a benchmark, which is tasks plus data, and from a metric, which is just the scorer. Confusing the three causes most tooling mistakes.

Names such as G-Eval and GPTScore are prompting techniques rather than frameworks or metrics, a distinction Evidently makes and most pages blur. The practical test for a framework: can you add a custom criterion in your own code and run it in CI without a hosted account? If not, it is a product. Price it as one.

How large does an LLM evaluation dataset need to be?

Roughly 246 samples for a metric expected around 80 percent, at 95 percent confidence with a 5 percent margin of error, from n = z²p(1−p)/e² with z = 1.96, as derived in the practitioner guide at arxiv.org/html/2506.13023v1. Tighter margins scale sharply. A ±3 point margin needs 683 samples, ±2 points needs 1,537.

Start below that anyway. Twenty to fifty real failures is the right week-one set, because it teaches you what your criteria should be. Size up to your target margin before you let any number gate a release, and stratify by segment so per-slice scores stay interpretable.

Can you trust LLM-as-a-judge scores?

Only after you measure the judge. The procedure: human-label a stratified sample of 100 to 300 outputs against the same rubric, run the judge on the identical items, compute Cohen’s kappa rather than raw agreement, inspect the confusion matrix for the direction of error, then revise the rubric and repeat.

The documented failure modes are position bias, verbosity bias, self-enhancement bias and poor numeric calibration, per Microsoft Learn’s metric reference. Named mitigations: Balanced Position Calibration, length control, heterogeneous juries, Multiple Evidence Calibration, and binary rubrics instead of 1-to-5 scales. Cite published reliability figures only when you have the primary URL.

What does it cost to run LLM evaluations?

Cost per judged sample equals input tokens times the published input rate, plus output tokens times the published output rate, multiplied by samples × criteria × orderings × repeats. A 300-sample, 5-criterion suite run with 2 orderings and 3 repeats is 9,000 judge calls, which the arithmetic above prices at roughly $30 without retrieved context.

Three levers dominate: retrieved-context length in the judge prompt, criteria count, and repetition. The tiered pattern controls spend. Free verifiers on every commit. Cheap-judge triage per PR. The full frontier suite nightly. Human review weekly. Quote every rate from the provider’s pricing page with an access date.

What is the difference between offline and online LLM evaluation?

Offline evaluation scores a fixed candidate against a curated dataset before exposure, with ground truth usually available and every run reproducible. Online evaluation scores live responses continuously, on traffic your test set never contained. Ground truth rarely exists there, so you substitute reference-free descriptors, sampled judge scoring, and user-signal proxies such as retries and escalations.

Each has a blind spot. Offline evaluation cannot cover inputs you did not anticipate. Online evaluation cannot tell you whether a different candidate would have done better, because you only observe the branch you shipped. The loop compounds only when production failures flow back into the offline set.


What this guide does not claim. Nothing here was measured first-hand. No tool was installed, no benchmark was run, no latency was timed. Every number is either arithmetic shown in full from stated inputs (the sample-size table, the cost model) or a third-party result cited to its source. Current agreement rates between specific frontier judges and human labels on a fixed rubric would require running an eval harness against human-labelled data, which this page has not done and does not estimate. That is work a publisher may choose to commission. Until it exists, only published results with primary URLs are cited.

Method note. The framework capability table was built to be populated from each project’s own documentation or repository, one URL and access date per row, with “not documented” recorded where the docs are silent rather than inferred. Recheck each row against the linked docs before publication. Vendor statements about their own products are labelled as such with their URLs and are never restated as independent findings. Rate assumptions in the cost model are placeholders, to be replaced with figures copied from each provider’s public pricing page.

Last reviewed: 29 September 2026. Change log: sample-size table recomputed from the published formula; cost model arithmetic rebuilt against stated token assumptions; unsourced judge-reliability figures removed; framework table flagged for per-cell verification.

Next action, in order of return: pull 20 failures from last week’s logs, write one pass/fail criterion each, and run them by hand. That set will tell you which of the LLM evaluation techniques in this guide you actually need, and it costs you an afternoon rather than a quarter.

Frequently Asked Questions

What are the best LLM evaluation tools?

There is no single best tool, because the category splits three ways. Benchmark harnesses such as lm-evaluation-harness, Stanford HELM and NVIDIA NeMo Evaluator compare models on standardised tasks. Application eval frameworks such as DeepEval, RAGAS, promptfoo, Evidently and Inspect AI test your system on your data. Tracing platforms such as MLflow, Arize Phoenix, LangSmith, Langfuse and Braintrust score production traffic. Choose on four verifiable criteria: licence, how easily you can author a custom criterion, whether it runs in CI without a hosted service, and whether the dashboard requires an account. The capability table above lists all fourteen against those columns. Any "best tool" claim sourced from a vendor's own page is marketing, including the two labelled in this guide.

What are the different techniques used in LLM testing?

The LLM evaluation techniques in use today fall into five families: deterministic reference matching (exact match, schema validation), statistical overlap (BLEU, ROUGE, METEOR, Levenshtein), model-based scoring (BERTScore, NLI entailment, fine-tuned classifiers), LLM-as-a-judge (direct scoring, pairwise, G-Eval, juries), and verifiers (unit tests, boxed-answer extraction, trace assertions). Red teaming and robustness perturbation layer across all five rather than sitting beside them. Both are ways of choosing inputs, and any family can score the outputs. Two facts about your task determine the family: whether ground truth exists, and whether a program can check the output.

What are the 5 basic types of evaluation?

That phrasing comes from programme evaluation, where the five types are formative, summative, process, outcome and impact evaluation. Each maps onto LLM work. Formative is offline iteration on a dev set. Summative is the pre-release gate. Process is component and span-level scoring. Outcome is task completion and online quality. Impact is business and user-satisfaction metrics. The mapping is worth keeping, because it exposes the stage most teams skip. Almost everyone does formative and summative evaluation. Few correlate offline scores against in-product satisfaction, which is the impact layer, and without it you cannot tell whether your eval set represents your users at all.

What are LLM evaluation frameworks?

A framework is three things together: a dataset specification, a set of scorers, and an execution and reporting harness that runs them and stores results. That distinguishes it from a benchmark, which is tasks plus data, and from a metric, which is just the scorer. Confusing the three causes most tooling mistakes. Names such as G-Eval and GPTScore are prompting techniques rather than frameworks or metrics, a distinction Evidently makes and most pages blur. The practical test for a framework: can you add a custom criterion in your own code and run it in CI without a hosted account? If not, it is a product. Price it as one.

How large does an LLM evaluation dataset need to be?

Roughly 246 samples for a metric expected around 80 percent, at 95 percent confidence with a 5 percent margin of error, from n = z²p(1−p)/e² with z = 1.96, as derived in the practitioner guide at [arxiv.org/html/2506.13023v1](https://arxiv.org/html/2506.13023v1). Tighter margins scale sharply. A ±3 point margin needs 683 samples, ±2 points needs 1,537. Start below that anyway. Twenty to fifty real failures is the right week-one set, because it teaches you what your criteria should be. Size up to your target margin before you let any number gate a release, and stratify by segment so per-slice scores stay interpretable.

Can you trust LLM-as-a-judge scores?

Only after you measure the judge. The procedure: human-label a stratified sample of 100 to 300 outputs against the same rubric, run the judge on the identical items, compute Cohen's kappa rather than raw agreement, inspect the confusion matrix for the direction of error, then revise the rubric and repeat. The documented failure modes are position bias, verbosity bias, self-enhancement bias and poor numeric calibration, per [Microsoft Learn's metric reference](https://learn.microsoft.com/en-us/ai/playbook/technology-guidance/generative-ai/working-with-llms/evaluation/list-of-eval-metrics). Named mitigations: Balanced Position Calibration, length control, heterogeneous juries, Multiple Evidence Calibration, and binary rubrics instead of 1-to-5 scales. Cite published reliability figures only when you have the primary URL.

What does it cost to run LLM evaluations?

Cost per judged sample equals input tokens times the published input rate, plus output tokens times the published output rate, multiplied by samples × criteria × orderings × repeats. A 300-sample, 5-criterion suite run with 2 orderings and 3 repeats is 9,000 judge calls, which the arithmetic above prices at roughly $30 without retrieved context. Three levers dominate: retrieved-context length in the judge prompt, criteria count, and repetition. The tiered pattern controls spend. Free verifiers on every commit. Cheap-judge triage per PR. The full frontier suite nightly. Human review weekly. Quote every rate from the provider's pricing page with an access date.

What is the difference between offline and online LLM evaluation?

Offline evaluation scores a fixed candidate against a curated dataset before exposure, with ground truth usually available and every run reproducible. Online evaluation scores live responses continuously, on traffic your test set never contained. Ground truth rarely exists there, so you substitute reference-free descriptors, sampled judge scoring, and user-signal proxies such as retries and escalations. Each has a blind spot. Offline evaluation cannot cover inputs you did not anticipate. Online evaluation cannot tell you whether a different candidate would have done better, because you only observe the branch you shipped. The loop compounds only when production failures flow back into the offline set. -- *What this guide does not claim. Nothing here was measured first-hand. No tool was installed, no benchmark was run, no latency was timed. Every number is either arithmetic shown in full from stated inputs (the sample-size table, the cost model) or a third-party result cited to its source. Current agreement rates between specific frontier judges and human labels on a fixed rubric would require running an eval harness against human-labelled data, which this page has not done and does not estimate. That is work a publisher may choose to commission. Until it exists, only published results with primary URLs are cited. *Method note. The framework capability table was built to be populated from each project's own documentation or repository, one URL and access date per row, with "not documented" recorded where the docs are silent rather than inferred. Recheck each row against the linked docs before publication. Vendor statements about their own products are labelled as such with their URLs and are never restated as independent findings. Rate assumptions in the cost model are placeholders, to be replaced with figures copied from each provider's public pricing page. *Last reviewed: 29 September 2026. Change log: sample-size table recomputed from the published formula; cost model arithmetic rebuilt against stated token assumptions; unsourced judge-reliability figures removed; framework table flagged for per-cell verification. Next action, in order of return: pull 20 failures from last week's logs, write one pass/fail criterion each, and run them by hand. That set will tell you which of the LLM evaluation techniques in this guide you actually need, and it costs you an afternoon rather than a quarter.

Explore More

Free Newsletter

Get the LLM Evals Newsletter

Platform comparisons, pricing changes and eval technique deep-dives. No spam.

Related Articles