How to Evaluate an LLM: Test Sets, Scorers and Model Judges
8 min readBytePatterns
How to evaluate an LLM: frozen test cases, normalised and deterministic scorers first, rubrics and model judges after, and enough cases to trust a difference.
You changed a prompt, swapped a model or added retrieval, and the demo looks better. Is it? Without an evaluation, nobody knows, and "it looks better" is how quality quietly regresses. Evaluating an LLM is less exotic than it sounds: it is testing, with two complications. Correct answers can be phrased many ways, and the system is often non-deterministic. Both have standard answers.
The problem it solves
An evaluation turns "is this better?" into a number you can compare across runs. It needs three things:
- Frozen inputs. A fixed set of cases, versioned like code, so two runs are comparable.
- An expected outcome per case. An exact answer, a pattern, a test the output must pass, or a rubric.
- A scorer. The function that turns one output into pass or fail, or a grade.
With those, every change becomes an experiment on the same cases. Without them, you are comparing anecdotes.
The intuition
Order the scorers from cheapest and most trustworthy to most flexible:
- Deterministic checks first. Normalised exact match for short answers, a regular expression for formats such as dates, "does the code pass its tests" for code, a schema check for JSON. They are fast, free and never disagree with themselves.
- Normalise before comparing. Case, surrounding spaces and a trailing full stop are not errors. Raw string equality scores them as failures, and a scorer that lies is worse than none.
- Rubrics for what is left. Tone, helpfulness and faithfulness to sources need criteria written down: "cites a source for every claim", "under 100 words". People grade against the rubric, or a model does.
- A model judge is a model. It has an error rate. Measure it against human labels on a sample before trusting it, and re-measure when the judge changes. The model-as-judge lesson covers its biases.
Then the statistics: an eval set of 40 cases cannot reliably tell 75% from 80%. Compare versions per case on the same inputs and ask whether the split of fixed and broken cases could be chance.
Watch it run
The animation is the lesson's scoreboard. An evaluation is three things: fixed inputs, an expected outcome for each, and a scorer. Three cases each carry the answer a correct system should give, and the model's actual answers are recorded character for character. Scored with plain string equality, case 1 fails because "Paris" is not " paris ", case 2 fails because a trailing full stop is still a different string, and case 3 fails as well, so raw scoring reports 0 / 3 and is wrong about two of them. So both sides are normalised: lower-case, trim, collapse spaces, drop a trailing stop. Case 1 now compares "paris" with "paris" and passes; case 2 passes too, since the model was right and only its punctuation differed. Case 3 still fails, as it should, because green is not blue. 2 / 3 is the honest number, against 0 / 3 before normalising. Deterministic checks come first: exact match, a pattern, or simply whether the code runs. Softer qualities need a rubric graded by people or by a model, and that judge has to be spot-checked against human labels, where in the final frame it agrees on only 2 of 3.
Evaluating LLMs
Step 1 of 14
An evaluation is three things: fixed inputs, an expected outcome for each, and a scorer.
The same interactive animation as the lesson — step through it with the controls.
The code
A toy model of an evaluation harness. Three deterministic scorers share one interface; the lesson's normaliser is the first:
import math
import random
import re
from fractions import Fraction
from itertools import combinations, product
def norm(s):
return " ".join(s.lower().strip().split()).rstrip(".")
def exact(want, got): # the lesson's normalised match
return norm(want) == norm(got)
def pattern(want, got): # want is a regular expression
return re.fullmatch(want, got.strip()) is not None
def runs(want, got): # want is a test the answer's code must pass
env = {}
try:
exec(got, env) # toy only: sandbox real model output
return eval(want, env) is True
except Exception:
return False
SCORERS = {"exact": exact, "pattern": pattern, "runs": runs}
cases = [
("exact", "Paris", " paris "),
("exact", "42", "42."),
("exact", "blue", "green"),
("pattern", r"\d{4}-\d{2}-\d{2}", "2026-10-02"),
("pattern", r"\d{4}-\d{2}-\d{2}", "October 2, 2026"),
("runs", "f(3) == 9", "def f(x):\n return x * x"),
("runs", "f(3) == 9", "def f(x):\n return x + x"),
]
verdicts = [SCORERS[kind](want, got) for kind, want, got in cases]
print([int(v) for v in verdicts], sum(verdicts), "/", len(cases)) # [1, 1, 0, 1, 0, 1, 0] 4 / 7
Now the question every eval exists to answer: is the new version better? In this toy, the old version passes each case with probability 0.75 and the new one with 0.80, so the new one really is better. A sign test looks only at the cases where the versions disagree and asks how likely that split would be from a fair coin:
def sign_test(wins, losses):
"""Exact two-sided p-value: could this split be a fair coin?"""
n, k = wins + losses, min(wins, losses)
return min(1.0, 2 * sum(math.comb(n, i) for i in range(k + 1)) / 2 ** n)
def compare(n, seed):
rng = random.Random(seed)
old = [rng.random() < 0.75 for _ in range(n)]
new = [rng.random() < 0.80 for _ in range(n)]
wins = sum(b and not a for a, b in zip(old, new)) # cases the new one fixed
losses = sum(a and not b for a, b in zip(old, new)) # cases it broke
return sum(new) - sum(old), sign_test(wins, losses)
small = [compare(40, seed) for seed in range(10)]
print([d for d, _ in small]) # [5, 1, -1, -2, 8, 7, 4, -2, 0, 0]
print(sum(p < 0.05 for _, p in small)) # 1
big = compare(2000, 0)
print(big[0], big[1] < 0.001) # 103 True
Ten 40-case runs of the same two versions: the "improvement" ranges from minus two to plus eight, half the runs say the new version is no better, and only one clears the usual 0.05 bar. With 2,000 cases the gain is unmistakable. Next, a judge that passes everything humans pass plus some they fail, and the pass@k estimate used for code generation (from memory: the unbiased form below is the standard one):
rng = random.Random(7)
human = [rng.random() < 0.6 for _ in range(50)]
judge = [h or rng.random() < 0.3 for h in human] # a lenient judge
agree, lenient = sum(map(lambda h, j: h == j, human, judge)), sum(map(lambda h, j: j > h, human, judge))
print(agree, lenient) # 47 3
def pass_at_k(n, c, k):
"""n samples, c correct: chance that k drawn without replacement include one."""
if n - c < k:
return Fraction(1)
return 1 - Fraction(math.comb(n - c, k), math.comb(n, k))
print(pass_at_k(10, 3, 1), pass_at_k(10, 3, 5)) # 3/10 11/12
The judge agrees on 47 of 50, and all three disagreements go the same way: it passed answers humans failed. Agreement alone hides that direction; count it. Finally, both formulas are checked by brute force on 300 seeded cases: pass@k against every k-subset of the samples, and the sign test against every way a fair coin could split the disagreements:
ok = True
for seed in range(300):
r = random.Random(seed)
n = r.randint(1, 10)
c, k = r.randint(0, n), r.randint(1, n)
picks = list(combinations([True] * c + [False] * (n - c), k))
ok &= pass_at_k(n, c, k) == Fraction(sum(any(p) for p in picks), len(picks))
w, l = r.randint(0, 7), r.randint(0, 7)
flips = list(product([0, 1], repeat=w + l))
extreme = sum(min(sum(f), w + l - sum(f)) <= min(w, l) for f in flips)
ok &= abs(sign_test(w, l) - min(1.0, extreme / len(flips))) < 1e-12
print(ok) # True
The complexity
- Deterministic scoring: linear in the size of the outputs, effectively free next to generating them.
- A run:
cases × samples per casegenerations; non-deterministic systems need several samples per case. - Model judging: one extra generation per graded output, plus the human labels needed to measure the judge.
Where it goes wrong
- Too few cases. As above, 40 cases cannot separate close versions. Report the number of cases with every score.
- Unnormalised scoring. Cosmetic differences become failures and the score lies.
- Contamination. Public benchmark questions may sit in a model's training data, inflating its score. Keep a private set built from your own traffic.
- Tuning on the test set. Iterate on a development set; keep a held-out set you open rarely, as in classic machine learning.
- Executing model output unsandboxed. The
runsscorer is a toy; real harnesses run code in an isolated sandbox. - An unmeasured judge. Re-check judge agreement when the judge, the rubric or the domain changes.
When it shows up in interviews
In AI engineering rounds: "how would you know your RAG system got better?", "how do you ship prompt changes safely?". Expect follow-ups on golden datasets, regression gates in CI, judge reliability and online metrics. It pairs with RAG, where retrieval and generation are scored separately, and with choosing between fine-tuning, prompting and RAG, which evals decide. As of October 2026, model-graded rubrics are common practice, and only as good as their agreement with human labels.
How to say it in an interview
"An eval is a frozen, versioned set of cases with an expected outcome per case and a scorer. I use deterministic scorers first: normalised exact match, patterns, running the code, schema checks. Softer qualities get a written rubric graded by people or a model judge, and I measure the judge's agreement with human labels, including which way it errs. To compare versions I run both on the same cases, look at per-case wins and losses, and make sure the set is large enough that the difference is not noise. Then it becomes a regression gate on every change."