LLM as a Judge: Position Bias, Length Bias and Agreement
9 min readBytePatterns
LLM as a judge explained: pairwise grading, why the verdict flips when you swap the order, how longer answers win, swap-and-tie, and kappa against human labels.
When an answer cannot be checked by a string match, a common shortcut is to have another model grade it: an LLM as a judge. It is fast, cheap compared with people, and scales to thousands of comparisons a night. It is also a model, and its failures lean in predictable directions. This article measures two of them, position and length, in a small simulation, then shows the cheap fixes and the number that says whether to trust the judge.
The problem it solves
Open-ended outputs, such as summaries or support replies, have no single correct string. Exact match fails, and human grading is slow and expensive. The LLM evaluation article puts deterministic scorers first and judges after; this one is about the judge itself.
Two set-ups are common:
- Pairwise: show two answers and ask which is better against a rubric.
- Pointwise: ask for a score, say 1 to 10, for one answer.
Pairwise is often preferred, because free scores cluster: most answers come back a 7 or an 8, which separates almost nothing.
The intuition
A judge's verdict should depend only on the content. In practice it also depends on things that should not matter:
- Position. The same pair, shown in the opposite order, can get the opposite verdict. Measure it by swapping and re-asking.
- Length and fluency. A longer, smoother answer often beats a shorter correct one.
- Self-preference. Judges tend to rate answers in their own house style highly.
The mitigations cost little. Judge every pair in both orders: if the verdicts agree, keep it; if not, call it a tie, so order can no longer decide anything. Hide which system wrote which answer. Ask for a short reason before the verdict. Above all, measure the judge like any other model: compare it with human labels on a held-out sample, and re-check whenever the judge or rubric changes.
Raw agreement is not enough for that. If humans prefer answer A in 53% of pairs, a judge that always says "A" agrees 53% of the time while knowing nothing. Cohen's kappa corrects for chance: observed agreement minus the agreement two independent raters with the same label frequencies would reach by luck, divided by the most that could be gained over luck. Zero is chance level; one is perfect.
Watch it run
The animation starts with two answers, one rubric, and no string match that could settle it, so a model grades. Answer A is the correct one, answer B the fluent one. A is shown first; the judge reads both and picks A. Then the order is swapped and exactly the same question is asked, with nothing else changed. B wins. The verdict tracked position, not quality. So every pair is graded twice, in both orders, and the two verdicts averaged, which turns the conflict into a tie and cancels the position bias. The other pull is length: the longer, more fluent answer wins even when it is the wrong one. Then the remaining fixes: hide which system produced which answer, and demand a reason before the verdict. Finally, the judge is measured like any other model, by its agreement with human labels, checked every release.
LLM as a Judge
Step 1 of 9
Two answers, one rubric, and no string match that could settle it. So a model grades.
The same interactive animation as the lesson — step through it with the controls.
The code
A toy model, with no real model inside: each answer has a hidden quality and a length, humans prefer the higher quality, and the "judge" sees quality through noise plus a pull towards the first slot and towards longer text. Every judgement is deterministic, so the numbers reproduce:
import random
def make_pairs(n, seed):
"""Toy data: each answer has a hidden quality and a length; humans label the better one."""
rng = random.Random(seed)
pairs = []
for i in range(n):
a = {"id": f"{i}a", "quality": rng.gauss(0, 1), "words": rng.randint(40, 400)}
b = {"id": f"{i}b", "quality": rng.gauss(0, 1), "words": rng.randint(40, 400)}
pairs.append((a, b, "a" if a["quality"] > b["quality"] else "b"))
return pairs
def judge(first, second, position=0.6, length=0.004):
"""Toy judge: sees quality through noise, plus a pull towards the first slot and longer text."""
noise = random.Random(first["id"] + second["id"]).gauss(0, 0.5)
margin = first["quality"] - second["quality"] + noise
margin += position + length * (first["words"] - second["words"])
return "first" if margin > 0 else "second"
def pick(a, b, **bias):
"""Show a first, b second; translate the slot back into a name."""
return "a" if judge(a, b, **bias) == "first" else "b"
def pick_swapped(a, b, **bias):
return "b" if judge(b, a, **bias) == "first" else "a"
pairs = make_pairs(400, seed=38)
human = [h for _, _, h in pairs]
once = [pick(a, b) for a, b, _ in pairs]
swapped = [pick_swapped(a, b) for a, b, _ in pairs]
agree = lambda xs: sum(x == h for x, h in zip(xs, human)) / len(human)
print(f"a shown first: {agree(once):.2f}", f"b shown first: {agree(swapped):.2f}")
print("flips when swapped:", sum(x != y for x, y in zip(once, swapped)))
print("first slot wins:", sum(x == "a" for x in once) + sum(y == "b" for y in swapped), "of 800")
# a shown first: 0.82 b shown first: 0.80
# flips when swapped: 119
# first slot wins: 519 of 800
Either order alone looks respectable at about 80% agreement, which is exactly why position bias goes unnoticed. Swapping exposes it: 119 of 400 verdicts flip, and the first slot wins 65% of the time. Next, judging in both orders, the length trap, and kappa:
def both_orders(a, b, **bias):
"""Judge twice, once per order. Agreement is a verdict; disagreement is a tie."""
x, y = pick(a, b, **bias), pick_swapped(a, b, **bias)
return x if x == y else "tie"
fair = [both_orders(a, b) for a, b, _ in pairs]
decided = [(v, h) for v, h in zip(fair, human) if v != "tie"]
print("ties:", fair.count("tie"), f"agreement when decided: {sum(v == h for v, h in decided) / len(decided):.2f}")
def longer_wrong(verdicts):
"""Pairs where the longer answer is the worse one: how often did it win anyway?"""
rows = [(v, a, b, h) for v, (a, b, h) in zip(verdicts, pairs) if v != "tie"]
traps = [(v, a, b) for v, a, b, h in rows if (a["words"] > b["words"]) != (h == "a")]
return sum((v == "a") == (a["words"] > b["words"]) for v, a, b in traps), len(traps)
print("longer but worse, still won:", longer_wrong(fair))
print("same judge, length=0:", longer_wrong([both_orders(a, b, length=0) for a, b, _ in pairs]))
def kappa(xs, ys):
"""Agreement corrected for the agreement two independent raters would reach by chance."""
n = len(xs)
observed = sum(x == y for x, y in zip(xs, ys)) / n
chance = sum((xs.count(k) / n) * (ys.count(k) / n) for k in set(xs) | set(ys))
return (observed - chance) / (1 - chance)
always_a = ["a"] * len(human)
print(f"always 'a': agreement {agree(always_a):.2f}, kappa {kappa(always_a, human):.2f}")
print(f"judge, one order: agreement {agree(once):.2f}, kappa {kappa(once, human):.2f}")
# ties: 119 agreement when decided: 0.94
# longer but worse, still won: (16, 112)
# same judge, length=0: (1, 129)
# always 'a': agreement 0.53, kappa 0.00
# judge, one order: agreement 0.82, kappa 0.63
Both orders turn the 119 flips into ties, and the remaining verdicts agree with humans 94% of the time. The ties are not waste: they are the close calls. Length still bites: when the longer answer is the worse one, it wins 16 of 112 decided pairs, against 1 of 129 for the same judge without the length pull. And kappa separates a useless rater, 53% agreement and kappa 0, from a real one.
The seeded check, on 300 random datasets and judges: the both-orders verdict must not depend on which answer is passed first, ties must equal flips, and kappa must match a brute-force chance rate averaged over every possible pairing of the two raters' labels:
from fractions import Fraction
mirror = {"a": "b", "b": "a", "tie": "tie"}
ok = True
for seed in range(300):
r = random.Random(seed)
data = make_pairs(r.randint(1, 40), seed)
bias = {"position": r.uniform(-1, 1), "length": r.uniform(0, 0.01)}
verdicts = [both_orders(a, b, **bias) for a, b, _ in data]
ok &= verdicts == [mirror[both_orders(b, a, **bias)] for a, b, _ in data] # order cannot matter
flips = sum(pick(a, b, **bias) != pick_swapped(a, b, **bias) for a, b, _ in data)
ok &= verdicts.count("tie") == flips
xs = [pick(a, b, **bias) for a, b, _ in data]
ys = [h for _, _, h in data]
n = len(xs)
chance = Fraction(sum(x == y for x in xs for y in ys), n * n) # every (rater 1, rater 2) pairing
if chance != 1:
observed = Fraction(sum(x == y for x, y in zip(xs, ys)), n)
ok &= abs(kappa(xs, ys) - float((observed - chance) / (1 - chance))) < 1e-9
print(ok) # True
The complexity
- Cost: one judge call per comparison; both orders double it. A round robin over
msystems ism(m - 1)/2pairs per test case, so compare against one baseline whenmgrows. - Human labels: a labelled sample per domain, refreshed when anything changes.
Where it goes wrong
- Trusting one order. Always measure the flip rate.
- Reporting raw agreement. A judge that always picks the majority label looks decent; report kappa too.
- Rewarding length. Put "longer is not better" in the rubric, and check wins against length.
- A judge grading its own family's output. Use a different judge, or measure self-preference directly.
- Never re-measuring. A new judge version is a new instrument.
When it shows up in interviews
In ML system design, as "how would you evaluate a chatbot or a summariser at scale?", and as the follow-up to proposing a judge: "how do you know the judge is right?". Have the biases, both-orders grading and agreement against human labels ready. As of October 2026, pairwise judging with order swaps is a common pattern in public model comparisons; studies of judges have reported position, length and self-preference effects (from memory).
How to say it in an interview
"For open-ended outputs I'd use a model as a judge, pairwise against a rubric rather than a 1 to 10 score, which clusters. Judges have known biases: position, length and self-preference. So I'd judge each pair in both orders and count disagreement as a tie, hide which system wrote which answer, and ask for a reason before the verdict. Then I'd measure the judge itself: agreement and Cohen's kappa against human labels on a held-out set, re-checked whenever the judge, rubric or domain changes."