LLM Temperature, Top-k and Top-p Sampling Explained
8 min readBytePatterns
LLM temperature, top-k and top-p explained with real numbers: how dividing the logits sharpens or flattens softmax, and how each filter trims the unlikely tail.
Ask a language model the same question twice and you can get two different answers. That is a setting, not a bug. At every step the model scores every possible next token, and a sampling procedure picks one. Temperature reshapes the scores before the pick; top-k and top-p cut away the unlikely tail. Interviews for LLM and machine learning roles ask what these knobs do to the numbers, because "higher temperature means more creative" is where most answers stop.
The problem it solves
A model's raw output at each step is a vector of logits, one real number per token in the vocabulary. Softmax turns them into probabilities: exponentiate each, divide by the total. Then something has to choose.
Always taking the most likely token, greedy decoding, is deterministic but tends to be repetitive. Sampling straight from the distribution gives variety, but the long tail occasionally produces nonsense, and over hundreds of steps "occasionally" becomes "somewhere in every answer". The knobs choose a point between those failures.
The intuition
Temperature divides every logit by t before softmax.
- With
t = 1, nothing changes: this is the model's own distribution. - With
tbelow 1, dividing by a small number stretches the gaps between logits, and exponentiation turns bigger gaps into much bigger probability ratios. The favourite takes nearly everything. Astapproaches zero, sampling becomes greedy decoding. - With
tabove 1, the gaps shrink and the distribution flattens toward uniform, so rare tokens get a real chance.
The ranking never changes; only how confident the distribution is. The ratio between two tokens' probabilities is exp((a - b) / t), which is the whole story in one formula.
Top-k keeps the k most likely tokens, zeroes the rest and renormalises. But k is fixed: when the model is confident, k = 50 still admits 49 bad options; when it is unsure, 50 may be too few.
Top-p, also called nucleus sampling, keeps the smallest set of tokens whose probabilities add up to at least p. The cut follows the shape of the distribution: one token when the model is sure, many when it is not. That adaptivity is why top-p is the common default companion to temperature.
Order matters in practice: temperature reshapes, filters trim, then one token is drawn. Hosted APIs expose these as request parameters, but the allowed ranges, the defaults and whether you may combine them differ by provider, so check the current documentation (as of September 2026). And temperature 0 is best read as "as deterministic as possible": batching and floating-point effects mean hosted models are not always bit-for-bit reproducible.
Watch it run
The animation uses the lesson's own softmax over the logits 2.0, 1.0 and 0.0. The model returns a score per token, and softmax turns them into probabilities that add up to one. At t = 1.0 that is the model's own distribution: 0.665, 0.245, 0.090. Temperature rescales the scores before that conversion, so the dial turns down: 0.8, then 0.5, where the favourite takes 0.867, since dividing by a small number widens every gap. Lower still and it is effectively deterministic, the same answer every time. Then the other way, back through the model's own setting and up to t = 2.0, where the field flattens to 0.506, 0.307 and 0.186; rare tokens get a real chance, and loosened too far, the lesson's tour guide starts inventing history. Sampling also trims the tail. Top-k keeps a fixed count, here the two highest, whatever they are. Top-p instead keeps the smallest set whose probabilities add up to p, so the cut follows the shape; at p = 0.8 that is two of the three. The last frames tie it to use: pulling a field into a fixed format wants a very low temperature, and fresh phrasing wants a higher one. One dial, two jobs.
Temperature and Sampling
Step 1 of 13
The model returns a score per token. These three are 2.0, 1.0 and 0.0.
The same interactive animation as the lesson — step through it with the controls.
The code
A numerically stable softmax with temperature. Subtracting the largest logit first changes nothing mathematically and stops exp from overflowing at small temperatures, which the naive version does:
import math
def softmax(scores, t=1.0):
top = max(scores) # subtract the max: exp never overflows
exp = [math.exp((s - top) / t) for s in scores]
total = sum(exp)
return [e / total for e in exp]
def show(probs):
return [round(p, 3) for p in probs]
logits = [2.0, 1.0, 0.0]
for t in (0.2, 0.5, 1.0, 2.0):
print(t, show(softmax(logits, t)))
# 0.2 [0.993, 0.007, 0.0]
# 0.5 [0.867, 0.117, 0.016]
# 1.0 [0.665, 0.245, 0.09]
# 2.0 [0.506, 0.307, 0.186]
def naive_softmax(scores, t):
exp = [math.exp(s / t) for s in scores]
return [e / sum(exp) for e in exp]
try:
naive_softmax(logits, 0.001) # exp(2000) does not fit in a float
except OverflowError:
print("OverflowError") # OverflowError
print(show(softmax(logits, 0.001))) # [1.0, 0.0, 0.0] the limit is plain argmax
Top-k and top-p as filters that zero out the tail and renormalise. At t = 2 the top two tokens hold about 0.814 of the mass, so p = 0.8 keeps two and p = 0.9 needs all three. On a confident distribution, top-p keeps one token where top-k still keeps two:
def renormalise(probs, keep):
mass = sum(probs[i] for i in keep)
return [probs[i] / mass if i in keep else 0.0 for i in range(len(probs))]
def top_k(probs, k):
keep = sorted(range(len(probs)), key=lambda i: -probs[i])[:k]
return renormalise(probs, keep)
def top_p(probs, p):
keep, running = [], 0.0
for i in sorted(range(len(probs)), key=lambda i: -probs[i]):
keep.append(i) # smallest set whose mass reaches p
running += probs[i]
if running >= p:
break
return renormalise(probs, keep)
hot = softmax(logits, 2.0)
print(show(top_k(hot, 2))) # [0.622, 0.378, 0.0]
print(show(top_p(hot, 0.8))) # [0.622, 0.378, 0.0]
print(show(top_p(hot, 0.9))) # [0.506, 0.307, 0.186]
peaked = softmax([6.0, 1.0, 0.5, 0.0])
print(show(peaked)) # [0.987, 0.007, 0.004, 0.002]
print(show(top_p(peaked, 0.9)), show(top_k(peaked, 2)))
# [1.0, 0.0, 0.0, 0.0] [0.993, 0.007, 0.0, 0.0]
What that means for the text: 10,000 seeded draws of the next word at two temperatures. At 0.5 the favourite wins about 86 per cent of the time; at 2.0, only about half:
import random
random.seed(25)
words = ["the", "a", "one"]
for t in (0.5, 2.0):
draws = random.choices(words, weights=softmax(logits, t), k=10_000)
print(t, {w: draws.count(w) for w in words})
# 0.5 {'the': 8618, 'a': 1229, 'one': 153}
# 2.0 {'the': 5021, 'a': 3110, 'one': 1869}
Checked on 2,000 seeded random logit vectors: the stable softmax must match the naive one, sum to one and keep the ranking; a higher temperature must never lower the entropy; and top-p must keep exactly as many tokens as a brute-force search over every subset for the smallest one holding mass p:
from itertools import combinations
def entropy(probs):
return -sum(p * math.log(p) for p in probs if p > 0)
def smallest_set_size(probs, p):
for size in range(1, len(probs) + 1):
if any(sum(c) >= p for c in combinations(probs, size)):
return size
ok = True
for _ in range(2_000):
scores = [random.uniform(-5, 5) for _ in range(random.randint(1, 6))]
t1, t2 = sorted(random.uniform(0.05, 5) for _ in range(2))
a, b = softmax(scores, t1), softmax(scores, t2)
ok &= abs(sum(a) - 1) < 1e-9 and all(x >= 0 for x in a)
ok &= max(abs(x - y) for x, y in zip(a, naive_softmax(scores, t1))) < 1e-9
ok &= sorted(range(len(scores)), key=lambda i: scores[i]) == sorted(range(len(a)), key=lambda i: a[i])
ok &= entropy(a) <= entropy(b) + 1e-9 # hotter never means more certain
p = random.uniform(0.05, 0.99)
ok &= sum(x > 0 for x in top_p(a, p)) == smallest_set_size(a, p)
k = random.randint(1, len(scores))
ok &= sum(x > 0 for x in top_k(a, k)) == k
print(ok) # True
The complexity
- Temperature: one division per logit,
O(V)for a vocabulary ofVtokens. - Top-k:
O(V log k)with a heap, or a partial sort. - Top-p: a sort,
O(V log V), then a prefix sum until the mass reachesp. - Relative to the model: negligible next to a forward pass.
Where it goes wrong
- Dividing by zero.
t = 0is special-cased as greedy decoding, never computed. - Unstable softmax. Exponentiating large logits, or logits over a small
t, overflows; subtract the maximum first. - Expecting temperature to fix facts. It changes how the model chooses among its options, not what it knows. Low temperature makes a wrong answer consistent.
- Tuning everything at once. Temperature and top-p both control diversity; change one at a time.
- High temperature for structured output. Extraction and code want low temperature; one stray token breaks a format.
When it shows up in interviews
Expect it in LLM application and ML engineering interviews: "why does the model give different answers?", "what would you set for a JSON extractor versus a brainstorming tool?", "what is the difference between top-k and top-p?". It connects to tokenization, since sampling happens over tokens rather than words, and to attention, which is what produces the logits in the first place.
How to say it in an interview
"The model outputs a logit per token; softmax turns them into probabilities, and sampling picks one. Temperature divides the logits before softmax: below one it widens the gaps, so the top token dominates, and near zero it becomes greedy; above one it flattens the distribution. It never changes the ranking. Top-k keeps a fixed number of tokens; top-p keeps the smallest set whose mass reaches p, so it adapts to how confident the model is. For extraction I use a low temperature; for varied drafts a moderate one with top-p."