LLM Context Window Explained: Token Budget and Truncation
8 min readBytePatterns
What an LLM context window is: one token budget shared by instructions, history and the reply, why old turns fall off, and how to trim, pin and summarise.
A language model has no memory between calls. Every time a chat application answers, it sends the instructions, the conversation so far, any retrieved documents and the new question, all over again. The context window is the hard limit on how much of that fits: the most tokens a model can attend to in one call. Anything outside it does not exist for the model, which is why long conversations "forget" their beginning.
The problem it solves
Understanding the window explains three everyday problems:
- Forgetting. A chatbot that ignores a rule you gave it an hour ago usually never saw it on this call.
- Cost and latency. You are billed per token sent, and the whole history is resent every turn, so long conversations get more expensive with every message.
- Failed calls. Input plus requested output over the limit is rejected or truncated, depending on the provider.
The engineering question is "what goes in, and what is dropped when it is full?"
The intuition
Think of the window as one budget shared by everything in the call: the system instructions, the conversation history, retrieved passages, tool results, and the reply the model is about to write. The reply is the part most often forgotten: if the input fills the window, there is no room left to answer.
Budgets are counted in tokens, not words or characters. A token is a chunk of text from the model's tokenizer, often a word or a piece of one, and the same text costs different numbers of tokens under different tokenizers. As of September 2026, advertised windows range from a few thousand tokens to around a million, and providers usually cap the reply length separately as well.
When history outgrows the budget, something must leave, and the only question is whether you choose it:
- Keep the newest turns, walking backwards until the budget is spent: the lesson's approach.
- Pin what must survive, above all the system instructions, before filling the rest.
- Reserve the reply's share before filling anything.
- Summarise older turns into a short note instead of dropping them.
- Retrieve instead of stuffing: store past material outside the window and fetch only the relevant pieces, which is what retrieval-augmented generation does.
A bigger window is not a free fix. Every token sent is paid for on every call, the attention computation grows with the length of the input, and research has found that models tend to use information at the start and end of a long context more reliably than information buried in the middle.
Watch it run
The animation walks the lesson's six turns against a 700-token window, a budget that instructions, history, retrieved passages and the reply all share. Six turns so far: together they are 1,270 tokens, nearly twice the window. So walk backwards from the newest turn and keep what still fits: 120. The assistant turn before it is 400, which still fits at 520. Add the 180-token user turn and the window is exactly full at 700. The next turn is 250 more; 950 is over budget, so the walk stops. Everything older is dropped, including the system turn, quietly, and for this call those turns do not exist: the model sees three of six. Then the catch: the reply is charged to the same budget, so reserving 120 tokens for it squeezes the history further, down to two turns. The fix is to choose what leaves: summarise older turns into 60 tokens instead of letting them fall off, which fills the window exactly with the reply's share intact. Last, a much bigger window is not free: you pay for every token on every call, and attention costs more too.
Context Windows
Step 1 of 12
The context window is the most tokens a model can attend to in one call.
The same interactive animation as the lesson — step through it with the controls.
The code
A toy model: turns are (role, tokens) pairs with given token counts. The lesson's newest-first loop drops the system turn; the second version pins it and reserves the reply's share first:
history = [("system", 20), ("user", 300), ("assistant", 250),
("user", 180), ("assistant", 400), ("user", 120)]
def newest_first(history, budget):
"""The lesson's loop: keep the longest run of newest turns that fits."""
kept, used = [], 0
for role, size in reversed(history):
if used + size > budget:
break
kept.append((role, size))
used += size
return list(reversed(kept)), used
def fit(history, window, reply_budget):
"""Pin the system turn, reserve room for the reply, then keep the newest turns."""
system, turns = history[0], history[1:]
budget = window - reply_budget - system[1]
if budget < 0:
raise ValueError("the instructions and the reply alone overflow the window")
kept, used = newest_first(turns, budget)
return [system] + kept, used + system[1]
kept, used = newest_first(history, 700)
print([r for r, _ in kept], used) # ['user', 'assistant', 'user'] 700
kept, used = fit(history, 700, reply_budget=120)
print([r for r, _ in kept], used) # ['system', 'assistant', 'user'] 540
The pinned version keeps the instructions and leaves 40 tokens unused, because the next-oldest turn (180) does not fit. Now the cost side. The model is stateless, so turn k resends all k - 1 earlier turns: the tokens sent over a conversation grow with the square of its length, unless older turns are folded into a summary:
def tokens_sent(n_turns, per_turn, summarise_every=None, summary=60):
"""Total input tokens over a conversation where every call resends the history."""
total, live = 0, 0
for k in range(1, n_turns + 1):
live += per_turn # the new message joins the history
total += live # ...and the whole history is sent again
if summarise_every and k % summarise_every == 0:
live = summary # fold everything so far into a short note
return total
for n in (10, 20, 40):
print(n, tokens_sent(n, 200), tokens_sent(n, 200, summarise_every=10))
# 10 11000 11000
# 20 42000 22600
# 40 164000 45800
Doubling the conversation quadruples the tokens; with summaries, growth is close to linear. Checked on 3,000 seeded random histories: the lesson's loop equals a brute force that tries every suffix and keeps the longest that fits, the pinned version never exceeds the window, and per_turn · n(n + 1) / 2 matches the simulation:
import random
def longest_fitting_suffix(turns, budget):
"""Brute force: try every suffix, keep the longest whose total fits."""
for start in range(len(turns) + 1):
if sum(size for _, size in turns[start:]) <= budget:
return turns[start:]
rng = random.Random(31)
ok = True
for _ in range(3_000):
turns = [("system", rng.randint(5, 80))] + [
(rng.choice(["user", "assistant"]), rng.randint(1, 400)) for _ in range(rng.randint(0, 12))]
window, reply = rng.randint(100, 1_500), rng.randint(0, 300)
kept, used = newest_first(turns, window)
ok &= kept == longest_fitting_suffix(turns, window) and used == sum(s for _, s in kept)
if turns[0][1] + reply <= window:
kept, used = fit(turns, window, reply)
ok &= kept[0] == turns[0] and used + reply <= window
n, per = rng.randint(1, 60), rng.randint(1, 500)
ok &= tokens_sent(n, per) == per * n * (n + 1) // 2
print(ok) # True
The complexity
- Fitting the history is one backwards pass:
O(t)fortturns, after counting tokens once per turn. - Tokens sent per conversation grow as
O(n²)in the number of turns without summarisation, and roughlyO(n)with periodic summaries. - Model-side cost grows with context length: standard self-attention does work proportional to the square of the input length, and the KV cache that speeds up generation grows linearly with it.
Where it goes wrong
- Dropping the instructions. Newest-first truncation silently removes the system prompt first. Pin it.
- No room for the answer. Filling the window with input leaves the reply truncated or the call rejected. Reserve the output budget up front.
- Counting words, not tokens. Code and non-English text often cost more tokens per word. Use the model's tokenizer.
- Assuming big windows mean perfect recall. Facts buried mid-context can be missed. Put key material near the start or end.
- Summaries that lose facts. A summary that drops "the user is on the paid plan" is a bug.
When it shows up in interviews
In machine learning interviews as "what is a context window?" or "how would you handle a conversation longer than the window?", and in system design prompts for chat assistants and RAG pipelines, where the follow-ups are token budgets, cost per conversation, and what to summarise or retrieve.
How to say it in an interview
"The context window is the maximum number of tokens the model can attend to in one call, and everything shares it: system prompt, history, retrieved context, and the output. Because the model is stateless, I resend what matters on every call, so I'd budget explicitly: pin the system prompt, reserve the reply's share, then fill with the newest turns. Older turns get summarised or moved to a store and retrieved when relevant. I'd count with the real tokenizer, and I wouldn't just buy a bigger window, because every token costs money and latency on every call and recall in the middle of a long context is weaker."