The KV Cache
AI & ML: lesson 24 of 32
Keep the past keys and values so each new token is cheap.
Lesson 24 of 32 · 5 min
The KV Cache
Step 1 of 9
Four tokens are in. To produce the fifth, attention needs every earlier token's key and value.
The Idea
Attention at each step looks back over every earlier token's key and value. Those depend only on the prefix, so they never change. Keeping them turns each step from "redo the whole prefix" into "project one new token and read the cache".
Real-World Example
A court stenographer who keeps the running transcript beside them. Each new line is added to the pile; nobody retypes the morning's testimony to record the afternoon's.
The Code
n = 100
redo = sum(range(1, n + 1)) # every step re-projects the whole prefix
cached = n # every step projects exactly one token
print(redo, cached, round(redo / cached, 1)) # 5050 100 50.5
layers, heads, dim, width = 32, 32, 128, 2 # width = bytes per number
per_token = 2 * layers * heads * dim * width # keys and values
print(per_token, round(per_token * 8192 / 1e9, 2)) # 524288 4.29Your turn
Fill in the blank.
n = 6
redo = sum(range(1, n + 1))
print(redo, n) # ___ 6Mini quiz
1 / 3