Skip to content
BytePatterns

KV Cache Memory Budget

EasyAI & ML#capacity-math#memory-estimate~10m

Problem

While a transformer generates text, it keeps a key vector and a value vector for every earlier token, in every layer and every key-value head, so that each new token does not have to recompute them. Given the number of layers, key-value heads, the size of each head, the number of cached tokens, the batch size and the bytes per stored number, return the memory the cache needs in MiB (2^20 bytes). Serving capacity is planned from exactly this number, since it grows with every token of every conversation in the batch.

Examples

Input:  layers = 32, kv_heads = 32, head_dim = 128, tokens = 4096, batch = 1, 2 bytes per number
Output: 2048.0
Why:    512 KiB per token, times 4,096 tokens
Input:  layers = 32, kv_heads = 8, head_dim = 128, tokens = 4096, batch = 1, 2 bytes per number
Output: 512.0
Why:    sharing each key-value head across 4 query heads cuts the cache by 4
Input:  layers = 32, kv_heads = 8, head_dim = 128, tokens = 0
Output: 0.0
Why:    edge case, an empty prompt has nothing cached yet

Hints

0 / 3

Stuck on the idea rather than the code? The KV Cache covers it.