Speculative Decoding
AI & ML: lesson 25 of 32
A small model guesses ahead; the big one checks in one pass.
Lesson 25 of 32 · 5 min
Speculative Decoding
Step 1 of 9
Producing one token at a time is sequential. Each one has to wait for the one before it.
The Idea
Producing tokens one at a time is sequential and slow. Checking a sequence you already have is parallel and fast. So a small draft model proposes several tokens cheaply, and the large model verifies them all in one forward pass — keeping the correct prefix and correcting the first mistake.
Real-World Example
A junior surveyor pegs out the whole boundary line by eye. The senior walks it once, agrees with the first four pegs, and moves the fifth. One walk, several pegs placed.
The Tradeoff
Your speedup is set by the acceptance rate and by how cheap the drafter is. A drafter too weak is rejected constantly and you pay for both models; a drafter strong enough to be always right is nearly as slow as the target. Both models sit in memory, which is exactly the resource serving is already short of.
Your turn
Put the steps in the right order.
- The target model scores all proposed tokens in one pass
- The small draft model proposes four tokens ahead
- Keep the accepted prefix and correct the first rejected token
- Discard everything after the correction and draft again
Mini quiz
1 / 3