Skip to content
BytePatterns

Speculative Decoding

AI & ML: lesson 25 of 32

A small model guesses ahead; the big one checks in one pass.

Lesson 25 of 32 · 5 min

Speculative Decoding

Step 1 of 9

Producing one token at a time is sequential. Each one has to wait for the one before it.

The Idea

Producing tokens one at a time is sequential and slow. Checking a sequence you already have is parallel and fast. So a small draft model proposes several tokens cheaply, and the large model verifies them all in one forward pass — keeping the correct prefix and correcting the first mistake.

Real-World Example

A junior surveyor pegs out the whole boundary line by eye. The senior walks it once, agrees with the first four pegs, and moves the fifth. One walk, several pegs placed.

The Tradeoff

Your speedup is set by the acceptance rate and by how cheap the drafter is. A drafter too weak is rejected constantly and you pay for both models; a drafter strong enough to be always right is nearly as slow as the target. Both models sit in memory, which is exactly the resource serving is already short of.

Your turn

Put the steps in the right order.

  1. The target model scores all proposed tokens in one pass
  2. The small draft model proposes four tokens ahead
  3. Keep the accepted prefix and correct the first rejected token
  4. Discard everything after the correction and draft again

Mini quiz

1 / 3

The speedup comes from the fact that verifying several tokens:

New lessons land every few weeks

Leave an address and we will tell you when the next one is up. That is the only reason we will use it.

One address, stored so we can email you. Nothing else, ever.