Skip to content
BytePatterns

Design a Job Scheduler

System Design Cases: lesson 11 of 20

The worker died mid-job. The ticket goes back on the rail.

Lesson 11 of 20 · 7 min

Design a Job Scheduler

Step 1 of 11

Five million jobs a day, and the one thing you cannot promise is that the worker survives the job.

The Idea

Accept a job now, run it later, and survive the worker that dies halfway through. Assume five million jobs a day, a spike on every hour boundary, and a promise that nothing quietly disappears.

Real-World Example

A kitchen ticket rail. A ticket stays clipped until a cook takes it down, and if that cook walks out mid-service the ticket returns to the rail instead of vanishing with them.

claim:  UPDATE jobs SET lease_until = now()+30s, owner = me
        WHERE run_at <= now() AND lease_until < now() LIMIT 10
run:    heartbeat extends lease_until every 10s
ok:     delete the row
fail:   attempts += 1, run_at = now() + 2^attempts
        attempts >= 5 -> move to dead_letter

The Tradeoff

A lease with a heartbeat makes a crashed worker recoverable and makes duplicate runs possible, so every handler has to be safe to run twice. Retries back off to spare whatever failed, and a job that exhausts them is parked rather than left to block the queue behind it.

Your turn

Put the steps in the right order.

  1. After repeated failures, back off and park the job in the dead-letter queue
  2. Store the job with its run-at time and hand back an id
  3. Heartbeat while the handler runs, then delete the row on success
  4. Claim due jobs by taking a short lease over them

Mini quiz

1 / 3

A worker's lease expires while its handler is still running, so the scheduler must assume:

New lessons land every few weeks

Leave an address and we will tell you when the next one is up. That is the only reason we will use it.

One address, stored so we can email you. Nothing else, ever.