Design a Job Scheduler
System Design Cases: lesson 11 of 20
The worker died mid-job. The ticket goes back on the rail.
Lesson 11 of 20 · 7 min
Design a Job Scheduler
Step 1 of 11
Five million jobs a day, and the one thing you cannot promise is that the worker survives the job.
The Idea
Accept a job now, run it later, and survive the worker that dies halfway through. Assume five million jobs a day, a spike on every hour boundary, and a promise that nothing quietly disappears.
Real-World Example
A kitchen ticket rail. A ticket stays clipped until a cook takes it down, and if that cook walks out mid-service the ticket returns to the rail instead of vanishing with them.
claim: UPDATE jobs SET lease_until = now()+30s, owner = me
WHERE run_at <= now() AND lease_until < now() LIMIT 10
run: heartbeat extends lease_until every 10s
ok: delete the row
fail: attempts += 1, run_at = now() + 2^attempts
attempts >= 5 -> move to dead_letter
The Tradeoff
A lease with a heartbeat makes a crashed worker recoverable and makes duplicate runs possible, so every handler has to be safe to run twice. Retries back off to spare whatever failed, and a job that exhausts them is parked rather than left to block the queue behind it.
Your turn
Put the steps in the right order.
- After repeated failures, back off and park the job in the dead-letter queue
- Store the job with its run-at time and hand back an id
- Heartbeat while the handler runs, then delete the row on success
- Claim due jobs by taking a short lease over them
Mini quiz
1 / 3