Message Queues Explained: At-Least-Once, Idempotency, Dead Letters
8 min readBytePatterns
Message queues explained: why producers stop waiting, how acks and redelivery give at-least-once delivery, why consumers must be idempotent, and what a DLQ is.
Almost every system design answer eventually says "put a queue in front of it". The phrase is cheap; the follow-up questions are not. What happens when a worker crashes halfway through a message? Can the same message be processed twice? What stops one malformed payload from jamming everything behind it? A queue is a small idea with three consequences, and a good answer names all three.
The problem it solves
Without a queue, the service that creates work calls the service that does it and waits. Every slow job becomes a slow request, every burst becomes timeouts, and the two services have to be up, scaled and deployed together.
A message queue sits between a producer and one or more consumers:
- The producer writes a message and returns as soon as the queue has accepted it.
- Consumers pull messages when they have capacity and process them at their own pace.
- A burst becomes a backlog instead of a wave of failures, and the backlog is a number you can graph and alert on.
- Either side can be restarted or scaled without the other noticing.
Typical uses: emails and notifications, resizing uploads, charging cards after checkout, analytics, and any work the page need not wait for.
The intuition
The design question is what happens between "a consumer received the message" and "the work is done". Queues answer it with an acknowledgement. Receiving a message does not delete it; it hides it from other consumers for a while. Only an explicit ack removes it. If no ack arrives in time, because the consumer crashed, hung or lost its connection, the message becomes visible again and another consumer gets it.
That rule is what makes a queue reliable, and it has a price. A consumer can finish the work and crash before its ack lands. The message is then delivered again, and the work is done twice. This is at-least-once delivery, the default guarantee of most queues. Exactly-once processing is not something the queue can give you on its own, because the queue cannot see whether your side effect happened.
So the consumer has to make a second delivery harmless. That property is idempotency: processing a message twice has the same effect as processing it once. Common ways to get it are recording processed message ids, using upserts keyed by a business id, or making the operation naturally repeatable, such as "set status to shipped" rather than "add 10 to balance".
The last consequence is the poison message, a payload that fails every time, for instance because it cannot be parsed. Redelivery would retry it forever. Queues count delivery attempts and, past a limit, move the message to a dead-letter queue, where a person or a separate process can look at it while the main queue keeps moving.
As of September 2026, the big managed and open-source brokers all expose these ideas under their own names: a visibility timeout or unacknowledged-message redelivery, a maximum receive count, and a dead-letter target.
Watch it run
The animation starts with a queue between the side that makes work and the side that does it. The producer writes a message and returns to its caller straight away. A worker pulls the message once it has spare capacity, not before; it finishes, acknowledges, and the message is removed for good. Then a burst arrives: four hundred jobs in ten seconds. Nothing times out; the backlog piles up visibly, and the producer never waited. Add workers and the queue drains at whatever rate you are willing to pay for. Then the guarantees: at-least-once delivery means the same message can arrive twice, so processing it twice has to be harmless, and consumers must be idempotent. One payload keeps failing; retried three times, it is still poison, so it is parked in a dead-letter queue and cannot block the pipeline forever. The last frame is the caveat: if the caller genuinely needs the answer now, a queue only hides the waiting somewhere less visible.
Message Queues
Step 1 of 12
A queue sits between the side that makes work and the side that does it.
The same interactive animation as the lesson — step through it with the controls.
The code
A toy model, not a broker: a ready list, an in-flight map, attempt counts and a dead-letter list are enough to show acks, redelivery and poison handling:
from collections import deque
class ToyQueue:
"""Toy model: at-least-once delivery with acks, redelivery and a dead-letter queue."""
def __init__(self, max_attempts=3):
self.ready = deque()
self.in_flight = {} # message id -> message, delivered but not acked
self.attempts = {}
self.dead = []
self.max_attempts = max_attempts
def publish(self, msg_id, body):
self.ready.append((msg_id, body)) # the producer returns right here
def receive(self):
if not self.ready:
return None
msg = self.ready.popleft()
self.in_flight[msg[0]] = msg # hidden, not deleted
self.attempts[msg[0]] = self.attempts.get(msg[0], 0) + 1
return msg
def ack(self, msg_id):
self.in_flight.pop(msg_id) # only now is it gone for good
def timeout(self, msg_id):
msg = self.in_flight.pop(msg_id) # no ack in time: deliver it again
if self.attempts[msg_id] >= self.max_attempts:
self.dead.append(msg) # poison: park it, keep the line moving
else:
self.ready.append(msg)
def run(q, handle, crash_after_work):
"""Deliver until empty. crash_after_work(msg_id) says the worker died before acking."""
while (msg := q.receive()) is not None:
try:
handle(msg)
except ValueError:
q.timeout(msg[0]) # a failing payload is never acked
continue
if crash_after_work(msg[0]):
q.timeout(msg[0]) # the work happened, the ack did not
else:
q.ack(msg[0])
Three payments, and the worker dies once, right after applying the second one. The naive consumer charges it twice; the idempotent one remembers what it has applied:
balance, seen = {"naive": 0, "idempotent": 0}, set()
def naive(msg):
balance["naive"] += msg[1]
def idempotent(msg):
if msg[0] in seen: # already applied: a redelivery is harmless
return
seen.add(msg[0])
balance["idempotent"] += msg[1]
crashed = set()
def crash_once_on_2(msg_id): # the worker dies once, after doing message 2
if msg_id == 2 and msg_id not in crashed:
crashed.add(msg_id)
return True
return False
for handler in (naive, idempotent):
q, crashed = ToyQueue(), set()
for i, amount in enumerate([10, 20, 30], start=1):
q.publish(i, amount)
run(q, handler, crash_once_on_2)
print(balance) # {'naive': 80, 'idempotent': 60}
def poison_aware(msg):
if msg[1] == "not-a-number":
raise ValueError("cannot parse")
idempotent(msg)
q = ToyQueue()
for msg in [(4, 5), (5, "not-a-number"), (6, 7)]:
q.publish(*msg)
run(q, poison_aware, lambda _: False)
print(q.dead, q.attempts[5], balance["idempotent"]) # [(5, 'not-a-number')] 3 72
def drain_seconds(backlog, workers, per_worker_per_second):
return backlog / (workers * per_worker_per_second)
print(drain_seconds(400, 2, 5), drain_seconds(400, 8, 5)) # 40.0 10.0
Checked on 500 seeded random runs with random crash rates and occasional poison payloads: the idempotent total must equal applying every good message once, every poison message must be dead-lettered, nothing may be left ready or in flight, and the naive consumer must double-count at least once:
import random
random.seed(27)
ok, redelivered_runs = True, 0
for _ in range(500):
msgs = [(i, "not-a-number" if random.random() < 0.05 else random.randint(1, 100))
for i in range(random.randint(1, 30))]
crash_rate = random.random() * 0.5
totals = {}
for name in ("naive", "idempotent"):
q, applied, done = ToyQueue(), set(), []
def handle(msg):
if msg[1] == "not-a-number":
raise ValueError
if name == "naive" or msg[0] not in applied:
applied.add(msg[0])
done.append(msg[1])
for m in msgs:
q.publish(*m)
run(q, handle, lambda _: random.random() < crash_rate)
totals[name] = sum(done)
ok &= not q.ready and not q.in_flight
ok &= {m for m in msgs if m[1] == "not-a-number"} <= set(q.dead)
exactly_once = sum(b for _, b in msgs if b != "not-a-number")
ok &= totals["idempotent"] == exactly_once
redelivered_runs += totals["naive"] != exactly_once
print(ok, redelivered_runs > 0) # True True
The complexity
- Publish, receive, ack:
O(1)each in the queue itself. - Drain time: backlog divided by total consumer throughput. Doubling workers halves it, until a downstream dependency becomes the bottleneck.
- Idempotency memory: one stored id per processed message, usually kept only for as long as redelivery is possible.
Where it goes wrong
- Assuming exactly-once. Design every consumer for duplicates.
- Acking before the work. Then a crash loses the message instead of duplicating it: at-most-once, usually the worse failure.
- A visibility timeout shorter than the work. Healthy slow consumers get their messages redelivered to someone else while still working.
- Expecting global order with parallel consumers. Order holds per queue or per partition key at best, not across workers.
- No dead-letter queue, or one nobody watches. Poison messages then either loop forever or vanish silently.
- An unbounded backlog. A queue absorbs bursts, not a permanent shortfall; see producer and consumer for why bounds matter.
When it shows up in interviews
Queues appear inside larger designs: the delivery pipeline of a notification system, video transcoding, order processing, webhooks. Follow-ups test the three consequences above, plus the difference between a work queue, where one consumer gets each message, and publish-subscribe, where every subscriber does, which SQS vs SNS vs EventBridge covers for one cloud.
How to say it in an interview
"The producer publishes and returns, and workers pull at their own pace, so bursts become a visible backlog instead of timeouts, and both sides scale independently. A received message is hidden, not deleted, until the consumer acks it; if the consumer dies first, it is redelivered. That gives at-least-once delivery, so consumers must be idempotent, for example by recording processed message ids. Messages that fail repeatedly go to a dead-letter queue after a few attempts. And if the caller needs the result synchronously, I would not use a queue at all."