Skip to content
BytePatterns

Don't Block the Event Loop: Find and Fix Blocking Calls

9 min readBytePatterns

Don't block the event loop: why one slow call delays every task, how debug mode names the culprit, and fixes: await, yield in slices, or move it off the loop.

An async service handles ten thousand connections on one thread, and then one request makes all ten thousand wait. The cause is almost always the same: a call that holds the thread instead of handing it back. The async/await article explains the loop itself. This one is the field guide to its most common failure: what the damage really is, how to find the call responsible, and the three ways to fix it.

The problem it solves

An event loop is cooperative. Control changes hands only at an await on something that is not ready yet. Between two awaits, the running task owns the only thread, and every other task, timer and incoming request waits.

That gives a precise rule for the damage: every task's latency includes the longest stretch any task runs without yielding. A handler that computes for 200 ms adds up to 200 ms to every request that arrives meanwhile, including the trivial health check, which may then time out and get a healthy server restarted.

The usual suspects:

  • Blocking I/O written for synchronous code: time.sleep, a synchronous HTTP client, a synchronous database driver, reading a large file.
  • CPU work: parsing a huge JSON payload, hashing a password, resizing an image, a long loop.
  • Hidden blocking: a logging handler that writes over the network, or a library call that does I/O internally.

The intuition

Three fixes, chosen by what the slow call is:

  1. Await the async version. For I/O, use an async client, so the wait happens with the thread handed back.
  2. Yield in slices. For a long loop you control, insert await asyncio.sleep(0) every few thousand iterations. The total work is unchanged, but nobody waits longer than one slice.
  3. Move it off the loop. asyncio.to_thread or run_in_executor runs the call on another thread while the loop keeps going. For pure-Python CPU work in standard CPython, a process pool also runs it in parallel, since threads share one interpreter lock (as of October 2026; free-threaded builds change this, from memory).

Finding the call comes first. asyncio has a debug mode that logs any step that holds the thread longer than slow_callback_duration, 0.1 seconds by default, naming the task responsible. In production, the same idea is a monitor that schedules a tick at a fixed interval and records how late it fires: event-loop lag.

Watch it run

The animation draws one lane, because there is one thread. Whatever is running owns it until it gives it back. Task A runs until it hits an await, then hands the thread straight back, so B starts while A is still waiting: that is where the concurrency comes from. Both waits overlap, so two one-tenth-second sleeps take one tenth of a second. Then the wait becomes a blocking one. Nothing yields, so A holds the thread, and B sits in the ready queue the whole time, unable to start, because there is no second thread to start it on. Only when A finishes does B get a turn, so the two sleeps add up to 0.2 seconds instead. A tight computation, here parsing 200 MB, does exactly the same, and the ready queue fills with B, a timer and another request. The fix: push that work to a thread or process pool, and the loop is free and responsive again.

Blocking the Event Loop

Step 1 of 9

An event loop is one thread. Whatever is running owns it until it gives it back.

The same interactive animation as the lesson — step through it with the controls.

The code

The lesson's race, printed as comparisons so the output does not depend on the machine, then debug mode catching the culprit by name:

import asyncio
import logging
import time

async def gauge(depth):
    await asyncio.sleep(0.1)              # yields the thread to the loop
    return depth

async def rude(depth):
    time.sleep(0.1)                       # holds the only thread
    return depth

async def race(worker):
    start = time.perf_counter()
    await asyncio.gather(worker(3), worker(7))
    return time.perf_counter() - start

print(asyncio.run(race(gauge)) < 0.15, asyncio.run(race(rude)) >= 0.2)   # True True

class Collect(logging.Handler):
    def __init__(self):
        super().__init__()
        self.lines = []
    def emit(self, record):
        self.lines.append(record.getMessage())

caught = Collect()
logging.getLogger("asyncio").addHandler(caught)

async def watched():
    asyncio.get_running_loop().slow_callback_duration = 0.05   # default is 0.1 s
    await asyncio.gather(rude(3), gauge(7))

asyncio.run(watched(), debug=True)
print(any("rude()" in m and "took" in m for m in caught.lines))   # True

The logged line reads "Executing", then the task running rude(), then "took 0.110 seconds". Now a CPU job, a checksum over four million numbers, run three ways while a ping task tries to answer every 5 ms. The measure is the ping's worst delay:

def checksum(n):                          # pure CPU: no await anywhere
    total = 0
    for i in range(n):
        total = (total * 31 + i) % 1_000_003
    return total

async def checksum_in_slices(n, every=20_000):
    total = 0
    for i in range(n):
        total = (total * 31 + i) % 1_000_003
        if i % every == 0:
            await asyncio.sleep(0)        # a deliberate yield point
    return total

async def worst_ping(job):
    """Run job while a ping task tries to answer every 5 ms; return its worst delay."""
    worst, done = 0.0, asyncio.Event()
    async def ping():
        nonlocal worst
        while not done.is_set():
            asked = time.perf_counter()
            await asyncio.sleep(0.005)
            worst = max(worst, time.perf_counter() - asked - 0.005)
    p = asyncio.create_task(ping())
    await asyncio.sleep(0.02)
    result = await job()
    done.set()
    await p
    return result, worst

N = 4_000_000

async def inline():
    return checksum(N)

async def sliced():
    return await checksum_in_slices(N)

async def offloaded():
    return await asyncio.to_thread(checksum, N)

a, slow = asyncio.run(worst_ping(inline))
b, fine = asyncio.run(worst_ping(sliced))
c, free = asyncio.run(worst_ping(offloaded))
print(a == b == c, slow > 0.1, fine < slow / 4, free < slow / 4)   # True True True True

Same answer three times. Inline, the ping waited the whole computation, about 180 ms on the test machine; sliced and offloaded, a few hundredths of a second at most. The thread still competes for the interpreter lock, which is why CPU-heavy work belongs in a process pool when it matters.

Finally a toy model of the latency rule, with a fake millisecond clock: a job runs CPU slices back to back and yields between them, and a quick request is served at the next yield point. One 200 ms slice makes the worst request wait 199 ms; twenty 10 ms slices, 9 ms. The seeded check compares the formula with a brute-force tick-by-tick simulation on 300 random jobs, and checks that no request ever waits longer than the longest slice:

import bisect
import itertools
import random

def waits(slices, arrivals):
    """Toy model of the loop thread: a job runs CPU slices (ms) back to back, yielding between
    them; a quick request that arrives at t is served at the first yield point at or after t."""
    yields = [0, *itertools.accumulate(slices)]
    out = []
    for t in arrivals:
        i = bisect.bisect_left(yields, t)
        out.append(yields[i] - t if i < len(yields) else 0)
    return out

every_ms = range(200)
print(max(waits([200], every_ms)), max(waits([10] * 20, every_ms)))   # 199 9
print(max(waits([5, 150, 5, 40], every_ms)))                           # 149

def waits_by_ticks(slices, arrivals):
    """Brute force: advance a clock one millisecond at a time."""
    todo, left, clock, served = list(slices), 0, 0, {}
    pending = sorted(set(arrivals))
    while pending:
        if left == 0:                             # a yield point, or the job is over
            while pending and pending[0] <= clock:
                t = pending.pop(0)
                served[t] = clock - t
            if todo:
                left = todo.pop(0)
        if left:
            left -= 1
        clock += 1
    return [served[t] for t in arrivals]

ok = True
for seed in range(300):
    r = random.Random(seed)
    slices = [r.randint(1, 30) for _ in range(r.randint(0, 12))]
    arrivals = [r.randint(0, sum(slices) + 5) for _ in range(r.randint(1, 40))]
    w = waits(slices, arrivals)
    ok &= w == waits_by_ticks(slices, arrivals)
    ok &= max(w) <= max(slices, default=0)        # no request waits longer than one slice
    if slices:                                    # requests every ms: the worst is one slice
        ok &= max(waits(slices, range(sum(slices)))) == max(slices) - 1
print(ok)                                         # True

The complexity

  • Latency: worst added delay equals the longest stretch between yields, across all tasks.
  • Slicing: total work unchanged, plus a little overhead per yield.
  • Offloading: a thread hop or a process round trip, with arguments and results pickled for processes.

Where it goes wrong

  • Sync libraries inside async handlers. One blocking client call per request serialises the server.
  • to_thread for CPU work, expecting a speed-up. It keeps the loop responsive; a process pool adds parallelism.
  • Slices too coarse. A yield every ten seconds of work is no yield at all.
  • Unbounded offloading. A pool has a size; queue behind it deliberately, with a semaphore if needed.
  • No lag metric. Without one, a blocked loop looks like a slow network.

When it shows up in interviews

As "what happens if you call a blocking function in an async handler?", "why is the health check timing out under load?", or the Node.js version, "how do you handle CPU-heavy work on the event loop?". Name the rule, the detection and the three fixes.

How to say it in an interview

"An event loop is one thread, and tasks only switch at an await, so any call that doesn't yield delays every task by its full duration, including health checks. I'd find it with debug mode or an event-loop lag metric. Then: await an async version for I/O, yield in slices if I own the loop, or move the call off the loop with to_thread, or a process pool for CPU-bound Python. The worst latency is bounded by the longest stretch between yields."