Async/Await and the Event Loop Explained: One Thread, Taking Turns
8 min readBytePatterns
Async/await and the event loop explained: one thread, cooperative turns at every await, why one blocking call freezes everything, and when threads fit better.
A server that handles ten thousand connections does not need ten thousand threads. Most connections spend their time waiting for bytes, and waiting needs no processor. Async/await is built on that observation: one thread, an event loop, and many tasks that take turns, each handing the thread back whenever it has to wait. Interviews ask about it to see whether you know what actually runs concurrently, and the one mistake that freezes an entire async service.
The problem it solves
The classic way to serve many clients is a thread per client. Each thread costs memory for its stack, the operating system switches between them, and shared data needs locks, with the race conditions that come with them. For work that is mostly waiting, such as calling APIs, querying databases or streaming from sockets, that is expensive machinery for doing nothing. Async runs all of those waits on one thread: while one task waits, the thread runs another that has something to do.
The intuition
A coroutine, defined with async def, is a function that can pause. It runs normally until it reaches an await on something that is not ready yet, such as a network read or a timer. At that point it suspends and hands control back to the event loop.
The event loop is a plain loop around two collections: a ready queue of tasks that can run now, and the timers and sockets being waited on. It runs a ready task until its next await; when a timer fires or a socket has data, the task waiting on it goes back into the ready queue.
That makes async cooperative, not preemptive. Threads are interrupted by the operating system at arbitrary moments; a coroutine loses the thread only at an await, visible in the source, so between two awaits nothing else runs.
The same fact is the danger. A coroutine that calls something slow without awaiting, like time.sleep, a synchronous HTTP client or a long CPU loop, keeps the one thread, and every other task stops. The fix is to use the async version of the call, or to push the blocking work onto a thread with asyncio.to_thread, which is in the standard library from Python 3.9 on; the Python details here are as of September 2026. JavaScript works the same way: one thread, an event loop, and await as the yield point.
Watch it run
The animation starts from the rule: async runs on one thread, and gather schedules both coroutines and hands that thread to the loop. A runs and prints step 1, then hits await, where it gives the thread back. The loop resumes whichever task is ready; that is B, so B prints step 1. B awaits in turn, so the loop comes back round to A for step 2. B finishes with step 2. The printed order is A, B, A, B on every single run, because without preemption there are no surprise interleavings: every yield point is visible in the source. Then a rude task arrives. time.sleep(5) never awaits, so it keeps the one thread. A has been ready the whole time and cannot run; one thread means one rude task freezes everything. The closing frames sum it up: waiting is nearly free here and computing is not, so async suits thousands of idle network calls, and cooperative scheduling is a bargain that needs no locks, provided every slow call is actually awaited.
Async and the Event Loop
Step 1 of 10
Async runs on one thread. gather schedules both coroutines and hands that thread to the loop.
The same interactive animation as the lesson — step through it with the controls.
The code
The lesson's two workers. Each await asyncio.sleep(0) hands the thread back, so the output alternates in the same order every time:
import asyncio
import time
async def worker(name, rounds):
for step in range(1, rounds + 1):
print(name, "step", step)
await asyncio.sleep(0) # yield: let the loop run someone else
async def main():
await asyncio.gather(worker("A", 2), worker("B", 2))
asyncio.run(main())
# A step 1
# B step 1
# A step 2
# B step 2
Why waiting is nearly free. A hundred simulated 0.2-second calls, gathered, finish in about 0.2 seconds; five awaited one after another take a second. Timings print as comparisons so the output does not depend on the machine:
async def fetch(i):
await asyncio.sleep(0.2) # stands in for a network call
return i
async def all_at_once():
t0 = time.perf_counter()
results = await asyncio.gather(*(fetch(i) for i in range(100)))
return len(results), time.perf_counter() - t0
async def one_by_one():
t0 = time.perf_counter()
for i in range(5):
await fetch(i)
return time.perf_counter() - t0
count, together = asyncio.run(all_at_once())
apart = asyncio.run(one_by_one())
print(count, together < 0.6, apart >= 1.0) # 100 True True
The rude task, measured. A heartbeat tries to tick every 20 ms while another task waits half a second, and the longest gap between ticks is recorded. time.sleep stops the heartbeat for the full half second; await asyncio.sleep, or asyncio.to_thread, keeps it ticking:
async def heartbeat(beats, stop):
while not stop.is_set():
beats.append(time.perf_counter())
await asyncio.sleep(0.02)
async def longest_gap(slow_call):
beats, stop = [], asyncio.Event()
hb = asyncio.create_task(heartbeat(beats, stop))
await asyncio.sleep(0.05)
await slow_call()
await asyncio.sleep(0.05)
stop.set()
await hb
return max(b - a for a, b in zip(beats, beats[1:]))
async def rude():
time.sleep(0.5) # blocks: never gives the thread back
async def polite():
await asyncio.sleep(0.5) # waits without holding the thread
async def offloaded():
await asyncio.to_thread(time.sleep, 0.5) # the blocking call runs on a worker thread
print(asyncio.run(longest_gap(rude)) >= 0.5) # True
print(asyncio.run(longest_gap(polite)) < 0.25) # True
print(asyncio.run(longest_gap(offloaded)) < 0.25) # True
A toy model of the loop, not how asyncio is implemented: generators play the coroutines, each yield is an await asyncio.sleep(delay), and the loop keeps a ready queue, a timer heap and a fake clock that jumps to the next timer when nothing is ready. Tasks finish in order of total sleep, not start order:
import heapq
from collections import deque
def toy_loop(tasks):
"""Toy model of an event loop: a ready queue, a timer heap and a fake clock."""
ready, timers, clock, seq, log = deque(tasks.items()), [], 0, 0, []
while ready or timers:
if not ready: # nothing runnable: jump to the next timer
clock, _, name, gen = heapq.heappop(timers)
ready.append((name, gen))
name, gen = ready.popleft()
try:
delay = next(gen) # run until the next "await"
except StopIteration:
log.append((clock, name, "done"))
continue
if delay == 0:
ready.append((name, gen)) # sleep(0): back of the queue
else:
seq += 1
heapq.heappush(timers, (clock + delay, seq, name, gen))
return log
def sleeper(delays):
for d in delays:
yield d # "await asyncio.sleep(d)"
print(toy_loop({"A": sleeper([3]), "B": sleeper([1]), "C": sleeper([1, 1])}))
# [(1, 'B', 'done'), (2, 'C', 'done'), (3, 'A', 'done')]
Checked on 200 seeded random sets of tasks. With sleep(0) only, real asyncio's step order must equal the toy loop's and a brute-force round-robin; with random fake sleeps, each toy task must finish exactly at the sum of its delays:
import random
async def recorder(name, steps, out):
for step in range(steps):
out.append((name, step))
await asyncio.sleep(0)
async def real_order(counts):
out = []
await asyncio.gather(*(recorder(n, k, out) for n, k in counts.items()))
return out
def toy_order(counts):
out = []
def gen(name, steps):
for step in range(steps):
out.append((name, step))
yield 0
toy_loop({n: gen(n, k) for n, k in counts.items()})
return out
def round_robin(counts):
"""Brute force: go round the tasks in order, one step each, skipping finished ones."""
out, step = [], 0
while any(k > step for k in counts.values()):
out += [(n, step) for n, k in counts.items() if k > step]
step += 1
return out
random.seed(24)
ok = True
for _ in range(200):
counts = {f"t{i}": random.randint(0, 5) for i in range(random.randint(1, 6))}
ok &= asyncio.run(real_order(counts)) == toy_order(counts) == round_robin(counts)
delays = {f"t{i}": [random.randint(1, 9) for _ in range(random.randint(0, 4))] for i in range(4)}
done = {name: t for t, name, _ in toy_loop({n: sleeper(d) for n, d in delays.items()})}
ok &= done == {n: sum(d) for n, d in delays.items()} # each task ends at its total sleep
print(ok) # True
The complexity
- Waiting: a suspended coroutine is a small object, not a thread stack, so tens of thousands of concurrent waits fit on one thread.
- Computing: one thread uses one core; async adds no parallelism.
- Latency: every task waits on the longest stretch any task runs between two awaits.
Where it goes wrong
- Blocking calls inside coroutines.
time.sleep, synchronous database drivers and file reads all freeze the loop. Use async libraries orasyncio.to_thread. - Forgetting
await. Calling a coroutine function without awaiting it creates a coroutine object that never runs. - CPU-heavy work on the loop. Parsing a huge file or hashing passwords holds the thread; move it to a process or thread pool. The trade-offs are in threads vs processes.
- Assuming no races at all. State is safe between awaits, but a read, an
await, then a write can still interleave with another task. - Losing exceptions. A task never awaited can fail silently; await it, or use
asyncio.TaskGroup, available since Python 3.11.
When it shows up in interviews
Backend rounds ask it as "concurrency or parallelism?", "why is Node.js fast with one thread?" or "what happens if you call a blocking function in an async handler?". System design answers lean on it for chat servers, websocket gateways and crawlers, where most connections are idle. Have semaphores vs mutexes ready too, since a semaphore is how async code caps concurrent requests.
How to say it in an interview
"Async is cooperative multitasking on one thread. A coroutine runs until it awaits something that is not ready, then suspends and the event loop runs another ready task; when the timer or socket completes, that task is ready again. Waiting is almost free, which suits thousands of idle network calls. The catch is that anything blocking or CPU-heavy holds the only thread and stalls every task, so I use async libraries throughout and push blocking or CPU work to a thread or process pool."