Skip to content
BytePatterns

Thread Pools Explained: Reuse Workers and Cap Concurrency

8 min readBytePatterns

Thread pools explained: a fixed crew of workers fed from a queue, why map returns input order, how pool size caps concurrency, and the starvation deadlock trap.

The first version of a concurrent program usually starts a thread per task. It works in the demo, and then a burst of ten thousand requests creates ten thousand threads, each with its own stack, all fighting over the same cores and the same database. A thread pool is the standard fix: create a fixed number of workers once, put tasks in a queue, and let the workers pull from it. The interview questions around it are about what the pool size really limits, what order results come back in, and how a pool can deadlock itself.

The problem it solves

Creating a thread costs time and memory: the operating system sets up a stack and scheduling state for each one. For short tasks that setup can cost more than the work. Worse, a thread per task means concurrency is set by the traffic rather than by you, and a traffic spike becomes a resource spike.

A pool separates the two concerns. The task queue absorbs however much work arrives; the worker count decides how much of it runs at once. Threads are made once and reused, so the setup cost is paid at start-up rather than per task.

The intuition

A pool has three parts:

  • Workers: long-lived threads, each looping "take a task from the queue, run it, repeat".
  • A queue: tasks waiting for a free worker.
  • Futures: a handle per submitted task that will eventually hold its result or its exception.

The worker count is a deliberate cap. It should match whatever is actually scarce. For work that mostly waits, such as HTTP calls or disk reads, many threads are cheap because a blocked thread is not using a core; for CPU-heavy work, about one per core is the ceiling. The follow-up lesson, sizing a thread pool, works through that choice. In CPython the global interpreter lock also means threads do not run Python bytecode in parallel on the standard build (as of September 2026, free-threaded builds exist but are not the default), so CPU-bound Python work goes to a process pool; see threads vs processes.

And a bigger pool only helps if the bottleneck is the pool. If every task needs one of four database connections, the fifth thread just waits in a different queue.

Watch it run

The animation is the lesson's example: two workers, created once, and four pages waiting in the queue, not four threads. Both workers take a task; the pool size is a deliberate ceiling, so only two run at once. Worker 1 finishes index and is reused, not destroyed; it picks up pricing, and the thread count is still two. about was the slow one, so worker 2 hands it back and takes faq, leaving one of the four running. Then all four are done, in completion order, which is not the order they went in. pool.map lines the results back up with the inputs before handing them over. Both threads are still alive between tasks, because building one per task is what a pool avoids. Doubling the pool does not help when the bottleneck is a database already at its limit. The closing frame sums it up: a pool is two decisions at once, reuse expensive workers and cap how much runs in parallel.

Thread Pools

Step 1 of 9

Two workers, created once. Four pages wait in the queue — not four threads.

The same interactive animation as the lesson — step through it with the controls.

The code

The standard library pool. Four tasks on two workers use at most two threads, and map returns results in input order however they finish:

import threading
import time
from concurrent.futures import ThreadPoolExecutor

def fetch_size(page):
    time.sleep(0.05 if page == "about" else 0.01)   # about is the slow one
    return page, len(page) * 100, threading.get_ident()

pages = ["index", "about", "pricing", "faq"]
with ThreadPoolExecutor(max_workers=2) as pool:
    results = list(pool.map(fetch_size, pages))

print([(page, size) for page, size, _ in results])
# [('index', 500), ('about', 500), ('pricing', 700), ('faq', 300)]
print(len({ident for *_, ident in results}) <= 2)   # True  never more than two threads

A pool from scratch, so the moving parts are visible: a queue, a fixed set of worker threads, one sentinel per worker to shut them down, and results stored by input position:

import queue

class TinyPool:
    def __init__(self, workers):
        self.tasks = queue.Queue()
        self.threads = [threading.Thread(target=self._work) for _ in range(workers)]
        for t in self.threads:
            t.start()                          # created once, reused for every task

    def _work(self):
        while True:
            item = self.tasks.get()
            if item is None:                   # sentinel: time to stop
                return
            fn, arg, slot, out = item
            try:
                out[slot] = fn(arg)
            except Exception as exc:           # keep the worker alive, like a future
                out[slot] = exc
            finally:
                self.tasks.task_done()

    def map(self, fn, args):
        out = [None] * len(args)
        for slot, arg in enumerate(args):
            self.tasks.put((fn, arg, slot, out))
        self.tasks.join()                      # wait until every task is done
        return out                             # slot order = input order

    def close(self):
        for _ in self.threads:
            self.tasks.put(None)
        for t in self.threads:
            t.join()

pool = TinyPool(3)
squares = pool.map(lambda x: x * x, list(range(10)))
pool.close()                                   # one sentinel per worker, then join
print(squares, len(pool.threads))
# [0, 1, 4, 9, 16, 25, 36, 49, 64, 81] 3

The trap: a task that waits for another task in the same pool. With one worker, the outer task holds the only thread while waiting for an inner task that can never start:

from concurrent.futures import TimeoutError as FutureTimeout

def outer(pool):
    inner = pool.submit(lambda: "inner done")
    try:
        return inner.result(timeout=0.5)       # waits for a worker that is... us
    except FutureTimeout:
        return "starved"

with ThreadPoolExecutor(max_workers=1) as p:
    print(p.submit(outer, p).result())         # starved

A toy model of the "bigger pool" question on a fake clock: 40 tasks of 10 ms, each needing one of four database connections. The makespan stops improving at four workers, because the scarce thing is the connection:

def makespan(tasks, workers, connections):
    lanes = min(workers, connections)          # a task needs a thread AND a connection
    finish = [0] * lanes
    for duration in tasks:
        i = finish.index(min(finish))          # next free lane takes the next task
        finish[i] += duration
    return max(finish)

jobs = [10] * 40
print([makespan(jobs, w, connections=4) for w in (1, 2, 4, 8, 16)])
# [400, 200, 100, 100, 100]

Checked on 300 seeded random workloads: the scratch pool must return exactly what a plain sequential loop returns, in input order, for any worker count, and the toy makespan must match a tick-by-tick brute-force simulation:

import random

def makespan_by_ticks(tasks, lanes):
    """Brute force: advance one millisecond at a time, hand tasks to idle lanes."""
    pending, busy, t = list(tasks), [0] * lanes, 0
    while pending or any(busy):
        for i in range(lanes):
            if busy[i] == 0 and pending:
                busy[i] = pending.pop(0)
        t += 1
        busy = [max(0, b - 1) for b in busy]
    return t

random.seed(26)
ok = True
for _ in range(300):
    args = [random.randint(-50, 50) for _ in range(random.randint(0, 25))]
    p = TinyPool(random.randint(1, 6))
    got = p.map(lambda x: x * 3 - 1, args)
    p.close()
    ok &= got == [x * 3 - 1 for x in args]
    durations = [random.randint(1, 9) for _ in range(random.randint(1, 12))]
    w, c = random.randint(1, 6), random.randint(1, 6)
    ok &= makespan(durations, w, c) == makespan_by_ticks(durations, min(w, c))
print(ok)                                      # True

The complexity

  • Per task: one queue put and one get, O(1), instead of creating and destroying a thread.
  • Memory: one stack per worker, fixed, plus whatever the queue holds. The standard executor's queue is unbounded.
  • Throughput: at most min(workers, scarce resource) tasks in flight; beyond that, extra workers add memory and context switches, not speed.

Where it goes wrong

  • Tasks that wait on the same pool. Nested submission can starve every worker, as above. Use a separate pool, or restructure so parents do not block on children.
  • An unbounded queue. Submitting millions of tasks holds all of them in memory. Bound it and push back on producers; backpressure covers how.
  • Swallowed exceptions. A failed task stores its exception in the future; if nobody calls result(), nobody sees it.
  • One pool for everything. A slow dependency can occupy every worker and block unrelated work. Separate pools per dependency isolate the damage.
  • Threads for CPU-bound Python. Use a process pool on the standard CPython build.

When it shows up in interviews

Concurrency rounds ask you to implement a small pool, like the one above, or to explain map against as_completed. System design rounds ask how many workers a service should run and what happens when the downstream database is the limit. Both connect to the producer-consumer bounded buffer, which is exactly the queue between the submitters and the workers, and to semaphores, the other way to cap how many things run at once.

How to say it in an interview

"A thread pool creates a fixed number of workers once and feeds them from a queue, so I pay thread creation once and I choose the concurrency instead of letting traffic choose it. The size should match the scarce resource: about the core count for CPU work, more for IO-bound work, and never more than the downstream can accept. I would bound the queue so overload pushes back instead of eating memory, and I would never have a task block on another task in the same pool, because with every worker waiting, nothing can run."