Skip to content
BytePatterns

AWS Lambda Cold Starts Explained: Init, Concurrency, Provisioning

9 min readBytePatterns

What a Lambda cold start is, what runs in the init phase, how concurrency is counted, and when reserved or provisioned concurrency actually helps with latency.

"Lambda is slow on the first request" is half true. What is slow is creating a new execution environment, and whether a request pays for that depends on concurrency, not on how long the function has existed. Every behaviour and number below comes from the AWS Lambda Developer Guide pages listed at the end, as of September 2026; quotas change, so check them before you rely on them.

The problem it solves

Lambda runs each invocation in an execution environment: an isolated sandbox with your code, a runtime and any extensions. When a request arrives and no initialised environment is free, Lambda builds one first, and the request waits. That wait is the cold start.

The documentation says cold starts typically occur in under 1% of invocations and last from under 100 ms to over a second. The average hides the problem: latency-sensitive endpoints care about the slowest requests.

The intuition

An environment goes through three phases:

  • Init. Lambda starts any extensions, bootstraps the runtime and runs your function's static code, everything outside the handler. On-demand functions get 10 seconds for this phase; if it does not finish, Lambda retries it at the first invocation, within the function's configured timeout. The init time is billed.
  • Invoke. The handler runs, bounded by the function timeout, at most 15 minutes. Afterwards the environment is frozen, and a later request can reuse it: a warm start, with no init.
  • Shutdown. Lambda keeps idle environments for a while, but it also terminates them every few hours for maintenance, even for functions invoked continuously. You cannot rely on an environment living forever.

Concurrency is the number of requests in flight at once, and each needs its own environment. The documented rule of thumb: concurrency equals requests per second times average duration in seconds. A cold start happens whenever concurrency rises above the number of idle warm environments, so steady load is almost all warm starts and a burst is a wave of cold ones.

Two controls follow:

  • Reserved concurrency is both a maximum and a minimum for one function. It does not pre-initialise anything, so it does not prevent cold starts. It costs nothing extra.
  • Provisioned concurrency keeps environments initialised in advance, so requests up to that number skip init. It is billed extra, and requests above it spill over to on-demand environments that can cold start.

Watch it run

The animation feeds upload events through an SQS queue. The first batch finds no environment, so Lambda builds one: the cold start. A burst brings two more cold starts and a concurrency of 3; reserved concurrency of 3 caps it there while the rest of the burst waits in the queue. Env 1 then takes the next batch warm. The last frames follow a failing message to the dead-letter queue and show provisioned concurrency keeping environments initialised ahead of time.

Lambda & Event-Driven Design

Step 1 of 11

Upload events land in a queue. Lambda runs a poller that reads the queue for you and invokes your function with batches.

The same interactive animation as the lesson — step through it with the controls.

The code

What "static code runs once per environment" means, simulated locally. Module-level objects are created during init and survive between warm invocations, which is why the documentation recommends creating SDK clients and connections outside the handler:

import time

STARTED = time.time()        # module level: runs once, in the init phase
calls = 0                    # survives between warm invocations

def handler(event, context):
    global calls
    calls += 1
    return {"calls_in_this_environment": calls}

print(handler({}, None))     # {'calls_in_this_environment': 1}
print(handler({}, None))     # {'calls_in_this_environment': 2}

Next, a toy model, not Lambda's scheduler, of the one rule that decides cold starts: a request reuses an idle environment if there is one and creates a new one otherwise. Times are in milliseconds:

import heapq

def simulate(requests, provisioned=0, reserved=None):
    """A toy model, not Lambda's scheduler. requests: (arrival, duration) pairs.
    A request reuses any idle environment, else creates one (a cold start),
    unless reserved concurrency is already fully busy (a throttle)."""
    idle = provisioned                  # pre-initialised, so never cold
    busy = []                           # heap of finish times
    cold = throttled = peak = 0
    for arrival, duration in sorted(requests):
        while busy and busy[0] <= arrival:
            heapq.heappop(busy)
            idle += 1                   # finished: environment stays warm
        if reserved is not None and len(busy) >= reserved:
            throttled += 1
            continue
        if idle:
            idle -= 1                   # warm start
        else:
            cold += 1                   # init phase before the handler
        heapq.heappush(busy, arrival + duration)
        peak = max(peak, len(busy))
    return {"cold": cold, "peak": peak, "throttled": throttled}

burst = [(0, 1000)] * 5 + [(500, 1000)] * 3 + [(1200, 1000)] * 4
print(simulate(burst))                  # {'cold': 8, 'peak': 8, 'throttled': 0}
print(simulate(burst, provisioned=5))   # {'cold': 3, 'peak': 8, 'throttled': 0}
print(simulate(burst, reserved=6))      # {'cold': 6, 'peak': 6, 'throttled': 2}

steady = [(i * 10, 500) for i in range(1000)]   # 100 requests/s, 500 ms each
print(simulate(steady))                 # {'cold': 50, 'peak': 50, 'throttled': 0}

The steady load matches the documented formula: 100 requests per second times 0.5 seconds is a concurrency of 50, and after 50 cold starts the remaining 950 requests are all warm. In the burst, provisioned concurrency of 5 removes five cold starts but not the three above it, and a reserve of 6 limits concurrency by refusing work: a throttle, not a speed-up. With a queue in front, as in the lesson, that refused work waits instead of failing.

The model's invariants, checked on 2,000 random workloads: peak concurrency equals a brute-force count of overlapping requests, cold starts equal the peak minus the provisioned environments, and every request is either served or throttled:

import random

def brute_peak(requests):
    return max((sum(a <= t < a + d for a, d in requests) for t, _ in requests), default=0)

random.seed(8)
ok = True
for _ in range(2000):
    reqs = [(random.randint(0, 50), random.randint(1, 20)) for _ in range(random.randint(0, 25))]
    p = random.randint(0, 5)
    r = simulate(reqs, provisioned=p)
    ok &= r["peak"] == brute_peak(reqs)
    ok &= r["cold"] == max(0, r["peak"] - p)
    cap = random.randint(1, 6)
    t = simulate(reqs, reserved=cap)
    ok &= t["peak"] <= cap and t["throttled"] <= len(reqs)
    ok &= (t["throttled"] == 0) == (brute_peak(reqs) <= cap)
print(ok)                               # True

The complexity

The costs are latency and money, not computation:

  • Init work is paid per environment. Large packages and slow connection setup lengthen every cold start; the documentation recommends importing only what the function uses and loading rarely used objects lazily.
  • Concurrency is shared. By default an account has 1,000 concurrent executions per Region across all functions, and 100 always stay unreserved. Reserved and provisioned concurrency both count against that pool.
  • SnapStart, for Java 11+, Python 3.12+ and .NET 8+ managed runtimes, runs init when you publish a version and resumes new environments from a snapshot. It works only on published versions and not with provisioned concurrency. Since one snapshot seeds many environments, unique IDs, secrets and random seeds must be generated after init.

Where it goes wrong

  • Reserving concurrency to fix cold starts. Reserved concurrency is a cap and a guarantee; it does not pre-warm anything.
  • Relying on warm state. Environments are recycled every few hours, and after a crash or timeout Lambda resets the environment, so the next request on it runs init again. Use module-level objects as a cache, never as the only copy of data.
  • Heavy work at module level "because it runs once". It runs once per environment, and a burst creates many environments at the same time, each paying the full init.

How to say it in an interview

"A cold start is the init phase of a new execution environment: extensions, runtime bootstrap, and code outside the handler. It happens when concurrency, requests per second times duration, rises above the idle warm environments, so bursts cause them and steady traffic mostly doesn't. I keep init light and create clients outside the handler so warm invocations reuse them. For strict latency I'd use provisioned concurrency, which costs money while configured. Reserved concurrency only caps and guarantees capacity. For spiky async work I'd put SQS in front so bursts wait instead of failing."

The queue side of the design, including partial batch failures and dead-letter queues, continues in SQS vs SNS vs EventBridge, and container-based alternatives for long jobs are in ECS vs EKS vs Fargate.

Sources