Kubernetes Rolling Updates Explained: maxSurge and maxUnavailable
8 min readBytePatterns
How a Deployment replaces v1 Pods with v2 without downtime, what maxSurge and maxUnavailable bound, how they round, and why a bad release stalls safely.
You change the image tag in a Deployment and, a minute later, every Pod runs the new version and nobody noticed. That is a rolling update, and two numbers decide how it goes: maxSurge and maxUnavailable. Knowing exactly what they bound — and what happens when the new version is broken — is what separates "I've used Kubernetes" from "I can run it". Every behaviour below is taken from the Kubernetes documentation listed at the end; defaults are as documented as of September 2026.
The problem it solves
Replacing a running service has two naive options, and both are bad. Stop everything, then start the new version: simple, but the service is down in between. That is the Recreate strategy, where all existing Pods are killed before new ones are created. Or start a complete second copy, then switch: no downtime, but you need double the capacity.
A rolling update sits between them. It replaces Pods a few at a time, keeping enough of them serving throughout and never creating too many extra. RollingUpdate is the default strategy for a Deployment.
The intuition
A Deployment does not manage Pods directly. It manages ReplicaSets, and each ReplicaSet keeps a number of identical Pods running. When you change the Pod template — the image, an environment variable, a label in the template — the Deployment creates a new ReplicaSet for the new template, scales it up, and scales the old one down, in steps. Only a template change starts a rollout; scaling the replica count does not.
The two knobs bound every intermediate state:
maxSurge: how many Pods may exist above the desired count. It is how much extra capacity the rollout may borrow.maxUnavailable: how many Pods may be missing from the desired count. It is how much capacity the rollout may give up.
Both accept a number or a percentage, and both default to 25%. Percentages become Pod counts with deliberate rounding: maxSurge rounds up, maxUnavailable rounds down. Both cannot be 0 at once, since then nothing could ever move. With four replicas and the defaults, the documentation's own example gives between 3 and 5 Pods at every moment.
"Available" is the key word. A new Pod counts only once it is ready — its readiness probe passes — and has stayed ready for minReadySeconds, which defaults to 0. So a new version that never becomes ready can never push the rollout forward.
Watch it run
The animation starts with the scheduler placing one Pod: filter out nodes that cannot fit its CPU request, score the rest, bind. Then it rolls four v1 Pods to v2 with maxSurge 1 and maxUnavailable 0: one v2 Pod starts, passes readiness, and only then does one v1 Pod terminate — surge, ready, shrink, four times. Finally it ships a v3 whose Pod never becomes ready. The rollout stalls with all four v2 Pods still serving, and kubectl rollout undo returns to the previous ReplicaSet.
Scheduling & Rolling Updates
Step 1 of 13
A new Pod requests 500m of CPU. Until the scheduler binds it to a node, it is Pending.
The same interactive animation as the lesson — step through it with the controls.
The code
First, the rounding, applied to the defaults for a few replica counts:
import math
def resolve(replicas, max_surge="25%", max_unavailable="25%"):
"""Percentages become Pod counts: surge rounds up, unavailable rounds down."""
def count(value, rounding):
if isinstance(value, str):
return rounding(replicas * int(value.rstrip("%")) / 100)
return value
return count(max_surge, math.ceil), count(max_unavailable, math.floor)
for n in (1, 4, 10):
surge, unavailable = resolve(n)
print(n, "replicas: at most", n + surge, "Pods, at least", n - unavailable, "available")
# 1 replicas: at most 2 Pods, at least 1 available
# 4 replicas: at most 5 Pods, at least 3 available
# 10 replicas: at most 13 Pods, at least 8 available
With a single replica, rounding down gives maxUnavailable 0, so the defaults never take the only Pod away before its replacement is ready.
Next, a toy model of the two bounds, not the controller's code: each tick, remove old Pods while enough stay available, then start new ones within the surge. A new Pod is ready one tick after it starts:
def roll(replicas, surge, unavailable, new_becomes_ready=True, max_ticks=50):
"""A toy model of the two documented bounds, not the controller's code.
Each tick: remove v1 Pods while enough Pods stay available, then start
v2 Pods within the surge. A v2 Pod started on one tick is ready on the next."""
old, new_ready, new_starting = replicas, 0, 0
history = []
for _ in range(max_ticks):
new_ready += new_starting if new_becomes_ready else 0
new_starting = 0 if new_becomes_ready else new_starting
floor = replicas - unavailable # the availability floor
old -= max(0, min(old, old + new_ready - floor))
room = replicas + surge - (old + new_ready + new_starting)
new_starting += max(0, min(room, replicas - new_ready - new_starting))
history.append((old, new_ready, new_starting))
if old == 0 and new_ready == replicas:
break
return history
for step in roll(4, surge=1, unavailable=0):
print(step) # (v1 Pods, v2 ready, v2 starting)
# (4, 0, 1)
# (3, 1, 1)
# (2, 2, 1)
# (1, 3, 1)
# (0, 4, 0)
print(roll(4, surge=1, unavailable=1))
# [(3, 0, 2), (1, 2, 2), (0, 4, 0)]
print(roll(4, surge=1, unavailable=0, new_becomes_ready=False)[-1])
# (4, 0, 1) -- v2 never gets ready: the rollout stalls, all four v1 Pods keep serving
The trade-off is visible: maxUnavailable 0 never dips below four available Pods but takes five ticks; the defaults finish in three by dropping to three available at the start.
The model is then checked against both documented bounds on 3,000 random configurations, for releases that become ready and releases that never do:
import random
random.seed(6)
ok, stalled_ok = True, True
for _ in range(3000):
replicas = random.randint(1, 20)
surge, unavailable = random.randint(0, 5), random.randint(0, 5)
if surge == unavailable == 0:
continue # not allowed: nothing could ever move
for ready in (True, False):
history = roll(replicas, surge, unavailable, new_becomes_ready=ready)
for old, new_ready, new_starting in history:
ok &= old + new_ready + new_starting <= replicas + surge # the surge cap
ok &= old + new_ready >= replicas - unavailable # the availability floor
if ready:
ok &= history[-1] == (0, replicas, 0) # it finishes
else:
stalled_ok &= history[-1][1] == 0 and history[-1][0] >= replicas - unavailable
print(ok, stalled_ok) # True True
The commands themselves, from the lesson, are illustrative — the image name is a placeholder:
kubectl set image deployment/api api=registry.example.com/api:2.0.0
kubectl rollout status deployment/api
kubectl rollout undo deployment/api
The complexity
The cost of a rollout is time and headroom, not computation:
maxSurgeabove 0 needs spare cluster capacity for the extra Pods. If the scheduler cannot fit them, they stay Pending and the rollout waits.maxUnavailableabove 0 needs no spare capacity but runs below full strength for part of the rollout.- Larger values finish faster; smaller values are gentler. Each step also waits for readiness, so slow-starting Pods make every step slow.
Where it goes wrong
- No readiness probe. Without one, the kubelet treats the readiness result as a success, so a broken version can be declared available and the old Pods removed.
- Expecting an automatic rollback. When a rollout stops progressing for
progressDeadlineSeconds(default 600), the Deployment reports aProgressingcondition with reasonProgressDeadlineExceeded. Kubernetes takes no other action; rolling back is your job or your tooling's. - Setting
revisionHistoryLimitto 0. Old ReplicaSets are the rollback history, 10 by default. With 0, the rollout cannot be undone. - Counting terminating Pods. Pods that are shutting down can take a while, so the real number of Pods can briefly exceed what the bounds suggest.
How to say it in an interview
"A Deployment rolls out by creating a new ReplicaSet and shifting Pods from the old one to the new one in steps. maxSurge caps how many extra Pods may exist, rounded up; maxUnavailable caps how many may be missing, rounded down; both default to 25%. A new Pod only counts once it's ready, so a bad release stalls instead of replacing healthy Pods. I'd use maxUnavailable 0 where capacity must not dip, give every container a readiness probe, and roll back with kubectl rollout undo — Kubernetes reports a stalled rollout but doesn't roll back on its own."
Readiness is what makes all of this safe, and probes and self-healing covers it; the objects being rolled are introduced in Pods, ReplicaSets and Deployments.
Sources
- Deployments — Kubernetes Documentation
- Kubernetes Scheduler — Kubernetes Documentation
- Resource management for Pods and containers — Kubernetes Documentation
- Liveness, readiness and startup probes — Kubernetes Documentation