Kubernetes StatefulSet vs Deployment: PVs, PVCs and Stable Storage
9 min readBytePatterns
StatefulSet vs Deployment in Kubernetes: stable Pod names, one PVC per replica, ordered rollout, why scale-down keeps data, and how reclaim policies delete it.
A Deployment treats its Pods as cattle: any replica can replace any other, and their names end in random suffixes. That is exactly wrong for a database, where replica 0 might be the primary and each replica has its own data on disk. Kubernetes answers with three objects that work together: PersistentVolumes, PersistentVolumeClaims and the StatefulSet. Everything below comes from the Kubernetes documentation pages listed at the end, as of September 2026.
The problem it solves
Run three replicas of a database on a cluster where Pods are rescheduled all the time. You need:
- Storage that outlives the Pod. When a Pod dies, its data must not die with it.
- One disk per replica. Replicas do not share a data directory.
- Stable identity. The replacement for replica 1 must come back as replica 1, with replica 1's disk, at an address clients already know.
- Order. Start the first replica before the others join it, and stop them in reverse.
A Deployment gives none of the last three. Put a claim in a Deployment's Pod template and every replica mounts the same claim.
The intuition
A PersistentVolume (PV) is a piece of storage whose lifecycle is independent of any Pod. An administrator can create it in advance, or a StorageClass can provision it on demand.
A PersistentVolumeClaim (PVC) is a request for storage: a size and an access mode. A control loop binds each claim to a matching volume, one to one and exclusively; a claim with no match stays unbound until one appears. Pods never name a volume directly, only a claim.
Access modes say how a volume may be mounted. ReadWriteOnce means read-write by a single node, which can still allow several Pods on that node; ReadWriteOncePod narrows it to a single Pod; ReadOnlyMany and ReadWriteMany allow many nodes.
A StatefulSet adds identity on top:
- Names: with N replicas, Pods get ordinals 0 to N - 1 and names like
db-0,db-1. - Addresses: a headless Service gives each Pod a stable DNS name such as
db-0.db, so clients can find the primary by name. - Storage: each entry in
volumeClaimTemplatesproduces one claim per Pod, named after the template and the Pod, such asdata-db-0. A replacement Pod with the same name mounts the same claim. - Order: by default Pods are created from 0 upwards, each only after its predecessors are Running and Ready, and removed from the highest ordinal down.
Deleting is deliberately conservative. Scaling down or deleting a StatefulSet does not delete its claims by default. What happens when you delete a claim yourself depends on the volume's reclaim policy: Retain keeps the volume for manual cleanup, while Delete removes the volume and the storage behind it. Dynamically provisioned volumes inherit the StorageClass's policy, which defaults to Delete.
Watch it run
The animation starts with three replicas that are not interchangeable: each needs its own name and its own data. The StatefulSet creates db-0 first, with a claim from its template, data-db-0. A StorageClass provisions a 10Gi volume and binds it: one claim, one volume. Only once db-0 is Running and Ready does db-1 start, then db-2; with ReadWriteOnce, every replica has a disk of its own. Through the headless Service, db-0 is reachable as db-0.db. Then db-1's node fails: the Pod is gone, its claim and volume are not, and the replacement comes back as db-1 and mounts data-db-1 again. Scaling to 2 stops db-2 first and keeps its claim. Deleting data-db-2 by hand, with the default Delete policy for dynamic volumes, takes the data with it. Stable names, one claim per replica, ordered start and stop, and cleanup that is deliberately yours.
Storage: PV, PVC & StatefulSet
Step 1 of 11
A database's replicas are not interchangeable: each needs its own name and its own data.
The same interactive animation as the lesson — step through it with the controls.
The code
The lesson's StatefulSet, one 10Gi claim per replica:
# illustrative: one claim per replica; serviceName is a headless Service
apiVersion: apps/v1
kind: StatefulSet
metadata: { name: db }
spec:
serviceName: db
replicas: 3
selector: { matchLabels: { app: db } }
template:
metadata: { labels: { app: db } }
spec:
containers:
- name: db
image: registry.example.com/db:1.0
volumeMounts: [{ name: data, mountPath: /data }]
volumeClaimTemplates:
- metadata: { name: data }
spec: { accessModes: [ReadWriteOnce], resources: { requests: { storage: 10Gi } } }
A toy model of the controller's storage behaviour, not the real controller: ordered scale-up, reverse scale-down, claims that survive both, and a reclaim policy applied when a claim is deleted:
class ToyStatefulSet:
"""Toy model of StatefulSet storage rules, as documented in September 2026."""
def __init__(self, name, template="data", reclaim="Delete"):
self.name, self.template, self.reclaim = name, template, reclaim
self.pods, self.claims, self.volumes, self.log = [], {}, set(), []
def scale(self, replicas):
while len(self.pods) < replicas: # create 0, 1, 2 ... in order
pod = f"{self.name}-{len(self.pods)}"
claim = f"{self.template}-{pod}"
if claim not in self.claims: # reuse a kept claim if present
self.claims[claim] = f"pv-{claim}"
self.volumes.add(f"pv-{claim}")
self.pods.append(pod)
self.log.append(f"+{pod}")
while len(self.pods) > replicas: # remove the highest ordinal first
self.log.append(f"-{self.pods.pop()}") # its claim is kept
def node_lost(self, pod):
self.log.append(f"{pod} rescheduled, mounts {self.template}-{pod}")
def delete_claim(self, claim):
volume = self.claims.pop(claim)
if self.reclaim == "Delete":
self.volumes.discard(volume) # the data goes with it
db = ToyStatefulSet("db")
db.scale(3)
db.node_lost("db-1")
db.scale(2)
print(db.log)
# ['+db-0', '+db-1', '+db-2', 'db-1 rescheduled, mounts data-db-1', '-db-2']
print(sorted(db.claims)) # ['data-db-0', 'data-db-1', 'data-db-2']
db.scale(3)
print(db.claims["data-db-2"]) # pv-data-db-2 the old volume, data intact
db.scale(2)
db.delete_claim("data-db-2")
print(sorted(db.volumes)) # ['pv-data-db-0', 'pv-data-db-1']
The model against a direct statement of the documented rules, over 2,000 random sequences of scaling and claim deletions: Pods are exactly ordinals 0 to N - 1, every Pod has its own claim, and a claim exists for every ordinal ever created unless someone deleted it:
import random
random.seed(17)
ok = True
for _ in range(2000):
s = ToyStatefulSet("db", reclaim=random.choice(["Delete", "Retain"]))
created, deleted = set(), set()
for _ in range(random.randint(1, 8)):
if random.random() < 0.75 or not s.claims:
n = random.randint(0, 5)
s.scale(n)
created |= set(range(n))
deleted -= set(range(n)) # a new Pod gets a fresh claim
else:
i = random.choice([int(c.rsplit("-", 1)[1]) for c in s.claims])
if f"db-{i}" in s.pods:
continue # a claim in use is protected
s.delete_claim(f"data-db-{i}")
deleted.add(i)
ok &= s.pods == [f"db-{i}" for i in range(len(s.pods))]
ok &= set(s.claims) == {f"data-db-{i}" for i in created - deleted}
ok &= all(f"data-{p}" in s.claims for p in s.pods)
if s.reclaim == "Retain":
ok &= len(s.volumes) == len(created) # nothing is ever wiped
print(ok) # True
The complexity
The costs are availability and cleanup, not time:
- Ordered startup is slow on purpose. With the default policy, replica N waits for all earlier replicas to be Running and Ready.
podManagementPolicy: Parallelstarts them together when order does not matter. - Rolling updates go one Pod at a time, highest ordinal first, and a
partitioncan hold the lower ordinals back as a canary. - Kept claims cost money. Every scale-down leaves a volume behind until you delete the claim, or set
persistentVolumeClaimRetentionPolicytoDelete.
Where it goes wrong
- Using a Deployment for a database. All replicas share one claim, or none has stable storage.
- Forgetting the headless Service. It is required for the Pods' network identity.
- Assuming
ReadWriteOncemeans one Pod. It means one node;ReadWriteOncePodmeans one Pod. - Deleting a claim with the
Deletepolicy while you still need the data. It removes the storage too. - Deleting the StatefulSet to stop it. There is no ordering guarantee on deletion; scale to 0 first.
When it shows up in interviews
It shows up in platform and backend interviews as "how would you run a database on Kubernetes?" or "StatefulSet or Deployment?". Expect follow-ups on what happens to a replica's data when its node dies, why scaling down does not free the disk, what ReadWriteOnce really restricts, and how clients find the primary. A good answer also names the alternative: a managed database outside the cluster, with Kubernetes kept for the stateless tier.
How to say it in an interview
"A PVC is a request for storage, a size and an access mode, and it binds one to one to a PV that an admin or a StorageClass provides. A StatefulSet gives each replica a stable name like db-0, a stable DNS name through a headless Service, and its own claim from a volumeClaimTemplate. Pods start in order and stop in reverse. If db-1's node dies, db-1 comes back with the same name and the same claim. Scaling down keeps the claims by default, and with the Delete reclaim policy, deleting a claim deletes its data."
Rolling out new versions of stateless Pods is covered in Kubernetes rolling updates, and how traffic finds Pods in Kubernetes Services vs Ingress.
Sources
- Persistent Volumes — Kubernetes Documentation
- StatefulSets — Kubernetes Documentation
- Storage Classes — Kubernetes Documentation