Pod vs ReplicaSet vs Deployment in Kubernetes, Explained
8 min readBytePatterns
Pod vs ReplicaSet vs Deployment: what each object owns, why a lost Pod is replaced, never moved, how labels decide ownership, and how rollbacks reuse history.
Many Kubernetes interviews start with "what is the difference between a Pod, a ReplicaSet and a Deployment?" The three nest inside each other, and the question is about responsibility. A Pod runs containers. A ReplicaSet keeps a count of Pods. A Deployment keeps the history of versions. Knowing which one to edit, and what happens when a node dies, is most of the answer. Everything below comes from the Kubernetes documentation pages listed at the end, as of September 2026.
The problem it solves
You want three copies of an api container running on whatever machines are healthy, and you want to change the version without downtime and roll it back if it breaks. Starting containers by hand gives you none of that. Something has to notice when a copy disappears, start a new one, and manage the switch between versions.
The intuition
A Pod is the unit Kubernetes runs. The documentation calls it the smallest deployable unit: one or more containers with shared storage and network resources, always co-located and co-scheduled. Containers in the same Pod share a network namespace, so they reach each other over localhost. Pods are designed to be relatively ephemeral and disposable, and a given Pod, identified by its UID, is never rescheduled to another node. When its node fails, that Pod is simply gone. That is why you rarely create Pods yourself, even singletons.
A ReplicaSet is a count. It ensures that a specified number of Pod replicas are running. It finds its Pods with a label selector and records ownership in each Pod's ownerReferences. When there are too few matching Pods it creates more from its template; when there are too many it deletes some. A replacement is a new Pod with a new name and a new UID, not the old one moved. The selector cuts both ways: a Pod with no controlling owner whose labels match is acquired immediately, even if it was not created from the template. The documentation strongly recommends that bare Pods do not carry labels matching a ReplicaSet's selector. Changing a Pod's labels so they no longer match is how you take it out of a ReplicaSet for debugging.
A Deployment is version history. It manages ReplicaSets and adds declarative updates. Each distinct Pod template gets its own ReplicaSet, told apart by a pod-template-hash label the controller adds. Only a change to the Pod template triggers a rollout; scaling does not. During a rollout the Deployment scales the new ReplicaSet up and the old one down, and it keeps old ReplicaSets at zero replicas, ten by default, so kubectl rollout undo has somewhere to go. In apps/v1 a Deployment's selector is immutable, and the documentation is explicit: do not manage ReplicaSets owned by a Deployment.
All three are held together by controllers, control loops that observe the actual state, compare it with the desired state, and act to close the gap. The lesson's picture is a thermostat for your cluster.
Watch it run
The animation starts with desired state: a Deployment that wants three replicas of api. The Deployment controller creates a ReplicaSet for this version of the Pod template. The ReplicaSet creates three Pods from the template, and the scheduler places each one on a node. A Pod is the unit Kubernetes runs: one or more containers sharing a network namespace, which means localhost, and storage. Then a node fails and takes one Pod with it. Pods are not moved, so that one is simply gone. The ReplicaSet sees 2 where 3 are wanted and creates a replacement: a new Pod, new name, new UID. That is every controller, observe, compare, act, forever, a thermostat for the cluster. Change the image in the template, and the Deployment creates a second ReplicaSet for v2, scaling it up as v1 scales down. The old ReplicaSet stays, at zero, so a rollback has somewhere to go. Pod: the unit. ReplicaSet: the count. Deployment: the version history. You only ever edit the Deployment.
Pods, ReplicaSets & Deployments
Step 1 of 10
You write desired state: a Deployment that wants three replicas of api.
The same interactive animation as the lesson — step through it with the controls.
The code
The lesson's manifest:
# illustrative — kubectl apply -f api.yaml, then kubectl get rs,pods -l app=api
apiVersion: apps/v1
kind: Deployment
metadata: { name: api }
spec:
replicas: 3
selector: { matchLabels: { app: api } }
template:
metadata: { labels: { app: api } }
spec:
containers:
- name: api
image: registry.example.com/api:1.4.2
ports: [{ containerPort: 8000 }]
A toy model of the three objects, not the Kubernetes controllers: no API server, no scheduler beyond round-robin, and a rollout that moves one Pod at a time. The ReplicaSet adopts matching orphans, fills the count with new UIDs, and trims extras:
import hashlib
import itertools
class ToyCluster:
"""Toy model of Pods, ReplicaSets and a Deployment, not the Kubernetes controllers."""
def __init__(self, nodes):
self.nodes, self.pods, self.uids = list(nodes), {}, itertools.count(1)
def create_pod(self, labels, image, owner=None):
uid = next(self.uids) # a new Pod always gets a new UID
node = self.nodes[uid % len(self.nodes)] # stand-in for the scheduler
self.pods[uid] = {"labels": dict(labels), "image": image, "owner": owner, "node": node}
return uid
def fail_node(self, node): # its Pods are gone, not moved
for uid in [u for u, p in self.pods.items() if p["node"] == node]:
del self.pods[uid]
class ToyReplicaSet:
def __init__(self, cluster, name, selector, image, replicas=0):
self.c, self.name, self.selector = cluster, name, dict(selector)
self.image, self.replicas = image, replicas
def owned(self):
return sorted(u for u, p in self.c.pods.items() if p["owner"] == self.name)
def reconcile(self): # observe, compare, act
for p in self.c.pods.values(): # adopt orphans that match
if p["owner"] is None and self.selector.items() <= p["labels"].items():
p["owner"] = self.name
mine = self.owned()
for _ in range(self.replicas - len(mine)):
self.c.create_pod(self.selector, self.image, owner=self.name)
for uid in mine[self.replicas:]: # too many: delete the extras
del self.c.pods[uid]
A node failure, a replacement, and a bare Pod that the selector quietly adopts in place of a real replica:
c = ToyCluster(["node-1", "node-2", "node-3"])
rs = ToyReplicaSet(c, "api-rs", {"app": "api"}, "api:1.4.2", replicas=3)
rs.reconcile()
print(rs.owned(), sorted({c.pods[u]["node"] for u in rs.owned()}))
# [1, 2, 3] ['node-1', 'node-2', 'node-3']
c.fail_node("node-2") # Pod 1 was on node-2
print(rs.owned()) # [2, 3]
rs.reconcile()
print(rs.owned()) # [2, 3, 4] a replacement, new UID
del c.pods[2] # kubectl delete pod
c.create_pod({"app": "api"}, "someone-elses:0.1") # a bare Pod with a matching label
rs.reconcile()
print(rs.owned(), [c.pods[u]["image"] for u in rs.owned()])
# [3, 4, 5] ['api:1.4.2', 'api:1.4.2', 'someone-elses:0.1']
The Deployment: one ReplicaSet per template hash, a step-by-step handover, and a rollback that reuses the old set:
class ToyDeployment:
def __init__(self, cluster, name, replicas):
self.c, self.name, self.replicas, self.sets = cluster, name, replicas, {}
def apply(self, image): # a changed template means a rollout
h = hashlib.sha256(image.encode()).hexdigest()[:6] # stand-in for pod-template-hash
if h not in self.sets:
self.sets[h] = ToyReplicaSet(self.c, f"{self.name}-{h}",
{"app": self.name, "pod-template-hash": h}, image)
new = self.sets[h]
old = [rs for rs in self.sets.values() if rs is not new]
while new.replicas < self.replicas or any(rs.replicas for rs in old):
if new.replicas < self.replicas:
new.replicas += 1 # one up...
for rs in old:
if rs.replicas:
rs.replicas -= 1 # ...one down
break
for rs in self.sets.values():
rs.reconcile()
return {rs.image: rs.replicas for rs in self.sets.values()}
d = ToyDeployment(ToyCluster(["node-1", "node-2"]), "api", replicas=3)
print(d.apply("api:1.4.2")) # {'api:1.4.2': 3}
print(d.apply("api:1.5.0")) # {'api:1.4.2': 0, 'api:1.5.0': 3}
print(d.apply("api:1.4.2")) # {'api:1.4.2': 3, 'api:1.5.0': 0}
print(len(d.sets)) # 2 the rollback reused the old set
The model against the documented rules over random event sequences: after every reconcile the owned count equals the spec, Pods with other labels are never touched, UIDs are never reused, and every rollout ends with exactly the desired number of Pods on the new image:
import random
random.seed(19)
ok = True
for _ in range(400):
c = ToyCluster(["n1", "n2", "n3"])
rs = ToyReplicaSet(c, "rs", {"app": "api"}, "img", replicas=random.randint(0, 5))
rs.reconcile()
ever, strangers = set(c.pods), set()
for _ in range(15):
event = random.choice(["delete", "fail", "orphan", "stranger", "scale"])
if event == "delete" and rs.owned():
del c.pods[random.choice(rs.owned())]
elif event == "fail":
c.fail_node(random.choice(c.nodes))
elif event == "orphan":
c.create_pod({"app": "api", "extra": "x"}, "img")
elif event == "stranger":
strangers.add(c.create_pod({"app": "web"}, "img"))
elif event == "scale":
rs.replicas = random.randint(0, 5)
rs.reconcile()
ok &= len(rs.owned()) == rs.replicas # reality matches the spec
ok &= all(c.pods[u]["owner"] is None for u in strangers if u in c.pods)
ok &= all(u > max(ever, default=0) for u in set(c.pods) - ever) # UIDs never reused
ever |= set(c.pods)
for _ in range(200):
d = ToyDeployment(ToyCluster(["n1", "n2"]), "api", replicas=random.randint(1, 5))
images = [random.choice(["v1", "v2", "v3"]) for _ in range(random.randint(1, 6))]
for img in images:
state = d.apply(img)
ok &= state[img] == d.replicas and sum(state.values()) == d.replicas
ok &= all(p["image"] == img for p in d.c.pods.values())
ok &= len(d.sets) == len(set(images))
print(ok) # True
The complexity
The costs are in recovery time and in what you can undo:
- A lost Pod is replaced only after the controller notices, so capacity dips briefly.
- A rollout runs old and new versions side by side; old ReplicaSets at zero run no Pods.
Where it goes wrong
- Running bare Pods in production. Nothing replaces them when they are deleted or their node fails.
- Editing a Deployment's ReplicaSets by hand. The documentation says not to; change the Deployment.
- Overlapping selectors. A stray Pod with matching labels is counted as a replica, and a real one is never created.
- Expecting a Pod to move. A replacement is a new Pod with a new UID.
- Expecting a scale change to roll out a new version. Only template changes do.
When it shows up in interviews
It opens most Kubernetes and DevOps interviews, often followed by "what happens when a node dies?" and "how do you roll back a bad release?". The rollout mechanics go deeper in Kubernetes rolling updates, and keeping replicas healthy in liveness, readiness and startup probes.
How to say it in an interview
"A Pod is the unit Kubernetes runs: one or more containers sharing a network namespace and storage, and it is disposable; a Pod is never moved to another node. A ReplicaSet keeps a set number of Pods matching its selector running, so when a node dies it creates a replacement with a new UID. A Deployment manages ReplicaSets: each template version gets its own, a rollout scales the new one up and the old one down, and old ones are kept at zero for rollbacks. Controllers reconcile desired against actual state continuously. In practice I only edit the Deployment and keep labels unique."
Sources
- Pods — Kubernetes Documentation
- ReplicaSet — Kubernetes Documentation
- Deployments — Kubernetes Documentation
- Controllers — Kubernetes Documentation