Containers vs Virtual Machines: Kernels, Namespaces, cgroups
8 min readBytePatterns
Containers vs virtual machines: one shared kernel or a guest kernel each, namespaces for what a process sees, cgroups for what it uses, and opt-in limits.
"What is the difference between a container and a virtual machine?" opens many Docker and Kubernetes interviews, and the weak answer is "containers are lightweight VMs". A virtual machine brings its own kernel; a container is a fenced-off process on the host's kernel. Two Linux kernel features build the fence, and saying which does what is most of a strong answer. Everything below comes from the Docker and Kubernetes documentation pages listed at the end, as of September 2026.
The problem it solves
You want to run several applications on one machine without them interfering: separate files, separate network ports, and no single application able to eat all the memory. There are two ways to get that isolation.
- A virtual machine runs an entire operating system with its own kernel, hardware drivers, programs and applications, on a hypervisor. The Docker documentation points out that spinning up a VM only to isolate a single application is a lot of overhead.
- A container is an isolated process with all of the files it needs to run. If you run several containers, they all share the same kernel, which lets you run more applications on less infrastructure.
The two are not rivals. In a cloud environment the machines you provision are typically VMs, and a VM with a container runtime runs many containerised applications instead of one.
The intuition
A container is a normal process with two kinds of fence around it.
Namespaces decide what it can see. Docker creates a set of namespaces for each container, and each aspect of the container runs in a separate namespace with access limited to it. A PID namespace gives the container its own process list, where its main process is PID 1. A network namespace gives it its own interfaces and ports, so two containers can each listen on port 80. Mounts get a namespace of their own as well.
Control groups decide what it can use. On Linux, cgroups constrain the resources allocated to processes. Docker's --memory and --cpus flags become limits on the container's cgroup, and the Kubernetes kubelet and container runtime use the same mechanism to enforce CPU and memory requests and limits for Pods. The crucial default: a container has no resource constraints unless you set them, and can use as much of a resource as the host kernel's scheduler allows.
The shared kernel is the trade. A container is only a process, so starting one boots no operating system, and many more fit on a machine. But every container depends on that one kernel, so its boundary is thinner than a VM's, which is a kernel of its own.
Watch it run
The animation shows two ways to isolate an app on one machine: a virtual machine, or a container. The VM carries a whole guest operating system with its own kernel, running on a hypervisor. A container is an isolated process, and every container on the host shares the one host kernel. Namespaces decide what it can see: inside api, docker exec api ps shows a process list of its own, starting at PID 1. A network namespace each, too, so api and worker can both listen on port 80 without a clash. cgroups decide what it can use: the memory and CPU flags of docker run --memory 256m --cpus 0.5 worker become limits on the container's control group. Leave the flags off and there is no limit; the container may use as much as the kernel's scheduler allows. One kernel for all of them is why more apps fit: a third container is one more process. The price is that every container leans on that one kernel. So namespaces are what it sees, cgroups are what it uses, and the shared kernel is both the density and the trade.
Containers vs VMs
Step 1 of 10
Two ways to isolate an app on one machine: a virtual machine, or a container.
The same interactive animation as the lesson — step through it with the controls.
The code
The lesson's commands:
# illustrative — one container, with its own limits
docker run -d --name api --memory 256m --cpus 0.5 -p 8080:80 nginx
docker stats api # live CPU and memory, against the limit
A toy model of one host kernel, not Linux: containers are processes with their own PID numbering and port table, and a memory limit per cgroup. A VM gets a separate kernel object of its own:
import itertools
class ToyKernel:
"""Toy model of namespaces and cgroups on one kernel, not Linux."""
def __init__(self, memory_mb):
self.memory_mb, self.host_pids = memory_mb, itertools.count(1000)
self.procs, self.ports, self.limit, self.used = {}, {}, {}, {}
def run(self, name, memory=None): # docker run [--memory]
self.procs[name], self.ports[name] = [], set()
self.limit[name], self.used[name] = memory, 0
return self.start(name)
def start(self, name): # PID namespace: numbering starts at 1
self.procs[name].append(next(self.host_pids))
return len(self.procs[name])
def ps(self, name):
return list(range(1, len(self.procs[name]) + 1))
def listen(self, name, port): # network namespace: one table each
if port in self.ports[name]:
return "address in use"
self.ports[name].add(port)
return "listening"
def alloc(self, name, mb): # cgroup: a limit only if one was set
cap = self.limit[name]
if cap is not None and self.used[name] + mb > cap:
return "killed: over its cgroup limit"
if sum(self.used.values()) + mb > self.memory_mb:
return "host out of memory"
self.used[name] += mb
return "ok"
The animation's two containers, one limited and one not, plus a VM with a kernel of its own:
host = ToyKernel(memory_mb=4096)
print(host.run("api"), host.run("worker", memory=256)) # 1 1 each sees PID 1
print(host.start("api"), host.ps("api")) # 2 [1, 2]
print(host.listen("api", 80), host.listen("worker", 80)) # listening listening
print(host.listen("api", 80)) # address in use
print(host.alloc("worker", 200), host.alloc("worker", 100))
# ok killed: over its cgroup limit
print(host.alloc("api", 3800)) # ok no limit was set
print(host.alloc("api", 200)) # host out of memory
vm_guest = ToyKernel(memory_mb=2048) # a VM boots its own kernel
print(vm_guest is host) # False
The model against the rules it encodes, over 2,000 random sequences of runs, starts, listens and allocations: every container numbers its processes from 1, host PIDs are never shared, a limited container never exceeds its limit, the host total never exceeds physical memory, and a port clashes only inside the same container:
import random
random.seed(20)
ok = True
for _ in range(2000):
k = ToyKernel(memory_mb=random.choice([512, 1024]))
names = [f"c{i}" for i in range(random.randint(1, 4))]
for n in names:
k.run(n, memory=random.choice([None, 128, 256]))
bound = set()
for _ in range(25):
n, op = random.choice(names), random.choice(["start", "listen", "alloc"])
if op == "start":
k.start(n)
elif op == "listen":
port = random.choice([80, 443, 8080])
ok &= (k.listen(n, port) == "address in use") == ((n, port) in bound)
bound.add((n, port))
else:
k.alloc(n, random.randint(1, 200))
ok &= all(k.ps(c) == list(range(1, len(k.procs[c]) + 1)) for c in names)
pids = [p for c in names for p in k.procs[c]]
ok &= len(pids) == len(set(pids))
ok &= all(k.limit[c] is None or k.used[c] <= k.limit[c] for c in names)
ok &= sum(k.used.values()) <= k.memory_mb
print(ok) # True
The complexity
The costs are density and blast radius, not steps:
- Density: a container is one more process on a running kernel; a VM carries a whole operating system.
- Isolation: a VM's boundary is a separate kernel; containers share one, so a kernel problem is every container's problem.
- Resources: unlimited by default. Without
--memory, one leaking container competes with everything else on the host.
Where it goes wrong
- Calling a container a lightweight VM. It has no kernel of its own; it is a process with namespaces and a cgroup.
- Assuming limits exist by default. They do not; a container can use as much as the host kernel's scheduler allows until you set them.
- Mixing up Linux namespaces and Kubernetes namespaces. Linux namespaces isolate what a process sees; a Kubernetes namespace is a scope for object names in a cluster.
- Treating the container boundary as a VM's. For untrusted workloads, raise the shared kernel.
- Swapping the two features. Namespaces limit visibility; cgroups limit consumption.
When it shows up in interviews
It is the usual first question in a Docker or Kubernetes round, and it leads into how limits work in a cluster, covered in Kubernetes requests, limits and the HPA, and into what an image is, covered in Docker image layers and build cache. A broader question bank is in Docker and Kubernetes interview questions.
How to say it in an interview
"A virtual machine runs a full guest operating system with its own kernel on a hypervisor. A container is an isolated process that shares the host's kernel. Linux namespaces limit what it can see, such as its own process list starting at PID 1, its own network interfaces and ports, and its own mounts; cgroups limit what it can use, such as CPU and memory. Limits are opt-in: without --memory or --cpus, a container can use as much as the host allows. Sharing the kernel is why containers are dense, and why their isolation is thinner than a VM's. In the cloud the two are combined: containers inside VMs."
Sources
- What is a container? — Docker Docs
- What is Docker? — Docker Docs
- Resource constraints — Docker Docs
- About cgroup v2 — Kubernetes Documentation
- Namespaces — Kubernetes Documentation