Skip to content
BytePatterns

Kubernetes Liveness vs Readiness vs Startup Probes Explained

9 min readBytePatterns

What each Kubernetes probe does when it fails, the default thresholds, how startup probes protect slow boots, and how a bad liveness probe causes cascades.

A container that is running is not necessarily a container that is working. It can be deadlocked, still warming a cache, or overloaded for a minute. Kubernetes lets the kubelet ask three different questions about each container, and the whole topic comes down to knowing which question triggers which action. Everything below comes from the Kubernetes documentation pages listed at the end, as of September 2026.

The problem it solves

Two failure modes look the same from outside and need opposite responses:

  • The process is stuck for good. A deadlock will not clear by itself. The fix is to restart the container.
  • The process is alive but cannot serve right now. It is loading files, waiting on a dependency, or recovering from overload. Restarting it throws away progress; the fix is to stop sending it traffic until it recovers.

A third case sits in front of both: an application that legitimately takes a long time to start, and must not be mistaken for a stuck one while it boots.

The intuition

Each probe answers one question and triggers one action:

  • Liveness: "should this container be restarted?" If it fails more times in a row than its failureThreshold, the kubelet restarts the container. The documentation's example is catching a deadlock, where the application runs but cannot make progress.
  • Readiness: "should this container get traffic?" If it fails, the Pod's IP is removed from the EndpointSlices of every Service that matches it. Nothing restarts, and the probe keeps running for the container's whole lifecycle, so the Pod rejoins once it passes again.
  • Startup: "has it finished starting?" While a startup probe is configured and has not yet succeeded, liveness and readiness probes do not run. If the startup probe fails, the kubelet kills the container and the restart policy applies.

Each probe uses exactly one mechanism: exec succeeds on exit code 0, httpGet on a status code from 200 up to but not including 400, tcpSocket when a connection opens, and grpc when the health check reports SERVING.

The defaults matter because most manifests rely on them: initialDelaySeconds 0, periodSeconds 10, timeoutSeconds 1, successThreshold 1, which must stay 1 for liveness and startup probes, and failureThreshold 3.

Watch it run

The animation follows one api container with its three probes as rows. While it boots, only the startup probe runs, with a budget of 30 × 2 s. Once startup succeeds, readiness and liveness take turns. A failing readiness check keeps the Pod out of the Service's endpoints with nothing restarted, and traffic returns when it passes. Then the process deadlocks: liveness fails repeatedly, the kubelet restarts the container, and repeated crashes back off into CrashLoopBackOff. The last frames show the risk of an over-eager liveness probe.

Probes & Self-Healing

Step 1 of 12

The container is booting. Only the startup probe runs; liveness and readiness wait for it.

The same interactive animation as the lesson — step through it with the controls.

The code

The lesson's three probes on one container. Liveness and readiness here use different endpoints, /healthz for "is the process alive" and /ready for "can it serve":

# illustrative: 30 × 2 s gives the app up to 60 s to start
startupProbe:
  httpGet:
    path: /healthz
    port: 8000
  failureThreshold: 30
  periodSeconds: 2
readinessProbe:
  httpGet:
    path: /ready
    port: 8000
  periodSeconds: 5
livenessProbe:
  httpGet:
    path: /healthz
    port: 8000
  periodSeconds: 10

A toy model, not the kubelet, of the arithmetic behind these settings. The documentation's rule of thumb: if a container usually needs longer than initialDelaySeconds + failureThreshold × periodSeconds to start, give it a startup probe:

def startup_budget(initial_delay=0, failure_threshold=3, period=10):
    """Toy model: seconds a slow starter gets before the kubelet gives up."""
    return initial_delay + failure_threshold * period

print(startup_budget(failure_threshold=30, period=2))      # 60
print(startup_budget())                                   # 30

With every default, a container has about 30 seconds of failed checks before liveness acts. An app that boots in 45 seconds would be restarted in a loop, which is exactly what a startup probe prevents.

Next, the same toy model for thresholds. One result per period; a probe acts only after failureThreshold failures in a row, so a short blip does nothing:

def run_probe(results, kind, failure_threshold=3, success_threshold=1):
    """Toy model of one probe. results: True/False per period.
    Returns the events a probe of this kind would cause."""
    events, fails, oks, ready = [], 0, 0, False
    for t, ok in enumerate(results):
        fails = 0 if ok else fails + 1
        oks = oks + 1 if ok else 0
        if kind == "liveness" and fails == failure_threshold:
            events.append((t, "restart container"))
            fails = 0                          # a fresh process starts over
        if kind == "readiness":
            if ready and fails == failure_threshold:
                ready = False
                events.append((t, "remove from endpoints"))
            elif not ready and oks >= success_threshold:
                ready = True
                events.append((t, "add to endpoints"))
    return events

blip = [True, True, False, False, True, False, False, False, True]
print(run_probe(blip, "readiness"))
# [(0, 'add to endpoints'), (7, 'remove from endpoints'), (8, 'add to endpoints')]
print(run_probe(blip, "liveness"))       # [(7, 'restart container')]

When a container keeps exiting, the kubelet waits longer before each restart. The documented default starts at 10 seconds, doubles, and is capped at 300 seconds, and the timer resets after 10 minutes of running without problems:

def backoff_delays(restarts, cap=300):
    """Documented default: 10 s, doubling each restart, capped at 300 s."""
    return [min(10 * 2 ** i, cap) for i in range(restarts)]

print(backoff_delays(8))       # [10, 20, 40, 80, 160, 300, 300, 300]

Both models against independently written references on 3,000 random probe histories: restarts must happen exactly when the last k results since the previous restart are all failures, and the closed-form back-off must match the doubling loop:

import random

def restart_times(results, k):
    """Reference: restart at t when the last k results since the previous
    restart are all failures."""
    out, since = [], 0
    for t in range(len(results)):
        window = results[max(since, t - k + 1): t + 1]
        if len(window) == k and not any(window):
            out.append(t)
            since = t + 1
    return out

def backoff_reference(restarts, cap=300):
    delays, d = [], 10
    for _ in range(restarts):
        delays.append(d)
        d = min(d * 2, cap)
    return delays

random.seed(9)
ok = True
for _ in range(3000):
    k = random.randint(1, 5)
    results = [random.random() < 0.6 for _ in range(random.randint(0, 40))]
    got = [t for t, e in run_probe(results, "liveness", failure_threshold=k)]
    ok &= got == restart_times(results, k)
    n = random.randint(0, 12)
    ok &= backoff_delays(n) == backoff_reference(n)
print(ok)                      # True

The complexity

The costs are operational, not asymptotic:

  • Detection time. A failure is acted on after roughly failureThreshold × periodSeconds. Lower values react faster and misfire more often.
  • Probe load. Every probe is a request, every period, on every container. A health endpoint that queries a database multiplies that load across the fleet.
  • Back-off. A crash-looping container spends more and more time waiting; with the defaults, the gap reaches five minutes. Newer kubelet settings can change the cap per node.

Where it goes wrong

  • An aggressive liveness probe. The documentation warns that misconfigured liveness probes can cause cascading failures: containers restart under high load, requests fail, and the remaining Pods take more load. Liveness should signal an unrecoverable failure, such as a deadlock.
  • Liveness that checks dependencies. If the database is down, restarting every web container does not bring it back. The documentation points readiness, not liveness, at the case of an app that depends on external services after startup.
  • Expecting readiness to wait for anything. Liveness and readiness do not depend on each other. Use initialDelaySeconds or a startup probe to hold liveness back.
  • No startup probe for a slow boot. Stretching liveness thresholds to cover startup also slows deadlock detection for the rest of the container's life.
  • Adding liveness by reflex. If the process exits on its own when it breaks, the restart policy already handles it, and a liveness probe may not be needed.

How to say it in an interview

"Liveness decides restarts, readiness decides traffic, and startup holds the other two back until a slow app is up. A failing readiness probe only removes the Pod from the Service's endpoints; a failing liveness probe gets the container restarted, with exponential back-off if it keeps crashing. I make liveness cheap and narrow, catching real deadlocks, with a higher failure threshold than readiness, because an aggressive liveness probe under load restarts healthy-but-slow containers and turns a slowdown into a cascading failure."

How readiness feeds a Service's endpoints is covered in Kubernetes Services vs Ingress, and rollouts that wait for ready Pods are in Kubernetes rolling updates.

Sources