Skip to content
BytePatterns

Kubernetes Requests vs Limits and the HPA Formula Explained

9 min readBytePatterns

What Kubernetes requests and limits do, why CPU is throttled but memory is OOM-killed, how the HPA computes replicas from the request, and key defaults.

Two lines in a container spec, requests and limits, decide where a Pod runs, what happens when it misbehaves, and whether it can autoscale at all. They are not one setting with two sizes: they are different mechanisms used by different components, and the HorizontalPodAutoscaler builds on one of them. Everything below comes from the Kubernetes documentation pages listed at the end, as of September 2026.

The problem it solves

A cluster packs many Pods onto a few nodes. Three questions follow:

  • Placement: which node has room for this Pod?
  • Protection: what stops one Pod from starving its neighbours?
  • Scaling: when the Pods run hot, how many more are needed?

Requests answer the first, limits the second, and the HPA the third, using the request as its yardstick.

The intuition

A request is a reservation. The scheduler places a Pod only on a node where the sum of the requests fits in the node's capacity, whatever the containers are actually using at that moment. If you set a limit but no request, Kubernetes copies the limit and uses it as the request.

A limit is a ceiling, enforced differently per resource. CPU is compressible: a container that tries to use more than its CPU limit is throttled, running slower but still running. Memory is not: the documentation says memory limits are enforced reactively, and a container that goes over can be OOM-killed by the kernel. So a CPU limit costs latency, and a memory limit that is too low costs restarts.

Requests and limits together set the QoS class. Every container with equal CPU and memory requests and limits makes the Pod Guaranteed; any request or limit short of that makes it Burstable; none at all makes it BestEffort, the first candidate for eviction when a node runs short.

The HPA measures utilization against the request. Pods requesting 250m CPU and using 225m are at 90%. The controller computes:

desiredReplicas = ceil(currentReplicas × currentMetricValue / desiredMetricValue)

With 4 replicas at 90% and a target of 60%, that is ceil(4 × 90 / 60) = 6. Three defaults keep it calm: it skips any change when the ratio is within a tolerance of 0.1 of 1.0; it re-evaluates every 15 seconds; and before scaling down it waits out a stabilization window of 300 seconds, acting on the highest recommendation from that window. If a container has no request for the metric, the documentation is explicit: utilization is not defined, and the autoscaler takes no action for it.

Watch it run

The animation draws one bar per Pod, the height being its CPU use as a percentage of its 250m request, against a 60% target. Traffic spikes to 90% and the HPA computes ceil(4 × 90 / 60) = 6. The new Pods need 250m each and no node has room, so they wait Pending until node autoscaling adds a node. With six Pods the load lands on target. A wobble to 64% is a ratio of 1.07, inside the tolerance, so nothing happens. In the evening use falls to 30%: the formula says 3, but minReplicas is 4, and even that waits out the 300-second window. Without a request, the percentage has no denominator.

Requests, Limits & HPA

Step 1 of 12

One bar per Pod: its CPU use as a percentage of its request, 250m. The HPA's target is 60%.

The same interactive animation as the lesson — step through it with the controls.

The code

The lesson's manifest: a memory limit, requests for both resources, and an HPA holding CPU at 60% of the request:

# illustrative: requests and a limit on the container...
resources:
  requests: { cpu: 250m, memory: 256Mi }
  limits: { memory: 512Mi }
---
# ...and an HPA that holds average CPU at 60% of the request
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata: { name: api }
spec:
  scaleTargetRef: { apiVersion: apps/v1, kind: Deployment, name: api }
  minReplicas: 4
  maxReplicas: 10
  metrics:
    - type: Resource
      resource: { name: cpu, target: { type: Utilization, averageUtilization: 60 } }

A toy model of the HPA's CPU rule, not the controller. It uses exact fractions so that a ratio of exactly 7 is never rounded up to 8 by floating point:

import math
from fractions import Fraction

def hpa_desired(replicas, usage_m, request_m, target_pct, min_r=1, max_r=10, tolerance=0.1):
    """Toy model of the HPA's CPU rule, not the controller. usage_m and
    request_m are millicores per Pod; utilization is usage over request."""
    if request_m is None:
        return replicas                        # no request: no utilization, no action
    ratio = Fraction(usage_m * 100, request_m * target_pct)   # exact: 7.0 stays 7
    if abs(ratio - 1) <= Fraction(str(tolerance)):
        return replicas                        # close enough: leave it alone
    return max(min_r, min(max_r, math.ceil(replicas * ratio)))

print(hpa_desired(4, 225, 250, 60, 4, 10))   # 6   90% of the request, ceil(4 x 1.5)
print(hpa_desired(6, 160, 250, 60, 4, 10))   # 6   64%: ratio 1.07, inside the tolerance
print(hpa_desired(6, 75, 250, 60, 4, 10))    # 4   ceil(6 x 0.5) = 3, raised to minReplicas
print(hpa_desired(4, 225, None, 60, 4, 10))  # 4   no request, nothing to divide by

The scale-down window as a toy model: every 15 seconds the controller applies the highest recommendation of the last 300, so a load drop at 60 seconds removes Pods only at 360. And the QoS rules for the lesson's container and two others:

def scale_down_with_window(recommendations, window=300, period=15):
    """Toy model: each sync, apply the highest recommendation from the last
    `window` seconds, so a brief dip does not remove Pods."""
    applied, seen = [], []
    for i, rec in enumerate(recommendations):
        seen.append((i * period, rec))
        now = i * period
        applied.append(max(r for t, r in seen if now - t <= window))
    return applied

recs = [6] * 4 + [4] * 22                     # load drops at t = 60 s
out = scale_down_with_window(recs)
print(out[3], out[4], out[23], out[24])       # 6 6 6 4
print((24 - 4) * 15)                          # 300  seconds from the dip to the scale-in

def qos_class(containers):
    """Toy model of the QoS rules for container-level resources."""
    if all(c.get("requests") == c.get("limits") and c.get("limits")
           and {"cpu", "memory"} <= c["limits"].keys() for c in containers):
        return "Guaranteed"
    if any(c.get("requests") or c.get("limits") for c in containers):
        return "Burstable"
    return "BestEffort"

print(qos_class([{"requests": {"cpu": "250m", "memory": "256Mi"},
                  "limits": {"memory": "512Mi"}}]))            # Burstable
print(qos_class([{"requests": {"cpu": "1", "memory": "1Gi"},
                  "limits": {"cpu": "1", "memory": "1Gi"}}]))  # Guaranteed
print(qos_class([{}]))                                         # BestEffort

With the tolerance off, the formula should equal "the fewest Pods that keep average utilization at or under the target". A reference that tries every count agrees on 5,000 random cases:

import random

def smallest_replicas(replicas, usage_m, request_m, target_pct, min_r, max_r):
    """Reference: the fewest Pods that keep average utilization at or under target."""
    total = replicas * usage_m                # all the CPU the Pods are using
    for n in range(min_r, max_r + 1):
        if total * 100 <= target_pct * n * request_m:
            return n
    return max_r

random.seed(16)
ok = True
for _ in range(5000):
    r, req = random.randint(1, 12), random.choice([100, 250, 500])
    use, tgt = random.randint(0, 2 * req), random.choice([50, 60, 70, 80])
    lo, hi = random.randint(1, r), random.randint(r, 20)
    ok &= hpa_desired(r, use, req, tgt, lo, hi, tolerance=0) == \
          smallest_replicas(r, use, req, tgt, lo, hi)
print(ok)                                     # True

The complexity

The costs are capacity, latency and reaction time:

  • Requests too high reserve node capacity nobody uses; too low and the scheduler packs Pods onto nodes that cannot serve them all at once.
  • Scale-up is fast but not instant. By default the HPA may add up to 100% of the current replicas or 4 Pods per 15-second period, whichever is more, and new Pods may still wait Pending for a node.
  • Scale-down is slow on purpose. The 300-second window trades a few minutes of spare Pods for no flapping.

Where it goes wrong

  • No CPU request. The HPA cannot compute utilization and does nothing for that metric, silently from the application's point of view.
  • A tight CPU limit. Throttling shows up as latency, not errors, which makes it hard to spot.
  • A memory limit below the real peak. The container is OOM-killed and restarts, and the HPA cannot help: more replicas do not change a per-container limit.
  • HPA and a fixed replicas in the same manifest. Re-applying the manifest resets the count the HPA chose; leave replicas out once an HPA owns it.

When it shows up in interviews

It shows up in platform, SRE and backend interviews, often as arithmetic: "Pods request 250m and use 225m; what utilization does the HPA see?" (90%), or "4 replicas at 90% with a 60% target?" (6). Then comes the design follow-up: why the new Pods are Pending (node autoscaling), and why a memory leak is not fixed by the HPA. In practice this is everyday cluster tuning: setting requests from observed usage so that both the scheduler and the autoscaler see the truth.

How to say it in an interview

"A request is what the scheduler reserves; a limit is the ceiling. Past its CPU limit a container is throttled; past its memory limit it can be OOM-killed. The HPA measures CPU as a percentage of the request, so without a request it will not scale on CPU. It computes ceil(replicas × current / target), ignores changes within a 0.1 tolerance, and waits out a 300-second window before scaling down. If the new Pods do not fit, they stay Pending until node autoscaling adds capacity."

Rolling out new versions of the same Deployment is covered in Kubernetes rolling updates, and the same scaling idea on EC2 is in EC2 Auto Scaling explained. Health checks that decide whether a Pod gets traffic are in probes and self-healing.

Sources