Requests, Limits & HPA
Docker & Kubernetes for Interviews: lesson 10 of 12
Reserve what a Pod needs, cap what it takes, and scale out when it runs hot.
Lesson 10 of 12 · 7 min
Requests, Limits & HPA
Step 1 of 12
One bar per Pod: its CPU use as a percentage of its request, 250m. The HPA's target is 60%.
The Idea
A request is what the scheduler reserves for a container; a limit is the ceiling. Past its CPU limit a container is throttled; past its memory limit it can be OOM-killed. The HorizontalPodAutoscaler adds or removes Pods to hold a metric, such as CPU utilization measured against the request.
Real-World Example
An api Deployment requests 250m CPU per Pod and targets 60%. Traffic spikes, the average hits 90%, and the HPA computes ceil(4 × 90 / 60) = 6 replicas. If the new Pods do not fit, they wait Pending until node autoscaling adds a node.
The Tradeoff
Without a request, CPU utilization has no denominator and the HPA will not act on it. Requests set too high waste nodes; set too low, the scheduler packs Pods onto nodes that cannot serve them all. Scale-down waits out a stabilization window, so replicas do not flap.
Hands-On
# illustrative — requests and a limit on the container…
resources:
requests: { cpu: 250m, memory: 256Mi }
limits: { memory: 512Mi }
---
# …and an HPA that holds average CPU at 60% of the request
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata: { name: api }
spec:
scaleTargetRef: { apiVersion: apps/v1, kind: Deployment, name: api }
minReplicas: 4
maxReplicas: 10
metrics:
- type: Resource
resource: { name: cpu, target: { type: Utilization, averageUtilization: 60 } }
Your turn
Put the steps in the right order.
- It computes ceil(4 × 90 / 60) = 6 and raises the Deployment's replicas
- The metrics pipeline reports each Pod's CPU use
- Two new Pods are scheduled, or wait Pending for a new node
- The HPA compares average utilization with its 60% target
Mini quiz
1 / 3
Pods request 250m CPU and use 225m on average. The HPA sees CPU utilization of:
Sources
- Resource management for Pods and containers — Kubernetes Documentation
- Horizontal Pod Autoscaling — Kubernetes Documentation
- Pod Quality of Service classes — Kubernetes Documentation
- Node autoscaling — Kubernetes Documentation