EC2 Auto Scaling Explained: Target Tracking, Health Checks, ALB
9 min readBytePatterns
How an EC2 Auto Scaling group holds min, max and desired capacity, how target tracking picks a size, when health checks replace instances, and ALB versus NLB.
"Put the servers in an Auto Scaling group behind a load balancer" is the default AWS answer to "how does it handle more traffic?" The follow-ups test what the group actually does: how it picks a size, what happens when an instance breaks, and why the load balancer and the group can disagree about health. Everything below comes from the Amazon EC2 Auto Scaling and Elastic Load Balancing guides listed at the end, as of September 2026.
The problem it solves
A fixed fleet is wrong most of the day: sized for the peak, you pay for idle instances at night; sized for the average, the peak takes you down. Instances also fail. You want one address for clients, a fleet size that follows the load within limits you choose, and broken instances replaced without anyone being paged.
The intuition
The Auto Scaling group launches instances from a launch template and has three numbers: minimum, maximum and desired capacity. It keeps the desired number running, and scaling policies only change the desired capacity, always between the minimum and the maximum. Across several Availability Zones it distributes the desired capacity and keeps the zones balanced as it scales, so losing a zone costs a share of the fleet, not all of it.
Target tracking is the usual policy. You pick a metric and a target value, such as average CPU at 50%, and the service manages the CloudWatch alarms; like a thermostat, it adds capacity above the target and removes it below. It rounds conservatively: if it computes that 1.5 more instances would bring CPU to the target, it adds 2; if removing 0.5 instances would push CPU above the target, it does not scale in. Scale-in is deliberately more gradual than scale-out.
Health checks come from several sources, and two matter here. Amazon EC2 status checks are always on and are the group's default. Elastic Load Balancing health checks are ignored by default: the load balancer stops routing to a failing target on its own, but the group only replaces an instance the load balancer calls unhealthy if you turn ELB health checks on for the group. A grace period keeps a booting instance from being judged too early; the guide gives 300 seconds as the console default and 0 for groups created with the CLI or an SDK.
The load balancer is the single address. An Application Load Balancer works at layer 7 and routes by path, host, headers and more; a Network Load Balancer works at layer 4 for TCP, UDP and TLS. Instances the group launches are registered with the target group automatically.
Watch it run
The animation starts with two instances in two zones behind an ALB, min 2, max 4. Traffic doubles; the policy, holding CPU near 50%, raises the desired capacity, and the group launches two instances, one per zone. The ALB routes to each once its health check passes. Then i-2 fails its health check: the ALB stops sending it traffic at once, and the group replaces it, which assumes ELB health checks are on for the group. At night the group scales in toward the minimum, never below it.
EC2, Auto Scaling & Load Balancers
Step 1 of 11
Two instances in two Availability Zones behind a load balancer. The group says: min 2, max 4.
The same interactive animation as the lesson — step through it with the controls.
The code
The lesson's policy, from the AWS CLI:
# illustrative: needs an AWS account and an existing group named web-asg
aws autoscaling put-scaling-policy \
--auto-scaling-group-name web-asg \
--policy-name cpu50 \
--policy-type TargetTrackingScaling \
--target-tracking-configuration \
'{"PredefinedMetricSpecification": {"PredefinedMetricType": "ASGAverageCPUUtilization"}, "TargetValue": 50.0}'
A toy model of the sizing, not the AWS algorithm: the smallest fleet that holds the target, clamped to the group's bounds. Rounding up reproduces the guide's two examples, a fractional instance added in full and a scale-in that would overshoot skipped:
def target_capacity(current, avg_cpu, target, min_size, max_size):
"""Toy model of target tracking on average CPU, not the AWS algorithm.
Load is treated as current * avg_cpu 'percent-instances'. Rounding up makes
a scale-out add enough and a scale-in stop before it would overshoot."""
needed = -(-current * avg_cpu // target) # integer ceiling of the load / target
return max(min_size, min(max_size, needed))
print(target_capacity(2, 90, 50, 2, 4)) # 4 ceil(3.6), the animation's scale-out
print(target_capacity(2, 90, 50, 2, 3)) # 3 capped by the maximum
print(target_capacity(4, 20, 50, 2, 4)) # 2 night: back to the minimum, not below
print(target_capacity(3, 40, 50, 1, 10)) # 3 removing one would push CPU to 60%
print(target_capacity(4, 60, 50, 2, 10)) # 5 4.8 instances' worth, rounded up
Two more toy models: zone balance on launch, and when an instance is replaced:
def launch_zones(counts, n):
"""Toy model of AZ balance: each launch goes to the zone with the fewest."""
counts = dict(counts)
for _ in range(n):
zone = min(sorted(counts), key=lambda z: counts[z])
counts[zone] += 1
return counts
print(launch_zones({"a": 1, "b": 1}, 2)) # {'a': 2, 'b': 2}
print(launch_zones({"a": 2, "b": 0, "c": 1}, 3)) # {'a': 2, 'b': 2, 'c': 2}
def replaces(elb_check_on, lb_healthy, ec2_running, secs_in_service, grace=300):
"""Toy model of whether the group replaces an InService instance."""
if not ec2_running:
return True # not running: replaced even in the grace period
if secs_in_service < grace:
return False # still booting: optional checks wait
return elb_check_on and not lb_healthy # ELB results are ignored unless turned on
print(replaces(False, False, True, 900)) # False the ALB skips it; the group keeps it
print(replaces(True, False, True, 900)) # True
print(replaces(True, False, True, 60)) # False inside the grace period
print(replaces(False, True, False, 10)) # True
The sizing rule against a reference that tries every fleet size, on 5,000 random groups:
import random
def smallest_fleet(current, avg_cpu, target, min_size, max_size):
"""Reference: try every size and keep the smallest that holds the target."""
load = current * avg_cpu
for n in range(min_size, max_size + 1):
if load <= n * target:
return n
return max_size
random.seed(16)
ok = True
for _ in range(5000):
lo = random.randint(0, 5)
hi = random.randint(max(lo, 1), 20)
cur = random.randint(max(lo, 1), hi)
cpu, tgt = random.randint(0, 100), random.choice([30, 40, 50, 60, 70])
ok &= target_capacity(cur, cpu, tgt, lo, hi) == smallest_fleet(cur, cpu, tgt, lo, hi)
print(ok) # True
The complexity
The costs here are time and money:
- Scaling takes minutes. An instance must launch, boot and pass health checks, and it is not counted in the group's metrics until its warmup time expires. For a predictable spike, add scheduled or predictive scaling.
- Metric resolution sets reaction time. EC2 metrics arrive every five minutes by default; the guide recommends one-minute metrics for target tracking.
- The minimum is paid for around the clock; the maximum caps the bill during a runaway.
Where it goes wrong
- Expecting the group to act on ALB health. Without ELB health checks turned on, a hung application that passes EC2 status checks gets no traffic but is never replaced.
- A grace period shorter than boot time. New instances are judged unhealthy and replaced in a loop.
- The wrong metric. The guide says a target tracking metric must change in proportion to the number of instances. Total
RequestCounton the load balancer does not;ALBRequestCountPerTargetdoes. - Several policies. The group scales out if any target tracking policy wants to, and in only if all of them do: availability first.
When it shows up in interviews
It shows up in AWS system design rounds as "the site gets ten times the traffic during a sale, what happens?" A strong answer names min for availability, max for cost, target tracking in between, several zones, and health checks wired to the load balancer. The other common question is ALB versus NLB: ALB when routing depends on the HTTP request, such as /api to one service and /images to another; NLB for raw TCP or UDP, millions of requests per second, or a static IP address per zone.
How to say it in an interview
"I run the web tier in an Auto Scaling group across at least two Availability Zones, with a minimum for availability and a maximum for cost. A target tracking policy holds average CPU near a target, 50% for example, by changing the desired capacity; it rounds up on scale-out and scales in gradually. An Application Load Balancer is the entry point and routes only to healthy targets, and I turn on ELB health checks for the group, with a grace period longer than boot time, so an unhealthy instance is replaced. For raw TCP or UDP I would use a Network Load Balancer."
The whole architecture this tier belongs to is in design a system on AWS, and the same idea for containers is in Kubernetes requests, limits and the HPA.
Sources
- Auto Scaling groups — Amazon EC2 Auto Scaling User Guide
- Target tracking scaling policies — Amazon EC2 Auto Scaling User Guide
- Health checks for instances in an Auto Scaling group and about the health checks — Amazon EC2 Auto Scaling User Guide
- Set the health check grace period — Amazon EC2 Auto Scaling User Guide
- What is an Application Load Balancer? — Elastic Load Balancing
- What is a Network Load Balancer? — Elastic Load Balancing