Skip to content
BytePatterns

AWS Cost Optimization: Spot vs Savings Plans vs Reserved

8 min readBytePatterns

AWS cost optimization in order: right-size, scale, commit the floor with Savings Plans or Reserved Instances, restartable work on Spot, and watch NAT traffic.

"How would you cut this AWS bill?" is a common architecture interview question, and the weak answer is a list of discounts. The strong answer is an order: stop paying for what you do not use, make what remains follow demand, and only then commit to the part that never goes away. Discounts on waste are still waste. Everything below comes from the AWS documentation pages listed at the end, as of September 2026.

The problem it solves

A service needs three instances all night and nine at its busiest hour. Sized for the peak and left running, every hour pays for nine, and most of that capacity idles. The problem is not the price of an instance but how many you pay for, and when.

The EC2 purchasing options give you different prices for different kinds of use:

  • On-Demand: pay by the second for what you launch, with no commitment.
  • Savings Plans: a commitment to a consistent amount of usage, measured in dollars per hour, for a one- or three-year term.
  • Reserved Instances: a commitment to a consistent instance configuration, including instance type and Region, for one or three years.
  • Spot Instances: unused EC2 capacity at a significant discount, which EC2 can take back.

The intuition

Pull the levers in the lesson's order.

One, switch off idle and right-size. Compute Optimizer flags over-provisioned instances. A smaller type is a saving with no commitment attached.

Two, scale with demand. An Auto Scaling group pays for the shape of the day instead of its peak.

Three, commit to the floor, the part of the curve that never goes away. Compute Savings Plans are the most flexible: they apply regardless of instance family, size, Region, operating system or tenancy, and also to Fargate and Lambda usage, with prices up to 66% below On-Demand. EC2 Instance Savings Plans go up to 72% in exchange for committing to one instance family in one Region. A commitment cannot be changed after purchase, so size it from what you measured, not from what you hope.

Four, send interruptible work to Spot. Spot is spare capacity, and EC2 gives a two-minute interruption notice before it stops or terminates the instance. The notice arrives as an EventBridge event and as an instance-action item in instance metadata, and AWS recommends checking every five seconds. Delivery is best effort, so the work must survive losing the machine: checkpointed batch jobs, CI, rendering.

Five, watch the wires. A NAT gateway is charged for every hour it is available and for every gigabyte it processes. A gateway endpoint for S3 or DynamoDB takes that traffic off it, and keeping resources in their NAT gateway's Availability Zone avoids cross-zone transfer.

Watch it run

The animation shows one service's day in twelve two-hour blocks, each bar being how many instances that block actually needs. Sized for the peak and left running, every block pays for 9, and the gap between the bars and the ceiling is idle. Lever one right-sizes, with Compute Optimizer flagging the over-provisioned. Lever two scales with demand, so an Auto Scaling group pays for the shape of the day. Then notice the floor: 3 instances run every hour of every day, and steady, predictable use is exactly what a commitment is for. Lever three covers the floor with a Savings Plan or Reserved Instances, and what rises above it stays On-Demand. A nightly batch job needs 3 more at night; it can restart from a checkpoint, so lever four runs it on Spot, which EC2 can reclaim with a two-minute notice. Lever five hides outside compute: data transfer, where a NAT gateway bills per GB processed and a gateway endpoint takes S3 traffic off it. The closing frame gives the order to say it in.

AWS Cost Levers

Step 1 of 10

One service's day in two-hour blocks. Each bar is how many instances that block actually needs.

The same interactive animation as the lesson — step through it with the controls.

The code

The lesson's commands, which need Compute Optimizer enabled and Cost Explorer access:

# illustrative
aws compute-optimizer get-ec2-instance-recommendations \
  --filters name=Finding,values=Overprovisioned
aws ce get-cost-and-usage \
  --time-period Start=2026-09-01,End=2026-10-01 \
  --granularity MONTHLY --metrics UnblendedCost \
  --group-by Type=DIMENSION,Key=SERVICE

The animation's day, counted in instance-blocks rather than dollars:

DEMAND = [3, 3, 3, 3, 5, 8, 9, 9, 8, 6, 4, 3]     # instances needed per two-hour block
peak, floor = max(DEMAND), min(DEMAND)

print("sized for the peak:", peak * len(DEMAND))   # sized for the peak: 108
print("scaled with demand:", sum(DEMAND))          # scaled with demand: 64
print("floor, every block:", floor * len(DEMAND))  # floor, every block: 36

BATCH = [3, 3, 3, 3] + [0] * 8                      # the nightly job, restartable
committed = floor * len(DEMAND)
on_demand = sum(d - floor for d in DEMAND)
spot = sum(BATCH)
print(committed, on_demand, spot)                   # 36 28 12

How much to commit is where people overspend. A toy model, not AWS billing: the model counts instances, while a Savings Plan commits dollars per hour, and the committed rate below is an illustrative ratio, not a price. A committed instance is paid in every block, used or not:

from fractions import Fraction

def cost(level, demand, rate):
    """Toy bill in On-Demand instance-blocks: `level` committed at `rate`, the rest On-Demand."""
    return level * len(demand) * rate + sum(max(0, d - level) for d in demand)

rate = Fraction(6, 10)                              # illustrative ratio, not an AWS price
for level in (0, 3, 4, 9):
    print(level, float(cost(level, DEMAND, rate)))
# 0 64.0
# 3 49.6
# 4 49.8
# 9 64.8

The rule behind those numbers: the k-th committed instance is worth it only if it would be busy in more than rate of the blocks. The fourth instance is busy in 7 of 12 blocks, just under 60%, so committing it costs more than it saves:

def best_by_rule(demand, rate):
    """Commit the k-th instance only if it would be busy in more than `rate` of the blocks."""
    return sum(1 for k in range(1, max(demand) + 1)
               if Fraction(sum(d >= k for d in demand), len(demand)) > rate)

print(best_by_rule(DEMAND, rate))                   # 3   the floor
print(best_by_rule(DEMAND, Fraction(1, 2)))         # 4   a deeper discount justifies more

The rule against trying every commitment level, on 3,000 random demand curves and rates:

import random

random.seed(19)
ok = True
for _ in range(3000):
    demand = [random.randint(0, 12) for _ in range(random.randint(1, 24))]
    r = Fraction(random.randint(1, 99), 100)
    costs = [cost(level, demand, r) for level in range(max(demand) + 1)]
    ok &= cost(best_by_rule(demand, r), demand, r) == min(costs)
print(ok)                                           # True

The complexity

The costs are commitment risk and interruption risk:

  • Committing above the floor pays for idle capacity in every hour the demand dips, for one or three years.
  • Spot trades price for the chance of losing the instance at two minutes' notice, so restart cost must be small.

Where it goes wrong

  • Buying commitments before right-sizing. You lock in a discount on instances you did not need.
  • Committing to the peak. The commitment runs whether you use it or not.
  • Stateful work on Spot without checkpoints, such as a primary database.
  • Relying on the notice. It is best effort, and hibernation starts without the two-minute warning.
  • Ignoring data transfer. NAT gateways bill per gigabyte processed.

When it shows up in interviews

It appears in AWS and architecture interviews as "our bill doubled, what do you look at?", and as a follow-up to any design: "how would you make this cheaper?" Interviewers listen for the order and for matching each option to a usage pattern. The scaling half is covered in EC2 Auto Scaling, and the NAT and endpoint layout in VPC subnets and security groups.

How to say it in an interview

"First I would stop paying for waste: switch off idle resources and right-size using Compute Optimizer's findings. Then scale with demand, so we pay for the curve rather than the peak. Then I would measure the floor and cover it with a Compute Savings Plan, or an EC2 Instance Savings Plan or Reserved Instances if the family is stable, and leave peaks On-Demand. Interruptible, checkpointed work goes on Spot, which can be reclaimed with a two-minute notice. Finally I would check data transfer: NAT gateways charge per gigabyte, so S3 and DynamoDB traffic should go through gateway endpoints."

Sources