Skip to content
BytePatterns

CloudFront Caching Explained: Cache Keys, TTLs and Invalidation

9 min readBytePatterns

How CloudFront caches at edge locations and regional edge caches, how the cache key and TTLs decide hits, and why versioned file names beat invalidations.

"Put a CDN in front of it" is an easy sentence in a system design answer. The follow-ups are harder: what is cached, for how long, and what happens on deploy? For Amazon CloudFront, the cache key, the TTLs and how you retire old files answer all three. Everything below comes from the Amazon CloudFront Developer Guide pages listed at the end, as of September 2026.

The problem it solves

The origin, an S3 bucket or a load balancer, lives in one Region; viewers are everywhere. Every trip to the origin costs the viewer latency and the origin load. A CDN keeps copies close to viewers, which raises three questions:

  • Where are the copies kept, and in how many layers?
  • When are two requests "the same object"?
  • How long does a copy stay, and how do you replace it early?

The intuition

Two layers of cache. DNS routes a request to a point of presence, an edge location, typically the nearest one in terms of latency. On a miss, the POP usually asks a regional edge cache, which sits between the POPs and the origin and has a larger cache, so less popular objects stay close to viewers for longer. Only if both miss does CloudFront go to the origin. The answer is cached in both layers, so every POP in that region shares one copy. There are documented exceptions: PUT, POST, PATCH, OPTIONS and DELETE go straight from the POP to the origin, dynamic requests skip the regional layer, and so does an S3 origin in the same Region as the regional edge cache.

The cache key is the unique identifier of an object in the cache. By default it holds the distribution's domain name and the URL path, nothing else: two requests for the same path with different query strings, cookies and User-Agent headers are a hit. A cache policy can add headers, cookies and query strings. Each one added splits the cache, and the guide's example is Accept-Language, where en-US,en and en,en-US mean the same thing but become different objects.

The TTL decides how long a copy lives. The cache policy sets minimum, maximum and default TTLs; the origin's Cache-Control header chooses a value inside those bounds. With no cache policy the default TTL is 24 hours. s-maxage, if present, is what CloudFront uses, while browsers use max-age. When a copy expires, CloudFront asks the origin again, and the origin answers 304 Not Modified if the cached version is still current.

Watch it run

The animation sends a request for /app.js through edge A, the regional cache and the origin, then shows a second viewer on edge B hitting the regional copy while the origin hears nothing. After a deploy the edges keep serving the old file until the TTL runs out; the last frames compare an invalidation with a versioned file name.

CloudFront & Caching Layers

Step 1 of 10

The origin lives in one Region. Viewers are everywhere, so CloudFront puts caches in edge locations near them.

The same interactive animation as the lesson — step through it with the controls.

The code

The lesson's commands, with placeholder names: a hashed bundle cached for a year, and an invalidation for the files that keep their names:

# illustrative: bucket and distribution ID are placeholders
aws s3 cp app.3f9a1c.js s3://site-bucket/js/ \
  --cache-control "public, max-age=31536000, immutable"
aws cloudfront create-invalidation --distribution-id E2EXAMPLE \
  --paths "/index.html" "/css/*"

A toy model of the guide's expiration table, not CloudFront itself. A header value is clamped into the policy's bounds; without one, the default applies:

def edge_ttl(min_ttl, default_ttl, max_ttl, max_age=None, s_maxage=None, no_store=False):
    """How long an edge keeps an object. no_store stands for Cache-Control
    no-cache, no-store and/or private."""
    if no_store:
        return min_ttl                           # 0 unless a minimum TTL overrides it
    header = s_maxage if s_maxage is not None else max_age
    if header is None:
        return max(min_ttl, default_ttl)
    return min(max(header, min_ttl), max_ttl)    # clamp into the policy's bounds

DAY, YEAR = 86400, 31536000
print(edge_ttl(0, DAY, YEAR))                            # 86400      no header: default TTL
print(edge_ttl(0, DAY, YEAR, max_age=YEAR))              # 31536000   hashed asset
print(edge_ttl(0, DAY, 3600, max_age=YEAR))              # 3600       capped by maximum TTL
print(edge_ttl(60, DAY, YEAR, max_age=5))                # 60         raised to minimum TTL
print(edge_ttl(0, DAY, YEAR, max_age=60, s_maxage=600))  # 600        edges use s-maxage
print(edge_ttl(0, DAY, YEAR, no_store=True))             # 0

The two layers and the cache key, simulated: a thousand requests for three files, each with a random ?ref= parameter. With the default key, the path, the origin sees each file once; with the whole URL, most requests:

import random

def origin_calls(requests, key_of):
    """Toy model: POP caches in front of one shared regional edge cache."""
    pops, regional, calls = {}, set(), 0
    for pop, url in requests:
        k = key_of(url)
        cache = pops.setdefault(pop, set())
        if k in cache:
            continue                             # POP hit
        if k not in regional:
            calls += 1                           # both layers missed
            regional.add(k)
        cache.add(k)
    return calls

random.seed(15)
paths = ["/app.js", "/app.css", "/logo.png"]
reqs = [(random.choice("ABC"), random.choice(paths) + "?ref=" + str(random.randint(1, 500)))
        for _ in range(1000)]
print(origin_calls(reqs, lambda url: url.split("?")[0]),
      origin_calls(reqs, lambda url: url))       # 3 722

Invalidation billing: the first 1,000 paths a month are free across all distributions in the account, and a wildcard path counts once however many files it removes. The guide's own example is three distributions with 600 paths each:

def billable_paths(paths_this_month):
    return max(0, len(paths_this_month) - 1000)

print(billable_paths(["/img/a.png"] * 1800), billable_paths(["/*"]))   # 800 0

The clamp against a reference that follows the table row by row, on 5,000 random policies and headers:

def table_ttl(mn, df, mx, max_age, s_maxage, no_store):
    if no_store:
        return 0 if mn == 0 else mn
    h = s_maxage if s_maxage is not None else max_age
    if h is None:
        return df if mn == 0 else max(mn, df)
    if mn == 0:
        return min(h, mx)                        # "the lesser of" the header or maximum TTL
    if h < mn:
        return mn
    if h > mx:
        return mx
    return h

ok = True
for _ in range(5000):
    mn, df, mx = sorted(random.choice([0, 0, 1, 60, 3600, DAY, YEAR]) for _ in range(3))
    ma, sm = (random.choice([None, 0, 5, 600, DAY, 10 * YEAR]) for _ in range(2))
    ns = random.random() < 0.1
    ok &= edge_ttl(mn, df, mx, ma, sm, ns) == table_ttl(mn, df, mx, ma, sm, ns)
print(ok)                                        # True

The complexity

The costs here are hit ratio and money:

  • Every value in the cache key multiplies the number of copies. User-Agent has thousands of variations and session cookies can be unique per user, so the guide calls both poor candidates. Values the origin needs but that do not change the response belong in an origin request policy, not the key.
  • A longer TTL means more hits and less origin load; a shorter one means fresher content.
  • Invalidations are billed per path beyond 1,000 a month, and each path in a request counts separately.
  • Versioned file names cost nothing to invalidate; you still pay to transfer the new file to the edges.

Where it goes wrong

  • Caching per user by accident. A session cookie in the key gives every user a private copy and a hit ratio near zero.
  • A minimum TTL above zero. CloudFront then caches for that minimum even when the origin says no-cache, no-store or private.
  • Invalidating instead of versioning. An invalidation clears CloudFront, but a viewer's browser or a corporate proxy may keep the old file until it expires. A new name is a new key everywhere.
  • Viewers bypassing the CDN. Keep the S3 bucket private and grant read access only to the distribution with origin access control; the bucket policy names the cloudfront.amazonaws.com service principal and the distribution's ARN. AWS recommends OAC over the legacy origin access identity.

How to say it in an interview

"CloudFront answers from the nearest edge location; on a miss it usually asks a regional edge cache, and only then the origin, so all POPs in a region share one copy. What counts as the same object is the cache key, by default the domain and path, and I add only the headers, cookies or query strings that really change the response, because each one splits the cache. The cache policy sets TTL bounds and the origin's Cache-Control picks a value inside them. For deploys I use content-hashed file names cached for a year, and invalidate only the few files that keep their names, like index.html. The S3 origin stays private behind origin access control."

Where CloudFront fits in a full design is in design a system on AWS, and caching as a general tool is in CDNs.

Sources