How a CDN Works: Edge Caching, Latency and Purging
9 min readBytePatterns
How a CDN works: why distance is latency you cannot tune away, what an edge hit and miss cost, origin shields, private pages, and purging vs versioned files.
You can make a server faster and tune the database, and a visitor on the other side of the world will still wait, because the bytes cross that distance at the speed of light in glass. A content delivery network attacks the one cost you cannot optimise: it puts copies of your files close to every visitor, so most requests never make the long trip.
The problem it solves
Three things go wrong when every request goes to one origin:
- Latency from distance. A page load is many round trips, each paying the full distance twice.
- Origin load. Every request for the same logo and script hits your servers.
- Spikes. A launch multiplies traffic that is mostly for the same few files.
A CDN answers all three with geographically spread caches, called edge locations or points of presence, each serving its region and asking your origin only for what it lacks.
The intuition
Start with physics. Light in optical fibre covers roughly 200 kilometres per millisecond, so a 6,000 km path costs at least 60 ms per round trip before routing detours or queueing. A fresh HTTPS connection needs several round trips before content arrives: the TCP handshake, a TLS 1.3 handshake, then the request. Three round trips across an ocean is 180 ms of pure distance; from an edge 60 km away, under 2 ms.
An edge is a pull-through cache. The visitor is routed, usually by DNS or anycast, to a nearby edge. On a miss, the edge fetches the file from the origin once, stores it, and serves it. Every later visitor nearby gets a hit, and the origin is never touched. The origin's load becomes the miss rate, not the request rate.
Response headers decide what may be cached and for how long. Cache-Control: public, max-age=... allows shared caches; s-maxage sets the lifetime for shared caches only; private or no-store keeps a response out of the edge. A page personalised for one signed-in user has nothing to reuse and must never reach someone else, so it passes straight through. The edge still helps: it terminates TLS nearby and keeps warm connections to the origin, so uncached requests skip the long handshakes.
Two refinements show up in any serious design:
- Origin shield. A new file is a miss at every edge, and each miss goes to the origin. A middle tier, one regional cache in front of the origin, turns those misses into one.
- Invalidation. Cached copies outlive a bad deploy. A purge must reach every edge; as of October 2026 that takes from under a second to several minutes, depending on the provider. The cleaner answer is versioned filenames: a hash of the content in the name, so new content is a new object.
Watch it run
The animation opens on the lesson's line: distance is latency you cannot tune away, so a CDN keeps copies near the visitor, with the origin 6,000 km off. The first visitor asks for the logo and is routed to the nearest edge, 60 km away. The edge has no copy yet; that is a miss. So it fetches the file once from the origin, across the ocean, and the latency readout shows 310 ms. It stores the file under its cache rules and serves the visitor. A second visitor nearby asks for the same file: a hit, served in 12 ms, and the origin is never touched. Every later visitor in the region is the same hit; at 97 of 98 the origin has seen one request. A page personalised for one signed-in user has nothing to reuse, so it skips the edge. And a bad deploy is now cached at fifty edges at once, which is why assets ship under a versioned filename: a new name is a new object, and no purge is needed.
Content Delivery Networks
Step 1 of 11
Distance is latency you cannot tune away. A CDN keeps copies near the visitor.
The same interactive animation as the lesson — step through it with the controls.
The code
A toy model. First the latency floor, then 100,000 requests over 50 edges, each holding the 200 most recently used of 2,000 objects whose popularity falls off with rank, with and without a shield in front of the origin:
import hashlib
import random
from collections import OrderedDict
KM_PER_MS = 200 # light in fibre: roughly 200,000 km per second
def rtt_ms(km):
return 2 * km / KM_PER_MS # there and back, before any queueing or routing detours
def new_connection_ms(km, round_trips=3):
return round_trips * rtt_ms(km) # TCP handshake + TLS 1.3 handshake + the request
print(rtt_ms(6_000), new_connection_ms(6_000), round(new_connection_ms(60), 1)) # 60.0 180.0 1.8
class LRU:
"""One cache: a fixed number of objects, least recently used evicted first."""
def __init__(self, capacity):
self.capacity, self.items = capacity, OrderedDict()
def get(self, key):
if key in self.items:
self.items.move_to_end(key)
return True
self.items[key] = True
if len(self.items) > self.capacity:
self.items.popitem(last=False)
return False # a miss: the caller fetched it from upstream
def simulate(requests, edges, edge_capacity, shield_capacity=None):
"""requests: (edge, object). Returns (edge hits, origin fetches)."""
caches = [LRU(edge_capacity) for _ in range(edges)]
shield = LRU(shield_capacity) if shield_capacity else None
hits = origin = 0
for edge, obj in requests:
if caches[edge].get(obj):
hits += 1
elif shield is None or not shield.get(obj):
origin += 1 # nobody upstream had it
return hits, origin
rng = random.Random(33)
objects = 2_000
weights = [1 / rank for rank in range(1, objects + 1)] # a few objects are very popular
traffic = [(rng.randrange(50), obj)
for obj in rng.choices(range(objects), weights, k=100_000)]
for label, shield in (("no shield", None), ("shield of 500", 500), ("shield of 2000", 2_000)):
hits, origin = simulate(traffic, 50, 200, shield)
print(label, round(hits / len(traffic), 3), origin)
# no shield 0.588 41200
# shield of 500 0.588 22958
# shield of 2000 0.588 2000
def versioned(name, content):
"""Put a hash of the bytes in the filename: new content, new object, no purge."""
stem, ext = name.rsplit(".", 1)
return "%s.%s.%s" % (stem, hashlib.sha256(content).hexdigest()[:8], ext)
print(versioned("logo.svg", b"<svg>v1</svg>"), versioned("logo.svg", b"<svg>v2</svg>"))
# logo.bb27c39a.svg logo.0b66dbb6.svg
The edges hit 58.8% of the time, yet the origin served 41,200 requests: each edge misses separately on the long tail. A shield holding the catalogue cuts that to 2,000. Then a check on 2,000 seeded random traces: edge hits against a brute-force list LRU, and with caches too large to evict, origin fetches equal the distinct (edge, object) pairs, or the distinct objects with a shield:
def lru_hits_brute(keys, capacity):
"""Brute force: a plain list, most recent at the end, scanned on every request."""
order, hits = [], 0
for k in keys:
if k in order:
hits += 1
order.remove(k)
order.append(k)
if len(order) > capacity:
order.pop(0)
return hits
ok = True
for _ in range(2_000):
edges, n_obj = rng.randint(1, 6), rng.randint(1, 30)
reqs = [(rng.randrange(edges), rng.randrange(n_obj)) for _ in range(rng.randint(0, 120))]
cap = rng.randint(1, 12)
hits, _ = simulate(reqs, edges, cap)
ok &= hits == sum(lru_hits_brute([o for e, o in reqs if e == k], cap) for k in range(edges))
big = len(reqs) + 1 # nothing is ever evicted
ok &= simulate(reqs, edges, big)[1] == len(set(reqs)) # one origin fetch per (edge, object)
ok &= simulate(reqs, edges, big, big)[1] == len({o for _, o in reqs}) # one per object
print(ok) # True
The complexity
- Per request at the edge:
O(1)average lookup; latency set by the distance to the nearest edge. - Origin load: the miss rate times traffic, growing with edges times distinct objects per edge, unless a shield collapses them.
- Storage: every edge copies what is popular in its region, so total storage is many times the catalogue.
Where it goes wrong
- Caching personalised responses. A shared cache can hand one user's page to another; mark them private.
- Purging as the release process. It is slow to propagate and easy to get wrong; version the filenames.
- Cache keys that include noise. A tracking query parameter or an unneeded header in the key splits one object into thousands of misses; CloudFront caching walks through cache keys and TTLs on one provider.
- Expecting a CDN to fix slow APIs. Uncached dynamic requests still go to the origin; the CDN shortens the handshakes, not the query.
When it shows up in interviews
As "what is a CDN?", and as a component in almost every read-heavy design: video streaming, image hosting and static sites. Expect follow-ups on invalidation, what must not be cached, and surviving a cold cache. The general patterns are in caching explained.
How to say it in an interview
"Distance is latency I can't optimise: light in fibre does about 200 km per millisecond, and a new HTTPS connection needs several round trips, so a cross-ocean visitor pays a lot before any work happens. A CDN puts pull-through caches near users: the first request at an edge misses and fetches from origin once; everyone nearby after that is a hit. Cache-Control decides what's shareable; personalised pages are private. I'd add an origin shield so many edges' misses become one fetch, and ship static assets under content-hashed filenames, so a deploy never needs a purge."