Skip to content
BytePatterns

Design Video Streaming: Transcoding, Segments and CDN Edges

9 min readBytePatterns

Design a video streaming service for a system design interview: transcoding into a bitrate ladder, segments and manifests, CDN edges, and adaptive bitrate.

"Design a video streaming service" sounds like a storage problem, but storing files is easy. The hard parts are that every viewer's connection is different and keeps changing, and that the bytes leaving your system dwarf everything else. Transcoding, segments, a CDN and a player that adapts are the four ideas that answer those two facts, and the interview is mostly about explaining why each one is there.

The problem it solves

Requirements you should state up front:

  • Start fast. Playback begins within a couple of seconds, anywhere in a two-hour film.
  • Never stall. On a weak or changing connection, lower the quality instead of buffering.
  • Serve huge egress. 20,000 concurrent viewers at 5 Mb/s is 100 Gb/s leaving the system, on the lesson's numbers.
  • Accept uploads of large files that arrive once and are watched many times.

The read-to-write ratio is extreme. That single fact justifies spending a lot of work once, at upload time, to make every one of the many reads cheap.

The intuition

Follow one video from upload to screen:

  1. Upload. The client sends the original in chunks to object storage, so a dropped connection resumes instead of restarting. A metadata record (title, owner, status) goes to a database.
  2. Transcode once, into a ladder. A pipeline, fed by a message queue, encodes the original into several renditions: in the lesson, 1080p at 6 Mb/s, 720p at 3 Mb/s and 360p at 0.8 Mb/s. Workers can split the video into chunks and encode them in parallel. It costs CPU and storage, paid once per upload.
  3. Segment and describe. Each rendition is cut into short segments, a few seconds each, and a manifest lists them. As of September 2026, HLS and MPEG-DASH are the two widespread formats built on this idea.
  4. Serve from the edge. Segments are plain, immutable files, so a CDN caches them. The first request at an edge location is a miss that pulls from origin; every nearby viewer after that is a hit. See caching for the general pattern.
  5. Adapt in the player. The player reads the manifest, measures how fast segments actually arrive, and picks the rendition for the next segment. This is adaptive bitrate streaming, and it is client-side because only the client can measure the link it is using.

Segments are what make the last step possible: quality can only change at a segment boundary, so short segments react faster, while longer ones are more efficient to serve. For live video that trade becomes latency, since the player sits a few segments behind real time.

Watch it run

The animation starts from the lesson's sizing: twenty thousand viewers at 5 Mb/s is 100 Gb/s of egress, and that decides the shape. The upload is transcoded once, into several bitrates rather than one: 1080p at 6 Mb/s, 720p at 3, 360p at 0.8. Each rendition is cut into short segments, listed in a manifest. A player reads the manifest and asks the nearest edge for segment one. It is a cold edge, so that one segment is pulled from origin and kept. Every viewer nearby after that is a hit, so origin serves roughly one copy. Then the train enters a cutting and the measured bandwidth drops to 1.4 Mb/s, under 720p's 3. So the player asks for the next segment at 360p: quality drops, playback does not. The switch is only possible at a segment boundary, which is what segments buy. The client decides, because only the client can measure the link it is actually using. Live is the same picture with a clock: longer segments are cheaper and later, here about ten seconds behind with 4-second segments.

Design Video Streaming

Step 1 of 11

Twenty thousand viewers at 5 Mb/s is 100 Gb/s of egress. That decides the shape.

The same interactive animation as the lesson — step through it with the controls.

The code

A toy model with illustrative numbers. First the sizing, then a player that picks the highest rendition fitting inside 80% of the last measured throughput, on a train journey where the link drops from 5 to 1.4 Mb/s for four segments:

LADDER = [("1080p", 6.0), ("720p", 3.0), ("360p", 0.8)]   # Mb/s, highest first
SEGMENT_S = 4

viewers, avg_mbps = 20_000, 5
print(viewers * avg_mbps / 1_000, "Gb/s of egress")          # 100.0 Gb/s of egress
per_hour_gb = sum(r for _, r in LADDER) * 3_600 / 8 / 1_000
print(round(per_hour_gb, 2), "GB stored per hour of video")  # 4.41 GB stored per hour of video
print(2 * 3_600 // SEGMENT_S * len(LADDER), "segments for a 2 h film")   # 5400 segments for a 2 h film

def pick(measured_mbps, safety=0.8):
    """Highest rendition that fits inside a safety margin of the measured link."""
    for name, rate in LADDER:
        if rate <= safety * measured_mbps:
            return name, rate
    return LADDER[-1]                                       # never refuse to play

def play(trace, adaptive, start=("720p", 3.0)):
    """trace[i] = link speed while segment i downloads. Returns (stall s, renditions)."""
    buffer, stall, measured, chosen = 0.0, 0.0, None, []
    for bw in trace:
        name, rate = pick(measured) if adaptive and measured else start
        download = rate * SEGMENT_S / bw                    # seconds for this segment
        if chosen:                                          # playback drains the buffer
            stall += max(0.0, download - buffer)
            buffer = max(0.0, buffer - download)
        buffer += SEGMENT_S
        measured = bw                                       # the player measures the link
        chosen.append(name)
    return round(stall, 1), chosen

train = [5, 5, 5, 5, 1.4, 1.4, 1.4, 1.4, 5, 5]              # Mb/s, a cutting in the middle
print(play(train, adaptive=False)[0])                       # 13.5
stall, chosen = play(train, adaptive=True)
print(stall, " ".join(name[:-1] for name in chosen))
# 0.0 720 720 720 720 720 360 360 360 360 720

A player stuck on 720p stalls for 13.5 seconds in the cutting. The adaptive one does not stall at all, but notice the fifth segment: it was requested at 720p because the drop had not been measured yet, and only the buffer built up earlier covered its slow download. That is why real players also watch their buffer level, not just throughput. Next, the edge: 20,000 viewers spread over 50 edge locations watch the first 30 segments, and a brute-force check on 2,000 seeded random cases confirms that origin reads equal the number of distinct edge, rendition and segment triples, and that pick equals trying every rendition:

from collections import defaultdict
import random

def origin_reads(requests):
    """requests: (edge, rendition, segment). A cold edge pulls once, then keeps it."""
    cached, reads = defaultdict(set), 0
    for edge, rendition, seg in requests:
        if (rendition, seg) not in cached[edge]:
            reads += 1
            cached[edge].add((rendition, seg))
    return reads

rng = random.Random(32)
crowd = [(rng.randrange(50), "720p", seg) for _ in range(20_000) for seg in range(30)]
print(len(crowd), origin_reads(crowd))                      # 600000 1500

ok = True
for _ in range(2_000):
    reqs = [(rng.randrange(4), rng.choice(LADDER)[0], rng.randrange(20))
            for _ in range(rng.randint(0, 200))]
    ok &= origin_reads(reqs) == len(set(reqs))              # brute force: distinct triples
    m = rng.uniform(0.1, 12)
    fits = [r for r in LADDER if r[1] <= 0.8 * m]
    ok &= pick(m) == (max(fits, key=lambda r: r[1]) if fits else LADDER[-1])
print(ok)                                                   # True

600,000 segment requests cost origin 1,500 reads, a quarter of one percent. The model's edges never evict; real ones do, so the long tail of rarely watched videos is where origin load comes from.

The complexity

  • Transcoding: proportional to video length times the number of renditions, paid once per upload and parallel across chunks.
  • Storage: the sum of the ladder's bitrates per hour of video, here about 4.4 GB, plus the original.
  • Serving: O(1) per segment request at the edge; origin load scales with edges times distinct segments, not with viewers.
  • Player decision: O(r) per segment for r renditions.

Where it goes wrong

  • Serving from origin. 100 Gb/s of egress is the bill the CDN exists to avoid.
  • One big file. No quality switching, and seeking is awkward to cache.
  • Server-side quality choice. Only the player can see its link.
  • Throughput-only adaptation. It reacts one segment late; the buffer is the safety net.
  • Live with long segments. Cheap to serve, but viewers see the goal after the neighbours cheer.

When it shows up in interviews

As "design a video streaming platform", "design a video upload and playback service" or "design live streaming". Expect follow-ups on the upload path, the transcoding queue, how view counts are aggregated without touching the hot path, and live latency.

How to say it in an interview

"Reads dwarf writes and egress dominates, so I'd do the work once at upload: store the original, then a queue feeds transcoding workers that produce a bitrate ladder, say 1080p, 720p and 360p, each cut into segments of a few seconds with a manifest. Segments are immutable, so a CDN serves them and origin only sees cache misses. The player reads the manifest, measures throughput and its buffer, and picks the rendition per segment, stepping down instead of stalling. For live, shorter segments trade serving efficiency for lower latency."