WebSockets vs Polling vs Server-Sent Events Explained
8 min readBytePatterns
WebSockets vs polling vs server-sent events: the HTTP upgrade handshake, frames and masking, the latency and request cost of each, and scaling open sockets.
Any feature that says "live", such as chat, prices, a shared cursor or a delivery on a map, needs the server to tell the browser something the browser did not ask for. HTTP was built the other way round: the client asks, the server answers. There are three standard ways around that, and interviews expect you to pick one for a reason, not by habit.
The problem it solves
A client wants updates the moment they happen, but it cannot know when that is. The options:
- Polling: ask on a timer. Simple, stateless, and mostly wasted: with a 5-second interval, an hour is 720 requests, even if only three updates happened. An update also waits for the next poll, 2.5 seconds on average.
- Long polling: ask, and let the server hold the request open until there is news or a timeout. Updates arrive promptly, but every update costs a fresh request.
- Server-sent events (SSE): one long HTTP response that the server keeps writing events into. One-way, server to client, plain HTTP, with automatic reconnection built into the browser API.
- WebSockets: one connection upgraded into a channel either side can write to at any time.
The intuition
A WebSocket starts life as an ordinary HTTP GET carrying Upgrade: websocket and a random Sec-WebSocket-Key. The server proves it understood by hashing that key with a fixed string from the protocol and answering 101 Switching Protocols with the result. From then on the TCP connection no longer carries HTTP. It carries frames: a byte with the final-fragment flag and an opcode (text, binary, close, ping, pong), a length, and the payload. Frames from the client are masked with a random four-byte key, which protects proxies that might misread the bytes as HTTP; frames from the server are not.
The design consequence is the one the lesson leads with. An open socket is state on one machine. The server that holds Bob's socket is the only one that can write to him. When Alice's message lands on another server, something has to carry it across: sticky routing alone does not help, because Alice and Bob are pinned to different machines. The usual answer is a pub/sub layer that every socket server subscribes to.
That is also the decision rule. If only the server talks, SSE is less to run. If updates are rare or seconds of delay are fine, polling is the cheapest to operate. WebSockets earn their cost when both sides talk often and latency matters.
Watch it run
The animation starts with polling, which asks "anything new?" on a timer whether or not there is news. Request one is a whole HTTP round trip, headers and all. Nothing has changed, so the answer is an empty 204 and the round trip bought nothing. Five seconds later it asks again, and an hour of that is 720 requests for three updates. When news does arrive, it waits up to a full interval before anyone hears it. Then the switch: a WebSocket starts as one ordinary HTTP request asking to upgrade, and the server answers 101 Switching Protocols. The channel is open. Now the server pushes a price the moment it changes, and nobody asked; then the client sends a bid over the same line, because either side writes whenever it has something to say. The cost comes next: that socket is state on one machine, so server 2 cannot answer for it, and scaling needs sticky routing or pub/sub. A close frame, or a dropped connection, ends the session.
WebSockets and Realtime
Step 1 of 11
Polling asks “anything new?” on a timer, whether or not there is news.
The same interactive animation as the lesson — step through it with the controls.
The code
The handshake's one computation, checked against the example in the protocol specification:
import base64
import hashlib
GUID = "258EAFA5-E914-47DA-95CA-C5AB0DC85B11" # fixed by the protocol
def accept_key(client_key):
"""What the server's 101 response must echo in Sec-WebSocket-Accept."""
digest = hashlib.sha1((client_key + GUID).encode()).digest()
return base64.b64encode(digest).decode()
print(accept_key("dGhlIHNhbXBsZSBub25jZQ==")) # s3pPLMBiTxaQ9kYGzzhZRbK+xOo=
A frame encoder and decoder. Lengths under 126 fit in the second byte; 126 and 127 announce a 16-bit or 64-bit length. Both printed frames match the specification's examples:
import struct
TEXT, BINARY, CLOSE, PING, PONG = 0x1, 0x2, 0x8, 0x9, 0xA
def encode_frame(payload, opcode=TEXT, mask=None):
"""One final frame. Clients must mask; servers must not."""
head = bytes([0x80 | opcode]) # FIN bit + opcode
n, m = len(payload), 0x80 if mask else 0
if n < 126:
head += bytes([m | n])
elif n < 65536:
head += bytes([m | 126]) + struct.pack("!H", n)
else:
head += bytes([m | 127]) + struct.pack("!Q", n)
if mask:
payload = bytes(b ^ mask[i % 4] for i, b in enumerate(payload))
head += mask
return head + payload
def decode_frame(data):
fin, opcode = data[0] >> 7, data[0] & 0x0F
masked, n, i = data[1] >> 7, data[1] & 0x7F, 2
if n == 126:
n, i = struct.unpack("!H", data[2:4])[0], 4
elif n == 127:
n, i = struct.unpack("!Q", data[2:10])[0], 10
key = data[i:i + 4] if masked else b"\0\0\0\0"
i += 4 if masked else 0
body = bytes(b ^ key[j % 4] for j, b in enumerate(data[i:i + n]))
return fin, opcode, bool(masked), body
print(encode_frame(b"Hello").hex(" ")) # from the server
# 81 05 48 65 6c 6c 6f
print(encode_frame(b"Hello", mask=bytes.fromhex("37fa213d")).hex(" ")) # from a client
# 81 85 37 fa 21 3d 7f 9f 4d 51 58
print(decode_frame(encode_frame(b"bid 43", mask=b"\x01\x02\x03\x04")))
# (1, 1, True, b'bid 43')
The polling bill from the animation, with three seeded update times in one hour:
import random
def polling_hour(events, interval):
"""Requests sent, and how late each event is seen, when polling every `interval` s."""
requests = 3600 // interval
delays = [interval - t % interval for t in events] # seen at the next poll
return requests, max(delays), sum(delays) / len(delays)
random.seed(29)
events = [random.uniform(0, 3600) for _ in range(3)]
req, worst, avg = polling_hour(events, 5)
print(req, "requests for", len(events), "updates; worst delay", round(worst, 2), "s")
# 720 requests for 3 updates; worst delay 3.54 s
A toy model of the scaling problem. Alice's server receives a message for Bob, whose socket lives on another server. Without a shared bus it is lost; with one, every server hears it and the one holding Bob's socket delivers:
class Server:
"""Toy model: the open sockets live in this process and nowhere else."""
def __init__(self, name, bus=None):
self.name, self.sockets, self.bus = name, {}, bus
if bus is not None:
bus.append(self)
def send(self, to, text):
for server in (self.bus if self.bus is not None else [self]):
server.deliver(to, text) # with a bus, every server hears it
def deliver(self, to, text):
if to in self.sockets:
self.sockets[to].append((self.name, text))
for bus in (None, []):
s1, s2 = Server("s1", bus), Server("s2", bus)
s1.sockets["alice"], s2.sockets["bob"] = [], []
s1.send("bob", "hi") # alice's server, bob's socket is elsewhere
print("bus" if bus is not None else "no bus", s2.sockets["bob"])
# no bus []
# bus [('s2', 'hi')]
Checked on 2,000 seeded random frames, including the 125, 126, 65,535 and 65,536-byte length boundaries: every frame must have the size the length rule predicts and decode back to its payload. A 20,000-event simulation must give a mean polling delay of half the interval:
random.seed(29)
ok = True
for _ in range(2_000):
n = random.choice([0, 1, 125, 126, 127, 65535, 65536, random.randint(0, 3000)])
payload = bytes(random.getrandbits(8) for _ in range(n))
mask = bytes(random.getrandbits(8) for _ in range(4)) if random.random() < 0.5 else None
frame = encode_frame(payload, random.choice([TEXT, BINARY, PING]), mask)
header = 2 + (2 if 126 <= n < 65536 else 8 if n >= 65536 else 0) + (4 if mask else 0)
ok &= len(frame) == header + n # brute force: the size rule
ok &= decode_frame(frame)[3] == payload # and the round trip
lots = [random.uniform(0, 3600) for _ in range(20_000)]
ok &= abs(polling_hour(lots, 5)[2] - 2.5) < 0.05 # mean delay: half the interval
print(ok) # True
The complexity
- Polling:
3600 / intervalrequests per client per hour, whatever happens; average delay half the interval. - Long polling: one request per update plus one per timeout; delay close to zero.
- WebSockets and SSE: one connection per client, a few bytes of framing per message, delay close to zero. The cost moves into memory and file descriptors per open connection, and into keeping them alive.
Where it goes wrong
- No heartbeat. A dead mobile connection can look open for minutes. Send pings and time out silent sockets.
- No reconnect plan. Clients must reconnect with backoff and resume from the last message they saw, or they miss updates.
- Pinning without a bus. Sticky sessions keep one user on one server; they do not deliver a message to a user on another.
- WebSockets for one-way feeds. A score ticker is SSE or polling, with far less to operate.
As of September 2026, WebSockets are also defined over HTTP/2 and HTTP/3, but support varies by browser, proxy and server; and SSE over HTTP/1.1 is limited by the browser's small per-host connection cap.
When it shows up in interviews
In design a chat app, live dashboards, collaborative editors and ride tracking, usually as "how does the server reach the client?" Follow-ups cover connection counts per server, how a message finds the server holding a socket, and what happens during a deploy. Load balancing long-lived connections is its own twist: least-connections beats round robin when connections last hours.
How to say it in an interview
"It depends on who talks and how often. For rare updates I'd poll: stateless and cheap to run. For a one-way feed I'd use server-sent events. For two-way, low-latency traffic like chat I'd use WebSockets: an HTTP request upgraded with a 101, then framed messages either way. The catch is that each socket is state on one server, so I'd put a pub/sub layer behind the socket servers so a message can reach whichever server holds the recipient, plus heartbeats and client reconnects with resume."