Skip to content
BytePatterns

LLM Quantization Explained: 8-Bit, 4-Bit, Scales and Outliers

8 min readBytePatterns

LLM quantization explained: how a scale maps weights to 8-bit or 4-bit integers, why fewer bytes speed up decoding, and why outlier weights force small groups.

A language model's weights are usually shipped as 16-bit floating-point numbers. Quantization stores them in 8 or 4 bits instead, and the surprise is not that the model gets smaller but that it gets faster, often with little visible change in quality. Both effects have one cause: generating text is limited by how quickly weights can be read from memory, not by arithmetic.

The problem it solves

A model with 7 billion parameters needs 14 GB for its weights at 2 bytes each, which decides what hardware can run it. And generation produces one token at a time, each forward pass reading essentially every weight once, so with a single request the arithmetic units mostly wait for bytes. Fewer bytes per weight means:

  • Less memory: 8 bits halves the footprint, 4 bits quarters it.
  • Faster decoding: fewer bytes cross the memory bus per token.
  • More room for everything else: the freed memory holds longer contexts and more concurrent requests, whose KV cache competes for the same space.

The price is precision. A 4-bit integer has 16 possible values. The question is how to spend them.

The intuition

Take a group of weights, a few dozen to a few hundred neighbours in one row. Map the largest magnitude to the largest integer available, 127 for signed 8-bit or 7 for signed 4-bit; the ratio is the group's scale. Every weight is divided by the scale and rounded. Stored: one scale per group, one small integer per weight. Read back: integer times scale.

Rounding moves each weight by at most half a step, and the step is the scale. So the error depends on two things:

  • Bits. From 8 to 4 bits the levels drop from 256 to 16, and each step here is about eighteen times wider (127 over 7). Error does not halve with the bits; it jumps.
  • The largest weight in the group. One large value stretches the scale for everyone, and tiny weights sharing a scale with an outlier all round to zero. Real models do contain such outliers, which is why schemes use small groups, keep sensitive layers in higher precision, or pick roundings that protect the channels that matter most.

Watch it run

The animation follows one weight, -0.82, stored in sixteen bits; every layer holds millions of these, and they all have to be read per token. Quantizing means keeping one scale for a group, here 0.95 / 127, and a small integer per weight. At eight bits the grid has 256 steps, so the value barely moves: it becomes -110, and the group's worst rounding error is 0.0035. Halve the bits again and the grid has sixteen steps; -0.82 becomes -6 and the rounding is visible, with a worst error of 0.037. Half the bits is a sixteenth of the levels, so the error does not halve, it jumps. What you bought is bandwidth: 0.5× the bytes per token at 8 bits and 0.25× at 4. That is why it speeds decoding up at all, since generation is memory-bound, not arithmetic-bound. Then the bill: a handful of weights are far larger than the rest, and one shared scale flattens the group around them. So quantize in small groups, keep the sensitive layers wide, and measure the task.

Quantization

Step 1 of 9

One weight, stored in sixteen bits. Every layer holds millions of these and they all have to be read per token.

The same interactive animation as the lesson — step through it with the controls.

The code

A toy model in plain Python, not a kernel: symmetric absmax quantization with one scale per group. It reproduces the lesson's numbers:

def quantize(weights, bits):
    """Toy model: symmetric absmax quantization of one group."""
    qmax = 2 ** (bits - 1) - 1                     # 127 for 8-bit, 7 for 4-bit
    scale = max(abs(w) for w in weights) / qmax or 1.0
    return [round(w / scale) for w in weights], scale

def dequantize(q, scale):
    return [v * scale for v in q]

def max_error(weights, bits):
    q, scale = quantize(weights, bits)
    return max(abs(a - b) for a, b in zip(weights, dequantize(q, scale)))

w = [-0.82, -0.11, 0.0, 0.37, 0.95]                 # the lesson's group
for bits in (8, 4):
    q, scale = quantize(w, bits)
    print(bits, q, round(max_error(w, bits), 4))
# 8 [-110, -15, 0, 49, 127] 0.0035
# 4 [-6, -1, 0, 3, 7] 0.0371

"Half a byte per weight" is literal: two 4-bit values share one byte, stored in two's complement and sign-extended when read back:

def pack4(q):                                       # two 4-bit values per byte
    nibbles = [v & 0xF for v in q] + [0] * (len(q) % 2)
    return bytes(lo | hi << 4 for lo, hi in zip(nibbles[::2], nibbles[1::2]))

def unpack4(data, n):
    out = []
    for byte in data:
        for nib in (byte & 0xF, byte >> 4):
            out.append(nib - 16 if nib > 7 else nib)   # back to signed
    return out[:n]

q4, _ = quantize(w, 4)
packed = pack4(q4)
print(len(packed), packed.hex(), unpack4(packed, len(q4)))
# 3 fa3007 [-6, -1, 0, 3, 7]

The memory arithmetic for 7 billion weights, in decimal gigabytes. The last line adds one 16-bit scale per group of 128 weights, which costs an eighth of a bit per weight:

params = 7_000_000_000
for label, bits_per_weight in [("16-bit", 16), ("8-bit", 8), ("4-bit", 4), ("4-bit + scales", 4 + 16 / 128)]:
    print(label, round(params * bits_per_weight / 8 / 1e9, 2), "GB")
# 16-bit 14.0 GB
# 8-bit 7.0 GB
# 4-bit 3.5 GB
# 4-bit + scales 3.61 GB

The outlier problem, measured on a seeded layer of 1,024 small weights plus one value 45 standard deviations out. With one scale for the layer, every small weight rounds to zero at 4 bits; smaller groups confine the damage:

import random

def grouped_error(weights, bits, group):
    total = 0.0
    for i in range(0, len(weights), group):
        chunk = weights[i:i + group]
        q, scale = quantize(chunk, bits)
        total += sum(abs(a - b) for a, b in zip(chunk, dequantize(q, scale)))
    return total / len(weights)

random.seed(28)
layer = [random.gauss(0, 0.02) for _ in range(1024)]
layer[300] = 0.9                                    # one outlier, 45 standard deviations out
for group in (1024, 128, 32):
    print(group, round(grouped_error(layer, 4, group), 5))
# 1024 0.01626
# 128 0.0037
# 32 0.00211

Checked on 2,000 seeded random groups at 2, 3, 4 and 8 bits, against a brute force that tries every integer level for every weight: rounding must pick a nearest level, stay in range, err by at most half a step, and 4-bit packing must round-trip:

random.seed(28)
ok = True
for _ in range(2_000):
    bits = random.choice([2, 3, 4, 8])
    ws = [random.uniform(-1, 1) * random.choice([0.01, 1, 5]) for _ in range(random.randint(1, 64))]
    q, scale = quantize(ws, bits)
    qmax = 2 ** (bits - 1) - 1
    ok &= all(-qmax <= v <= qmax for v in q)
    for x, v in zip(ws, q):                          # brute force: try every level
        best = min(range(-qmax, qmax + 1), key=lambda lvl: abs(x - lvl * scale))
        ok &= abs(x - v * scale) <= abs(x - best * scale) + 1e-12
        ok &= abs(x - v * scale) <= scale / 2 + 1e-12
    if bits == 4:
        ok &= unpack4(pack4(q), len(q)) == q
print(ok)                                           # True

The complexity

  • Storage: b bits per weight plus one scale per group of g, so about b + 16/g bits per weight with 16-bit scales.
  • Rounding error per weight: at most half the scale, and the scale is the group's largest magnitude divided by 2^(b-1) - 1.
  • Decoding speed: with a single request, roughly proportional to bytes read per token. Real speedups fall short of 4× because of dequantization work, activations and the KV cache.
  • Batching: when many requests share each weight read, arithmetic matters more and the gain shrinks.

Where it goes wrong

  • One scale for too many weights. A single outlier flattens everything around it, as the measurement shows.
  • Quantizing everything equally. Some layers are far more sensitive than others; mixed precision keeps them wider.
  • Trusting perplexity alone. A small change in an average score can hide a large drop on one task. Measure the task you care about.

As of September 2026, common practice is weight-only 8-bit or 4-bit storage with per-group scales, applied after training, with calibration that picks roundings to minimise each layer's output error or protects channels with large activations, plus 8-bit floating-point formats where hardware supports them. Details vary by library and hardware.

When it shows up in interviews

In ML engineering and inference-system interviews: "how would you serve this model on one smaller GPU?", "why is decoding memory-bound?", or "what does 4-bit cost you?". It sits next to the KV cache as the other big memory lever, and next to LoRA, since fine-tuning small adapters on top of a quantized base is a common way to adapt large models cheaply.

How to say it in an interview

"Quantization stores each group of weights as small integers plus one scale: divide, round, and multiply back when reading. At 8 bits the error is tiny; at 4 bits there are sixteen levels, so it jumps. The win is memory and speed: decoding reads every weight per token and is bandwidth-bound, so a quarter of the bytes means much faster tokens at small batch sizes. The risk is outliers, which stretch the scale and crush their neighbours, so I'd use small groups, keep sensitive layers wider, and evaluate on the actual task."