Skip to content
BytePatterns

Transformer Architecture Explained: One Block, Stacked

9 min readBytePatterns

The transformer architecture explained: token vectors, positions, a block of attention plus a per-position network with residuals, stacked, then vocab scores.

A transformer diagram can look like a circuit board, but the architecture behind most language models is one idea repeated. Turn tokens into vectors, add where each one sits, then pass the sequence through a stack of same-shaped blocks, each letting positions exchange information and then processing each position on its own. Once that one block is clear, depth is just repetition.

The problem it solves

A language model must read a sequence and produce, for each position, a prediction of what comes next. That requires two kinds of computation:

  • Mixing across positions. The meaning of "bank" depends on words before it, sometimes far before it.
  • Thinking per position. Once a position has gathered context, it needs a nonlinear transformation to turn that context into useful features.

Earlier recurrent networks mixed by reading strictly left to right, carrying a state forward: cheap per token, but sequential, so training cannot process positions in parallel, and distant information must survive many steps. The transformer, introduced in 2017 for translation, replaced recurrence with attention over all positions at once.

The intuition

Follow one sequence through:

  1. Tokens. Text is split into tokens, usually subword pieces from a tokenizer.
  2. Embeddings. Each token id looks up a learned vector of width d, the model dimension. The sequence is now a matrix of shape n × d. See embeddings for what those numbers mean.
  3. Positions. Attention is a weighted sum, and a sum does not care about order: shuffle the inputs and the outputs shuffle with them. So position information is added to the vectors, or applied inside attention.
  4. The block, repeated. Each block has two sublayers. Attention lets every position build its own blend of the others. A feed-forward network, usually about four times wider than d in its middle layer, then transforms each position independently. Each sublayer is wrapped in a residual connection, adding its output to its input rather than replacing it, and a normalisation step.
  5. The output. After the last block, a final linear layer maps each position's vector to one score per vocabulary entry. Softmax and sampling turn the last position's scores into the next token.

"Identical blocks" means identical in shape, not weights: every block has its own parameters. Each maps n × d to n × d, so they stack freely. A rule of thumb: with a feed-forward width of 4d, a block has about 12d² weights, 4d² in attention and 8d² in the feed-forward network, ignoring biases and norms.

Text generators are decoder-only with a causal mask: position i may attend only to positions up to i, so training can predict every next token in parallel without cheating. As of October 2026, most widely used large language models follow this decoder-only design, commonly with normalisation before each sublayer and rotary position encodings applied inside attention; the 2017 original was an encoder-decoder for translation.

Watch it run

A transformer is a stack of identical blocks refining one sequence, so the animation starts with the text and splits it into tokens, nine here; the model never sees characters or words. It looks up a vector for each token, 768 dimensions, so meaning is now a list of numbers. Then it adds position information, because nothing in a weighted sum records which token came first. Those vectors enter the first block, and every block has the same two parts: attention mixes across positions, each one building its own blend of the others, then a small network transforms each position on its own, in parallel. Thirty-two identical desks, and no single one rewrites the manuscript. The last layer scores every entry in the vocabulary, 50,257 scores. More blocks, 96 in the next frame, build more abstract features and cost more time and memory. And attention scores every pair of positions, so cost grows with the square of the length: 1,000 tokens make a million pairs, 8,000 make 64 million. Recurrent models were cheaper per token but had to read strictly in order; that bill buys parallel training.

Transformers: Big Picture

Step 1 of 12

A transformer is a stack of identical blocks refining one sequence. Start with the text.

The same interactive animation as the lesson — step through it with the controls.

The code

A toy model in plain Python: width 8, three blocks, a ten-token vocabulary and random weights. It predicts nothing useful but has the real structure: pre-normalisation, residuals, a causal mask, a per-position network. It traces shapes and counts parameters:

import math
import random

rng = random.Random(33)
D, HIDDEN, VOCAB, MAX_LEN = 8, 32, 10, 16           # toy sizes; HIDDEN = 4 * D as is common

def matrix(rows, cols):
    return [[rng.gauss(0, 1 / math.sqrt(rows)) for _ in range(cols)] for _ in range(rows)]

def matmul(X, W):
    return [[sum(x * w for x, w in zip(row, col)) for col in zip(*W)] for row in X]

def add(X, Y):
    return [[a + b for a, b in zip(r, s)] for r, s in zip(X, Y)]

def layer_norm(X):
    out = []
    for row in X:                                       # each position on its own
        mean = sum(row) / len(row)
        var = sum((x - mean) ** 2 for x in row) / len(row)
        out.append([(x - mean) / math.sqrt(var + 1e-5) for x in row])
    return out

def softmax(v):
    m = max(v)
    e = [math.exp(x - m) for x in v]
    return [x / sum(e) for x in e]

class Block:
    """Same shape as every other block, its own weights."""
    def __init__(self):
        self.wq, self.wk, self.wv, self.wo = (matrix(D, D) for _ in range(4))
        self.w1, self.w2 = matrix(D, HIDDEN), matrix(HIDDEN, D)

    def attention(self, X, causal):                     # mixes ACROSS positions
        Q, K, V = matmul(X, self.wq), matmul(X, self.wk), matmul(X, self.wv)
        out = []
        for i, q in enumerate(Q):
            seen = range(i + 1) if causal else range(len(X))   # causal: no peeking ahead
            w = softmax([sum(a * b for a, b in zip(q, K[j])) / math.sqrt(D) for j in seen])
            out.append([sum(wj * V[j][d] for wj, j in zip(w, seen)) for d in range(D)])
        return matmul(out, self.wo)

    def mlp(self, X):                                   # transforms EACH position alone
        return matmul([[max(0.0, h) for h in row] for row in matmul(X, self.w1)], self.w2)

    def __call__(self, X, causal=True):
        X = add(X, self.attention(layer_norm(X), causal))   # residual: add, never replace
        return add(X, self.mlp(layer_norm(X)))

    def params(self):
        return sum(len(m) * len(m[0]) for m in (self.wq, self.wk, self.wv, self.wo, self.w1, self.w2))

EMBED, POS, UNEMBED = matrix(VOCAB, D), matrix(MAX_LEN, D), matrix(D, VOCAB)
BLOCKS = [Block() for _ in range(3)]

def transformer(tokens, causal=True, positions=True, trace=None):
    X = [EMBED[t][:] for t in tokens]                   # look up a vector per token
    if positions:
        X = add(X, POS[:len(tokens)])                   # otherwise order is invisible
    for block in BLOCKS:
        X = block(X, causal)
        if trace is not None:
            trace.append((len(X), len(X[0])))
    return matmul(layer_norm(X), UNEMBED)               # a score per vocabulary entry

shapes = []
logits = transformer([3, 1, 4, 1, 5], trace=shapes)
print(shapes, (len(logits), len(logits[0])))            # [(5, 8), (5, 8), (5, 8)] (5, 10)
print(BLOCKS[0].params(), 12 * D * D)                   # 768 768
d = 768
print(12 * d * d, 12 * 12 * d * d)                      # 7077888 84934656

The shape never changes inside the stack; only the final layer widens each position to the vocabulary. At width 768 the 12d² rule gives about 7.1 million weights per block, 85 million for twelve, before the embedding table. Next, three properties on 300 seeded random sequences: changing a token never changes an earlier output, bit for bit; without positions or a mask, shuffling the input only shuffles the output; with positions, the same shuffle changes the result:

def max_diff(A, B):
    return max(abs(a - b) for r, s in zip(A, B) for a, b in zip(r, s))

ok = True
for _ in range(300):
    n = rng.randint(2, 6)
    tokens = rng.sample(range(VOCAB), n)                # distinct tokens
    out = transformer(tokens)
    # 1. causal: changing token k never changes the outputs before k, bit for bit
    k = rng.randrange(n)
    changed = tokens[:k] + [(tokens[k] + 1) % VOCAB] + tokens[k + 1:]
    ok &= transformer(changed)[:k] == out[:k]
    # 2. no positions, no mask: shuffling the input just shuffles the output
    perm = rng.sample(range(n), n)
    plain = transformer(tokens, causal=False, positions=False)
    shuffled = transformer([tokens[p] for p in perm], causal=False, positions=False)
    ok &= max_diff(shuffled, [plain[p] for p in perm]) < 1e-9
    # 3. with positions, the same shuffle is no longer invisible
    if perm != list(range(n)):
        with_pos = transformer([tokens[p] for p in perm], causal=False)
        ok &= max_diff(with_pos, [transformer(tokens, causal=False)[p] for p in perm]) > 1e-6
print(ok)                                               # True

The complexity

  • Attention: O(n² · d) per block for n positions, from scoring every pair; memory for the scores grows as n² unless computed in tiles.
  • Feed-forward: O(n · d²) per block, linear in length.
  • Whole model: times the number of blocks; short inputs are dominated by the feed-forward part, long ones by attention.
  • Generation: one token at a time, reusing earlier keys and values from the KV cache, within the context window.

Where it goes wrong

  • "Identical blocks share weights." They share shape; each has its own parameters.
  • Forgetting position. Without it, the model sees a bag of tokens.
  • Feed-forward mixing positions. It does not; only attention moves information between positions.
  • Dropping the residuals. They let each block make a small edit and gradients reach early layers.
  • Missing the mask. Without a causal mask, training lets each position read the answer it must predict.

When it shows up in interviews

As "explain the transformer architecture", "why do transformers need positional encodings?" and "why is attention quadratic?". In ML engineering rounds it leads to parameter counting, the KV cache and long-context cost.

How to say it in an interview

"Tokens become vectors from an embedding table, and position information is added because attention alone is order-blind. Then a stack of blocks, same shape but separate weights. Each block has self-attention, which mixes information across positions, and a feed-forward network applied to each position independently, both wrapped in residual connections with normalisation, so a block maps n by d to n by d and they stack. A decoder-only model masks attention causally so each position sees only the past. A final linear layer scores the vocabulary. Attention is quadratic in sequence length, which is the price of parallel training."