Skip to content
BytePatterns

Attention, Intuitively: What Does the Word 'Bank' Listen To?

7 min readBytePatterns

Attention is a weighted average whose weights are recomputed for every input. Queries, keys, values and the square-root scaling, built from scratch in Python.

"Attention" is the word most often used to explain modern language models and the word least often explained. Diagrams of stacked blocks and arrows make it look like machinery. Underneath, one attention step is something you already know: a weighted average. The only unusual part is where the weights come from — and that part is the whole idea.

The problem it solves

Take the word "bank". In "the river bank" it is about water; in "the bank loan" it is about money. A model that turns each word into a fixed vector gives "bank" the same vector in both sentences, somewhere between the two meanings. Nothing downstream can tell which one was meant.

What you want is for each word's representation to be rebuilt from its context: "bank" should come out of the layer pulled toward water in the first sentence and toward money in the second. And it has to work for any sentence, with any word in any position, without a hand-written rule for every ambiguous word in the language.

The intuition

Let every word ask the rest of the sentence a question, and take a blend of the answers it likes.

Each token carries three vectors, each produced by multiplying its embedding by a learned matrix:

  • a query — what this token is looking for,
  • a key — what this token offers, to be matched against other tokens' queries,
  • a value — what this token hands over if someone listens to it.

For one token, score its query against every key with a dot product: large when they point the same way, as in cosine similarity. Turn the scores into weights that are positive and sum to 1 with a softmax. The output is the weighted average of the values.

The weights are not parameters. They are computed fresh from the input, every time. The parameters only decide how queries and keys are made.

That is why the same model treats "bank" differently in two sentences: the learned matrices are identical, but the keys it is matched against are not. Nothing is selected and nothing is thrown away — every value contributes in proportion to its weight. Attention blends; it does not choose.

Watch it run

Each position is scored against every other — three tokens, nine scores. The animation isolates the row for "bank", turns its scores into weights of 0.1, 0.7 and 0.2, and blends the three value vectors into an output pulled toward "river". Then it changes the input and the whole desk is rebalanced.

Attention, Intuitively

Step 1 of 12

Attention lets every position decide which of the others to listen to.

The same interactive animation as the lesson — step through it with the controls.

The code

Keys and values here are written by hand in tiny dimensions so every number is readable. The value vectors have two components, which you can read as "water-ness" and "money-ness".

import math

def softmax(scores):
    top = max(scores)                          # subtract the max: no overflow
    exps = [math.exp(s - top) for s in scores]
    total = sum(exps)
    return [e / total for e in exps]

def attend(query, keys, values):
    d = len(query)
    scores = [sum(q * k for q, k in zip(query, key)) / math.sqrt(d) for key in keys]
    weights = softmax(scores)                  # how much to listen to each token
    return weights, [sum(w * v[i] for w, v in zip(weights, values))
                     for i in range(len(values[0]))]

#        what each token offers (key)   what it hands over (value: water, money)
KEY = {"the": [0, 0, 1], "river": [3, 0, 0], "loan": [0, 3, 0], "bank": [1, 1, 0]}
VAL = {"the": [0, 0],    "river": [1, 0],    "loan": [0, 1],    "bank": [0.5, 0.5]}
BANK_ASKS = [1, 1, 0]                          # "bank" looks for water or money

for sentence in (["the", "river", "bank"], ["the", "bank", "loan"]):
    w, out = attend(BANK_ASKS, [KEY[t] for t in sentence], [VAL[t] for t in sentence])
    print(sentence, [round(x, 2) for x in w], [round(x, 2) for x in out])
# ['the', 'river', 'bank'] [0.1, 0.58, 0.32] [0.74, 0.16]
# ['the', 'bank', 'loan'] [0.1, 0.32, 0.58] [0.16, 0.74]

Same query, same tables, two sentences. In the first, "bank" gives 58% of its attention to "river" and comes out as [0.74, 0.16] — mostly water. In the second, 58% goes to "loan" and the output flips to mostly money. "the" gets a small, non-zero share in both: it matches nothing, but softmax never gives exactly zero.

Now the mysterious / math.sqrt(d). With real dimensions — hundreds of components — dot products between unrelated vectors grow in size roughly with √d. Feed big scores to a softmax and it saturates: one weight near 1, the rest near 0, and gradients that barely flow during training. Dividing by √d keeps the scores in a range where the softmax is still a blend.

import random
random.seed(9)
d = 512
q = [random.gauss(0, 1) for _ in range(d)]
ks = [[random.gauss(0, 1) for _ in range(d)] for _ in range(8)]
raw = softmax([sum(a * b for a, b in zip(q, k)) for k in ks])
scaled = softmax([sum(a * b for a, b in zip(q, k)) / math.sqrt(d) for k in ks])
print(round(max(raw), 3), round(max(scaled), 3))           # 1.0 0.341

Unscaled, one of eight random keys takes essentially all the weight. Scaled, the largest share is about a third.

Real implementations do every token at once as matrix products. To check that the loop above is the same computation, compare it with the matrix form on random inputs:

import numpy as np
rng = np.random.default_rng(0)
ok = True
for _ in range(200):
    n, dk, dv = rng.integers(1, 9), rng.integers(1, 17), rng.integers(1, 5)
    Q, K, V = rng.normal(size=(n, dk)), rng.normal(size=(n, dk)), rng.normal(size=(n, dv))
    S = Q @ K.T / np.sqrt(dk)                  # all n × n scores at once
    W = np.exp(S - S.max(axis=1, keepdims=True))
    W /= W.sum(axis=1, keepdims=True)
    rows = [attend(list(Q[i]), [list(k) for k in K], [list(v) for v in V])[1]
            for i in range(n)]
    ok &= np.allclose(W @ V, rows) and np.allclose(W.sum(axis=1), 1)
print(ok)                                                  # True

Two hundred random cases, every row identical, every row of weights summing to 1.

The complexity

For a sequence of n tokens with dimension d, every token scores every token: n² dot products of length d, so O(n² · d) time and an n × n weight matrix per attention head. Doubling the input length quadruples that work. This quadratic term is the reason long context is expensive, and why so much engineering goes into avoiding recomputation — see the KV cache.

Where it goes wrong

  • Thinking attention picks one word. The output is a blend. A weight of 0.58 still leaves 42% of the output coming from elsewhere.
  • Treating the weights as learned constants. They change with every input. What is learned is the projections that produce queries, keys and values.
  • Dropping the scale factor. Without √d, softmax saturates in high dimensions and training suffers.
  • A naive softmax. exp of a large score overflows. Subtracting the maximum first changes nothing mathematically and keeps it finite.
  • Forgetting order. Attention on its own treats the input as a set: shuffle the tokens and each token's output is unchanged, just reordered. Position information has to be added separately.

For how these steps stack into a full model, see transformers, the big picture.

How to say it in an interview

"Each token produces a query, a key and a value through learned projections. For a given token, I dot its query with every key, scale by √d so the softmax doesn't saturate, and softmax the scores into weights that sum to one. The output is the weighted average of the values. The weights depend on the input, which is how the same word gets different representations in different contexts. The cost is quadratic in sequence length, because every token is scored against every other."

If you can add that attention by itself ignores word order, you have covered the follow-up before it is asked.