Skip to content
BytePatterns

Cosine Similarity Explained: Why Search Ignores Length

7 min readBytePatterns

The measure behind vector search, taken apart: what the angle means, why a dot product alone favours long documents, and when to normalise once instead.

Every semantic search box, every retrieval step in front of a language model, every "related articles" strip ends up computing the same small number. It is not a neural network and it is not learned. It is one line of arithmetic with a specific reason for each of its three parts.

The problem it solves

An embedding turns a piece of text into a list of numbers — a point in a space with a few hundred or a few thousand dimensions, positioned so that related meanings land near each other. Fine. Now you have a query as one such list and a million documents as a million others, and you need "near each other" to be an actual number.

The obvious candidate is the dot product: multiply the vectors element by element and add it up. It is fast, and it is the right shape — it grows when the two vectors agree. It also has a flaw that ruins search:

def dot(a, b):
    return sum(x * y for x, y in zip(a, b))

short = [1, 2, 0]          # a short paragraph about cheap flights
long_same = [4, 8, 0]      # the same topic, four times as much text

print(dot(short, long_same), dot(long_same, long_same))   # 20 80

long_same is more similar to itself than short is to it — fine — but notice why. Nothing about the topic changed between those two numbers. The second is larger because the vector is longer. Rank by raw dot product and the longest documents win every query, regardless of what they are about.

The intuition

Split a vector into the two things it is carrying: a direction and a length. For an embedding, the direction is the meaning and the length is mostly an artefact — how much text there was, how emphatic it was, how the model happened to scale it.

Cosine similarity is the dot product after both lengths have been divided out. What remains is the angle, and the angle is the meaning.

Divide the dot product by both magnitudes and you get the cosine of the angle between the two vectors, which lands on a fixed scale that means the same thing for every pair:

  • 1 — same direction. Same topic, whatever the size.
  • 0 — perpendicular. Unrelated; they share no signal at all.
  • -1 — opposite directions.

Those endpoints are not conventions someone chose. They are what the cosine function does between 0 and 180 degrees, which is exactly why this measure is comparable across pairs while a raw dot product is not. In practice many embedding models cluster their scores in a narrow positive band, so cosine gives you a stable ordering rather than a calibrated probability.

Watch it run

Two vectors, one angle. Change a vector's length and watch the score stay put; change its direction and watch the score move.

Cosine Similarity

Step 1 of 11

Cosine similarity compares direction and deliberately ignores magnitude.

The same interactive animation as the lesson — step through it with the controls.

That invariance is the whole design: the measure was built to be blind to exactly the quantity you do not want to rank on.

The code

import math

def dot(a, b):
    return sum(x * y for x, y in zip(a, b))

def norm(v):
    return math.sqrt(dot(v, v))          # a vector's length is its own dot product

def cosine(a, b):
    return dot(a, b) / (norm(a) * norm(b))

short = [1, 2, 0]          # "cheap flights to rome"
long_same = [4, 8, 0]      # same topic, four times as much text
other = [0, 0, 3]          # "how to bake sourdough"
opposite = [-1, -2, 0]

print(round(cosine(short, long_same), 4))   # 1.0
print(round(cosine(short, other), 4))       # 0.0
print(round(cosine(short, opposite), 4))    # -1.0

The length that broke the dot product is gone: short and long_same score a perfect 1.0 because they point the same way, and the four-times difference in magnitude is divided out.

For a real index you would not do that division a million times per query. Normalise once, at write time, and the expensive part disappears:

def unit(v):
    n = norm(v)
    return [x / n for x in v]            # length 1, direction unchanged

u, w = unit(short), unit(long_same)
print(round(dot(u, w), 4))               # 1.0 — a plain dot product now

def euclid(a, b):
    return math.sqrt(sum((x - y) ** 2 for x, y in zip(a, b)))

print(round(euclid(u, w), 4), round(euclid(unit(short), unit(other)), 4))
# 0.0 1.4142
print(round(math.sqrt(2 - 2 * cosine(short, other)), 4))   # 1.4142

Two facts fall out of that, and both are worth having ready. On unit vectors the dot product is the cosine, so storing normalised embeddings turns every query into a matrix multiply with no divisions. And straight-line distance between unit vectors is sqrt(2 - 2·cos) — a decreasing function of the cosine, which is why a vector database built on Euclidean distance returns the same ranking as one built on cosine, provided the vectors were normalised on the way in.

The complexity

Time: O(d) per pair, where d is the number of dimensions — one multiply-add per dimension, plus two square roots. For 1,536 dimensions that is a few thousand floating-point operations, and modern hardware does it in microseconds.

Searching a corpus: O(n·d) for a brute-force scan of n documents. That is the number that decides your architecture. At ten thousand documents, scan everything and stop worrying. At ten million it is too slow per query, and the answer is approximate nearest neighbour search, which trades a small amount of recall for a very large amount of speed.

Space: O(n·d). At four bytes per dimension, a million 1,536-dimensional embeddings is about six gigabytes — which is why quantisation exists.

Where it goes wrong

  • Comparing vectors from different models. Two embedding models produce two unrelated coordinate systems. The arithmetic still runs and the number still looks like a similarity; it means nothing.
  • Normalising twice, or forgetting once. If the index stores unit vectors and the query does not, the dot product is scaled by the query's length — harmless for ranking a single query, and quietly wrong the moment you compare scores across queries or apply a fixed threshold.
  • Treating the score as a probability. 0.9 is not "90% relevant". Thresholds must be calibrated per model, on your own data, or not used.
  • Ignoring magnitude when magnitude is the signal. Cosine discards length by design; for count vectors where volume genuinely matters, the dot product is the right tool.
  • Assuming a high score means the answer is in there. Retrieval returns the nearest chunks, always — even when nothing in the corpus answers the question.

How to say it in an interview

For a system design round, the framing matters more than the formula:

"I would embed the documents once and store the vectors normalised to unit length. Then relevance is cosine similarity — the dot product divided by both magnitudes — which measures the angle and so ignores document length; a raw dot product would rank long documents highest regardless of topic. Since the vectors are already normalised, the query is a plain dot product, O(d) per document. A brute-force scan is fine up to roughly the low millions; past that I would put an approximate nearest-neighbour index in front of it and accept a small recall loss for a large latency win."

Length, angle, normalisation, and the point where brute force stops working — four sentences, and the whole answer.