Skip to content
BytePatterns

What Are Embeddings? Meaning as a List of Numbers

8 min readBytePatterns

Embeddings explained from scratch: how text becomes a vector, why similar meanings land close, mean pooling and its trap, and why two models never mix.

Search engines, recommendation feeds, retrieval for chatbots, duplicate detection: a large share of modern machine learning runs on one idea. Turn each item into a fixed-length list of numbers, arranged so that items with similar meaning get similar numbers. That list is an embedding. Once meaning is a position, "find related things" stops being a language problem and becomes geometry, which computers are very good at.

The problem it solves

Computers compare strings by characters. "Kitten" and "cat" share no letters in the right places, while "car" and "cat" differ by one. Keyword search inherits that blindness: a query for "cheap flights" misses a page about "low-cost airfare". What we want is a representation where:

  • Related items are close and unrelated items are far, whatever their spelling.
  • Every item has the same shape, so a sentence, a paragraph or a product photo can be compared with a single operation.
  • Comparison is cheap, so millions of items can be scanned or indexed.

A fixed-length vector gives all three. Closeness is usually measured with cosine similarity, the angle between two vectors.

The intuition

Nobody hand-picks the coordinates. A model is trained on a task where similar things must produce similar outputs, and the vectors it learns along the way are the embeddings. The oldest version of that signal is still the easiest to see: a word is characterised by the words around it. "Cat" and "kitten" both appear near "chased" and "milk"; "car" appears near "fuel" and "road". Count those neighbours and the counts already form vectors in which cat and kitten point the same way.

Modern embedding models are neural networks trained on far richer signals, such as pairs of texts that answer each other, but the contract is the same: a fixed-length vector, distances that mean something, and axes that mean nothing on their own. As of September 2026, widely used text embedding models output a few hundred to a few thousand dimensions, and some are trained so a shorter prefix of the vector still works.

Two consequences matter in practice. A sentence embedding is often built by pooling token vectors, commonly by averaging, which works when the words agree and blurs when they do not. And every model invents its own axes, so vectors from two different models cannot be compared, even if both have the same length.

Watch it run

The animation uses the lesson's three words in two dimensions; real embeddings have hundreds or thousands. An embedding turns text into a fixed-length list of numbers, a position in space. "cat" becomes [0.9, 0.1], so it lands there. "kitten" is [0.8, 0.2], almost the same place, because training put it there. "car" is [0.1, 0.9] and lands nowhere near either. Nobody chose those coordinates: training pushed related items into neighbouring regions, and once meaning is geometry, searching and grouping become arithmetic; cat and kitten are 0.14 apart. Average cat and kitten and you get [0.85, 0.15], still squarely in animal territory. Average cat and car and you get [0.5, 0.5], a meaningless midpoint that means neither. Nothing in the picture stores the text; only the position survives the trip. Each model invents its own axes, so a distance only means something within one space. The closing frame is the lesson's paint matcher: close in numbers means indistinguishable to your eye.

Embeddings

Step 1 of 11

An embedding turns text into a fixed-length list of numbers — a position in space.

The same interactive animation as the lesson — step through it with the controls.

The code

A toy model trained on ten sentences: each word's vector counts the other words it shares a sentence with. Very common words are skipped because they sit next to everything:

import heapq
import math
from collections import Counter

sentences = [s.split() for s in """the cat chased the mouse | the kitten chased the mouse
the cat drank the milk | the kitten drank the milk | the cat slept on the sofa
the car needs fuel | the truck needs fuel | the truck carried the load
the car drove down the road | the truck drove down the road""".replace("\n", "|").split("|")]
SKIP = {"the", "on", "down"}                        # too common to say anything

def train(sentences):
    """Toy model: a word's vector counts the other words it shares a sentence with."""
    vocab = sorted({w for s in sentences for w in s} - SKIP)
    near = {w: Counter() for w in vocab}
    for s in sentences:
        words = [w for w in s if w not in SKIP]
        for i, w in enumerate(words):
            near[w].update(words[:i] + words[i + 1:])
    return {w: [near[w][c] for c in vocab] for w in vocab}

def cosine(a, b):
    dot = sum(x * y for x, y in zip(a, b))
    return dot / (math.sqrt(sum(x * x for x in a)) * math.sqrt(sum(y * y for y in b)))

vecs = train(sentences)
print(len(vecs["cat"]), "dimensions")                           # 16 dimensions
print(round(cosine(vecs["cat"], vecs["kitten"]), 3),
      round(cosine(vecs["cat"], vecs["car"]), 3))               # 0.816 0.0

def nearest(query, space, k=2):
    ranked = sorted(((cosine(query, v), w) for w, v in space.items()), reverse=True)
    return [w for _, w in ranked[:k]]

print(nearest(vecs["kitten"], vecs, 2), nearest(vecs["truck"], vecs, 2))
# ['kitten', 'cat'] ['truck', 'car']

"Kitten" never appears in a sentence with "cat", yet it is cat's nearest neighbour, because they keep the same company. Sentence vectors by averaging, and the midpoint trap:

def embed(text):
    """Mean pooling: a sentence is the average of its word vectors."""
    words = [vecs[w] for w in text.split() if w in vecs]
    return [sum(col) / len(words) for col in zip(*words)]

docs = {d: embed(d) for d in ["kitten drank milk", "truck needs fuel", "cat chased mouse"]}
print(nearest(embed("cat drank milk"), docs, 1), nearest(embed("car needs fuel"), docs, 1))
# ['kitten drank milk'] ['truck needs fuel']
mix = embed("cat car")
print(round(cosine(mix, vecs["cat"]), 2), round(cosine(mix, vecs["car"]), 2))   # 0.77 0.63

A second "model" with the same geometry on different axes, made by rotating every vector. Inside each space the answers are identical; across spaces, the same word is unrecognisable. Then 300 seeded random collections, checking a normalise-then-dot-product search (what vector indexes do) and the rotated space against a brute force that computes the cosine with every item:

import random

def other_model(space, seed):
    """A second toy model: the same geometry on different axes (a random rotation)."""
    rng = random.Random(seed)
    dim = len(next(iter(space.values())))
    out = {w: list(v) for w, v in space.items()}
    for _ in range(4 * dim):                         # a chain of random plane rotations
        i, j = rng.sample(range(dim), 2)
        t = rng.uniform(0, 2 * math.pi)
        c, s = math.cos(t), math.sin(t)
        for v in out.values():
            v[i], v[j] = c * v[i] - s * v[j], s * v[i] + c * v[j]
    return out

vecs_b = other_model(vecs, seed=30)
print(round(cosine(vecs_b["cat"], vecs_b["kitten"]), 3),
      nearest(vecs_b["kitten"], vecs_b, 2))          # 0.816 ['kitten', 'cat']
print(round(cosine(vecs["cat"], vecs_b["cat"]), 3))   # -0.402  the same word, across models

def top_k(query, space, k):
    """What a vector search does: normalise once, then rank by dot product."""
    norm = lambda v: [x / math.sqrt(sum(y * y for y in v)) for x in v]
    q = norm(query)
    scored = [(sum(a * b for a, b in zip(q, norm(v))), w) for w, v in space.items()]
    return [w for _, w in heapq.nlargest(k, scored)]

rng = random.Random(30)
ok = True
for _ in range(300):
    dim, n = rng.randint(2, 12), rng.randint(1, 40)
    space = {f"d{i}": [rng.gauss(0, 1) for _ in range(dim)] for i in range(n)}
    query, k = [rng.gauss(0, 1) for _ in range(dim)], rng.randint(1, n)
    brute = nearest(query, space, k)                  # brute force: cosine against every item
    ok &= top_k(query, space, k) == brute
    rotated = other_model({**space, "q": query}, seed=rng.random())
    q2 = rotated.pop("q")
    ok &= nearest(q2, rotated, k) == brute            # same model, new axes: same answer
print(ok)                                            # True

The complexity

  • Embedding a text costs one model forward pass, roughly linear in its token count for the encoder models typically used; do it once at write time and store the vector.
  • Comparing two vectors is O(d) for d dimensions. Normalise once and cosine becomes a plain dot product.
  • Exact search over n items is O(n·d) per query. At millions of items, vector databases trade a little recall for approximate indexes that look at a small fraction.
  • Memory is n·d numbers: a million 768-dimensional float32 vectors is about 3 GB before any index.

Where it goes wrong

  • Mixing models. Vectors from different models, or different versions of one model, live in unrelated spaces. Changing the model means re-embedding everything.
  • Averaging too much. One vector for a long, mixed document is a midpoint of its topics and close to none of them. Split long text into chunks, as retrieval-augmented generation pipelines do.
  • Treating closeness as truth. "Similar" means "used in similar contexts". Antonyms like "hot" and "cold" often land close.
  • Assuming the text is gone. An embedding is not a copy of the text, but research has shown that short texts can be partly reconstructed from some embeddings, so vectors of private text deserve the same care as the text.

When it shows up in interviews

In machine learning engineering interviews as "what is an embedding?", "how would you build semantic search?" or "design a RAG system", and in system design as the storage and indexing side of those. Expect follow-ups on the similarity metric, the dimension trade-off, chunking and re-embedding cost when the model changes. The attention layer that produces contextual token vectors is covered in the attention mechanism.

How to say it in an interview

"An embedding is a fixed-length vector a trained model assigns to an input, arranged so semantically similar inputs land close together. The individual axes mean nothing; only distances within one model's space do. I'd embed documents once at write time, normalise them, and search by cosine similarity, exact for small collections and an approximate index for large ones. For long documents I'd chunk before embedding, because an average of many topics is close to none of them, and I'd never compare vectors from two different models: switching models means re-embedding the corpus."