RAG Chunking and Reranking: Overlap, Recall and Rerankers
9 min readBytePatterns
RAG chunking and reranking explained: how a bad cut makes a fact unanswerable, what overlap guarantees, and why no reranker can recover a missed passage.
When a retrieval-augmented system gives a bad answer, the generator gets the blame. Often the damage was done earlier: when the documents were cut into chunks, or when the first search stage fetched the wrong passages. As the lesson puts it, retrieval quality is decided before the model reads a word. This article measures the two levers that decide it, how you chunk and whether you rerank. The whole pipeline is in RAG explained.
The problem it solves
- What is a passage? Too long, and one vector blurs several topics; too short, or cut in the wrong place, and it loses the context that made it relevant.
- Which passages reach the model? The first stage must be fast over millions of chunks, so it compares stored vectors, which costs precision. Yet the generator reads only a few passages, so those few must be right.
The intuition
Chunking happens once and decides everything afterwards. If a cut separates a subject from its fact, "the glaze kiln" at the end of one chunk and "reaches cone ten" at the start of the next, no chunk contains the answer, and no later stage can recover it. Overlap makes seams survivable: consecutive windows share a few words, so a short span near a boundary lands whole in at least one of them. Splitting on sentence or section boundaries helps for the same reason. The price is storage, and near-duplicate chunks that crowd the shortlist.
Two stages beat one. The first stage, a vector or keyword index, fetches a generous shortlist, twenty rather than three, because it is cheap and its job is recall. A reranker then scores each candidate together with the question, which a stored vector cannot do: it can notice that the words appear in the right order and relation, not merely that they appear. That pass costs a model call per candidate, so it only runs on the shortlist.
The limit follows directly: a reranker only reorders what it is given. Whatever the first stage missed stays missed, so first-stage recall caps the whole system. Measure recall at twenty before buying a reranker.
Watch it run
The animation starts offline: everything the model can possibly say is decided at the cut, long before a prompt exists. Documents are split into chunks, and each chunk is embedded once. A bad cut is fatal: the subject is in c1 and the fact is in c2, so neither can answer. Overlapping the cuts by a few sentences, two here, makes the seam survivable. The embedded chunks go into the index, and the offline half is done. A question arrives and is embedded by the same model, never a different one. The first stage fetches generously, the top 20, and the leading candidates are c7 at 0.71, c2 at 0.68 and c9 at 0.66, with recall at twenty high. The reranker scores question and passage together, which a stored vector cannot do, and the order changes: c2 jumps to 0.94, c9 drops to 0.55, c7 to 0.31. Only the survivors, two passages, reach the model, so the context stays short and relevant. And the limit: c4 was missed by the first stage, and a passage it never returned cannot be reordered into existence. The ceiling is first-stage recall.
Chunking and Reranking
Step 1 of 10
Everything the model can possibly say is decided here, at the cut — long before a prompt exists.
The same interactive animation as the lesson — step through it with the controls.
The code
A toy model: words stand in for tokens, and the "embedding" is a bag of content words; real systems use learned embeddings and a learned reranker. First the cut, eight-word windows with no, a little, and enough overlap:
import math
import re
from collections import Counter
def chunk(words, size, overlap):
"""Windows of `size` words; each starts `size - overlap` words after the last."""
step = size - overlap
return [words[i:i + size] for i in range(0, max(len(words) - overlap, 1), step)]
def holds(piece, span):
return any(piece[i:i + len(span)] == span for i in range(len(piece) - len(span) + 1))
manual = ("Load shelves on Thursday . The glaze kiln reaches cone ten after nine hours . "
"Let it cool for two days before opening .").split()
fact = "The glaze kiln reaches cone ten".split()
print([" ".join(p) for p in chunk(manual, 8, 0)][:2])
# ['Load shelves on Thursday . The glaze kiln', 'reaches cone ten after nine hours . Let']
for overlap in (0, 2, 5):
print(overlap, len(chunk(manual, 8, overlap)), any(holds(p, fact) for p in chunk(manual, 8, overlap)))
# 0 3 False
# 2 4 False
# 5 7 True
Overlap 2 still split the fact; overlap 5 could not, and the check at the end proves why: with overlap o, every span of at most o + 1 words lies whole inside some window. That guarantee costs 7 chunks instead of 3.
Next, two stages. The first stage ranks every passage by cosine similarity of stored vectors. The reranker reads question and passage together and rewards question word pairs that appear in the same order in the passage, something no precomputed vector of single words can see:
STOP = {"a", "an", "and", "are", "at", "be", "before", "by", "can", "do", "does", "for", "how",
"is", "it", "its", "many", "may", "of", "on", "the", "to", "up", "what", "when", "which",
"with", "you", "after", "each", "every", "until", "than", "take", "should"}
def terms(text):
return [w for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in STOP]
def cosine(a, b):
dot = sum(a[w] * b[w] for w in a)
norm = math.sqrt(sum(v * v for v in a.values()) * sum(v * v for v in b.values()))
return dot / norm if norm else 0.0
passages = [
"Loaded on Thursday, the glaze kiln climbs slowly overnight and reaches cone ten after nine hours, then holds for twenty minutes.",
"The bisque kiln is loaded on Mondays and fired to cone 04.",
"Cone ten glazes need a slow cool in the kiln to grow crystals.",
"Kiln shelves must be coated with kiln wash before each firing.",
"Glaze kiln hours: see the glaze kiln log.",
"Students may not open the kiln until it drops below 150 degrees.",
"The studio closes at nine on weekdays and at six on weekends.",
"Glaze buckets are stirred before every use to mix settled glaze.",
"Reduction firing starves the kiln of oxygen to change glaze colour.",
"The electric one gets to top temperature faster than the gas one.",
]
index = [Counter(terms(p)) for p in passages] # embedded once, offline
def first_stage(question, n):
q = Counter(terms(question))
return sorted(range(len(passages)), key=lambda i: (-cosine(q, index[i]), i))[:n]
def rerank(question, shortlist):
"""Reads question and passage TOGETHER: shared words, plus shared word pairs in order."""
q = terms(question)
pairs = set(zip(q, q[1:]))
def score(i):
p = terms(passages[i])
return len(set(q) & set(p)) + 2 * len(pairs & set(zip(p, p[1:])))
return sorted(shortlist, key=lambda i: (-score(i), shortlist.index(i)))
q = "How many hours until the glaze kiln reaches cone ten?"
print(first_stage(q, 3), rerank(q, first_stage(q, 3))) # [4, 0, 2] [0, 4, 2]
The short passage 4 wins the vector comparison by saying "glaze kiln" twice; the reranker sees that only passage 0 says "glaze kiln reaches cone ten". Now six labelled questions, and top-1 accuracy with and without reranking a shortlist of size n:
questions = {
"How many hours until the glaze kiln reaches cone ten?": 0,
"When can students open the kiln?": 5,
"What goes on kiln shelves before firing?": 3,
"How should cone ten glazes cool?": 2,
"Which glaze firing changes the colour?": 8,
"Which kiln heats up quickest?": 9,
}
for n in (1, 3, 5):
first = sum(want in first_stage(q, n)[:1] for q, want in questions.items())
again = sum(want in rerank(q, first_stage(q, n))[:1] for q, want in questions.items())
ceiling = sum(want in first_stage(q, n) for q, want in questions.items())
print(n, first, again, ceiling)
# 1 4 4 4
# 3 4 5 5
# 5 4 5 5
Reranking a shortlist of three lifts top-1 from 4 of 6 to 5 of 6, exactly the first stage's recall at three. The last question never gets there: passage 9 says "the electric one" and "top temperature", sharing no word with the question, so no shortlist contains it. That miss needs a better first stage, such as learned embeddings or a keyword-plus-vector hybrid. Finally the seeded check: the overlap guarantee on every span of every text up to 24 words, for every window up to 8; and on 2,000 random corpora, the reranker returns a permutation of its shortlist and never beats the first stage's recall:
import random
ok = True
for size in range(1, 9):
for overlap in range(size):
for n in range(1, 25):
words = list(range(n)) # distinct words, so spans are unique
pieces = chunk(words, size, overlap)
ok &= sorted({w for p in pieces for w in p}) == words
for length in range(1, min(overlap + 1, n) + 1):
ok &= all(any(holds(p, words[s:s + length]) for p in pieces)
for s in range(n - length + 1))
rng = random.Random(16)
vocab = ["kiln", "glaze", "cone", "ten", "cool", "shelf", "wash", "fire", "log", "hours"]
saved = passages[:]
for _ in range(2_000):
passages[:] = [" ".join(rng.choices(vocab, k=rng.randint(1, 8))) for _ in range(rng.randint(1, 12))]
index[:] = [Counter(terms(p)) for p in passages]
q, want, n = " ".join(rng.choices(vocab, k=rng.randint(1, 5))), rng.randrange(len(passages)), rng.randint(1, 6)
short = first_stage(q, n)
again = rerank(q, short)
ok &= sorted(again) == sorted(short)
ok &= (want in again[:1]) <= (want in short)
full = sorted(range(len(passages)), key=lambda i: (-cosine(Counter(terms(q)), index[i]), i))
ok &= short == full[:n]
passages[:] = saved
print(ok) # True
The complexity
- Chunking: with window
wand overlapo, aboutN / (w - o)chunks forNwords, so overlap multiplies storage and embedding cost byw / (w - o). - First stage: one query embedding plus a nearest-neighbour search, usually approximate, as in vector databases.
- Reranking: one scoring pass per candidate,
nmodel calls of question plus passage; the reasonnis tens, not thousands.
Where it goes wrong
- Fixed cuts with no overlap. Seams split facts from their subjects.
- Huge chunks. One vector averages several topics and matches none well, and the context window fills with filler.
- Near-duplicates crowding the shortlist. Heavy overlap returns three copies of one paragraph; deduplicate before taking the top few.
- A different embedding model for questions. Distances between two models' vectors mean nothing.
- Reranking to fix poor recall. Measure recall at the shortlist size first; if the passage is not there, the reranker is the wrong fix.
When it shows up in interviews
In AI system design questions, "build a question-answering bot over our documentation", and as follow-ups to RAG: how big are your chunks, why overlap, why two stages, and how would you know whether retrieval or generation is failing. As of October 2026, rerankers are commonly cross-encoder models that read the question and passage in one pass (from memory).
How to say it in an interview
"I'd chunk along sentence or section boundaries with some overlap, so a fact near a seam lands whole in at least one chunk, and accept the extra storage and near-duplicates. Retrieval is two-stage: a fast vector or hybrid search pulls a generous shortlist, maybe twenty, for recall; then a reranker scores each candidate together with the question for precision, and only the top few go to the model. The reranker only reorders, so first-stage recall is the ceiling. I'd measure recall at twenty on labelled questions before adding one."