Chunking and Reranking
AI & ML: lesson 16 of 32
Retrieval quality is decided before the model reads a word.
Lesson 16 of 32 · 5 min
Chunking and Reranking
Step 1 of 10
Everything the model can possibly say is decided here, at the cut — long before a prompt exists.
The Idea
Cutting documents into chunks happens once and decides everything afterwards. A cut that separates a definition from its subject leaves a passage that can never answer the question. So the cuts overlap, the first stage fetches generously, and a slower scorer reorders that shortlist.
Real-World Example
A library pulls thirty books off the shelves by catalogue keyword in a minute. A librarian then skims each index and hands you the three that actually address your question.
The Tradeoff
Overlap multiplies storage and lets near-duplicates crowd the shortlist, pushing out the one distinct passage you needed. Reranking adds a second model call to every query and only earns its latency once first-stage recall is already good. Measure recall at twenty before you buy a reranker.
Your turn
Put the steps in the right order.
- Score each candidate against the question with the slower model
- Cut the documents into overlapping chunks and embed each one
- Send the top few reordered passages to the generator
- Fetch a generous shortlist by vector similarity
Mini quiz
1 / 3