Skip to content
BytePatterns

Chunking and Reranking

AI & ML: lesson 16 of 32

Retrieval quality is decided before the model reads a word.

Lesson 16 of 32 · 5 min

Chunking and Reranking

Step 1 of 10

Everything the model can possibly say is decided here, at the cut — long before a prompt exists.

The Idea

Cutting documents into chunks happens once and decides everything afterwards. A cut that separates a definition from its subject leaves a passage that can never answer the question. So the cuts overlap, the first stage fetches generously, and a slower scorer reorders that shortlist.

Real-World Example

A library pulls thirty books off the shelves by catalogue keyword in a minute. A librarian then skims each index and hands you the three that actually address your question.

The Tradeoff

Overlap multiplies storage and lets near-duplicates crowd the shortlist, pushing out the one distinct passage you needed. Reranking adds a second model call to every query and only earns its latency once first-stage recall is already good. Measure recall at twenty before you buy a reranker.

Your turn

Put the steps in the right order.

  1. Score each candidate against the question with the slower model
  2. Cut the documents into overlapping chunks and embed each one
  3. Send the top few reordered passages to the generator
  4. Fetch a generous shortlist by vector similarity

Mini quiz

1 / 3

Overlapping chunks exist mainly to:

New lessons land every few weeks

Leave an address and we will tell you when the next one is up. That is the only reason we will use it.

One address, stored so we can email you. Nothing else, ever.