Skip to content
BytePatterns

RAG Explained: Retrieval-Augmented Generation Step by Step

9 min readBytePatterns

How retrieval-augmented generation works: chunk, embed, retrieve, then prompt. A runnable toy pipeline, where it fails, and how to explain RAG in an interview.

A language model answers from what it absorbed in training. That knowledge stops at a cutoff, knows nothing about your company's documents, and cannot say where a claim came from. Retrieval-augmented generation, RAG, fixes this by putting a search step in front of the model: find the relevant passages first, then ask the model to answer from them. The term comes from a 2020 paper by Lewis and colleagues, listed at the end; the pipeline below is the general pattern, described as of September 2026, and details vary by system.

The problem it solves

You want answers grounded in a specific, changing body of text: a product manual, support articles, internal policies. Two alternatives fall short:

  • Ask the model directly. It may not know the answer, and when it does not, it can still produce a fluent, confident one.
  • Retrain or fine-tune on the documents. Slow to repeat every time a document changes, and the answer still comes without a citation.

The Lewis et al. paper frames the difference as two kinds of memory: parametric memory, the knowledge stored in a model's weights, and non-parametric memory, an external index that can be searched. In their system the index was a dense vector index of Wikipedia, accessed by a neural retriever. RAG combines the two, so the knowledge can be updated by updating the index, without retraining anything.

The intuition

A RAG system has two phases that meet at a vector store.

Indexing, done ahead of time.

  • Split every source document into chunks, small enough that one chunk is about one thing.
  • Turn each chunk into a vector with an embedding model, and store the vectors next to the text.

Answering, done per question.

  • Embed the question with the same embedding model. Distances between vectors from two different models mean nothing.
  • Find the chunks whose vectors are closest to the question's, typically by cosine similarity.
  • Put those passages into the prompt with an instruction: answer only from these, and cite them.

The retriever decides what the model gets to read, so retrieval quality caps the whole system. A perfect model given the wrong passages gives a well-written wrong answer.

Watch it run

The animation shows both phases on one stage. Along the top, documents are chunked, embedded and stored in the vector database, once and offline. Along the bottom, a question arrives, is embedded with the same model, and the store returns the three closest chunks. Those passages go into the prompt with the question and a single instruction, and only then does the model write. The last step shows the failure the lesson warns about: when retrieval returns nothing relevant, the model can answer anyway, just as confidently.

Retrieval-Augmented Generation

Step 1 of 12

RAG puts a search step in front of the model. First, index your own corpus.

The same interactive animation as the lesson — step through it with the controls.

The code

A toy pipeline, not a production one. Its "embedding" counts content words, so it can only match exact words; real systems use a learned embedding model that places paraphrases near each other. Everything else, chunking with overlap, cosine similarity, top-k and the prompt, has the real shape:

import math
import re
from collections import Counter

STOP = {"a", "an", "and", "are", "at", "be", "can", "do", "for", "from", "how",
        "i", "if", "in", "is", "it", "my", "of", "on", "the", "to", "what", "when"}

def chunk(text, size=12, overlap=4):
    words = text.split()
    step = size - overlap
    return [" ".join(words[i:i + size])
            for i in range(0, max(len(words) - overlap, 1), step)]

def embed(text):
    """A toy embedding: counts of content words. Real systems use a learned model."""
    words = re.findall(r"[a-z0-9]+", text.lower())
    return Counter(w.rstrip("s") for w in words if w not in STOP)

def cosine(a, b):
    dot = sum(a[w] * b[w] for w in a)
    norm = math.sqrt(sum(v * v for v in a.values())) * math.sqrt(sum(v * v for v in b.values()))
    return dot / norm if norm else 0.0

docs = {
    "refunds.md": "Refunds are issued to the original payment method within 5 business days "
                  "of the return being received at our warehouse.",
    "shipping.md": "Standard shipping takes 3 to 7 days. Express shipping arrives in 1 to 2 days "
                   "and costs extra at checkout.",
    "returns.md": "Items can be returned within 30 days if unused. Start a return from the "
                  "orders page and print the prepaid label.",
}

index = []                                        # built once, offline
for name, text in docs.items():
    for i, piece in enumerate(chunk(text)):
        index.append((f"{name}#{i}", piece, embed(piece)))
print(len(index), index[0][1])
# 7 Refunds are issued to the original payment method within 5 business days

Retrieval with a score threshold. When nothing clears it, the pipeline returns no passages instead of weak ones:

import heapq

def retrieve(question, k=2, min_score=0.25):
    q = embed(question)                           # same embedding as the index
    scored = [(cosine(q, vec), cid, text) for cid, text, vec in index]
    top = heapq.nlargest(k, scored)
    return [(round(s, 2), cid) for s, cid, text in top if s >= min_score]

print(retrieve("How many business days until my refund is issued?"))
print(retrieve("How do I start a return?"))
print(retrieve("What is the capital of France?"))
print(retrieve("When do I get my money back?"))
# [(0.54, 'refunds.md#0'), (0.27, 'refunds.md#1')]
# [(0.53, 'returns.md#1'), (0.5, 'returns.md#0')]
# []
# []

The last two empty results are different stories. The France question is correctly out of scope. The money-back question is a retrieval miss: the answer is in refunds.md, but the toy embedding cannot see that "money back" means "refund". Closing that gap is exactly what learned embeddings are for, and why retrieval is evaluated on its own. Karpukhin et al. showed the effect directly: a dense retriever, with separate learned encoders for questions and passages, beat a strong keyword-based BM25 baseline on top-20 passage retrieval accuracy across the open-domain question answering datasets they tested.

Assembling the prompt. With no passages, the right move is to not call the model at all, or to answer "I don't know":

def build_prompt(question, hits):
    if not hits:
        return None                               # nothing relevant: do not guess
    text = dict((cid, t) for cid, t, _ in index)
    sources = "\n".join(f"[{cid}] {text[cid]}" for _, cid in hits)
    return ("Answer only from the sources below and cite them by id. "
            "If they do not contain the answer, say you do not know.\n\n"
            f"{sources}\n\nQuestion: {question}")

q = "How do I start a return?"
print(build_prompt(q, retrieve(q, k=1)))
# Answer only from the sources below and cite them by id. If they do not contain the answer, say you do not know.
#
# [returns.md#1] unused. Start a return from the orders page and print the prepaid
#
# Question: How do I start a return?

Notice the passage stops at "prepaid": a chunk boundary cut the sentence, which is why chunks overlap. The mechanics checked on 2,000 random inputs: chunks rebuild the original text exactly once the overlaps are removed, the heap's top-k equals a full sort's, and cosine stays between 0 and 1 for count vectors:

import random

random.seed(3)
vocab = ["refund", "return", "ship", "day", "label", "order", "cost", "item"]
ok = True
for _ in range(2000):
    words = [random.choice(vocab) for _ in range(random.randint(1, 40))]
    size = random.randint(2, 10)
    overlap = random.randint(0, size - 1)
    pieces = [p.split() for p in chunk(" ".join(words), size, overlap)]
    rebuilt = pieces[0] + [w for p in pieces[1:] for w in p[overlap:]]
    ok &= rebuilt == words                        # nothing lost, nothing added
    ok &= all(len(p) <= size for p in pieces)

    vecs = [Counter(random.choices(vocab, k=random.randint(1, 6))) for _ in range(12)]
    q = Counter(random.choices(vocab, k=3))
    scored = [(cosine(q, v), i) for i, v in enumerate(vecs)]
    k = random.randint(1, 5)
    ok &= heapq.nlargest(k, scored) == sorted(scored, reverse=True)[:k]
    ok &= all(-1e-9 <= s <= 1 + 1e-9 for s, _ in scored)
    ok &= abs(cosine(q, q) - 1) < 1e-9
print(ok)                                         # True

The complexity

  • Indexing embeds every chunk once. It is paid again only for documents that change.
  • Exact search compares the question with all N chunk vectors of dimension d: O(N · d) per question, plus O(N log k) to keep the top k with a heap. Beyond small corpora, vector databases use approximate nearest-neighbour indexes that trade a little recall for much less work.
  • The prompt grows with k times the chunk size, and it has to fit in the model's context window alongside the question and the answer.

Where it goes wrong

The lesson's list of failure points, one per stage:

  • Chunking. Chunks that are too large blur several topics into one vector; chunks that are too small, or cut mid-sentence, lose the context that made them relevant.
  • Retrieval. The one passage that mattered is not in the top k, as the money-back question shows. Measure retrieval separately from answers.
  • Mismatched embeddings. Indexing with one model and querying with another makes every distance meaningless.
  • Answering anyway. Nothing in generation forces the model to notice the evidence is missing. Use a score threshold, tell it to say when the sources do not contain the answer, and test that it does.

How to say it in an interview

"RAG adds a retrieval step before generation. Offline, I chunk the documents, embed each chunk and store the vectors. Per question, I embed it with the same model, retrieve the nearest chunks by cosine similarity, and put them in the prompt with an instruction to answer only from them and cite them. It gives fresh, citable answers without retraining, and retrieval can respect who may see which document. The weak points are chunking, retrieval misses, and a model that answers even when nothing relevant came back, so I'd evaluate retrieval on its own and add a threshold with an 'I don't know' path."

The pieces each have their own lesson: embeddings, cosine similarity, vector databases and chunking and reranking.

Sources