Skip to content
BytePatterns

Fine-Tuning vs Prompting vs RAG: Which One to Use

8 min readBytePatterns

Fine-tuning vs prompting vs RAG: what each changes, what each costs, a break-even calculation for tuning, and the order to try them in, with runnable code.

A language model gives you answers in the wrong format, misses your company's policies, or sounds nothing like your product. There are three levers: change the prompt, add retrieval so the right documents arrive with the question, or fine-tune so the behaviour moves into the weights. They fix different problems at very different costs, and picking the wrong one is an expensive way to stay stuck.

The problem it solves

Each lever changes a different thing:

  • Prompting changes the instructions. The model is frozen; you supply rules, context and a few worked examples ("few-shot") at request time. A change ships in seconds and reverts just as fast.
  • Retrieval-augmented generation changes the context. A search step finds relevant documents and puts them in the prompt, so answers can cite facts that were never in training. The details are in RAG explained.
  • Fine-tuning changes the weights. You continue training on curated input-output pairs until the model produces that behaviour without being told. Parameter-efficient methods such as LoRA train a small add-on instead of every weight, which makes this cheaper but does not change what it is good for.

The rule of thumb that falls out: facts go in retrieval, behaviour goes in weights, and everything starts as a prompt.

The intuition

Ask what kind of gap you are closing.

  1. Missing or changing knowledge. This week's prices, a customer's order history, a policy updated on Monday. Weights are a snapshot; retraining every time a document changes is slow and costly, and a tuned model can still blend old and new versions of a fact. Retrieval reads the current document on every request.
  2. Behaviour and shape. A strict JSON layout, a house style, a narrow classification task. These are hard to hold with instructions alone across thousands of calls and easy to teach with a few hundred good examples. Fine-tuning shines here, and shortens every prompt as a side effect.
  3. Neither, yet. Very often the honest answer is that the prompt was vague. A clear instruction plus three examples fixes a surprising share of problems for almost no cost.

The cost sides are asymmetric. Prompting costs tokens on every call, forever. Fine-tuning costs a dataset, a training run, an evaluation, and a model version you now own and must redo when the base model is upgraded. It pays back only when the per-call saving, multiplied by real traffic, beats that one-off bill.

Watch it run

The animation runs the two lanes. Prompting steers a frozen model with instructions supplied at request time, an 880-token prompt, and every single call carries those rules and worked examples with it. Change the wording and it takes effect on the very next request: instant and reversible, measured in seconds. The weights never move; nothing about the base model has been altered. Fine-tuning takes the other road: curate examples of the behaviour you want, 1,200 rows. A training run continues training the base model on them, four hours, one-off. The result is a model of your own with the behaviour baked into its weights, and now the prompt shrinks to just the input, 60 tokens where it was 880: the instructions and examples moved into the weights. Fewer tokens and a consistent output shape add up: prompt tokens per call fall 93%. Then the bill: dataset curation, a training run and a model version to maintain, and the whole run has to be redone on every base-model upgrade. Weights are a poor place to park facts; those belong in retrieval, which updates when the document does. The last frame is the order to try things in: measure the gap first, and treat prompting and retrieval as the cheaper first attempt.

Fine-Tuning vs Prompting

Step 1 of 13

Prompting steers a frozen model with instructions supplied at request time.

The same interactive animation as the lesson — step through it with the controls.

The code

A toy model of the economics, in made-up price units (not any provider's real prices). A call pays for prompt tokens and output tokens; tuning removes 820 prompt tokens per call and costs a one-off 5,000 units:

from fractions import Fraction as F
from math import ceil

def per_call(prompt_tok, output_tok, price_in, price_out):
    """Cost of one call in price units; output tokens are paid either way."""
    return prompt_tok * price_in + output_tok * price_out

def breakeven(base_call, tuned_call, one_off):
    """First call count at which tuning has paid for itself, or None if never."""
    saving = base_call - tuned_call
    return None if saving <= 0 else ceil(one_off / saving)

IN, OUT = F(1, 1000), F(4, 1000)           # illustrative: units per token
base = per_call(880, 200, IN, OUT)         # rules + 3 examples on every call
tuned = per_call(60, 200, IN, OUT)         # just the input
print(float(base), float(tuned), round(1 - 60 / 880, 3))           # 1.68 0.86 0.932
print(breakeven(base, tuned, one_off=5000))                         # 6098

pricier = per_call(60, 200, 2 * IN, 2 * OUT)   # same tokens, tuned model at 2x the rate
print(float(pricier), breakeven(base, pricier, one_off=5000))      # 1.72 None

Prompt tokens fall 93.2%, but the cost per call only falls from 1.68 to 0.86, because the 200 output tokens are paid either way. If the tuned model is billed at twice the base rate, it never pays back at all. Check how your provider prices tuned models before promising savings; as of October 2026 this varies between providers and between hosted and self-run models (from memory, not a quoted price list).

Facts are the other trap. A tuned model keeps whatever was true when it was trained:

docs = {"refund window": "30 days"}
weights = dict(docs)                       # facts baked in at training time
docs["refund window"] = "14 days"          # the policy changes on Monday
print(weights["refund window"], "|", docs["refund window"])      # 30 days | 14 days

The break-even formula is then checked against brute force on 500 seeded scenarios, with tuned prices from half to three times the base rate. The simulation pays for calls one at a time until the tuned side is no longer behind:

import random

rng = random.Random(35)
ok = True
for _ in range(500):
    p_in, p_out = F(rng.randint(1, 5), 1000), F(rng.randint(1, 20), 1000)
    mult = F(rng.randint(5, 30), 10)                      # tuned price 0.5x to 3x
    b = per_call(rng.randint(100, 3000), 200, p_in, p_out)
    t = per_call(rng.randint(10, 300), 200, p_in * mult, p_out * mult)
    one_off = rng.randint(1, 200)
    n, spent_base, spent_tuned = 0, F(0), F(one_off)
    while spent_tuned > spent_base and n <= 10_000:       # brute force: call by call
        n += 1
        spent_base += b
        spent_tuned += t
    got = breakeven(b, t, one_off)
    if spent_tuned <= spent_base:
        ok &= got == n
    else:                                                 # not paid back in 10,000 calls
        ok &= got is None or got > 10_000
print(ok)                                                  # True

The complexity

  • Prompting: zero setup; cost and latency grow with prompt length on every call, bounded by the context window.
  • RAG: an index to build and keep fresh, plus a search per request; the retrieved text also costs prompt tokens.
  • Fine-tuning: one-off data and training cost, then shorter prompts; repeated per base-model upgrade.

Where it goes wrong

  • Tuning facts into weights. They go stale, and the model can still answer confidently from the old version.
  • No evaluation set. Without a fixed set of test inputs and expected outputs, nobody can say whether the prompt, the retrieval or the tuned model actually helped.
  • Bad training data. A fine-tuned model reproduces its examples faithfully, inconsistencies included.
  • Tuning to fix retrieval. If the right document never reaches the prompt, no amount of tuning will make answers current.

When it shows up in interviews

In AI engineering and ML system design rounds: "our assistant gives outdated answers, what do you do?", "when would you fine-tune?", or "how would you cut inference cost?". Interviewers want the diagnosis before the tool: measure where outputs fall short, then pick the lever that matches the gap. Combining levers is normal; a tuned model behind retrieval is a common production shape.

How to say it in an interview

"I start with prompting, because it is instant and reversible, and I build an evaluation set first so I can measure the gap. If failures are missing or stale facts, I add retrieval, since it updates the moment the document does. If the gap is behaviour, a strict format, style or a narrow task, and it survives good prompting, I fine-tune on curated examples. That buys shorter prompts and consistency, but costs a dataset, a training run and a model version I have to redo on every base-model upgrade, so I check the break-even against real traffic first."