LoRA Explained: Low-Rank Adapters for Cheap Fine-Tuning
8 min readBytePatterns
LoRA explained: freeze the weights, train two thin matrices B and A, why their product has the layer's shape, what rank buys, and what merging does to serving.
Fine-tuning a large language model the classic way updates every weight, so each task needs its own full copy of the model and enough memory to hold gradients and optimizer state for all of it. LoRA, low-rank adaptation, trains a tiny correction beside each chosen weight matrix instead. The base stays frozen, and one base can serve many tasks.
The problem it solves
Take one square weight matrix W of width d = 4096. A full fine-tune trains all d² = 16,777,216 of its numbers, in every layer where it is applied, and optimizers such as Adam keep extra state per trained number, so memory for training is several times the model's own size. Storing a separate fine-tuned copy per customer or per task multiplies that again.
Two observations make a cheaper route possible. First, the change a task needs, ΔW, does not have to be an arbitrary matrix; the method's premise, supported by its experiments, is that task-specific updates have low intrinsic rank. Second, a low-rank matrix can be stored as a product of two thin ones.
The intuition
Freeze W. Learn two matrices: B, which is d × r, and A, which is r × d, with the rank r small, often single or low double digits. Their product BA is d × d, the same shape as W, so the layer computes:
h = W x + (α / r) · B A x
A few details matter in practice:
- Parameter count:
BandAhold2 · d · rnumbers. Ford = 4096andr = 8that is 65,536, about 0.39% of the full matrix. Halvingrhalves it. - Start from zero:
Ais initialised randomly andBto zeros, soBA = 0and training begins exactly at the base model. - Scale: the update is multiplied by
α / r, which keeps the effective step size stable when you changer. - Rank ceiling:
BAcan never have rank abover. That is the regulariser and the limit in one: too low and the adapter cannot express the task, too high and the savings shrink toward a full fine-tune. - Where it goes: usually on the attention projection matrices, sometimes on the feed-forward layers too.
At serving time there are two options. Keep the adapter separate and compute W x + B(A x): one base in memory, adapters swapped per request, a small extra cost per layer. Or merge: replace W with W + (α / r) BA. The merged model is exactly as fast as the original, but that copy now serves one task; switching means subtracting BA back out and adding another adapter.
As of September 2026, LoRA and its variants, including training adapters on top of a base quantized to 4 bits, are standard options in open fine-tuning libraries, and several serving systems batch requests for different adapters over one shared base.
Watch it run
The animation begins with the expensive baseline: a full fine-tune rewrites every weight, which means a whole copy of the model for every task, 16.8 million numbers for this one matrix. Instead, the base is frozen entirely; nothing in it is going to move. Two thin matrices are trained instead, d by r, then r by d, with r = 8: about 65 thousand numbers, four tenths of one percent of the weights the full tune would touch. Their product is d by d, the layer's own shape, so it can simply be added. Serving runs the sum, so the task-tuned model is one ordinary matrix. A second task swaps the adapter and reuses the same frozen base. Merged into W, inference costs nothing extra, but that copy serves one task until BA is subtracted again. And rank is the real dial: too low and it cannot learn the task, too high and you are back to fine-tuning.
Adapters and LoRA
Step 1 of 9
A full fine-tune rewrites every weight — which means a whole copy of the model for every task.
The same interactive animation as the lesson — step through it with the controls.
The code
The parameter arithmetic from the lesson, for three ranks:
def trained_params(d, r):
return d * d, 2 * d * r # full fine-tune vs B (d×r) plus A (r×d)
for r in (1, 8, 64):
full, lora = trained_params(4096, r)
print(r, lora, round(100 * lora / full, 2))
# 1 8192 0.05
# 8 65536 0.39
# 64 524288 3.12
A toy model, not a training loop: exact fractions on a 6 × 6 layer, so every equality is exact. With B at zero the adapter changes nothing; after "training", BA has the layer's shape but rank 2, the merged matrix gives the same output as the adapter path, and subtracting BA recovers the base:
from fractions import Fraction as F
import random
def matmul(X, Y):
return [[sum(X[i][k] * Y[k][j] for k in range(len(Y))) for j in range(len(Y[0]))]
for i in range(len(X))]
def add(X, Y, sign=1):
return [[a + sign * b for a, b in zip(rx, ry)] for rx, ry in zip(X, Y)]
def rank(M):
M, r = [row[:] for row in M], 0
for c in range(len(M[0])):
pivot = next((i for i in range(r, len(M)) if M[i][c] != 0), None)
if pivot is None:
continue
M[r], M[pivot] = M[pivot], M[r]
for i in range(len(M)):
if i != r and M[i][c] != 0:
factor = M[i][c] / M[r][c]
M[i] = [a - factor * b for a, b in zip(M[i], M[r])]
r += 1
return r
def lora_forward(W, B, A, x, scale=1):
base = matmul(W, x)
delta = matmul(B, matmul(A, x)) # x -> r numbers -> d numbers, never d×d
return add(base, [[scale * v for v in row] for row in delta])
random.seed(27)
d, r = 6, 2
W = [[F(random.randint(-9, 9)) for _ in range(d)] for _ in range(d)]
A = [[F(random.randint(-3, 3)) for _ in range(d)] for _ in range(r)] # r×d, random
B = [[F(0)] * r for _ in range(d)] # d×r, zeros
x = [[F(random.randint(-5, 5))] for _ in range(d)]
print(lora_forward(W, B, A, x) == matmul(W, x)) # True: B = 0, so training starts at the base
B = [[F(random.randint(-3, 3)) for _ in range(r)] for _ in range(d)] # after "training"
BA = matmul(B, A)
merged = add(W, BA)
print(len(BA), len(BA[0]), rank(BA)) # 6 6 2
print(lora_forward(W, B, A, x) == matmul(merged, x)) # True: merge changes nothing
print(add(merged, BA, sign=-1) == W) # True: subtracting BA recovers the base
Checked on 300 seeded random layers of width 2 to 8 with ranks 1 to 4 and random α / r scales: the adapter path must equal a brute-force dense update computed entry by entry, the update's rank must never exceed r, and unmerging must restore W exactly:
ok = True
for _ in range(300):
d, r = random.randint(2, 8), random.randint(1, 4)
W = [[F(random.randint(-9, 9)) for _ in range(d)] for _ in range(d)]
A = [[F(random.randint(-3, 3)) for _ in range(d)] for _ in range(r)]
B = [[F(random.randint(-3, 3)) for _ in range(r)] for _ in range(d)]
x = [[F(random.randint(-5, 5))] for _ in range(d)]
scale = F(random.choice([1, 2, 16]), r) # the alpha / r scaling
dense = [[sum(B[i][k] * A[k][j] for k in range(r)) for j in range(d)] for i in range(d)]
merged = [[W[i][j] + scale * dense[i][j] for j in range(d)] for i in range(d)]
ok &= lora_forward(W, B, A, x, scale) == matmul(merged, x)
ok &= rank(dense) <= min(r, d)
ok &= [[merged[i][j] - scale * dense[i][j] for j in range(d)] for i in range(d)] == W
print(ok) # True
The complexity
- Trained parameters per matrix:
2drinstead ofd², linear inr. - Extra forward cost, unmerged:
A xthenB(·), about2drmultiply-adds againstd²forW x. - Extra forward cost, merged: zero; the merge itself is one
d × daddition, done once. - Storage per task: the adapters only, megabytes where a full copy would be gigabytes.
Where it goes wrong
- Getting the shapes backwards.
Bisd × randAisr × d;A Bwould ber × rand could not be added toW. - Initialising both matrices randomly. Training then starts from a perturbed model instead of the base.
- Treating rank as free. Too small underfits; raising it trades back the savings.
- Assuming it matches a full fine-tune everywhere. It often comes close on narrow tasks; large shifts in knowledge or behaviour can need more capacity.
When it shows up in interviews
ML engineering and LLM infrastructure interviews ask how to customise a model per customer without a copy per customer, what the memory saving comes from, and what merging costs at serving time. It pairs with the KV cache in serving questions, and with attention, since the matrices LoRA usually targets are the query, key, value and output projections.
How to say it in an interview
"LoRA freezes the base weights and learns a low-rank update for selected matrices: B is d by r and A is r by d, so BA has the layer's shape and is added, scaled by alpha over r. That is 2dr trained numbers instead of d squared, well under one percent at rank 8 for a 4096-wide layer. B starts at zero, so training starts from the base model. At serving time I can keep adapters separate and share one base across tasks, or merge BA into W for zero extra latency, at the cost of that copy serving one task."