Skip to content
BytePatterns

Adapters and LoRA

AI & ML: lesson 17 of 32

Train a thin correction instead of the whole weight matrix.

Lesson 17 of 32 · 5 min

Adapters and LoRA

Step 1 of 9

A full fine-tune rewrites every weight — which means a whole copy of the model for every task.

The Idea

Full fine-tuning rewrites every weight, which means a full copy of the model per task. A low-rank adapter instead learns two thin matrices whose product has the same shape as the layer, and adds it. The base weights never move, so one base can serve many tasks.

Real-World Example

A tailor does not reweave a suit. They take in the waist and shorten the sleeves — a handful of stitched corrections, reversible, and a different set for each wearer.

The Code

d, r = 4096, 8                  # layer width, adapter rank
full = d * d                    # a full fine-tune rewrites all of these
adapter = 2 * d * r             # B is d×r, A is r×d
print(full, adapter)            # 16777216 65536
print(round(100 * adapter / full, 2), "%")   # 0.39 %
# merged weight = W + B @ A -> same shape, so serving cost is unchanged

Python

Your turn

Fill in the blank.

d, r = 1024, 4
adapter = ___ * d * r
print(adapter)   # 8192

Mini quiz

1 / 3

A low-rank adapter trains:

New lessons land every few weeks

Leave an address and we will tell you when the next one is up. That is the only reason we will use it.

One address, stored so we can email you. Nothing else, ever.