🔌

LoRA — Don't Change All 66M, Just Plug In a 1M Adapter

Efficient alternative to fine-tuning — freeze the base model, train just two small matrices

The Problem with Fine-tuning

Fine-tuning adjusts all 66M numbers. Problems: 7B models need dozens of GB GPU memory, 5 models = 5× storage, long training.

LoRA's Idea

"Don't change all 66M. Just plug in a small adapter."

Fine-tuning: adjust all 66M → save 66M (260MB)
LoRA: freeze 66M + train adapter 1M → save 1M only (4MB)

Analogy: fine-tuning = modifying the phone itself. LoRA = putting on a case.

How It Works — Matrix Decomposition

Instead of changing W (768×768 = 589K numbers), LoRA adds two small matrices:

W(frozen) + A(768×8) × B(8×768) = adapter is only 12K numbers

Code

from peft import LoraConfig, get_peft_model
lora_config = LoraConfig(r=8, lora_alpha=16, target_modules=["q_lin", "v_lin"])
model = get_peft_model(model, lora_config)
model.print_trainable_parameters()  # → 0.44% trainable
trainer.train()  # same as fine-tuning
model.save_pretrained("./lora-adapter")  # saves only adapter (few MB)

Fine-tuning vs LoRA

Fine-tuning LoRA
Train All 66M 300K (0.44%)
Save size 260 MB Few MB
Quality Baseline Nearly identical

Key Concepts

1

Freeze base model — lock all 66M numbers, set as non-trainable

2

Attach small matrices A(768×8), B(8×768) to each layer — adapter

3

Train only A and B — adjust only 0.44% of total

4

Inference: output = W×input + (A×B)×input — original weights + adapter correction

5

Save adapter only (few MB) — swap adapters per use case

Use Cases

Large model fine-tuning — train 7B/13B models on a single GPU Multi-purpose model — 1 base model + N adapters for multiple tasks Deployment size reduction — ship few MB adapter instead of 260MB