LoRA & QLoRA - Rank, Alpha, Target Modules, NF4 & Adapter Merging
LoRA math (W + alpha/r BA), rank/alpha/target modules, rsLoRA/DoRA notes, QLoRA NF4 + double quantization + paged optimizers, merging vs keeping adapters, when PEFT is the wrong choice; runnable LoRA/NF4 demo + memory budget calculator.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What does LoRA learn?
Answer
A low-rank update B A, scaled by alpha / r, on top of a frozen W.
L2
Why initialise B to zero?
Answer
So at step 0 the adapted model equals the base model.
L3
What does rank control?
Answer
The capacity of the update.
L4
How do alpha and rank interact?
Answer
With alpha / r, doubling r at fixed alpha halves the per-direction scale; rsLoRA uses alpha / sqrt(r).
L5
Which modules should get adapters?
Answer
All linear layers were needed in QLoRA's ablations to match full fine-tuning.
L6
What are QLoRA's three ideas?
Answer
NF4, double quantisation of block scales, and paged optimizers.
L7
When is LoRA the wrong choice?
Answer
Big knowledge or domain shifts, continued pretraining, or when you need the last bit of quality.
Failure modes
Attention-only adapters for a new domain
Training plateaus because the MLP layers carry much of the knowledge.
Merging into the 4-bit base directly
Compounding quantisation error looks like a bad fine-tune.
Wrong base revision at load time
The adapter quietly degrades on a base it was not trained on.
Misconceptions
Rank does not matter.
Rank sweeps without adjusting alpha or learning rate give confusing results.
LoRA prevents forgetting.
It limits drift, but a high learning rate and many epochs still overwrite behaviour.
QLoRA is just faster LoRA.
It is slower per step; it trades speed for fitting in memory.
Interviewer traps
Using QLoRA when bf16 LoRA already fits.
Use bf16 LoRA; QLoRA only slows each step.
Changing the tokenizer and saving only the adapter.
Save the resized embeddings too.
Design scenario
Same prompt for every reader.
Requirements
Fit training on the available GPUs, keep quality close to full fine-tuning, and decide which adapters to merge.
Failure assumptions
- Activations push memory over budget on long sequences.
- The serving base is quantised.
- Some adapters were trained at rank 256.
Constraints
- One 80 GB GPU per training job.
- Dozens of customer variants.
Prompt
Fine-tune a 70B model for a brand tone and also serve per-customer variants of an 8B model.
API
Which adapter config (rank, alpha, target modules) does each job use?
Data
How are adapters versioned against the exact base revision and tokenizer?
Architecture
Which jobs use LoRA vs QLoRA, and which adapters are merged vs served separately?
Overview
LoRA (Low-Rank Adaptation) freezes the pretrained weight matrix W and learns its update as the product of two thin matrices: W' = W + (alpha / r) * B A, with A of shape r x d_in and B of shape d_out x r, and r (the rank) typically 8-64 against hidden sizes in the thousands. Only A and B get gradients and optimizer state, so trainable parameters drop to well under 1 percent and the optimizer memory nearly vanishes. QLoRA goes further: it stores the frozen base in 4-bit NormalFloat (NF4) with double-quantised block scales, dequantises on the fly for each matmul, and trains bf16 LoRA adapters on top, with paged optimizers absorbing memory spikes. The result is fine-tuning a 70B model on one 80 GB GPU. This page explains rank, alpha, target modules, NF4, merging, and when parameter-efficient fine-tuning is the wrong choice.
The knobs
| Knob | What it does | Typical values | If too small | If too large |
|---|---|---|---|---|
| Rank r | Capacity of the update | 8-64 (behaviour), 64-256 (heavier domain shift) | Underfits, plateaus early | More memory, slower, can overfit small data |
| Alpha | Scales the update by alpha / r | alpha = r or 2r is common | Update too weak, needs higher LR | Unstable, effectively a higher LR |
| Target modules | Which linear layers get adapters | Attention q,v (original paper) up to all linear layers (q,k,v,o,gate,up,down) | Misses capacity in MLP layers | More params, usually still small |
| Learning rate | Step size for A and B | About 1e-4 to 2e-4 (higher than full FT) | Slow | Diverges, forgets |
| Dropout on the adapter input | Regularisation | 0-0.1 | - | Slower convergence |
Alpha and rank interact. With the alpha / r scale, doubling r at fixed alpha halves the per-direction scale, which is why sweeping r without adjusting alpha or LR gives confusing results. rsLoRA proposes alpha / sqrt(r) so the update's magnitude stays stable as rank grows (visible in the runnable output above).
Target modules. QLoRA's ablations found applying adapters to all linear layers was needed to match full fine-tuning quality; attention-only adapters are cheaper but leave the MLP blocks, where much of the model's knowledge lives, untouched.
bf16 LoRA or QLoRA on a 70B base?
Prefer
QLoRA (NF4 base + bf16 adapters)
Store the frozen base in 4-bit NF4 with double-quantised scales.
- 70B with r=16 all-linear needs 39.7 GB in the calculator.
- Double quantisation cut storage from 4.500 to 4.127 bits per parameter.
- Slower per step: each block is dequantised before its matmul.
Alternative
bf16 LoRA
Keep the frozen base in bf16 and train adapters on top.
- 70B with r=16 all-linear still needs 145 GB for the base.
- Faster steps when it fits, as on 8B (17 GB).
- Full fine-tuning 70B would need about 1130 GB.
The LoRA lifecycle
Diagram 1 condensed: train, then merge or keep separate, with the failure path.
- 1
Configure
Rank, alpha, target modules, learning rate. - 2
Train adapters
Base frozen (bf16 or NF4); only A and B learn. - 3
Evaluate
Compare against the base and the target evals. - 4
Merge or keep separate
Merge for one deployment; keep separate for many tenants. - 5
Merge into 4-bit base
Compounding quantisation error; dequantise to bf16 first.
Why low-rank works
The LoRA paper's hypothesis, supported empirically: the change a fine-tune makes to a pretrained weight has low intrinsic rank, even though the weight itself is full rank. Fine-tuning mostly re-weights directions the model already has. A rank-16 update can express a lot of behavioural change; it cannot easily express a whole new body of knowledge.
Initialisation matters: A random, B = 0, so at step 0 the adapted model is exactly the base model and training starts from a known-good point.
"""LoRA forward pass, merge, parameter counts, and NF4 vs uniform INT4 quantisation (numpy).
LoRA: keep W (d_out x d_in) frozen; learn B (d_out x r) and A (r x d_in); the layer becomes
y = W x + (alpha / r) * B (A x). B starts at zero, so training starts exactly at the base model.
QLoRA: store W in 4-bit NormalFloat (NF4) with one absmax scale per block of 64 weights, dequantise
on the fly to bf16 for the matmul, and train only A and B in higher precision.
"""
import numpy as np
rng = np.random.default_rng(0)
d_out, d_in, r, alpha = 1024, 1024, 16, 32
W = rng.normal(0, 0.02, (d_out, d_in)).astype(np.float32) # frozen pretrained weight
A = rng.normal(0, 1 / np.sqrt(d_in), (r, d_in)).astype(np.float32) # random init
B = np.zeros((d_out, r), dtype=np.float32) # zero init -> delta W = 0 at step 0
x = rng.normal(0, 1, (d_in,)).astype(np.float32)
scale = alpha / r
print("step 0 identical to base:", np.allclose(W @ x, W @ x + scale * (B @ (A @ x))))
B = rng.normal(0, 0.01, (d_out, r)).astype(np.float32) # pretend training updated B
y_lora = W @ x + scale * (B @ (A @ x)) # two skinny matmuls, W untouched
W_merged = W + scale * (B @ A) # merge for zero-overhead serving
print("merged == unmerged output:", np.allclose(y_lora, W_merged @ x, atol=1e-5))
full, lora = d_out * d_in, r * (d_out + d_in)
print(f"trainable params per layer: full={full:,} LoRA r={r}: {lora:,} ({lora/full:.2%})")
print(f"rank of the learned update B@A: {np.linalg.matrix_rank(B @ A)} (<= r by construction)")
# Why alpha/r: with alpha fixed, raising r would otherwise change the update's scale.
# Averaged over 20 random draws so the trend is visible: raw update norm grows ~sqrt(r).
for rr in (4, 16, 64):
norms = []
for _ in range(20):
Ar = rng.normal(0, 1 / np.sqrt(d_in), (rr, d_in)); Br = rng.normal(0, 0.01, (d_out, rr))
norms.append(np.linalg.norm(Br @ (Ar @ x)))
n = float(np.mean(norms))
print(f" r={rr:>2}: |BAx| raw={n:6.3f} x alpha/r={n*alpha/rr:6.3f} x alpha/sqrt(r) (rsLoRA)={n*alpha/np.sqrt(rr):6.3f}")
# NF4: 16 levels placed at quantiles of a standard normal (values from the QLoRA paper / bitsandbytes).
NF4 = np.array([-1.0, -0.6961928, -0.5250731, -0.3949175, -0.2844414, -0.1847734, -0.0910500, 0.0,
0.0795803, 0.1609302, 0.2461123, 0.3379152, 0.4407098, 0.5626170, 0.7229568, 1.0])
INT4 = np.linspace(-1, 1, 16) # uniform 4-bit grid on the same [-1, 1] range
def quant_dequant(w, levels, block=64):
w = w.reshape(-1, block)
absmax = np.abs(w).max(axis=1, keepdims=True) # one scale per block
normed = w / absmax
idx = np.abs(normed[..., None] - levels).argmin(axis=-1) # nearest level (a 4-bit code)
return (levels[idx] * absmax).reshape(-1)
flat = W.reshape(-1)
for name, lv in (("uniform INT4", INT4), ("NF4", NF4)):
err = flat - quant_dequant(flat, lv)
print(f"{name:<13} RMSE={np.sqrt((err**2).mean()):.6f} rel={np.linalg.norm(err)/np.linalg.norm(flat):.3%}")
bits = 4 + 32 / 64 # 4-bit codes + one fp32 absmax per 64 weights
bits_dq = 4 + 8 / 64 + 32 / (64 * 256) # double quantisation: 8-bit absmax, fp32 scale per 256 blocks
print(f"storage: {bits:.3f} bits/param with fp32 scales, {bits_dq:.3f} with double quantisation")Output:
step 0 identical to base: True
merged == unmerged output: True
trainable params per layer: full=1,048,576 LoRA r=16: 32,768 (3.12%)
rank of the learned update B@A: 16 (<= r by construction)
r= 4: |BAx| raw= 0.571 x alpha/r= 4.571 x alpha/sqrt(r) (rsLoRA)= 9.142
r=16: |BAx| raw= 1.240 x alpha/r= 2.480 x alpha/sqrt(r) (rsLoRA)= 9.919
r=64: |BAx| raw= 2.523 x alpha/r= 1.261 x alpha/sqrt(r) (rsLoRA)=10.092
uniform INT4 RMSE=0.002013 rel=10.058%
NF4 RMSE=0.001840 rel=9.192%
storage: 4.500 bits/param with fp32 scales, 4.127 with double quantisationMemory: where the bytes go
// Training-memory budget: full fine-tune vs LoRA vs QLoRA for Llama-3-shaped models.
// Simplified model: counts weights, gradients and AdamW states only. Activations, KV for
// long sequences, and framework overhead are EXTRA (use gradient checkpointing to tame them).
type Cfg = { name: string; params: number; layers: number; hidden: number; inter: number; kvDim: number };
const models: Cfg[] = [
{ name: "8B (Llama-3-8B shape)", params: 8.03e9, layers: 32, hidden: 4096, inter: 14336, kvDim: 1024 },
{ name: "70B (Llama-3-70B shape)", params: 70.6e9, layers: 80, hidden: 8192, inter: 28672, kvDim: 1024 },
];
// LoRA params per layer = r * (d_in + d_out) for each targeted linear layer.
function loraParams(m: Cfg, r: number, targets: "attn-qv" | "all-linear"): number {
const h = m.hidden, i = m.inter, kv = m.kvDim;
const mods: Array<[number, number]> = targets === "attn-qv"
? [[h, h], [h, kv]] // q_proj, v_proj (original LoRA paper style)
: [[h, h], [h, kv], [h, kv], [h, h], [h, i], [h, i], [i, h]]; // q,k,v,o,gate,up,down
return m.layers * mods.reduce((s, [a, b]) => s + r * (a + b), 0);
}
const GB = 1e9;
const ADAM_MIXED = 16; // bf16 weight 2 + bf16 grad 2 + fp32 master 4 + Adam m 4 + v 4 bytes
const NF4_BYTES = 4.127 / 8; // NF4 codes + double-quantised block scales
for (const m of models) {
console.log(`\n${m.name}`);
const full = m.params * ADAM_MIXED;
console.log(` full fine-tune : ${(full / GB).toFixed(0).padStart(5)} GB (every weight trainable)`);
for (const [r, t] of [[16, "attn-qv"], [16, "all-linear"], [64, "all-linear"]] as Array<[number, "attn-qv" | "all-linear"]>) {
const p = loraParams(m, r, t);
const lora = m.params * 2 + p * ADAM_MIXED;
const qlora = m.params * NF4_BYTES + p * ADAM_MIXED;
console.log(` r=${String(r).padEnd(2)} ${t.padEnd(10)} adapters=${(p / 1e6).toFixed(1).padStart(6)}M (${((100 * p) / m.params).toFixed(2)}%) LoRA bf16 base: ${(lora / GB).toFixed(0).padStart(4)} GB QLoRA NF4 base: ${(qlora / GB).toFixed(1).padStart(5)} GB`);
}
}
console.log("\nRead it as: full FT of 8B already needs 2+ GPUs with ZeRO/FSDP sharding; LoRA on 8B fits one 80 GB GPU,");
console.log("but LoRA on 70B still needs ~145 GB for the bf16 base. QLoRA fits 8B on a 24 GB card and 70B on one");
console.log("80 GB card, before activations. Adapter size barely matters;");
console.log("the frozen base dominates, which is why 4-bit base storage is the big lever.");Output:
8B (Llama-3-8B shape)
full fine-tune : 128 GB (every weight trainable)
r=16 attn-qv adapters= 6.8M (0.08%) LoRA bf16 base: 16 GB QLoRA NF4 base: 4.3 GB
r=16 all-linear adapters= 41.9M (0.52%) LoRA bf16 base: 17 GB QLoRA NF4 base: 4.8 GB
r=64 all-linear adapters= 167.8M (2.09%) LoRA bf16 base: 19 GB QLoRA NF4 base: 6.8 GB
70B (Llama-3-70B shape)
full fine-tune : 1130 GB (every weight trainable)
r=16 attn-qv adapters= 32.8M (0.05%) LoRA bf16 base: 142 GB QLoRA NF4 base: 36.9 GB
r=16 all-linear adapters= 207.1M (0.29%) LoRA bf16 base: 145 GB QLoRA NF4 base: 39.7 GB
r=64 all-linear adapters= 828.4M (1.17%) LoRA bf16 base: 154 GB QLoRA NF4 base: 49.7 GB
Read it as: full FT of 8B already needs 2+ GPUs with ZeRO/FSDP sharding; LoRA on 8B fits one 80 GB GPU,
but LoRA on 70B still needs ~145 GB for the bf16 base. QLoRA fits 8B on a 24 GB card and 70B on one
80 GB card, before activations. Adapter size barely matters;
the frozen base dominates, which is why 4-bit base storage is the big lever.Expected8B (Llama-3-8B shape) full fine-tune : 128 GB (every weight trainable) r=16 attn-qv adapters= 6.8M (0.08%) LoRA bf16 base: 16 GB QLoRA NF4 base: 4.3 GB r=16 all-linear adapters= 41.9M (0.52%) LoRA bf16 base: 17 GB QLoRA NF4 base: 4.8 GB r=64 all-linear adapters= 167.8M (2.09%) LoRA bf16 base: 19 GB QLoRA NF4 base: 6.8 GB 70B (Llama-3-70B shape) full fine-tune : 1130 GB (every weight trainable) r=16 attn-qv adapters= 32.8M (0.05%) LoRA bf16 base: 142 GB QLoRA NF4 base: 36.9 GB r=16 all-linear adapters= 207.1M (0.29%) LoRA bf16 base: 145 GB QLoRA NF4 base: 39.7 GB r=64 all-linear adapters= 828.4M (1.17%) LoRA bf16 base: 154 GB QLoRA NF4 base: 49.7 GB Read it as: full FT of 8B already needs 2+ GPUs with ZeRO/FSDP sharding; LoRA on 8B fits one 80 GB GPU, but LoRA on 70B still needs ~145 GB for the bf16 base. QLoRA fits 8B on a 24 GB card and 70B on one 80 GB card, before activations. Adapter size barely matters; the frozen base dominates, which is why 4-bit base storage is the big lever.
Press Run. Snippets must be self-contained — no network, files, or native modules.
The pattern: once you stop training the base, the frozen base weights dominate memory, so the next lever is storing them in fewer bits. That is QLoRA's contribution.
QLoRA's three ideas
- NF4 (4-bit NormalFloat). Pretrained weights are roughly normally distributed. NF4 places its 16 levels at quantiles of a normal distribution (with an exact zero), scaled per block of 64 weights by the block's absmax. For normal data it has lower error than a uniform 4-bit grid at the same storage.
- Double quantisation. The per-block scales themselves are quantised (8-bit, with a second-level scale per 256 blocks), cutting scale overhead from 0.5 to about 0.127 bits per parameter.
- Paged optimizers. Optimizer state can page between GPU and CPU memory (unified memory) to survive the memory spikes of long sequences.
Compute still happens in bf16: each 4-bit block is dequantised just before its matmul. QLoRA is therefore slower per step than bf16 LoRA; you trade speed for fitting at all.
The LoRA lifecycle
Diagram 1: train, merge or keep separate, and the failure path.
Decisions
- 1
Step 1: load frozen base (bf16 for LoRA, NF4 for QLoRA)
- nextStep 2: insert adapters A, B on target modules, B initialised to zero
- 2
Step 2: insert adapters A, B on target modules, B initialised to zero
- nextStep 3: train only A and B with the SFT or DPO loss of your choice
- 3
Step 3: train only A and B with the SFT or DPO loss of your choice
- nextStep 4: evaluate adapter on task and regression sets
- ?
Step 4: evaluate adapter on task and regression sets
- nextStep 5: one tenant or many?
- nextFailure path: plateau below target quality
- ?
Step 5: one tenant or many?
- nextStep 6a: merge W + (alpha/r) B A into the base, serve as a normal model
- nextStep 6b: keep adapters separate, serve many on one base (multi-LoRA serving)
- 6
Step 6a: merge W + (alpha/r) B A into the base, serve as a normal model
- nextMerge pitfall: merging into a 4-bit base re-quantises and can lose quality, merge into a bf16 copy
- 7
Step 6b: keep adapters separate, serve many on one base (multi-LoRA serving)
- 8
Failure path: plateau below target quality
- nextRaise rank, target all linear layers, tune alpha and LR; if still short, full fine-tune
- 9
Raise rank, target all linear layers, tune alpha and LR; if still short, full fine-tune
- nextStep 2: insert adapters A, B on target modules, B initialised to zero
- 10
Merge pitfall: merging into a 4-bit base re-quantises and can lose quality, merge into a bf16 copy
Lesson map
LoRA & QLoRA - Rank, Alpha, Target Modules, NF4 & Adapter Merging
LoRA math (W + alpha/r BA), rank/alpha/target modules, rsLoRA/DoRA notes, QLoRA NF4 + double quantization + paged optimizers, merging vs keeping adapters, when PEFT is the wrong choice; runnable LoRA/NF4 demo + memory budget calculator.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB s1["Step 1: load frozen base (bf16 for LoRA, NF4 for QLoRA)"] s2["Step 2: insert adapters A, B on target modules, B initialised to zero"] s3["Step 3: train only A and B with the SFT or DPO loss of your choice"] s4["Step 4: evaluate adapter on task and regression sets"] s5["Step 5: one tenant or many?"] s6["Step 6a: merge W + (alpha/r) B A into the base, serve as a normal model"] s7["Step 6b: keep adapters separate, serve many on one base (multi-LoRA serving)"] f1["Failure path: plateau below target quality"] f2["Raise rank, target all linear layers, tune alpha and LR if still short, full fine-tune"] f3["Merge pitfall: merging into a 4-bit base re-quantises and can lose quality, merge into a bf16 copy"] s1 -->|continues| s2 s2 -->|continues| s3 s3 -->|continues| s4 s4 -->|continues| s5 s5 -->|continues| s6 s5 -->|continues| s7 s4 -->|continues| f1 f1 -->|continues| f2 f2 -->|continues| s2 s6 -->|continues| f3
Merge vs keep separate
| Choice | Serving cost | Flexibility | When |
|---|---|---|---|
| Merge into base | Zero overhead, normal model | One fine-tune per deployment | A single product variant |
| Keep adapter separate | Small extra matmuls per layer | Many adapters on one base, hot-swap | Multi-tenant, A/B tests, per-customer tunes |
Merging a QLoRA adapter: dequantise the base to bf16, merge there, then re-quantise for serving if needed, and re-run evals because quantisation error now applies to the merged weights.
Variants worth knowing
- DoRA: decomposes weights into magnitude and direction and applies LoRA to the direction; often closes part of the gap to full fine-tuning.
- rsLoRA: rank-stabilised scaling
alpha / sqrt(r). - LoRA+: different learning rates for A and B.
- Prefix and prompt tuning, adapters (bottleneck layers): older PEFT methods; LoRA dominates because it can be merged with zero inference cost.
Decision chart: full fine-tune, LoRA or QLoRA?
Decisions
- 1
D1
- nextFull fine-tune if you can shard it (FSDP or ZeRO)
- nextStep 2: does the bf16 base plus activations fit your GPU?
- 2
Full fine-tune if you can shard it (FSDP or ZeRO)
- nextServe as its own model
- ?
Step 2: does the bf16 base plus activations fit your GPU?
- nextLoRA on all linear layers, r 16-64
- nextQLoRA: NF4 base, double quantisation, paged optimizer
- 4
LoRA on all linear layers, r 16-64
- nextStep 3: will you serve many variants of the same base?
- 5
QLoRA: NF4 base, double quantisation, paged optimizer
- nextStep 3: will you serve many variants of the same base?
- ?
Step 3: will you serve many variants of the same base?
- nextKeep adapters separate, multi-LoRA serving
- nextMerge into a bf16 base, then quantise for serving and re-evaluate
- 7
Keep adapters separate, multi-LoRA serving
- 8
Merge into a bf16 base, then quantise for serving and re-evaluate
- 9
Serve as its own model
What happens if you choose otherwise
- Full fine-tune for a tone change: 10-30x more training memory and one more full model to store and serve, for quality LoRA usually matches.
- LoRA on attention-only for a new domain: plateaus below target; the MLP layers carry much of the knowledge.
- QLoRA when bf16 LoRA fits: slower steps for no benefit.
- Merge into the 4-bit base directly: compounding quantisation error; quality drops that look like a bad fine-tune.
Pitfalls
- Comparing ranks without adjusting alpha or LR, then concluding "rank does not matter."
- Forgetting that the adapter must be loaded with the exact base revision it was trained on.
- Saving only the adapter but changing the tokenizer (new special tokens) without saving resized embeddings.
- Assuming LoRA prevents forgetting entirely; it limits drift but a high LR and many epochs still overwrite behaviour.
Interview Q&A
Explain LoRA in one breath.
Answer
Freeze W, learn the update as B times A with a small rank r, scale by alpha over r, initialise B to zero so training starts at the base model, and optionally merge B A into W after training for zero inference overhead.
Why does QLoRA fit a 70B model on one GPU?
Answer
The frozen base is stored in 4-bit NF4 with double-quantised scales (about 4.1 bits per parameter, roughly 37 GB for 70B), only small bf16 adapters get gradients and Adam state, and paged optimizers absorb spikes. Activations still need memory, so gradient checkpointing helps.
How do you choose rank and alpha?
Answer
Start at r 16 with alpha 16-32 on all linear layers; raise r if the task plateaus, keep alpha/r or use rsLoRA scaling so magnitude stays comparable, and tune LR alongside.
When is LoRA the wrong choice?
Answer
Big knowledge or domain shifts, continued pretraining on large corpora, or when you need the last bit of quality and have the compute for full fine-tuning.
Merge the adapter or not?
Answer
Merge for a single deployment (no overhead). Keep separate to serve many adapters on one base or to swap them dynamically.
How do you merge a QLoRA adapter safely?
Answer
Dequantise the base to bf16, merge there, re-quantise for serving if needed, and re-run evals.
What do DoRA and LoRA+ change?
Answer
DoRA applies LoRA to the direction of a magnitude-direction decomposition; LoRA+ uses different learning rates for A and B.
Why does NF4 beat a uniform 4-bit grid on weights?
Answer
Weights are roughly normal, and NF4 places its 16 levels at normal quantiles; in the demo NF4 RMSE was 0.001840 vs 0.002013 for uniform INT4.
Check yourself
For a model you might fine-tune, compute the full fine-tune memory (about 16 bytes per parameter), the bf16 LoRA memory, and the QLoRA memory, then decide which fits your GPUs and whether you would merge.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: Distributed PyTorch — DDP, FSDP, Tensor & Pipeline Parallelism, TensorRT-LLM — Engine Build, In-Flight Batching & Quantization, Multi-LoRA Serving - S-LoRA/Punica Batching, Adapter Hot-Swap & Memory, LLM Weight Quantization - GPTQ, AWQ, SmoothQuant, FP8 & INT4/INT8 Trade-offs.