Multi-LoRA Serving - S-LoRA/Punica Batching, Adapter Hot-Swap & Memory
Merged-per-tenant vs merge-in-place vs multi-LoRA batching; Punica SGMV/BGMV gathered low-rank kernels; S-LoRA unified paging; GPU/host/disk adapter tiers, dynamic loading, adapter-affinity routing, rank caps; runnable batching + adapter cache sims.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
How can one batch serve 30 fine-tunes?
Answer
Shared base matmul plus a gathered low-rank matmul per request (SGMV/BGMV).
L2
What does it cost per request?
Answer
About 2r/d extra FLOPs per adapted layer plus kernel inefficiency and swap latency on misses.
L3
Why not merge and unmerge in place?
Answer
Each switch rewrites full weight matrices and a batch can hold only one adapter.
L4
What are the adapter tiers?
Answer
GPU slots for the hot set, host memory for the working set, disk or object storage for the long tail.
L5
Why does S-LoRA use unified paging?
Answer
Adapters of different ranks and KV would fragment memory; one paged pool keeps utilisation high.
L6
Why cap ranks?
Answer
Slots reserve the maximum rank, so one rank-256 adapter wastes memory for every slot.
L7
When would you still merge?
Answer
When one variant has heavy steady traffic.
Failure modes
Too few GPU slots for the working set
Constant swaps and head-of-line waits.
Cold tenants on the disk path
Cold-start latency for whoever lands on a cold adapter.
One tenant's burst
Fills the batch and slows everyone without fairness controls.
Misconceptions
Multi-LoRA overhead shows at batch 1.
Cost and benefit both show at realistic batch sizes with many distinct adapters.
Prefix caches are shared across adapters.
Cache keys include the adapter id.
Adapters load safely on any base revision.
Outputs degrade quietly on a different base.
Interviewer traps
Adding GPU slots to fix disk loads.
Make the host tier hold the working set.
No per-adapter metering.
Log per-adapter rate, slot hit rate, swap latency and tokens.
Design scenario
Same prompt for every reader.
Requirements
Low added latency, no cold-start spikes for hot customers, fair sharing, and per-customer billing.
Failure assumptions
- Most adapters are cold.
- A few adapters were trained at high rank.
- The base is quantized to FP8.
Constraints
- A fixed GPU budget.
- Adapters arrive daily from the fine-tuning platform.
Prompt
Serve 500 customer fine-tunes of one 8B base with Zipf-like traffic.
API
How are adapters loaded, validated and addressed per request?
Data
Which tiers, slot counts and per-adapter metrics are used?
Architecture
How do routing, fairness and rank caps fit together across replicas?
Overview
LoRA fine-tunes are small: a rank-16 adapter on an 8B model is tens of megabytes against a 16 GB base. If you merge each fine-tune into its own model, you pay for a full deployment per customer, per experiment, per language, and most of those GPUs idle. Multi-LoRA serving keeps one base model resident and applies many adapters within the same continuous batch, each request carrying its own adapter id. Making that fast needed new kernels (Punica's SGMV/BGMV gathered low-rank matmuls), new memory management (S-LoRA's unified paging of adapter weights and KV), and a tiered adapter cache (GPU slots, host memory, disk or object storage) with hot-swap APIs. This page compares serving strategies, explains the kernel trick, sizes memory, and shows where it breaks: adapter thrash, cold starts, rank heterogeneity and noisy neighbours. For how LoRA is trained (rank, alpha, NF4), see the LoRA and QLoRA page.
Strategies compared
| Strategy | GPUs for N fine-tunes | Per-request overhead | Swap cost | Good for |
|---|---|---|---|---|
| Merge each adapter, one deployment each | N x replicas (each sized for peak) | None | Full model load (minutes) | A few high-traffic variants |
| One base, merge or unmerge adapter in place per batch | 1 x replicas | None while active | Full d x d update per switch, batch can hold only one adapter | Rare switching, single tenant at a time |
| Multi-LoRA batching (Punica, S-LoRA, vLLM, SGLang, LoRAX, TensorRT-LLM) | Sized to total load | Extra low-rank matmuls, about 2r/d of the base FLOPs per adapted layer | Copy a few MB into a GPU slot | Many variants with uneven traffic |
Bigger GPU slot pool or a host tier that holds the working set?
Prefer
Host tier sized to the working set
Turn disk loads into memory copies.
- 32 GPU slots with 500 host slots cut disk loads from 13187 to 500.
- Mean load wait fell from 32.22 ms to 2.54 ms.
- A shared base needed 6 replicas instead of 500 dedicated ones.
Alternative
More GPU slots only
Raise the GPU hit rate.
- 8 to 32 GPU slots raised GPU hits from 29.6% to 54.5%.
- Disk loads stayed about 13,000 with a 100-slot host tier.
- Load wait barely moved (32.46 ms to 32.22 ms).
Lifecycle of a multi-LoRA request
Diagram 1 condensed: route, load, batch, and the failure path.
- 1
Request arrives with an adapter id
The model name maps to the base plus one adapter. - 2
Check GPU slot
Hit: join the batch immediately. - 3
Load from host or disk
PCIe copy from host; cold storage is slow. - 4
Gathered batch
Shared base plus SGMV/BGMV low-rank matmuls. - 5
Slot thrash
Too few slots for the working set causes constant swaps.
The kernel trick: gather, do not merge
For an adapted linear layer, request i with adapter a(i) needs y_i = x_i W^T + scale * (x_i A_a^T) B_a^T. The base part is one big shared matmul for the whole batch. The adapter part is a batched gather of tiny matmuls: each row uses its own A and B. BGMV (batched gather matrix-vector, one adapter per decode token) and SGMV (segmented gather, for contiguous segments of tokens sharing an adapter, as in prefill) do this in one kernel launch without materialising merged weights. vLLM, SGLang and LoRAX ship kernels in this family (Punica-derived, Triton or CUTLASS based).
"""Multi-LoRA serving: one base model, many adapters, ONE batched forward pass (numpy).
Each request i in a batch names an adapter a_i. The layer output is
y_i = W x_i + scale * B[a_i] (A[a_i] x_i)
The base matmul X @ W^T is shared by the whole batch (the expensive part). The adapter part is a
gather + two skinny matmuls per row, which S-LoRA/Punica-style kernels (BGMV / SGMV) fuse so a
mixed batch runs in one launch instead of one forward pass per adapter.
"""
import numpy as np, time
rng = np.random.default_rng(0)
d, r, n_adapters, batch, scale = 2048, 16, 50, 64, 2.0
W = rng.normal(0, 0.02, (d, d)).astype(np.float32)
A = rng.normal(0, 0.02, (n_adapters, r, d)).astype(np.float32)
B = rng.normal(0, 0.02, (n_adapters, d, r)).astype(np.float32)
X = rng.normal(0, 1, (batch, d)).astype(np.float32)
ids = rng.integers(0, n_adapters, batch) # adapter per request (mixed batch)
# Option 1: merge-per-adapter (what you would do with ONE adapter): group requests, merge W + BA, run.
t = time.perf_counter()
Y1 = np.empty_like(X)
for a in np.unique(ids):
rows = ids == a
Y1[rows] = X[rows] @ (W + scale * B[a] @ A[a]).T # a d x d merge per adapter per step!
t1 = time.perf_counter() - t
# Option 2: unmerged + gathered (BGMV-style): shared base matmul + per-row low-rank path.
t = time.perf_counter()
base = X @ W.T
shrink = np.einsum("brd,bd->br", A[ids], X) # (batch, r) "shrink"
expand = np.einsum("bdr,br->bd", B[ids], shrink) # (batch, d) "expand"
Y2 = base + scale * expand
t2 = time.perf_counter() - t
print("same outputs:", np.allclose(Y1, Y2, atol=1e-3))
print(f"distinct adapters in this batch: {len(np.unique(ids))}")
print(f"merge-per-adapter: {t1*1e3:7.1f} ms gathered low-rank: {t2*1e3:6.1f} ms (numpy CPU, relative only)")
flops_base = 2 * batch * d * d
flops_lora = 2 * batch * r * d * 2
print(f"adapter FLOPs overhead vs base matmul: {flops_lora/flops_base:.2%} (2r/d = {2*r/d:.2%})")
mem_adapter = 2 * r * d * 2 # A and B, fp16, ONE layer
print(f"one adapter, one {d}x{d} layer: {mem_adapter/1e6:.2f} MB vs base layer {d*d*2/1e6:.1f} MB")
print("\nFailure path: merging per adapter costs a full d x d update every time the adapter mix changes,")
print("and serving each fine-tune as its own merged deployment multiplies GPUs by the number of tenants.")Output:
same outputs: True
distinct adapters in this batch: 34
merge-per-adapter: 117.3 ms gathered low-rank: 5.9 ms (numpy CPU, relative only)
adapter FLOPs overhead vs base matmul: 1.56% (2r/d = 1.56%)
one adapter, one 2048x2048 layer: 0.13 MB vs base layer 8.4 MB
Failure path: merging per adapter costs a full d x d update every time the adapter mix changes,
and serving each fine-tune as its own merged deployment multiplies GPUs by the number of tenants.Memory: slots, tiers and paging
- GPU slots: the engine reserves room for a fixed number of adapters at a maximum rank (vLLM's
max_lorasandmax_lora_rank). Requests whose adapter is not in a slot wait for a swap. - Host tier: a larger LRU cache of adapters in CPU memory (vLLM
max_cpu_loras), so a swap is a PCIe copy, not a download. - Cold tier: disk or object storage; loading may take hundreds of milliseconds to seconds and should be off the request path where possible.
- Unified paging (S-LoRA): adapter weights of different ranks and KV blocks share one paged memory pool, so rank heterogeneity does not fragment memory, and thousands of adapters can be served from host memory with prefetch.
- Rank caps: reserving slots at the maximum rank wastes memory if most adapters are rank 8 and one is rank 256; cap ranks in your fine-tuning platform.
// Adapter hot-swap as a two-tier cache: GPU slots (fast, few) <- host RAM (slower, many) <- disk/object store.
// Traffic over tenants' adapters is skewed (Zipf). We measure how many requests must wait for a load,
// for different GPU slot counts, and compare fleet size against one deployment per fine-tune.
// Load latencies are EXAMPLE values for a rank-16 adapter of an 8B model (tens of MB).
function lcg(seed: number): () => number {
let s = seed >>> 0;
return () => ((s = (Math.imul(s, 1664525) + 1013904223) >>> 0) / 2 ** 32);
}
const rnd = lcg(11);
const N_ADAPTERS = 500, REQUESTS = 50_000, ZIPF_S = 1.1;
const HOST_MS = 3, DISK_MS = 120; // example: host->GPU copy vs fetch from storage
// Zipf sampler via cumulative weights.
const weights = Array.from({ length: N_ADAPTERS }, (_, i) => 1 / Math.pow(i + 1, ZIPF_S));
const total = weights.reduce((a, b) => a + b, 0);
const cdf: number[] = []; weights.reduce((acc, w, i) => (cdf[i] = acc + w / total), 0);
const sample = () => { const u = rnd(); let lo = 0, hi = N_ADAPTERS - 1; while (lo < hi) { const m = (lo + hi) >> 1; if (cdf[m] < u) lo = m + 1; else hi = m; } return lo; };
const trace = Array.from({ length: REQUESTS }, sample);
class LRU {
private m = new Map<number, true>();
constructor(private cap: number) {}
touch(k: number): boolean { // returns hit?, inserts on miss, evicts LRU
const hit = this.m.delete(k);
this.m.set(k, true);
if (this.m.size > this.cap) this.m.delete(this.m.keys().next().value as number);
return hit;
}
}
console.log("GPU slots host slots GPU hit host hit disk loads mean load wait/request");
for (const [gpuSlots, hostSlots] of [[8, 100], [32, 100], [32, 500], [128, 500]]) {
const gpu = new LRU(gpuSlots), host = new LRU(hostSlots);
let g = 0, h = 0, d = 0;
for (const a of trace) {
if (gpu.touch(a)) { g++; continue; }
if (host.touch(a)) h++; else d++;
}
const wait = (h * HOST_MS + d * DISK_MS) / REQUESTS;
console.log(`${String(gpuSlots).padStart(9)}${String(hostSlots).padStart(12)}${(100 * g / REQUESTS).toFixed(1).padStart(9)}%${(100 * h / REQUESTS).toFixed(1).padStart(9)}%${String(d).padStart(12)}${wait.toFixed(2).padStart(14)} ms`);
}
// Fleet comparison: each merged fine-tune needs >= 1 replica even when idle; shared base scales with load.
const replicasForLoad = 6; // what total traffic actually needs (example)
console.log(`\nfleet: dedicated merged deployments = ${N_ADAPTERS} replicas minimum (one per fine-tune, mostly idle)`);
console.log(` shared base + multi-LoRA = ${replicasForLoad} replicas sized to total load`);
console.log("Failure path: too few GPU slots for the working set -> constant swapping and head-of-line waits;");
console.log("cold tenants hit the disk path (cold-start latency). Pin hot adapters, prefetch on schedule, cap rank.");Output:
GPU slots host slots GPU hit host hit disk loads mean load wait/request
8 100 29.6% 44.4% 12968 32.46 ms
32 100 54.5% 19.1% 13187 32.22 ms
32 500 54.5% 44.5% 500 2.54 ms
128 500 78.2% 20.8% 500 1.83 ms
fleet: dedicated merged deployments = 500 replicas minimum (one per fine-tune, mostly idle)
shared base + multi-LoRA = 6 replicas sized to total load
Failure path: too few GPU slots for the working set -> constant swapping and head-of-line waits;
cold tenants hit the disk path (cold-start latency). Pin hot adapters, prefetch on schedule, cap rank.ExpectedGPU slots host slots GPU hit host hit disk loads mean load wait/request 8 100 29.6% 44.4% 12968 32.46 ms 32 100 54.5% 19.1% 13187 32.22 ms 32 500 54.5% 44.5% 500 2.54 ms 128 500 78.2% 20.8% 500 1.83 ms fleet: dedicated merged deployments = 500 replicas minimum (one per fine-tune, mostly idle) shared base + multi-LoRA = 6 replicas sized to total load Failure path: too few GPU slots for the working set -> constant swapping and head-of-line waits; cold tenants hit the disk path (cold-start latency). Pin hot adapters, prefetch on schedule, cap rank.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Reading the output: GPU hit rate grows with slots, but the big win is making the host tier hold the whole working set, which turns disk loads into memory copies. Traffic is usually Zipf-like: a few adapters are hot, a long tail is cold.
Diagram 1: lifecycle of a multi-LoRA request
Decisions
- 1
Step 1: request arrives with model = base and adapter id (or model name mapped to an adapter)
- nextStep 2: adapter in a GPU slot?
- ?
Step 2: adapter in a GPU slot?
- nextStep 4: join the continuous batch with its adapter index
- nextStep 3: adapter in host cache?
- nextFailure path: working set exceeds GPU slots, adapters thrash, requests queue behind swaps
- 3
Step 4: join the continuous batch with its adapter index
- nextStep 5: base matmul for the whole batch plus gathered low-rank matmuls per request (SGMV for prefill, BGMV for decode)
- ?
Step 3: adapter in host cache?
- nextStep 3a: copy weights into a free or LRU-evicted GPU slot
- nextStep 3b: load from disk or object store, then into host cache and a GPU slot
- 5
Step 3a: copy weights into a free or LRU-evicted GPU slot
- nextStep 4: join the continuous batch with its adapter index
- 6
Step 3b: load from disk or object store, then into host cache and a GPU slot
- nextStep 4: join the continuous batch with its adapter index
- 7
Step 5: base matmul for the whole batch plus gathered low-rank matmuls per request (SGMV for prefill, BGMV for decode)
- nextStep 6: prefix cache keys include the adapter id, so KV is reused only within the same adapter
- 8
Step 6: prefix cache keys include the adapter id, so KV is reused only within the same adapter
- nextStep 7: stream tokens, meter usage per adapter and tenant
- 9
Step 7: stream tokens, meter usage per adapter and tenant
- 10
Failure path: working set exceeds GPU slots, adapters thrash, requests queue behind swaps
- nextRaise slots or host cache, pin hot adapters, route tenants to replicas by adapter affinity, cap max rank
- 11
Raise slots or host cache, pin hot adapters, route tenants to replicas by adapter affinity, cap max rank
- nextStep 1: request arrives with model = base and adapter id (or model name mapped to an adapter)
Lesson map
Multi-LoRA Serving - S-LoRA/Punica Batching, Adapter Hot-Swap & Memory
Merged-per-tenant vs merge-in-place vs multi-LoRA batching; Punica SGMV/BGMV gathered low-rank kernels; S-LoRA unified paging; GPU/host/disk adapter tiers, dynamic loading, adapter-affinity routing, rank caps; runnable batching + adapter cache sims.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB s1["Step 1: request arrives with model = base and adapter id (or model name mapped to an adapter)"] s2["Step 2: adapter in a GPU slot?"] s5["Step 4: join the continuous batch with its adapter index"] s3["Step 3: adapter in host cache?"] s4["Step 3a: copy weights into a free or LRU-evicted GPU slot"] s4b["Step 3b: load from disk or object store, then into host cache and a GPU slot"] s6["Step 5: base matmul for the whole batch plus gathered low-rank matmuls per request (SGMV for prefill, BGMV for decode)"] s7["Step 6: prefix cache keys include the adapter id, so KV is reused only within the same adapter"] s8["Step 7: stream tokens, meter usage per adapter and tenant"] f1["Failure path: working set exceeds GPU slots, adapters thrash, requests queue behind swaps"] f2["Raise slots or host cache, pin hot adapters, route tenants to replicas by adapter affinity, cap max rank"] s1 -->|continues| s2 s2 -->|continues| s5 s2 -->|continues| s3 s3 -->|continues| s4 s3 -->|continues| s4b s4 -->|continues| s5 s4b -->|continues| s5 s5 -->|continues| s6 s6 -->|continues| s7 s7 -->|continues| s8 s2 -->|continues| f1 f1 -->|continues| f2 f2 -->|continues| s1
Operating it
- Dynamic loading: vLLM exposes runtime load and unload endpoints (enabled by a flag) and resolver plugins; LoRAX and SGLang support loading adapters on demand by id. Treat loading as a deployment: validate base-model match, rank, target modules and checksum.
- Routing with adapter affinity: with many replicas, route each adapter preferentially to replicas that already hold it (the same idea as prefix-aware routing), bounded by load.
- Fairness: one tenant's burst can fill the batch and slow everyone; apply per-tenant quotas and fair scheduling at the gateway (see the fairness and noisy neighbours page).
- Quantized base: adapters trained on a BF16 base usually work on an FP8 base; evaluate each important adapter on the quantized base, and for QLoRA-trained adapters remember they were trained against NF4 weights.
- Observability: per-adapter request rate, slot hit rate, swap latency, queue wait and token counts.
Decision chart: how to serve N fine-tunes
Decisions
- 1
D1
- nextSeparate deployments, or distill or convert to LoRA next time
- nextStep 2: few variants, each with heavy steady traffic?
- 2
Separate deployments, or distill or convert to LoRA next time
- ?
Step 2: few variants, each with heavy steady traffic?
- nextMerge each into its own deployment, zero adapter overhead
- nextStep 3: does the hot working set fit in GPU slots?
- 4
Merge each into its own deployment, zero adapter overhead
- ?
Step 3: does the hot working set fit in GPU slots?
- nextMulti-LoRA serving with static slots, preload hot adapters
- nextMulti-LoRA with host cache for the working set, dynamic loading, adapter-affinity routing
- 6
Multi-LoRA serving with static slots, preload hot adapters
- nextStep 4: one adapter dominates traffic?
- 7
Multi-LoRA with host cache for the working set, dynamic loading, adapter-affinity routing
- nextStep 4: one adapter dominates traffic?
- ?
Step 4: one adapter dominates traffic?
- nextGive it a merged dedicated deployment, keep the long tail on the shared base
- nextKeep everything shared, enforce per-tenant quotas
- 9
Give it a merged dedicated deployment, keep the long tail on the shared base
- 10
Keep everything shared, enforce per-tenant quotas
What happens if you choose otherwise
- One merged deployment per customer: hundreds of replicas sized for peak, mostly idle; cost scales with customer count instead of traffic.
- Merge and unmerge in place per batch: every adapter switch rewrites full weight matrices and batches cannot mix adapters, so throughput collapses with many tenants.
- Too few GPU slots for the working set: constant swaps, latency spikes for whoever lands on a cold adapter.
- No rank cap: one high-rank adapter forces every slot to reserve its size.
Pitfalls
- Forgetting that prefix-cache keys include the adapter id, then expecting cross-adapter cache hits on the same system prompt.
- Loading an adapter trained on a different base revision; outputs degrade quietly rather than failing.
- Measuring multi-LoRA overhead at batch 1; the cost and the benefit both show at realistic batch sizes with many distinct adapters.
- No per-adapter metering, so you cannot bill or debug tenants.
Interview Q&A
How can one batch contain requests for 30 different fine-tunes?
Answer
The base matmul is shared; the adapter contribution is a gathered batch of small matmuls where each request uses its own A and B (SGMV/BGMV kernels), so no merged weights are materialised.
What does multi-LoRA serving cost per request?
Answer
Roughly 2r/d extra FLOPs per adapted layer (low single-digit percent for typical ranks), plus kernel inefficiency with many distinct adapters, plus swap latency on slot misses.
How do you design the adapter cache?
Answer
Tiers: GPU slots for the hot set, host memory for the working set, object storage for the long tail, with LRU eviction, pinning of hot adapters, prefetch, and adapter-affinity routing across replicas.
When would you still merge an adapter?
Answer
When one variant has heavy steady traffic: a merged deployment avoids adapter overhead and slot contention.
Why does S-LoRA use unified paging?
Answer
Adapters of different ranks and dynamic KV caches would fragment GPU memory; one paged pool for both keeps utilisation high and lets thousands of adapters stream from host memory.
How do vLLM settings map to the tiers?
Answer
max_loras and max_lora_rank size the GPU slots; max_cpu_loras sizes the host LRU tier.
How should dynamic adapter loading be treated?
Answer
As a deployment: validate base-model match, rank, target modules and checksum.
What should you watch when serving adapters on a quantized base?
Answer
Evaluate each important adapter on the quantized base; QLoRA-trained adapters were trained against NF4 weights.
Check yourself
List your fine-tunes with their traffic share and rank. Decide which to merge, how many GPU slots and host slots you need for the working set, and the rank cap you would enforce.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: PagedAttention & Continuous Batching, vLLM vs SGLang — Runtime Choice (comparative), Fairness, Quotas & Noisy Neighbors, LoRA & QLoRA - Rank, Alpha, Target Modules, NF4 & Adapter Merging.