Operating systems
Part 5 of 6 · Virtual Memory — Paging, Swapping & CachesScaling Memory — Huge Pages, NUMA, Bandwidth, and Overcommit
Memory does not scale by adding gigabytes alone. Four walls appear as working sets grow. First, translation: with 4 KiB pages the TLB covers only a few megabytes, so large random-access heaps pay page walks constantly; huge pages (2 MiB, 1 GiB) multiply TLB reach but bring compaction stalls, bloat, and copy-on-write amplification. Second, locality: multi-socket servers are NUMA, so memory attached to the other socket is slower and its interconnect is shared. Third, bandwidth: many workloads, including LLM decode, are limited by bytes per second, not FLOPs or cores, so the fix is moving fewer bytes. Fourth, promises: Linux overcommits virtual memory, and when the bet fails the OOM killer picks a victim, often your largest cache. This lesson covers each wall with tradeoff tables, a TLB-reach and bandwidth calculator, an overcommit simulator, and the container and Kubernetes angle.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What is TLB reach?
Answer
How many bytes the TLB can translate without a walk: entry count times page size. About 6 MiB with 4 KiB pages and a 1536-entry second-level TLB.
L2
When do huge pages help?
Answer
Big, dense, long-lived, randomly accessed memory: buffer pools, JVM heaps, large hash tables.
L3
When do they hurt?
Answer
Fork-heavy or sparse memory. Redis BGSAVE and many small processes. Transparent huge pages copy 2 MiB and khugepaged compacts.
L4
What is first-touch NUMA policy?
Answer
Linux allocates the physical page on the node of the CPU that first writes it. One initializer thread can pin a whole pool to node 0.
L5
Why does a scan stop scaling with cores?
Answer
Memory channels saturate. More threads queue on the same bandwidth. The fix is fewer bytes or a closer NUMA node.
L6
What do the three overcommit modes do?
Answer
0 refuses obviously impossible requests. 1 never refuses. 2 enforces a commit limit of swap plus RAM times the ratio and returns ENOMEM early.
L7
Why is LLM decode a memory-system problem?
Answer
Each step reads the weights and the batch's KV. Tokens per second is roughly bandwidth divided by those bytes. Paging exists so more sequences fit per weight read.
Failure modes
THP always on a Redis host
Copy-on-write during BGSAVE copies 2 MiB pages. Latency spikes and the OOM risk rises.
First-touch on the wrong socket
Every later reader pays remote latency and shares the interconnect. The laptop benchmark hid it.
Limit equal to the heap
Page cache, thread stacks, and allocator overhead still charge the cgroup. The kill arrives under 100 percent of RSS.
Strict overcommit under a fat reservation
JVMs, Go, and fork snapshots reserve address space they may not touch. Mode 2 returns ENOMEM and the fork fails.
Misconceptions
More cores raise scan throughput without a limit.
Once the memory channels are full, extra cores add contention. The ceiling is bandwidth over bytes per row.
Huge pages are free speed.
They cost bloat, compaction, and larger copy-on-write. Explicit huge pages also reserve RAM others cannot use.
OOM means the node is out of RAM.
A cgroup limit kills inside the container while the node still has free pages.
Interviewer traps
Quoting a GPU token rate without the bytes in the denominator.
State weights plus KV, name the bandwidth, and point the KV formula at the prefill lesson.
Recommending overcommit mode 2 for Redis BGSAVE.
Redis wants mode 1 so the virtual duplication of fork is allowed. Most pages are never copied.
Design scenario
Same prompt for every reader.
Requirements
Assign a huge-page policy per workload, a NUMA placement, a bandwidth explanation for the thread ceiling, and an overcommit mode per process. Size the cgroup with headroom.
Traffic / scale
Random lookups in the pool, fork snapshots on Redis, and decode on the GPU.
Latency
Remote NUMA and THP copy-on-write show up in the tail. Bandwidth shows up as a flat throughput curve.
Consistency
A strict commit limit must fail the reservation, not a later random kill, if you choose mode 2.
Availability
The OOM victim under mode 1 is the largest RSS, often the cache. Protect it on purpose.
Failure assumptions
- Transparent huge pages may be always.
- The initializer runs on one CPU.
- The GPU's PCIe link attaches to socket 0.
Constraints
- Do not benchmark only on a laptop and call the server identical.
- Do not set the limit equal to the heap.
Prompt
A two-socket host runs one Redis with BGSAVE, one large buffer pool initialized by a single thread, and a GPU on socket 0. A container limit equals the process heap. vm.overcommit_memory is 2 because someone read that it is safer. Throughput stopped scaling past 16 threads.
API
Which sysctl or cgroup isolates one socket?
Data
What is TLB reach at 4 KiB versus 2 MiB for 1536 entries?
Architecture
Which process gets mode 1, and which memory is interleaved?
Which wall you are actually hitting
Prefer
Name reach, place, bandwidth, or promises
A random heap wants a bigger page. A two-socket box wants local memory. A scan wants fewer bytes. A fork wants an overcommit mode that allows the virtual copy.
- Prefer transparent huge pages in madvise mode and let the application opt in.
- One process per socket, or interleave only the structures every thread hits.
- Set maxmemory and the language heap below the cgroup limit.
Alternative
Add RAM and cores and leave the policy at the default
THP stays always, the pool is first-touched on node 0, the container limit equals the heap, and strict overcommit refuses the fork.
- The TLB still covers a few megabytes of a 64 GiB heap.
- Half the cores read remote memory.
- The largest RSS dies at the spike.
Four walls, in the order they show up
Gigabytes do not move a workload off the wall it is already on.
- 1
Translation
4 KiB pages make a large random heap walk the page tables. Huge pages multiply reach and grow the cost of a copy. - 2
Locality
The socket that first writes the page owns it. Remote access is slower and shares a link. A GPU on the far socket halves the effective copy. - 3
Bandwidth
Little compute per byte means throughput is bandwidth divided by bytes moved. More cores queue. - 4
Promises
Overcommit lets reservations exceed RAM. The failure is either an early ENOMEM or a late OOM kill.
Wall 1: huge pages
| Option | How you get it | Pros | Cons | Typical users |
|---|---|---|---|---|
| 4 KiB | The default | Fine-grained, low waste, cheap copy-on-write | Small TLB reach. Big page tables | Everything, until it hurts |
| Transparent huge pages | The kernel promotes aligned 2 MiB regions in always or madvise mode | No application change. Large TLB reach | khugepaged and compaction stalls. Bloat when sparse. Copy-on-write copies 2 MiB | JVMs and some databases in madvise mode |
| Explicit huge pages | Reserved at boot or via sysctl. The app maps them | Predictable. Not swapped. No compaction at fault time | Pre-sized. Reserved RAM is unusable by others | Oracle, Postgres with huge_pages on, DPDK, some KV stores |
| 1 GiB pages | Gigantic hugetlb pages, usually reserved at boot | Maximum reach for a static arena | Few TLB entries for them. All-or-nothing reservation | Large in-memory databases, VM backing for guest RAM |
Rules of thumb:
- Big, dense, long-lived, randomly accessed memory benefits. Buffer pools, JVM heaps, large hash tables.
- Fork-heavy or sparse memory suffers. Redis with BGSAVE, and many small processes. Redis and MongoDB recommend disabling transparent huge pages or using madvise.
- Prefer
madvisesystem-wide and let applications that know better opt in.
Wall 2: NUMA
A two-socket server is two memory systems. Each socket has its own controllers and DIMMs. The other socket is a hop across UPI or Infinity Fabric, often about 1.3x to 2x the latency, and the link is shared.
Flow
- 1
1 Cores 0-31 on socket 0
- next2 Local DRAM node 0
- 5 remote - slower shared link4 Local DRAM node 1
- 2
2 Local DRAM node 0
- 3
3 Cores 32-63 on socket 1
- next4 Local DRAM node 1
- 6 remote - slower shared link2 Local DRAM node 0
- 4
4 Local DRAM node 1
- 5
7 GPU or NIC on the PCIe root of socket 0
- next1 Cores 0-31 on socket 0
Lesson map
Scaling Memory — Huge Pages, NUMA, Bandwidth, and Overcommit
Memory does not scale by adding gigabytes alone. Four walls appear as working sets grow. First, translation: with 4 KiB pages the TLB covers only a few megabytes, so large random-access heaps pay page walks constantly; huge pages (2 MiB, 1 GiB) multiply TLB reach but bring compaction stalls, bloat, and copy-on-write amplification. Second, locality: multi-socket servers are NUMA, so memory attached to the other socket is slower and its interconnect is shared. Third, bandwidth: many workloads, including LLM decode, are limited by bytes per second, not FLOPs or cores, so the fix is moving fewer bytes. Fourth, promises: Linux overcommits virtual memory, and when the bet fails the OOM killer picks a victim, often your largest cache. This lesson covers each wall with tradeoff tables, a TLB-reach and bandwidth calculator, an overcommit simulator, and the container and Kubernetes angle.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB c0["1 Cores 0-31 on socket 0"] m0["2 Local DRAM node 0"] c1["3 Cores 32-63 on socket 1"] m1["4 Local DRAM node 1"] c0 -->|1 Cores 0-31 on socket 0| m0 c1 -->|3 Cores 32-63 on socket 1| m1 c0 -->|5 remote -| m1 c1 -->|6 remote -| m0
What goes wrong:
- First touch. Linux allocates a page on the node of the CPU that first writes it. One initializer puts a buffer pool on node 0. Threads on node 1 read remotely forever.
- Migration. The scheduler moves a thread across sockets and leaves its memory behind. Automatic NUMA balancing migrates pages back, which costs faults and TLB shootdowns.
- One node fills. With a strict local policy a process can reclaim or OOM on one node while the other is free.
- Device locality. A GPU or NIC hangs off one socket. Feeding it from the far node cuts host-to-device bandwidth.
Strategies: one process or container per socket, pinned with numactl or a cpuset. Interleave memory that every thread hits. Keep data-loader threads on the socket nearest the GPU. Watch numastat for remote allocations.
Wall 3: bandwidth
For a loop with little compute per byte, throughput is memory bandwidth divided by bytes moved per unit of work. Extra cores do not help once the channels are saturated.
| Workload | Bytes per unit of work | Bound by | Lever |
|---|---|---|---|
| Random hash lookup | One or more cache lines, plus a page walk | Latency and the TLB | Huge pages, prefetch, batched lookups |
| Columnar scan | The column bytes | Bandwidth | Compression, column pruning, NUMA-local placement |
| Redis GET of a small value | A few cache lines | Network and syscalls more than RAM | Pipelining, I/O threads |
| LLM decode step | All weights, plus KV for the batch | HBM bandwidth | A bigger batch per weight read, quantization, grouped-query attention, prefix sharing |
| LLM prefill | Weights once for many tokens | Compute | Chunked prefill, better kernels |
ProblemCompute TLB reach and leaf page-table bytes for a 64 GiB heap at 4 KiB, 2 MiB, and 1 GiB. Then estimate a socket scan rate and a rough decode ceiling.
Expected1536 entries of 4 KiB pages cover 6 MiB. 2 MiB pages cover 512 times more and use 512 times less leaf page-table memory. Remote bandwidth at 0.6x lowers the scan rate by the same factor.
Edge cases
- 1 GiB TLB entries are tens, not thousands.
- Achievable bandwidth is below the channel peak.
- Test: 4 KiB reach is 6 MiB
reach_4k == 6 * 1024 * 1024 - Test: 2 MiB reach is 512 times larger
reach_2m == reach_4k * 512 - Test: leaf page table shrinks 512x
pte_4k == pte_2m * 512 - Test: remote scan is 0.6 of local
abs(remote_scan - local_scan * 0.6) < 1
Press Run. Snippets must be self-contained — no network, files, or native modules.
4 KiB pages cover a rounding error of a 64 GiB heap. 2 MiB pages cover gigabytes and shrink the leaf table 512 times. That is the case for huge pages. The decode estimate shows KV bytes per step growing with batch times context, so cache size caps both concurrency and tokens per second. The KV formula in its serving context is Prefill vs decode. You do not escape a bandwidth wall with more cores. You move fewer bytes, or you put the bytes closer.
Wall 4: overcommit and the OOM killer
malloc and mmap reserve address space. Frames arrive on first touch. Linux can let reservations exceed RAM plus swap.
| vm.overcommit_memory | Behavior | Pros | Cons |
|---|---|---|---|
| 0, heuristic, the default | Refuse only obviously impossible requests | Works for most software. Sparse allocations are fine | Failures arrive late, as OOM kills, not as ENOMEM |
| 1, always | Never refuse | fork of a big process works. Redis BGSAVE needs this | Highest OOM risk |
| 2, strict | Commit limit is swap plus RAM times overcommit_ratio | Fails early with ENOMEM | Breaks software that reserves big and touches little: JVMs, Go, some allocators, fork snapshots |
ProblemOn an 8 GiB box with no swap, let Redis touch 5 GiB and let an API reserve a large arena. Compare overcommit mode 1 with mode 2 at ratio 0.9.
ExpectedMode 1 admits both, then the spike kills Redis, the largest resident set. Mode 2 refuses the API reservation and leaves Redis alive. sshd is protected by a large negative oom_score_adj.
Edge cases
- A reservation does not increase used memory.
- Mode 0 still rejects a single mapping larger than RAM plus swap.
- Test: mode 1 kills redis
alwaysMode.killed.indexOf('redis') >= 0 - Test: mode 1 keeps api and sshd
alwaysMode.survivors.indexOf('api') >= 0 && alwaysMode.survivors.indexOf('sshd') >= 0 - Test: mode 2 kills nobody
strictMode.killed.length === 0 - Test: mode 2 still has redis
strictMode.survivors.indexOf('redis') >= 0 && strictMode.survivors.indexOf('api') >= 0
Press Run. Snippets must be self-contained — no network, files, or native modules.
Under always-overcommit the failure is late and lands on the largest resident process, often the cache. Under strict accounting the failure is an early allocation error the application can handle. Real kernels score RSS, swap, and page tables, then add oom_score_adj.
Containers
- A container memory limit is a cgroup limit,
memory.maxon cgroup v2. Crossing it reclaims inside the cgroup and then issues a cgroup OOM kill, even if the node has free memory. - The charge includes anonymous memory, page cache the container faulted in, kernel memory such as socket buffers, and tmpfs. A pod that reads large files can hit pressure on cache alone. It is often reclaimable.
memory.highstalls are still real. - Requests schedule the pod. Limits kill it. Pods without limits are first in line under node pressure. The Kubernetes view is requests, limits, and QoS.
- The runtime must see the limit: a container-aware JVM,
GOMEMLIMIT, Node's old-space cap, and Redis maxmemory below the limit with headroom for fork and fragmentation.
Where the walls show up
Redis wants transparent huge pages off or in madvise, overcommit_memory=1 so BGSAVE fork succeeds, maxmemory well below the cgroup limit, and one instance per NUMA node on a big box. Eviction once that cap is hit is Redis eviction.
Buffer pools want huge pages to cut TLB misses, and an interleave or per-node placement on multi-socket hosts. Do not size the pool so the OS and connections push the box into reclaim. Page layout stays on B-tree internals.
CDN and edge caches are often bandwidth-bound, not capacity-bound. The same ceiling applies to the NIC and the disk.
GPU memory is the scarce tier for KV cache. HBM bandwidth caps decode. Engines reserve a fixed fraction of GPU memory for the block pool up front, closer to hugetlbfs than to demand allocation. Tensor parallelism splits KV across devices. Host-to-GPU tiers make NUMA and PCIe locality matter again. That split is inference parallelism. Allocation profiles when the heap itself is the leak are memory profiling.
Interview Q&A
What are the tradeoffs of huge pages?
Answer
TLB reach grows 512 times for 2 MiB pages. Walks shorten, page tables shrink, and random-access workloads speed up. The costs are internal fragmentation and bloat for sparse use, compaction and khugepaged stalls with transparent huge pages, 2 MiB copy-on-write after fork, harder swap, and the reservation burden of hugetlbfs. Use them for big dense long-lived memory. Avoid transparent huge pages for fork-heavy or sparse workloads.
What is NUMA and how does it bite a backend service?
Answer
Each socket has local memory. Remote memory is slower and shares an interconnect. It bites through first-touch placement, thread migration away from the data, one node exhausting while the other is free, and devices on the far socket. Fix it with per-socket pinning, interleaving for shared structures, and locality-aware threads.
The service is CPU-light but throughput stops scaling after 16 threads on a 64-core box. Why?
Answer
Likely memory-bandwidth saturation or cross-socket traffic. The workload streams data with little compute, so more threads queue on the memory channels. Check bandwidth counters and numastat. Reduce bytes moved, improve locality, or split into per-socket instances.
Explain overcommit and the OOM killer. How do you protect a critical process?
Answer
The kernel allows reservations beyond RAM plus swap and allocates on touch. When touch-time demand exceeds supply and reclaim fails, the OOM killer picks the highest badness, mostly footprint adjusted by oom_score_adj. Protect a process with that adjustment, a cgroup with guaranteed memory, limits on noisy neighbors, and memory.high so you throttle before you kill.
Why is LLM decode memory-bandwidth bound, and what does that imply for KV cache?
Answer
Each decode step reads the weights and the KV cache of every sequence in the batch to produce one token per sequence, with little compute per byte. Tokens per second are roughly bandwidth divided by bytes per step. Batch more sequences per weight read, which needs KV capacity and is why paging avoids waste. Shrink KV with grouped-query attention or quantization. Share KV with a prefix cache. The byte formula is Prefill vs decode.
Why does Redis want overcommit mode 1?
Answer
BGSAVE fork duplicates the address space in the commit charge. Most pages are never written, so copy-on-write will not allocate them. Strict mode refuses the fork anyway. Mode 1 lets the fork succeed. You still keep maxmemory under the cgroup limit so the copies that do happen have room.
What is the difference between transparent huge pages and hugetlbfs?
Answer
Transparent huge pages are created by the kernel, in always or madvise mode, and can be split, compacted, and copied at 2 MiB. hugetlbfs pages are reserved ahead of time, mapped explicitly, and do not depend on compaction at fault time. The reservation cannot be used by anyone else.
The container limit equals the heap size and the process is still OOMKilled. Why?
Answer
The cgroup also charges page cache, thread stacks, allocator metadata, and kernel memory. Leave headroom. Tell the runtime the limit so it does not reserve a heap the cgroup cannot hold.
Pitfalls
- Enabling transparent huge pages in always mode on a Redis or other fork-heavy host.
- Benchmarking on a single-socket laptop and deploying to a two-socket server without pinning.
- Setting a container limit equal to the heap and forgetting page cache, stacks, and allocator overhead.
- Using overcommit mode 2 under software that reserves large address ranges, then chasing ENOMEM.
Take the last memory incident you saw. Say whether it was TLB reach, a remote node, a full memory channel, or a promise the kernel could not keep. If you cannot tell, the next graph to open is numastat, a bandwidth counter, or PSI, not another RAM quote.