Operating systems
Part 4 of 6 · Virtual Memory — Paging, Swapping & CachesThe Page Cache — OS Cache vs Application Cache vs Buffer Pool
Linux caches file contents in otherwise idle RAM: the page cache. Every buffered read and write, every mmap of a file, and every database that does not bypass it goes through this cache, keyed by file and offset and evicted by kernel policy. On top of it you usually stack more caches: a database buffer pool keyed by page id, an in-process LRU keyed by object, a shared Redis keyed by business identity, a CDN keyed by URL, and in LLM serving a KV cache keyed by token prefix. Each layer exists because it knows something the layer below cannot: object boundaries, transaction visibility, invalidation events, network locality, or the fact that a KV block can be recomputed. Stacking them carelessly double-caches the same bytes, wastes RAM, and lets one layer's eviction sabotage another. This lesson explains the read and write paths, mmap versus read versus direct I/O, why Postgres and InnoDB made opposite choices, why Redis is not the page cache, and how to pick the layer for each kind of data.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What does the page cache key?
Answer
A file and a page offset. Two processes reading that file share one copy.
L2
When is a write durable?
Answer
Not when write returns. When fsync, or an equivalent flush, has pushed the dirty pages and metadata to stable storage.
L3
Why do databases fear mmap?
Answer
The kernel chooses eviction and writeback, errors arrive as signals, faults block at random, and unmap shoots down the TLB. WAL order is no longer yours.
L4
What is double caching?
Answer
The same page lives in the buffer pool and in the page cache. Under a fixed RAM budget the hot set fits about half as well.
L5
How do scans get stopped from flushing the hot set?
Answer
Probation lists, midpoint insertion, a small ring for sequential scans, or frequency admission. A second touch promotes. One touch does not.
L6
Why is Redis not the page cache?
Answer
Different key, different scope, manual invalidation, and a miss that runs a query. It skips work the page cache cannot see.
L7
How is a prefix cache unlike the page cache?
Answer
The miss recomputes attention. The key is a hash of the whole prefix, so one early token change invalidates every later block. Capacity is GPU memory.
Failure modes
shared_buffers at three quarters of RAM
The page cache is squeezed, so a page evicted from Postgres is not caught by the OS tier. Checkpoints and clock sweeps get heavier too.
Redis, app, and database on one host
Redis eats the RAM the buffer pool and the page cache needed. The query path gets slower while the cache hit rate looks fine.
Dirty ratio stalls writers
write blocks inside the syscall once dirty pages pass vm.dirty_ratio. The latency looks like an application bug.
Scan flushes a plain LRU
One backup or one analytics pass touches every page once and evicts the hot set.
Misconceptions
buff/cache in free means the machine is out of memory.
That RAM is reclaimable page cache. Available memory is the number to page on.
write is durable.
The bytes are dirty in RAM until writeback or fsync.
A full table scan is harmless if you have a big cache.
Under plain LRU it is a one-touch flood. Scan resistance is a second chance, not a bigger cap.
Interviewer traps
Explaining Redis TTL jitter when the question is the page cache.
Say the page cache is coherent on local writes, then point at cache-aside for application invalidation.
Treating a prefix-cache miss as a disk read.
Name recompute on the GPU, and send radix trees and batching to their pages.
Design scenario
Same prompt for every reader.
Requirements
Explain the Postgres sizing miss, the stolen page cache, scan resistance, and why a prefix cache is not a bigger shared_buffers.
Traffic / scale
OLTP plus a sequential backup, and a separate GPU tier with repeated prefixes.
Latency
Query latency follows buffer and page-cache hits. Prefill latency follows prefix hits. They are not the same miss.
Consistency
Local file reads stay coherent. Redis and the prefix cache do not.
Availability
A container can hit memory.high on page cache it faulted in, even when the cache is reclaimable.
Failure assumptions
- Buffered I/O is on.
- The backup is a one-shot scan.
- Prompts differ in an early token.
Constraints
- Do not call write durable.
- Do not move the stampede design onto this page.
Prompt
A 64 GiB database host has shared_buffers at 48 GiB and got slower. Redis on the same box has a high hit rate. A nightly backup makes the hot queries miss for minutes. An LLM router wants to reuse long system prompts.
API
Which call makes the write durable?
Data
What fraction of RAM is a typical shared_buffers starting point?
Architecture
Which layer owns file bytes, query results, and KV prefixes?
Who should hold the hot page
Prefer
One inclusive cache for the bytes, another only if it skips work
The page cache or the buffer pool holds file pages. Redis holds the result of a query the kernel cannot see. Do not pay RAM twice for the same page.
- Postgres keeps shared_buffers modest and lets the page cache be the second tier.
- InnoDB often uses direct I/O and owns the pool.
- A probation list keeps a one-shot scan from becoming the active set.
Alternative
Split a fixed RAM budget across two LRUs of the same pages
The hot page is stored twice. Disk traffic does not fall, and the database's cache starves so Redis can look fast.
- effective cache capacity drops.
- A backup scan under plain LRU evicts the working set.
- write is treated as durable and a crash loses the commit.
Pick the layer by the key you actually reuse
File bytes, query results, HTTP responses, and attention state are different caches.
- 1
If the reuse is raw file or page bytes on this host
Leave it in the page cache or the buffer pool. Do not copy it into a second inclusive LRU. - 2
If the reuse is a computed result across app servers
A shared cache such as Redis can skip the query. You now own TTL and invalidation. - 3
If it is a public HTTP response or a repeated token prefix
A CDN caches the response. An engine prefix cache reuses KV blocks. A miss on the second one recomputes. - 4
If none of those are true
Do not cache. Make the query or the compute cheaper.
Read path and write path
Sequence
- 1
App → Kernel VFS
1 read fd, offset, length
- 2
Kernel VFS → Page cache
2 lookup file and page index
- 3
Page cache → App
3a copy bytes to the user buffer
- 4
Page cache → Disk
4a read the page plus readahead
- 5
Disk → Page cache
4b fill pages
- 6
Page cache → App
4c copy bytes to the user buffer
- 7
App → Kernel VFS
5 write fd, offset, data
- 8
Kernel VFS → Page cache
6 copy into the page and mark it dirty
- 9
Page cache
7 flusher threads write dirty pages later
- 10
App → Kernel VFS
8 fsync fd
- 11
Kernel VFS → Disk
9 force dirty pages and metadata out
- 12
Disk → App
10 durable only now
Lesson map
The Page Cache — OS Cache vs Application Cache vs Buffer Pool
Linux caches file contents in otherwise idle RAM: the page cache. Every buffered read and write, every mmap of a file, and every database that does not bypass it goes through this cache, keyed by file and offset and evicted by kernel policy. On top of it you usually stack more caches: a database buffer pool keyed by page id, an in-process LRU keyed by object, a shared Redis keyed by business identity, a CDN keyed by URL, and in LLM serving a KV cache keyed by token prefix. Each layer exists because it knows something the layer below cannot: object boundaries, transaction visibility, invalidation events, network locality, or the fact that a KV block can be recomputed. Stacking them carelessly double-caches the same bytes, wastes RAM, and lets one layer's eviction sabotage another. This lesson explains the read and write paths, mmap versus read versus direct I/O, why Postgres and InnoDB made opposite choices, why Redis is not the page cache, and how to pick the layer for each kind of data.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB app["App"] vfs["Kernel VFS"] pc["Page cache"] disk["Disk"] app -->|1 read fd,| vfs vfs -->|2 lookup file| pc pc -->|3a copy bytes to| app pc -->|4a read the page| disk disk -->|4b fill pages| pc pc -->|4c copy bytes to| app
writereturning is not durability. Data sits dirty in RAM until writeback or fsync. Databases fsync the WAL on commit for this reason. Group commit and the latency of that flush are fsync and group commit.- Readahead detects sequential reads and fetches ahead. Scans are fast. Random reads on cold data are not.
- If dirty pages exceed
vm.dirty_ratio, writers stall inside the write call. - The page cache is shared. Two processes reading the same file share one copy, so a weight file mmapped by several workers costs RAM once.
read, mmap, and direct I/O
| Method | How it works | Pros | Cons | Who uses it |
|---|---|---|---|---|
| Buffered read and write | Copy between the page cache and a user buffer | Simple, readahead, kernel caching | An extra copy. Double caching if the app also caches | Postgres, most services |
| mmap of a file | Map page-cache pages into the address space | No copy. Lazy. Shared across processes | Faults hide I/O. Little control of eviction or write order. Errors as signals. TLB shootdowns on unmap | LMDB, model weights, an older MongoDB engine |
| Direct I/O | Bypass the page cache. DMA into aligned buffers | No double caching. The app schedules I/O | The app must cache, readahead, and align | InnoDB in a typical config, ScyllaDB, many engines |
| io_uring, with or without direct I/O | Async submission and completion rings | High IOPS, fewer syscalls | Complexity. Still needs a caching strategy | Newer storage engines |
The CMU argument against mmap in a database is the canonical one: you lose control of when pages are evicted and written, which breaks transactional guarantees and makes latency unpredictable. mmap is a good fit for read-mostly immutable data.
Double caching
If a database reads with buffered I/O and also keeps a buffer pool, a hot page can live twice. Under a fixed RAM budget that halves the room for distinct hot pages.
| Engine | Strategy | Consequence |
|---|---|---|
| PostgreSQL | Buffered I/O. shared_buffers often around 25 percent of RAM. The page cache is the second tier | Some duplication is accepted. effective_cache_size tells the planner about the OS cache |
| MySQL InnoDB | Direct I/O is the typical flush method. The buffer pool is often half to 80 percent of RAM | No duplication. InnoDB owns the policy |
| Redis | All data in its heap. Persistence files go through the page cache | The page cache matters for RDB and AOF, not for keys |
| Kafka | Little in-process caching. The page cache holds the log | Recent reads are served from RAM. sendfile can be zero-copy |
| Elasticsearch and Lucene | mmap index files. Leave about half of RAM to the OS | The JVM heap stays small so the page cache can hold segments |
On a 64 GiB host, shared_buffers at 48 GiB squeezes the page cache. Pages evicted from the buffer pool are no longer caught by the OS tier, and the duplicated portion wastes memory. A huge shared_buffers also makes checkpoints and clock-sweep scans heavier. A common starting point is around 25 percent, with effective_cache_size reflecting the rest.
ProblemSplit 1000 page slots between an application LRU and a page-cache LRU. Run a Zipf trace. Compare disk rate and how many keys sit in both caches.
ExpectedAn empty application cache duplicates nothing. A 50-50 split duplicates hot keys. A larger application cache raises its hit rate and does not magically erase disk reads of the same bytes.
Edge cases
- A zero-capacity application cache never puts.
- The intersection counts keys, not bytes of distinct files.
- Test: no app cache means no app hits and no duplicates
rows[0][0] == 0 and rows[0][3] == 0 - Test: a split stores keys twice
rows[500][3] > 0 - Test: a larger app cache hits more often
rows[950][0] >= rows[250][0] - Test: page cache still sees misses the app missed
rows[0][1] > 0
Press Run. Snippets must be self-contained — no network, files, or native modules.
The shape is the lesson in the source experiment, which used 20,000 keys and 60,000 accesses on a 1,000-page budget. This run uses a smaller trace so the sandbox finishes. Splitting a fixed budget across two inclusive LRUs stores hot items twice. Disk rate stays roughly flat or gets worse compared with putting the budget in one layer. The application cache still wins on CPU when a hit skips syscalls, parsing, and query execution. Its RAM should not come out of a starved page cache or buffer pool on the same box.
Scan resistance
A backup, an analytics query, or a log grep can read gigabytes once. Under plain LRU those pages push out the hot set.
- Linux puts new pages on the inactive list. A second access promotes them. MGLRU generalizes this into generations.
- InnoDB uses midpoint insertion. A new page lands in the old sublist and must survive
innodb_old_blocks_timebefore promotion. - PostgreSQL uses a small ring buffer for large sequential scans so they do not flood shared_buffers. Replacement is a clock sweep with usage counts.
- Redis samples LRU, or uses LFU. allkeys-lfu resists a one-off scan because one touch does not make a key popular. The policy page is Redis eviction.
- Caffeine and similar caches use a W-TinyLFU admission filter.
ProblemWarm 60 hot pages, scan 5000 cold pages once, then touch the hot set again. Compare a plain LRU with a two-list cache on a 100-slot budget.
ExpectedThe two-list hit rate on the hot set after the scan is higher than plain LRU. The scan churns the inactive list.
Edge cases
- The score ignores hits during the warmup and the scan.
- Demotion moves a cold active page to inactive instead of dropping it.
- Test: two-list beats plain LRU after the scan
rates['two-list'] > rates['plain LRU'] - Test: plain LRU lost most of the hot set
rates['plain LRU'] < 0.5 - Test: two-list kept a majority of the hot set
rates['two-list'] > 0.5
Press Run. Snippets must be self-contained — no network, files, or native modules.
Why Redis is not the page cache
| Question | Page cache | Redis |
|---|---|---|
| What is cached | File byte ranges | Application objects |
| Key | Inode and page index | A business key you choose |
| Scope | One host, one kernel | Shared by app servers over the network |
| Invalidation | Automatic on a file write | Your job: TTL, events, versioned keys |
| Eviction | Kernel policy across all files | maxmemory-policy on the instance |
| Miss cost | A disk read | A query, serialization, and the network |
| Survives process restart | Yes, until reclaimed | Only with persistence. Usually rebuilt |
| Consistency risk | None for local readers of the file | Stale reads, races between fill and invalidate |
The page cache cannot cache the result of a query. A join that filters and serializes JSON touches many pages. The page cache can make each page read fast. Only an application cache skips the query. Redis cannot make the database's own index reads faster. They are complementary. The classic mistake is to run both on one host and starve the database. Fill and invalidation mechanics stay on cache-aside. Coherence across a near cache is distributed caching consistency.
Choosing the layer
Decisions
- 1
1 What is being reused
- next2 Raw file bytes on this host
- ?
2 Raw file bytes on this host
- 2a yes3 Page cache or buffer pool
- 2b no4 Computed result across app servers
- 3
3 Page cache or buffer pool
- ?
4 Computed result across app servers
- 4a yes5 Shared cache - plan TTL and invalidation
- 4b no6 Hot tiny set, seconds of staleness ok
- 5
5 Shared cache - plan TTL and invalidation
- ?
6 Hot tiny set, seconds of staleness ok
- 6a yes7 In-process LRU or near cache
- 6b no8 Public cacheable HTTP response
- 7
7 In-process LRU or near cache
- ?
8 Public cacheable HTTP response
- 8a yes9 CDN with cache-control and purge
- 8b no10 Attention state for a repeated prefix
- 9
9 CDN with cache-control and purge
- ?
10 Attention state for a repeated prefix
- 10a yes11 Engine prefix cache
- 10b no12 Do not cache
- 11
11 Engine prefix cache
- 12
12 Do not cache
The coherence the page cache gets for free, one copy per host invalidated on write, is what a near cache rebuilds with events and versions. A CDN is a page cache for HTTP, keyed by URL, with prefetch as a cousin of readahead and shielding in front of the origin: CDN cache hierarchy.
A buffer pool holds pages that may contain many row versions. A long transaction keeps old versions alive, so the same hot table occupies more pages and pushes others out. MVCC bloat is a cache-capacity problem. See MVCC and snapshot isolation and VACUUM and bloat. The page itself, the split, and the pin stay on B-tree internals.
A prefix cache is a page cache for attention state, keyed by a hash of the token prefix instead of a file offset. There is no file to reread, so a miss recomputes prefill, and the cache can be tiered from GPU memory to CPU RAM to a remote store. Agent-level prompt and semantic caches are caching for agents. The block allocator is PagedAttention. This page stops at the layer choice.
Interview Q&A
Why might a database prefer direct I/O?
Answer
To avoid double caching and to control eviction, write ordering, and I/O scheduling. The database knows which pages are index roots, which belong to a scan, and which are dirty with WAL dependencies. The kernel does not. The cost is implementing caching, readahead, and alignment in the engine.
Why is mmap risky for a database?
Answer
The kernel decides when to evict and when to write back dirty pages, which can violate WAL ordering. I/O errors arrive as signals. Page faults are invisible blocking I/O on any thread. Unmapping triggers TLB shootdowns. It is a reasonable fit for read-mostly immutable data such as LMDB pages or model weights.
Explain why Redis is not the page cache.
Answer
Different keys, business objects versus file offsets. Different scope, a cluster versus one host. Different invalidation, manual versus automatic. Different miss cost, a query versus a disk read. The page cache speeds storage reads. Redis skips computation and network hops. They stack. Stampede control is cache-aside.
The box has 64 GiB, Postgres shared_buffers is 48 GiB, and performance got worse. Why?
Answer
The page cache was squeezed to a few GiB, so pages evicted from shared_buffers are no longer caught by the OS tier, and the duplicated portion wastes memory. A huge shared_buffers also makes checkpoints and clock-sweep scans heavier. A typical starting point is around 25 percent, with effective_cache_size reflecting the rest.
How does a full table scan hurt a cache, and how do systems defend?
Answer
It touches every page once and, under plain LRU, evicts the hot set. Defenses are probation lists, the Linux inactive list, InnoDB midpoint insertion, Postgres ring buffers for scans, and frequency admission such as LFU or TinyLFU.
How does an LLM prefix cache differ from the OS page cache?
Answer
Both keep reusable state keyed by content. The backing store of a prefix is computation. The unit is a KV block of tokens across layers. An entry must match the exact prefix, because a different earlier token invalidates everything after it. Capacity lives in scarce GPU memory, so eviction changes batch size and throughput. Radix structure and batching stay on the engine pages, including PagedAttention.
Why does Kafka lean on the page cache?
Answer
Consumers of recent data are served from RAM, and the broker keeps little of the log in its own heap. sendfile can hand pages to the socket without a copy into user space. That only works if the page cache is allowed to hold the hot log.
A file-heavy pod hits its memory limit while anonymous RSS looks fine. What happened?
Answer
The cgroup is charged for page cache it faulted in, plus kernel memory such as socket buffers. The cache is usually reclaimable, but memory.high throttling and stalls are real. Do not read buff/cache as a leak, and do not ignore the cgroup charge.
Pitfalls
- Reading buff/cache as used memory and buying RAM you do not need.
- Running Redis, the application, and the database on one host and letting Redis eat the database cache.
- Calling write durable. fsync, O_DSYNC, or a correctly configured battery-backed controller is the durability point.
- Forgetting that cgroups charge page cache to the container that faulted it in.
For one piece of data you cached last week, say the key, the backing store, who evicts it, and the cost of a miss. If you cannot, it is probably the wrong layer.