Performance
Part 4 of 6 · Performance EngineeringMemory Profiling — Allocations, Leaks & GC Pressure
Allocation rate vs retained heap; leaks vs caches; GC pressure on p99.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
RSS is climbing or p99 tracks GC
Prefer
Name the question first
Allocation rate and retained set are different measurements. The symptom picks which one you take.
- A high rate means you are creating garbage. Sample the allocation sites.
- A growing retained set means something still points at the objects. Diff a heap.
- Re-measure p99 and RSS under the same load after the fix.
Alternative
Turn the GC off, or take a prod dump first
Disabling collection hides the pause and then you OOM. A production heap dump is a pause, a privacy review, and a last resort.
- Continuous allocation samples are the always-on signal.
- Staging diffs catch most retainers before you dump production.
- A dump still cannot tell you the rate. Take the profile that matches the question.
From symptom to a second measurement
RSS, a GC pause, and an OOM are the alarm. They are not the diagnosis.
- 1
Symptom
RSS rising, GC pause on the tail, or an OOM kill. - 2
Split the question
Is the allocation rate high, or is the retained set growing under steady traffic? - 3
Capture
Rate: allocation sampling or tracemalloc. Retained: a heap snapshot diff across time. - 4
Fix
Fewer, smaller, or pooled objects. Or drop the retainer, bound the cache, and break the cycle. - 5
Re-measure
p99 and RSS, same load shape. A memory win that slows the CPU path is a new profile, not a victory lap.
Overview
Memory problems show up as rising RSS, OOM kills, or GC pauses that inflate p99 while the CPU flame graph looks like a runtime. You need two ideas that people collapse into "we have a leak."
- Allocation rate is how fast you create objects. That is garbage, and it is GC pressure.
- Retained heap is bytes still reachable from GC roots.
- A leak, in a managed runtime, is a retained set that grows without bound under steady traffic. It is usually an unintended retainer: an unbounded cache, a global list, a closure, a listener, a detached DOM node. It is rarely a forgotten
malloc. - A cache is intentional retention. It is still a leak when the key includes a request id, when there is no TTL or LRU, or when the bound is larger than the machine.
The cluster map is the hub. How a wide bar gets drawn is Flame Graphs. How that pause lands on a user-facing percentile is Latency Budgets.
Flow
- 1
1. RSS up, GC pause, or OOM
- next2. High rate or growing set
- 2
2. High rate or growing set
- next3. Sample allocs or diff heap
- 3
3. Sample allocs or diff heap
- next4. Cut objects or drop retainer
- 4
4. Cut objects or drop retainer
- next5. Re-measure p99 and RSS
- 5
5. Re-measure p99 and RSS
Lesson map
Memory Profiling — Allocations, Leaks & GC Pressure
Allocation rate vs retained heap; leaks vs caches; GC pressure on p99.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. RSS up, GC pause, or OOM"] b["2. High rate or growing set"] c["3. Sample allocs or diff heap"] d["4. Cut objects or drop retainer"] a -->|1. RSS up, GC pause, or OOM| b b -->|2. High rate or growing set| c c -->|3. Sample allocs or diff heap| d
GC pressure
Python. Reference counting plus a cyclic collector. Huge object churn shows up as pauses in generational collections. gc.get_stats and tracemalloc tell you about the Python heap. Native extensions allocate outside it, so a growing RSS with a flat tracemalloc graph is a clue to look at C.
V8 and Node. Young and old spaces. Allocation spikes become scavenges and major GC. Chrome DevTools heap snapshots and allocation timelines are the interactive tools. --inspect in production is a deliberate choice, with a locked-down port.
Why p99 cares. Collection is stop-the-world, or close to it. It sits on the critical path of whichever request was in that process when the collector ran. GC burns CPU during the collection and stalls request threads, so utilization and p99 rise together. "Disable GC" removes the pause and keeps the garbage. The fix is less garbage, or a shorter lifetime, or a bounded cache.
Tools
| Runtime | Allocation sites | Retained set | Heavy, last |
|---|---|---|---|
| Python | tracemalloc, memray, filprofiler | objgraph, heapy, snapshot diff | A full dump on a big worker |
| Node / Chrome | allocation sampling, Clinic HeapProfiler | heap snapshot, retainer path | An unbounded prod inspect |
| Native | jemalloc profiling, heaptrack | same tools, live allocs | Valgrind massif |
| JVM | async-profiler alloc mode, JFR | MAT on a dump | A heap dump as the first step |
| Method | Use it when | Cost |
|---|---|---|
| Heap dump or snapshot | You need the exact retainer path | Pause, size, privacy. Poor as an always-on signal |
| Allocation sampling | You need hot allocation sites, including in production | Statistical. A rare retainer can hide |
Leaks that are caches
RSS that rises and then flatlines can be a cache warming to its cap. A leak keeps growing under constant traffic. Diff heap snapshots hours apart, with the same request mix, before you call it a leak.
Caches become leaks when keys are unbounded, when nothing evicts, or when the key contains a unique request id. A Map from id to payload retains every value for the life of the map. A WeakRef does not keep the value alive by itself. The sandbox shows the difference in structure. A real browser leak still wants a heap snapshot, because a WeakRef whose key is strongly held is not a fix.
A detached DOM node is the frontend version: the node is gone from the tree and still referenced by a listener. The retainer path in a snapshot names the listener. Guessing "React leaked" is not the path.
A hot json.loads in an allocation sample means you are parsing too often, parsing too much, or copying into too many objects. Parse less, use smaller payloads, reuse buffers, or use a parser that allocates less. The collector is downstream of that choice.
If a memory fix makes CPU worse, the pooling or the arena got more expensive than the GC it replaced. Re-profile both axes. The CPU page is the second measurement.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
The Python cache holds 200 chunks of 10,000 bytes. tracemalloc attributes the growth to the line that appended them. That is an allocation-rate finding and a retainer finding at once, because the list is the root. Drop the list, or bound it, and the next diff should be flat.
Interview Q&A
RSS grows and then flatlines. Is it a leak?
Answer
It may be a cache warming to its capacity. A leak keeps growing under steady traffic. Diff heap snapshots hours apart at constant load. A plateau at the configured max is a capacity choice. A line that never bends is a retainer.
Allocation sampling is hot in json.loads. What do you change?
Answer
Reduce how often you parse, shrink the payload, reuse buffers, or switch to a parser that allocates less. Disabling the GC keeps the garbage and then you run out of memory. The sample already named the site.
Why can GC raise CPU and latency together?
Answer
The collector burns CPU while it runs, and it stalls the threads that were serving requests. Utilization and p99 move in the same window. A CPU profile full of GC frames and a latency histogram with pauses are the same incident.
How is a Python leak different from a C leak?
Answer
Python: an unexpected retainer. The allocator will reuse the memory once the reference is gone. C or any native extension: a true unfreed malloc. heaptrack, ASan, and jemalloc stats are the native tools. tracemalloc will not see bytes it does not own.
When do you take a heap dump in production?
Answer
After continuous allocation profiles and a staging diff have failed, and after you have a privacy story for the contents of the heap. A dump pauses the process and copies retained data. It is a last resort, not the daily signal.
How do caches become leaks?
Answer
Unbounded keys, no TTL or LRU, or a key that includes a unique request id so every call is a new entry. Bound the size, expire entries, and keep identifiers that repeat out of the key.
What is a detached DOM or closure leak?
Answer
A node left the document and a JavaScript listener still points at it, or a closure captured a large object the caller thinks it dropped. The heap snapshot retainer path names the edge. The fix is to drop that edge, not to restart the tab on a timer.
The memory fix made CPU worse. What happened?
Answer
Pooling, reuse, or a large arena cost more than the allocations it saved. Re-profile CPU and allocation rate. A win on RSS with a new hot path in the pool is a trade you have to price, not a free fix.
Which profile do you take for p99 pauses versus a slow OOM?
Answer
Pauses with a stable RSS point at allocation rate and GC. A slow climb to OOM points at the retained set. Take both if you are unsure, and let the one that moves across an hour decide.
Pitfalls
- Calling every RSS increase a leak, including a cache that has reached its cap.
- Disabling GC to "fix" pauses.
- A production heap dump as the first tool, with customer payloads in the file.
- Optimizing a wide GC frame on a flame graph without an allocation site.
- An LRU keyed by request id.
- A pool so expensive that the CPU profile got worse and nobody re-measured.
Steady traffic. RSS climbs for six hours. A CPU flame graph is mostly the collector. Say which capture you take next, what retainer pattern you expect if the keys are request ids, and which two numbers you compare after the bound ships.