Performance
Part 3 of 6 · Performance EngineeringFlame Graphs & Stack Collapse — Reading Hot Paths
Width = sample share, not one slow call; how to read plateaus and avoid common misreads.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
From samples to a picture
You cannot read a flame graph until the stacks have been merged. The merge is the whole trick.
- 1
Sample under load
The load shape has to match the incident. An idle box collapses to a runtime poll. - 2
Symbolicate
JIT, DWARF, or sourcemaps. A hex address is not a function you can fix. - 3
Collapse
Identical stacks become one row with a count. That count is the width. - 4
Draw rectangles
Width tracks samples. The y-axis is stack depth. The x-axis groups names. It is not a timeline. - 5
Read and re-profile
A plateau is a hypothesis. Change the code, then sample the same way again.
Where did the time go?
Prefer
Flame graph
Merged stacks. One glance shows which call chain owns the samples.
- Language-agnostic once you have stacks.
- Wide means often, or slow, or both. You still check the caller.
- Weak across service hops. That is a trace waterfall.
Alternative
Trace waterfall
Spans in time order across hops. The critical path is visible. The inside of a hot function is not.
- Right for fan-out, queues, and downstream waits.
- A top-N table gives self time and cumulative time with less hierarchy.
- Use both. The waterfall picks the process. The flame graph picks the function.
Overview
A flame graph is a picture of collapsed stacks. Identical call chains merge. Each merged chain is a rectangle whose width is proportional to the sample count, which is the relative time spent on that stack. It is not a timeline. It is not a call graph whose edges are the only weights that matter.
The y-axis is stack depth. The x-axis is alphabetical, or otherwise sorted by name among siblings, so identical frames sit together. Color is usually cosmetic. Search and zoom in Speedscope, the Firefox Profiler, or Brendan Gregg's SVG are how you land in one package.
How you collected the samples, on-CPU versus wall-clock, is CPU & Wall-Clock Profiling. A wide GC or allocator frame is a reason to open Memory Profiling, not a reason to micro-optimize the collector's leaf.
Flow
- 1
1. Sample stacks under load
- next2. Symbolicate the frames
- 2
2. Symbolicate the frames
- next3. Collapse identical stacks
- 3
3. Collapse identical stacks
- next4. Width is sample share
- 4
4. Width is sample share
- next5. Read plateaus and re-profile
- 5
5. Read plateaus and re-profile
Lesson map
Flame Graphs & Stack Collapse — Reading Hot Paths
Width = sample share, not one slow call; how to read plateaus and avoid common misreads.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Sample stacks under load"] b["2. Symbolicate the frames"] c["3. Collapse identical stacks"] d["4. Width is sample share"] a -->|1. Sample stacks under load| b b -->|2. Symbolicate the frames| c c -->|3. Collapse identical stacks| d
Folded stacks
The file format is one line per merged stack: frames joined by semicolons, then a count.
main;handler;parse;json_decode 42
main;handler;db;query 10
main;handler;db;wait 30json_decode owns more samples than query. On a wall-clock graph, wait can dominate even when the CPU graph looks like the parser. The classic pipeline is perf script, then stackcollapse-perf.pl, then flamegraph.pl. Inferno and Speedscope are the modern equivalents. The sandboxes below collapse the same three stacks so the counts are obvious before any SVG exists.
| Shape | How to read it |
|---|---|
| Wide plateau | Many samples. First optimization candidate |
| Tall thin tower | Deep stack, few samples. Usually the wrong place to start |
| Plateau under epoll, park, or IO | Wait time. Fix IO, pooling, or fan-out |
| Many small siblings | Scattered work. Look at the algorithm or at batching |
On-CPU, off-CPU, and icicles
- An on-CPU flame graph shows where cores run. It is the right picture for a compute burn.
- An off-CPU or wakeup graph shows where threads block and who wakes them. Lock contention and IO live here. Brendan Gregg's off-CPU analysis is the reference. An on-CPU graph often misses the lock entirely.
- Some tools have an explicit wall-clock mode:
py-spy --idle, async-profiler wall mode.
An icicle chart is the same data drawn downward from the root. The preference is UI. The width rule does not change.
Misreads
- Width is the sample share of that stack. It is not the duration of a single invocation. A fast function called millions of times is wide.
- Optimizing a leaf such as
strcmpormemcpywhile the caller does a quadratic number of comparisons just makes the same mistake faster. - Wide GC or allocator frames mean memory pressure. Pair them with a heap or allocation profile.
- The x-axis is not chronological. A trace waterfall is the timeline. Two frames side by side did not happen "left then right" in time.
- Two graphs from different load shapes, CPU counts, inlining decisions, library versions, JIT warmup, or sample durations do not have comparable absolute widths. Compare fractions, or compare profiles taken the same way.
- Inlining and tail calls delete frames. A missing function can mean the compiler was good at its job.
Interview soundbite: "That bar is wide, so that function is slow" is the mistake. The bar is wide, so that stack was on the samples. Ask whether it is slow once or cheap and frequent, and ask who called it.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
json_decode is the widest on-CPU-style stack in this toy. epoll_wait is the one that would own a wall-clock graph of the same handler. Which file you collapsed decides which fix you propose.
Interview Q&A
What does width mean in a flame graph?
Answer
Relative sample count for that stack, which is approximately the fraction of time on that stack. It is not one call's latency. A 2 microsecond function called on every request can be wider than a 2 millisecond function called once.
Why is the x-axis alphabetical?
Answer
So identical names merge into stable plateaus. Chronological order is what a trace is for. Reading left to right as a timeline will invent a sequence the profiler never claimed.
How do you get from perf to a flame graph?
Answer
perf script, then stackcollapse-perf.pl, then flamegraph.pl. Inferno and Speedscope do the same job with a less fragile pipeline. If symbols are missing, stop and symbolicate before you rank frames.
The widest frames are GC or the allocator. What next?
Answer
Leave the CPU micro-opts. Open an allocation profile or a heap diff. Reduce allocations, shrink objects, or shorten lifetimes. The collector is wide because the program is handing it garbage. That is the memory page.
Can a flame graph show lock contention?
Answer
On-CPU often cannot. Off-CPU, wall-clock, and wakeup graphs can. If you only have an on-CPU capture of a lock-bound service, take another capture in the mode that includes blocked time.
Why do two environments draw different graphs for the same code?
Answer
Different load, different CPU count, different inlining, different library versions, JIT warmup, or a different sample duration. Compare percentages from profiles taken under the same shape, or treat the difference as a question about the environment.
Flame graph versus icicle chart?
Answer
Same merged stacks. Icicles grow downward from the roots. Pick the UI your team can read. Do not change the width rule when the picture flips.
A leaf is wide. Do you rewrite the leaf?
Answer
Read the caller first. A wide memcpy, strcmp, or hash often means the parent is copying or comparing too much. Fix the call pattern, then re-profile. Rewriting the leaf in assembly is the reward you give yourself for skipping the parent.
When do you stop staring at the flame graph?
Answer
When the plateau is outside this process: a downstream, a queue, or a fan-out. Switch to the trace and the latency budget. Come back when a single hop is the one burning samples.
Pitfalls
- Saying "this function is slow" because its bar is wide.
- Optimizing a libc leaf and leaving a quadratic caller in place.
- Treating x position as time.
- Comparing two SVGs from different load tests as if the pixels were milliseconds.
- Ignoring inlined frames and then claiming the function "does not show up, so it is free."
- A CPU flame graph full of GC, and a sprint spent on the parser.
You have folded counts: json_decode 42, query 10, epoll_wait 30. Say which bar you open on an on-CPU graph, which bar you open on a wall-clock graph, and what you would be wrong to conclude about a single call's duration.