Performance
Part 1 of 6 · Performance EngineeringPerformance Engineering — Profiling, Flame Graphs, Latency Budgets & Hot Paths
Measure→hot-path→fix→re-measure loop; latency vs throughput vs capacity; when to profile CPU vs memory vs wait/IO.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
A user-facing API is missing its latency SLO
Prefer
Latency-first
Name the percentile, find the critical path, and spend the fix on the hop that owns the tail.
- p99 and p999 are the numbers users feel.
- A flame graph or a trace says which hop to open.
- Headroom can stay unused. That is the cost of the SLO.
Alternative
Throughput-first
Batching and bigger chunks raise completed units per second and often stretch the tail.
- Right for pipelines whose SLO is jobs per hour.
- A higher mean throughput can still fail a p99 UX SLO.
- Capacity you add on top of the wrong hot path mostly buys coordination.
The performance loop
Evidence comes before a rewrite. A fix that you do not re-measure just moves the bottleneck.
- 1
Observe
Latency, errors, and saturation. Name the metric and the percentile. - 2
Hypothesize
CPU, memory, wait or IO, or coordination across hops. - 3
Profile
Sampling or instrumentation on that axis, under a load shape that matches production. - 4
Visualize
Flame graph, heap diff, or a per-hop latency budget. - 5
Fix and re-measure
Algorithm, allocation, concurrency, or the budget. If the SLO holds, write the guardrail. If a new bottleneck appears, start again.
Overview
"It's slow, rewrite it" is how teams ship feel-good diffs that miss the SLO. Performance engineering is a closed loop: measure, find the hot path, fix, re-measure. Before you propose a change, name the metric, the percentile, the environment, and the evidence (a profile, a trace, or a load curve).
Latency, throughput, and capacity are related. They are not interchangeable. Optimizing the wrong one raises a dashboard and leaves the user-facing tail where it was.
This page is the map. Sampling versus instrumentation, on-CPU versus wall-clock, and the GIL / event-loop traps are CPU & Wall-Clock Profiling. How collapsed stacks become plateaus is Flame Graphs. Allocation rate versus retained heap is Memory Profiling. Per-hop milliseconds and fan-out tails are Latency Budgets. Lab wins that vanish in production are Microbenchmarks vs Load Tests.
Latency, throughput, and capacity
| Goal | Question it answers | Failure mode |
|---|---|---|
| Latency | How long does one unit of work take? | Tuning the mean while p99 stays bad |
| Throughput | How many units finish per second? | Batching that blows up the tail |
| Capacity | What load is sustainable before saturation? | Buying machines for a hot path you never found |
- Latency is time for one unit of work. Care about p50, p99, and p999. The mean is a poor SLO for a user-facing API. Histograms, and why averages hide the tail, are histogram vs average latency.
- Throughput is completed units per second. High throughput with a terrible p99 still fails a UX SLO.
- Capacity is the max sustainable load before a resource saturates: CPU, memory, connections, disk, or a downstream quota.
A latency-first change protects the SLO and can leave headroom unused. A throughput-first change fits batch and pipelines and can mask tail latency. A capacity-first change is cost planning, and it is meaningless until a latency budget says what "sustainable" means.
Flow
- 1
1. Observe latency and saturation
- next2. Hypothesize the axis
- 2
2. Hypothesize the axis
- next3. Profile CPU, wall, or memory
- 3
3. Profile CPU, wall, or memory
- next4. Visualize flame, heap, budget
- 4
4. Visualize flame, heap, budget
- next5. Fix the hot path
- 5
5. Fix the hot path
- next6. Re-measure matching load
- 6
6. Re-measure matching load
- next7. Document budget and guardrail
- 7
7. Document budget and guardrail
Lesson map
Performance Engineering — Profiling, Flame Graphs, Latency Budgets & Hot Paths
Measure→hot-path→fix→re-measure loop; latency vs throughput vs capacity; when to profile CPU vs memory vs wait/IO.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Observe latency and saturation"] b["2. Hypothesize the axis"] c["3. Profile CPU, wall, or memory"] d["4. Visualize flame, heap, budget"] a -->|1. Observe latency and saturation| b b -->|2. Hypothesize the axis| c c -->|3. Profile CPU, wall, or memory| d
Interview soundbite: never jump from "it's slow" to "rewrite it in another language." Name the metric, the percentile, the environment, and the evidence.
Which profiler axis
| What you see | First tool | Then |
|---|---|---|
| CPU hot, utilization high | Sampling CPU profile and a flame graph | Confirm it is on-CPU, not a blocked poll |
| Wall slow, CPU low | Wall-clock or off-CPU profile, plus traces | The time is in wait, locks, or a downstream |
| Memory or GC churn | Allocation sampling or a heap snapshot | Retained set versus allocation rate |
| Coordination across services | Critical-path budget and span traces | The slow hop, then profile inside it |
| Unknown | RED and USE first | Then pick the axis |
RED and USE tell you that something is wrong and where the signal sits. Profiling tells you why inside a process. Start from the observability triad and exemplars, RED, and USE when you do not yet know the axis. A flame graph of the wrong service is a confident answer to the wrong question.
Sampling and instrumentation, in one pass
- Sampling (perf, py-spy, async-profiler, pprof) interrupts on a timer and records stacks. Overhead stays low enough for production. Very short functions can be missed. The full comparison is the CPU page.
- Instrumentation (cProfile, APM spans, OpenTelemetry) records exact call counts and spans. Use it surgically. The probe can move the bottleneck.
What the rest of the cluster adds
- CPU & Wall-Clock Profiling — on-CPU versus wall-clock, and why a Python GIL or a Node event loop makes "CPU%" a trap.
- Flame Graphs & Stack Collapse — width is sample share. A wide plateau is many samples on that stack, including a fast function called millions of times.
- Memory Profiling — allocation rate, retained heap, leaks versus caches, and GC pauses on the p99.
- Latency Budgets & Critical Path — milliseconds per hop. Sequential work sums. Parallel fan-out waits on the slowest child.
- Microbenchmarks vs Load Tests — dead-code elimination and warmup in the lab; coordinated omission and closed-loop load in the system test.
Budgets also set timeout headroom. A timeout with no budget is how retries pile up. The mechanics of deadlines live in timeouts, budgets, and deadline propagation. This cluster owns the millisecond allocation. That page owns propagation and retry storms.
Measure both axes
A CPU burn and a sleep can share a wall-clock number and still be different problems. The burn occupies a core. The sleep is off-CPU. The sandboxes below time both so the split is visible before you pick a profiler.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
The TypeScript runner returns before timers fire, so the wait is the 50 ms you would have awaited. The CPU loop is real. On Node, that same loop blocks the event loop and shows up in p99 even when average CPU% looks calm. That trap is the next page.
Interview Q&A
How do you start when a service is slow?
Answer
Define the SLO metric and the percentile. Check recent deploys and config. Look at saturation. Then choose an axis: a CPU sample, a wall-clock or off-CPU profile, memory, or a distributed trace. Measure before you rewrite. "Rewrite it" is a hypothesis, and it needs a profile.
Latency improved and users still complain. Why?
Answer
You optimized the mean or the p50, and p99 or p999 still hurts. Or you optimized a hop that is off the critical path, so the response time never moved. Ask which percentile the SLO uses, and which hop owns that percentile.
Sampling versus instrumentation, in one sentence?
Answer
Sampling estimates where time goes, with low overhead. Instrumentation records exact events and spans, and the probe itself costs time. Production defaults to sampling plus a few high-value spans.
When is a flame graph the wrong first tool?
Answer
When the bottleneck is outside the process. Start with traces and with USE or RED, find the slow hop, then profile that process. A flat CPU flame graph on a network-bound service is a result: the time is in wait.
Throughput went up and p99 got worse. What happened?
Answer
Batching, queueing under load, GC pressure, or lock contention. More completed units per second can sit on a longer queue. Read the tail and the saturation signals together.
How do latency budgets relate to timeouts?
Answer
A budget allocates milliseconds per hop. A timeout for that hop sits above the expected p99, with headroom, and retry time is its own line. Copying the full client deadline into every hop invites a retry storm. Deadline propagation is the resilience page. The numbers you hand it come from the budget page.
Why re-measure after a fix?
Answer
Fixes move bottlenecks. A CPU win can expose IO wait, or it can raise allocation pressure and hand the tail to the GC. The same load shape, before and after, is the only comparison that counts.
Which number do you refuse as a user-facing SLO?
Answer
The mean. It is pulled around by rare outliers, and it hides a brutal p99 when you truncate those outliers. Ship p99 or p999 for interactive APIs, and keep the mean as a diagnostic.
What do you say in the first minute?
Answer
Measure, find the hot path, fix, re-measure. Latency, throughput, and capacity answer different questions. CPU-hot means a sampling flame graph. Wall-slow with idle CPU means wait, locks, or a downstream. Memory and GC own a different profile. Microbenchmarks pick algorithms. Load tests defend the SLO.
Pitfalls
- Rewriting the service before you have a percentile and a profile.
- Celebrating a mean-latency drop while p99 is the SLO.
- Maximizing throughput with batching, then acting surprised that the tail moved.
- Adding capacity on a hot path you never identified.
- Opening a CPU flame graph on a service that is waiting on the network.
- Treating a microbenchmark win as proof the production SLO moved.
A staff engineer says the checkout API is slow. Answer with the percentile you would ask for, the saturation check, the profiler axis you would pick for high CPU versus low CPU, and the re-measure you would require after the fix.