Performance
Part 2 of 6 · Performance EngineeringCPU & Wall-Clock Profiling — Sampling vs Instrumentation
Sampling vs instrumentation; on-CPU vs wall-clock; GIL/event-loop traps that make “CPU” misleading.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
You need to know where time goes, in production
Prefer
Sampling
Interrupt on a timer, record a stack, aggregate. Overhead stays in the low single digits.
- perf, py-spy, async-profiler, Go and Java pprof, continuous profilers.
- Finds hot functions you did not think to instrument.
- Statistical. Short functions and safepoints can skew the picture.
Alternative
Instrumentation
A hook or a span on every call. Counts and attributes are exact. The probe is part of the cost.
- cProfile, yappi, OpenTelemetry spans, APM agents, manual timers.
- Right for a distributed critical path and for exact call counts.
- Easy to over-collect. Cardinality and probe cost move the bottleneck.
Pick the axis, then confirm on the other one
A single method is a hypothesis. The second method is the check.
- 1
Name the symptom
High latency, high CPU, or both. Write down the percentile. - 2
Choose on-CPU or wall-clock
Cores busy means on-CPU. Cores idle and the request slow means wall-clock, off-CPU, or a trace. - 3
Sample or instrument
Sampling stacks for the hot function. Spans when you need the cross-service path. - 4
Aggregate
Stacks become a flame graph. Spans become a waterfall. - 5
Confirm
If sampling says CPU and the wall-clock profile says epoll, believe the wait.
Overview
CPU profiling answers "where does the processor spend time?" Wall-clock profiling answers "where does the request spend time, including waits?" Those are different clocks. A service can burn a core on JSON parsing, or it can sit in epoll while CPU% looks bored. The fix is different.
Sampling profilers (perf, py-spy, async-profiler, Go and Java pprof) interrupt periodically and record stacks. Overhead is low, and the result is slightly statistical. Instrumentation (cProfile, APM auto-instrumentation, manual timers) records every call or span. It is precise, and it can distort the thing you measure.
The map of the cluster is the hub. The picture you draw from these stacks is Flame Graphs.
Flow
- 1
1. Symptom is latency or CPU
- next2. Pick on-CPU or wall-clock
- 2
2. Pick on-CPU or wall-clock
- next3. Sample stacks or record spans
- 3
3. Sample stacks or record spans
- next4. Flame graph or waterfall
- 4
4. Flame graph or waterfall
- next5. Confirm on the other axis
- 5
5. Confirm on the other axis
Lesson map
CPU & Wall-Clock Profiling — Sampling vs Instrumentation
Sampling vs instrumentation; on-CPU vs wall-clock; GIL/event-loop traps that make “CPU” misleading.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Symptom is latency or CPU"] b["2. Pick on-CPU or wall-clock"] c["3. Sample stacks or record spans"] d["4. Flame graph or waterfall"] a -->|1. Symptom is latency or CPU| b b -->|2. Pick on-CPU or wall-clock| c c -->|3. Sample stacks or record spans| d
Sampling
Tools: Linux perf, py-spy, async-profiler on the JVM, go tool pprof, and continuous profilers.
- Overhead around 1–5% is the usual reason it is safe in production.
- It finds unexpected hot functions, because you did not have to name them first.
- It can miss very short functions. You need enough samples, which means duration under real load.
- Wall-clock sampling and on-CPU sampling are different switches. Turning on the wrong one gives you a confident, empty graph.
- Bias: short functions are under-represented unless the rate is high enough. Signal handlers and JVM or Python safepoints skew which stacks you are allowed to see.
Sample rate is a resolution versus cost trade. Start around 49–99 Hz for CPU. Raise the rate when stacks are too sparse to read. A higher rate costs overhead and storage.
Instrumentation
Tools: cProfile and yappi, OpenTelemetry spans, APM agents, and manual perf_counter or performance.now wrappers.
- You get exact call counts, timings, and attributes. That is what a critical-path budget wants from each hop.
- Spans correlate with logs and trace ids. That is how you join this page to the observability triad.
- The observer effect is real. The act of measuring changes timings. Auto-instrumentation noise and unbounded attributes are how a trace bill becomes the incident.
- Refuse always-on instrumentation in production when probe cost or cardinality would change latency more than the insight is worth. Prefer continuous sampling, plus sparse high-value spans.
On-CPU versus wall-clock
| Clock | What it includes | What a high number means |
|---|---|---|
| On-CPU | Threads actually running on a core | Algorithmic or compute hot path |
| Wall-clock | Run time plus sleep, IO, locks, scheduling | The request is waiting |
| Off-CPU | The blocked part of wall time | IO, lock contention, or a parked runtime |
A CPU flame graph of a network-bound service looks empty, or it is dominated by poll and epoll. You need a wall-clock profile or a trace. High on-CPU plus high latency points at compute. High wall-clock and low on-CPU points at wait, a lock, or a downstream. CPU at 20% with a bad p99 is the second case. Do not start by optimizing a tight loop you have not seen in a sample.
GIL and the event loop
Python. The GIL means several threads can look active while only one runs Python bytecode. Process-level sampling with py-spy is the tool that matches that model. C extensions often release the GIL during IO, so "the process is in a C frame" can be a wait, not a compute burn. py-spy and top disagree when they are measuring different clocks, different windows, or native frames the other tool cannot see.
Node. process.cpuUsage and a modest CPU% can look fine while a synchronous JSON.parse blocks the event loop and spikes p99. Profile with Clinic.js, 0x, or Chrome DevTools. Treat synchronous CPU on the main thread as latency poison: one request's parse is every other request's queue.
Go, in production. Expose net/http/pprof, capture CPU and heap profiles under load, and visualize with go tool pprof or a flame graph. Keep the scrape authenticated. A public pprof endpoint is a profile of your service and a free look at its traffic.
| Runtime | First profiler | Trap |
|---|---|---|
| Linux native | perf | Missing symbols, so you optimize a hex address |
| Python | py-spy | Reading CPU% as if the GIL were not there |
| JVM | async-profiler | Safepoint bias |
| Go | pprof | An unauthenticated debug server |
| Node | Clinic.js, 0x, DevTools | A sync parse that never moves average CPU |
A small instrumentation sketch
cProfile counts every Python call. A handful of wall-clock samples only tells you how long the whole workload sat there. Use the first when you need counts. Use the second when you need a cheap "is this even slow?" check. Neither one is a production sampling profiler. py-spy is.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Interview Q&A
When would you refuse always-on instrumentation in production?
Answer
When probe cost or attribute cardinality would change latency or spend more than the insight is worth. Prefer a continuous sampling profiler and a small set of high-value spans. Turn the heavy agent off and confirm the latency comes back before you blame the code.
CPU is 20% and p99 is bad. What is the next step?
Answer
A wall-clock or off-CPU profile, or a distributed trace. The request is waiting: IO, a lock, a downstream, or a brief sync spike that never moves average CPU. An on-CPU flame graph will look quiet, and that quiet is the finding.
What is the difference between perf and cProfile?
Answer
perf samples OS-visible stacks with tiny overhead. cProfile instruments Python function calls, so you get exact counts and a higher overhead. perf can see native frames. cProfile sees Python frames. Use perf or py-spy in production. Use cProfile when you want counts on a workstation.
Why might py-spy disagree with top?
Answer
They are different clocks and different windows. top is process CPU. py-spy can sample on-CPU or include idle time. The GIL, native extensions, and a short sample window all move the stacks. Reconcile the question each tool answers before you call one of them wrong.
How do you profile a Go service in production?
Answer
Expose net/http/pprof on an authenticated port. Capture a CPU profile and a heap profile under load. Visualize with go tool pprof or a flame graph. Do not leave the default debug server on a public interface.
Node shows high latency and low CPU. What are the causes?
Answer
Awaited IO, DNS, GC pauses, or occasional synchronous spikes that do not move average CPU%. A sync JSON parse blocks every other request on the event loop for the duration of the parse. Profile the main thread. Average CPU will not confess.
How do you avoid the observer effect?
Answer
Prefer sampling. Keep an instrumentation budget. Compare a run with profiling off. Validate the fix under the same load generator you used for the baseline. If the only fast version is the version with the agent detached, the agent was the workload.
How do you choose a sampling rate?
Answer
Higher rate, better resolution, more overhead and storage. Start near 49–99 Hz for CPU. Raise it only when the stacks are too sparse to name a function. A rate you cannot afford in production is a lab tool.
Why confirm with a second method?
Answer
Sampling can miss a short function or a safepoint-biased stack. Instrumentation can invent cost. If both point at the same frame, you have a hot path. If they disagree, you have a clock mismatch or an observer effect, and that is the next question.
Pitfalls
- Reading a CPU flame graph of a service that is stuck in poll, then optimizing the runtime.
- Turning on every APM integration and calling the new latency "the code."
- Trusting process CPU% on a GIL language or an event-loop language.
- A five-second sample of an idle box, then a confident ranking of functions.
- Publishing pprof, or any profiler port, without auth.
- Comparing a profile taken at 10 requests per second with one taken at saturation.
CPU is 15% and checkout p99 is 800 ms. Say which profile you capture first, and what you will believe if the on-CPU stacks are mostly epoll.