Performance
Part 6 of 6 · Performance EngineeringMicrobenchmarks vs Load Tests — Lies, Variance & Capacity
DCE/warmup/noise lies in microbench; open vs closed load; coordinated omission; capacity vs regression.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Is this faster?
Prefer
Match the question
An algorithm question gets a microbenchmark with stats. An SLO question gets an open-loop load test.
- Keep the result alive so the compiler cannot delete the work.
- Warm up, then report a median or an interval, not the mean of three runs.
- Confirm the winner again where contention, IO, and GC exist.
Alternative
One number from the wrong harness
A 10× lab win that disappears in production, or a closed-loop test that hides the queue.
- Dead-code elimination, a cold JIT, and an L1-sized input will flatter you.
- Closed-loop users slow their arrivals as latency rises, so overload looks polite.
- Coordinated omission records service time and forgets the wait the user felt.
Pick the experiment, then keep it honest
The same tools can answer two questions. The design of the run is the difference.
- 1
Name the question
Faster function, or a service that still meets the SLO? - 2
Name the scope
One function on a pinned core, or the service under a production-shaped mix. - 3
Run the matching test
Microbenchmark with warmup and a kept-alive result. Or an open-loop load test that records queue wait. - 4
Confirm
The lab winner still has to win in integration. The load test has to watch saturation and errors, not only a green RPS. - 5
Gate what is stable
A tight, relative threshold in CI. A heavy capacity run on a schedule, with a statistical gate.
Overview
A microbenchmark asks whether this function is faster in isolation. A load test asks whether the system meets its SLOs under concurrent, production-shaped traffic. Mixing them up is how a 10× "win" ships and then disappears. The hub already said to re-measure under a realistic shape. This page is why the first measurement lies.
Use microbenchmarks for algorithm choices, with statistical rigor. Use load tests for capacity, regression gates, and p99 under contention. The percentile you are defending comes from Latency Budgets.
Flow
- 1
1. Algorithm or system SLO
- next2. Name the scope
- 2
2. Name the scope
- next3. Microbench or open-loop load
- 3
3. Microbench or open-loop load
- next4. Confirm under contention
- 4
4. Confirm under contention
- next5. Gate regressions or size capacity
- 5
5. Gate regressions or size capacity
Lesson map
Microbenchmarks vs Load Tests — Lies, Variance & Capacity
DCE/warmup/noise lies in microbench; open vs closed load; coordinated omission; capacity vs regression.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Algorithm or system SLO"] b["2. Name the scope"] c["3. Microbench or open-loop load"] d["4. Confirm under contention"] a -->|1. Algorithm or system SLO| b b -->|2. Name the scope| c c -->|3. Microbench or open-loop load| d
How microbenchmarks lie
- Dead-code elimination. If the result is unused, the compiler deletes the work. A benchmark of nothing is very fast. Keep a sink: xor the result into a variable the process prints.
- Warmup and JIT. The first iterations include compilation. Report the steady state after warmup. A single cold start is a startup test, and only if you meant to measure startup.
- Noise. Noisy neighbors, frequency scaling, and GC. Take many iterations and a robust statistic. Distrust n = 3.
- Cache effects. A tiny working set fits in L1. Production data does not. A win on 100 keys can lose on 100 million.
- Toy inputs. Happy-path JSON is not an adversarial payload. Benchmark the size you actually serve.
- The wrong summary. Prefer intervals and tools that already know this: pytest-benchmark, hyperfine, JMH, Criterion, Go's
testing.B. Benchmark.js needs the same DCE discipline.
How load tests lie
- Closed loop. N users send the next request when the previous one returns. As latency rises, arrival rate falls. Overload hides, because the test backs off for you.
- Open loop. Arrivals follow a schedule that does not wait for completions. Queues show up. So does true capacity.
- Coordinated omission. You record only the service time of requests that were accepted, and you drop the time spent waiting to be accepted. Percentiles look better than what users experience. Gil Tene's talk is the short version of this failure. wrk2 and a constant-throughput model exist because of it.
- The wrong shape. Identical payloads, no cache-miss mix, no diurnal pattern. You measured a demo.
- A shared staging dependency. A saturated database makes your service look worse, or a warm shared cache makes it look better. Isolate the thing you claim to have measured.
| Experiment | Design | Question |
|---|---|---|
| Regression gate | Fixed load shape. Fail when p99 or errors cross a threshold | Did this change make the SLO worse? |
| Capacity plan | Raise load until the SLO breaks. Record max RPS and the saturated resource | How many replicas, with headroom? |
k6, wrk, wrk2, vegeta, Locust, and Gatling can run either experiment. wrk2 and k6's open models are the ones to reach for when the question is latency under a chosen arrival rate. A closed-loop Locust script is a different question: "what do N users do?" Both are valid. Name which one you ran.
| Tool | Role |
|---|---|
| pytest-benchmark, hyperfine, JMH, Criterion, testing.B | Microbenchmarks with warmup and stats |
| k6 | Scripted scenarios, thresholds, a modern open model |
| wrk and wrk2 | Raw HTTP. wrk2 holds a target rate and uses HdrHistogram |
| vegeta, Locust, Gatling | Load. Check whether the model is open or closed |
CI performance tests flake when the runner is noisy. Pin CPU and memory, warm up longer, use relative thresholds, and isolate the machine. Move the heavy capacity run to a scheduled job with a statistical gate. A red mainline build on a 3% jitter is how teams delete the test.
A microbenchmark that deletes its own work
The naive loop times work and drops the return value. A smart compiler, or a later inlining pass, can make that loop cheaper than the function you think you measured. The better pattern warms up, keeps a sink, and reports a median and a high percentile.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Python will not delete this loop the way a C or JVM compiler will. The sink still matters the moment the same function is ported, and the median still matters today because the first iterations and the noisy ones are in that list. Print the sink so the work stays live.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Ten closed-loop users at 50 ms can finish about 200 requests in one second, because each user waits. An open loop at 200 requests per second attempts all 200 in that same second. One worker at 50 ms can finish 20. The other 180 are the queue coordinated omission forgets if you only time the requests that reached the handler.
Interview Q&A
When is a microbenchmark enough to ship?
Answer
Almost never by itself. It is enough to pick an algorithm. Then confirm under integration and under load, on the critical path, with the percentile from the budget. A function that is off the request path can be twice as fast and change nothing a user sees.
What is coordinated omission?
Answer
Measuring only the service time of requests that were accepted, and ignoring the time those requests waited in a queue. The histogram looks healthy because the omitted waits are the tail the user felt. Record arrival time at the client, not only time inside the handler.
Closed loop versus open loop?
Answer
A closed loop fixes the number of users. The next arrival waits on the previous response, so offered load falls as latency rises. An open loop fixes the arrival process. It is the one that shows queues and the capacity knee. Say which one you ran before you quote a percentile.
How many iterations does a microbenchmark need?
Answer
Enough for a stable median or a confidence interval after warmup. The tools above automate that. Three runs and a mean is a coin flip. If the interval is wide, you do not have a winner yet.
k6 versus wrk?
Answer
k6 scripts scenarios and thresholds, and it can hold an open arrival model. wrk is raw HTTP throughput. wrk2 aims at a constant rate and records an HdrHistogram so the tail is harder to round away. Pick from the question, not from which binary was already installed.
CI performance tests are flaky. What now?
Answer
Pin CPU and memory, lengthen warmup, use a relative threshold, and isolate the runner. Move heavy capacity tests to a scheduled job with a statistical gate. Deleting the test because it was red on a noisy laptop is how the regression comes back in production.
The microbenchmark says 2× and production did not move. Why?
Answer
You were off the critical path. IO, contention, or GC dominates the request. The production compiler already did the optimization. Or the lab input fit in L1 and the production input does not. Re-measure on the path the budget named.
How do load tests relate to latency budgets?
Answer
The budget names the hop and the percentile. The load test checks that percentile under concurrency and finds the saturation knee. A load test with no percentile is a throughput demo. A budget with no load test is a spreadsheet.
What is the difference between a regression gate and a capacity plan?
Answer
A regression gate holds the load shape still and fails the build when p99 or the error rate gets worse. A capacity plan increases load until the SLO breaks, then sizes replicas with headroom. Running one and announcing the other's number is how capacity reviews go wrong.
Pitfalls
- Shipping a microbenchmark as an SLO result.
- A closed-loop test quoted as "we handled the peak."
- Service-time histograms that start at accept.
- n = 3, no warmup, result discarded.
- A shared staging database in the critical path of a test you call isolated.
- A CI gate so tight it fails on neighbor noise, then gets ignored.
A teammate says wrk showed 2 ms p99 at 50,000 requests per second, and the lab function is 10× faster than last week. Say which result might be coordinated omission, which might be dead-code elimination, and which experiment you would run before changing the replica count.