Performance
Part 5 of 6 · Performance EngineeringLatency Budgets & Critical Path — Where the p99 Lives
Budget each hop; find the critical path; fan-out p99 ≈ max of children; averages lie.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Build the budget from the SLO, then spend it
A budget that adds up to the entire SLO has no room for a retry or a GC pause.
- 1
Start from the percentile
Example: 300 ms at p99 for this API, not the average of every endpoint. - 2
List the hops
Edge, auth, handler, cache, database, fan-out, serialization. - 3
Allocate
Give each hop a number from measured p99, and leave headroom. - 4
Find the critical path
Sequential hops add. Parallel children contribute their max, plus merge cost. - 5
Revisit
Budgets rot when the architecture changes and nobody remeasures.
Three downstream calls on the request
Prefer
Parallel fan-out
Wall time tracks the slowest child, plus whatever it costs to merge.
- The parent p99 is worse than any single child's p99.
- A straggler dominates. Hedging is a later, expensive mitigation.
- Overlap is the win. You still budget the max, not the average child.
Alternative
Sequential calls
Wall time tracks the sum. Every extra hop is added to the user.
- Right when the next call needs the previous result.
- The fix is to remove serial work, not to tune the fastest hop.
- A trace waterfall shows the sum. A flame graph will not.
Overview
A latency budget splits an end-to-end SLO into per-hop allowances. The critical path is the chain whose delays determine the response. Sequential hops add. Parallel fan-out contributes the max of the children, plus merge cost. Averages lie: a healthy mean can hide a p99 driven by GC, retries, a noisy neighbor, or one slow dependency.
The hub says to measure before you rewrite. This page says which hop is worth measuring. Inside that hop, a flame graph explains the samples. Across hops, a trace explains the path. Use both.
Flow
- 1
1. Start from the p99 SLO
- next2. Budget each hop
- 2
2. Budget each hop
- next3. Sequential hops add
- 3
3. Sequential hops add
- next4. Fan-out waits on the max
- 4
4. Fan-out waits on the max
- next5. Compare the path to the SLO
- 5
5. Compare the path to the SLO
Lesson map
Latency Budgets & Critical Path — Where the p99 Lives
Budget each hop; find the critical path; fan-out p99 ≈ max of children; averages lie.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Start from the p99 SLO"] b["2. Budget each hop"] c["3. Sequential hops add"] d["4. Fan-out waits on the max"] a -->|1. Start from the p99 SLO| b b -->|2. Budget each hop| c c -->|3. Sequential hops add| d
A 300 ms example
These numbers are a miss path. A cache hit skips the database. Adding the miss-path caps lands on 300 ms, which means the example has no slack. A budget you can operate leaves headroom for a retry and a GC pause. If you publish the table below as the SLO, any hop that misses its cap breaks the promise.
| Hop | Budget at p99 |
|---|---|
| Edge and load balancer | 10 ms |
| AuthZ | 20 ms |
| Handler logic | 30 ms |
| Cache lookup | 15 ms |
| Primary database, on a miss | 80 ms |
| Fan-out, max of three children | 100 ms |
| Serialization and the rest | 45 ms |
If the database p99 is 150 ms, you are over budget before the fan-out runs. Fix the query, add a cache that actually hits, denormalize, move the read, or renegotiate the SLO. Do that in an order driven by the trace, not by the framework you wanted to try. One fat endpoint must not share a budget with a tiny health check. Average them together and the fat one spends the shared number.
Fan-out and the tail
Sequential latency is about the sum. Parallel latency is about the max. If each of 20 calls has a 1% chance of a slow spike, and the spikes are independent, the chance that at least one child is slow is 1 - 0.99^20, about 18%. The parent is slow on roughly one request in five even though each child looks fine at p99. That is why "our dependency is within SLO" can still blow the caller's p99.
Hedged requests and partial results are how you fight a known long-tail child. They cost extra load and they weaken a consistency story. Budget them on purpose. They are not the first fix.
Retries multiply load and can add another full hop to the tail. Put retry time on its own line. A budget that ignores retries describes the happy path only.
Percentiles
| Stat | What it is for |
|---|---|
| p50 | The typical request. Incomplete as an SLO |
| p99 | About 1 in 100. GC, retries, and lock waits show up here |
| p999 | The extreme tail. Capacity and noisy neighbors |
| Mean | Pulled by rare outliers, or blind to them if you drop them. A weak user-facing SLO |
Histograms are how you keep these numbers honest. The observability lesson is histogram versus average latency.
Timeouts, lightly
A hop budgeted at 80 ms might time out around 150–200 ms, with retries budgeted separately. Copying the full client deadline into every hop is how a slow dependency becomes a retry storm. Propagation, cancellation, and the storm itself are timeouts, budgets, and deadline propagation. This page's job is the number you hand that design.
Budgets also rot. If nobody remeasures after a new dependency ships, the table is fiction. Over-tight budgets push caching and hedging before the trace says you need them.
The sandbox draws each child's delay and combines them. Sequential adds the three draws. Fan-out keeps the max. Real concurrency overlaps the waits so wall time tracks that max. The arithmetic is the interview point. An asyncio.gather of the same three sleeps is the production-shaped version of the fan-out function.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
The production-shaped overlap looks like this. The page runner finishes before those timers would, so the runnable model above uses the same numbers without sleeping:
async def fanout() -> float:
delays = await asyncio.gather(
child("a", 0.02, 0.03),
child("b", 0.02, 0.03),
child("c", 0.02, 0.05),
)
return max(delays)Interview Q&A
How do you set a latency budget?
Answer
Start from the product SLO and the percentile. Leave headroom. Allocate from the measured p99 of each hop, not from a guess that sums to a round number. Revisit the table when a dependency is added or a cache is removed. A budget nobody remeasures is a slide.
Why is the p99 of a fan-out worse than the p99 of one call?
Answer
The parent waits for the slowest child. Tails combine. Twenty independent 1% spikes make the parent slow about 18% of the time. Quote the parent's percentile, not the child's.
Mean latency looks fine. Can we ship?
Answer
Only if the SLO is on the mean. A user-facing API SLO on p99 can still be broken. Means hide multimodal distributions and long tails, especially if you drop outliers before you average. Look at the histogram.
Critical path versus a flame graph?
Answer
The critical path is cross-service and span-level: which hop the user waits on. A flame graph is inside one process. The trace picks the process. The flame graph picks the function. Stopping at either one leaves half the story.
The database is over budget. What are the options?
Answer
Query and index work, a cache, denormalizing a read, a replica, moving the work off the request, or renegotiating the SLO. Evidence from the trace and the query plan decides the order. A new cache in front of an unindexed scan copies the slow query into another store.
How do retries affect the budget?
Answer
Each retry can add another full hop to the tail and multiplies load on the thing that is already slow. Budget retry time explicitly. A single 80 ms line that quietly includes two retries is how the table lies.
What is a latency SLO anti-pattern?
Answer
One number averaged across heterogeneous endpoints. The fat endpoint burns the shared budget and the cheap endpoints look virtuous. Budget the endpoint the user actually waits on.
Where does GC sit in the budget?
Answer
On the critical path of a random request, as a pause inside the hop that allocated. If you did not draw a box for it, it still spends the handler's milliseconds. The memory page is how you shrink it. The budget is where you notice it.
What do you say about headroom?
Answer
A table that sums to 100% of the SLO breaks as soon as one hop hits its cap on the same request as another. Leave room for the retry line and for pauses. Tightness is something you earn with measurements, not something you start with.
Pitfalls
- Budgeting the mean, then paging on p99.
- Adding sequential hops and calling the fan-out "parallel enough."
- Quoting a dependency's p99 as if it were the parent's.
- A 300 ms table with no slack, presented as a plan.
- Copying the client deadline into every downstream call.
- One SLO averaged across every route.
Give the miss path a budget that sums to less than 300 ms. Say which hop you would cut first if the database p99 is 150 ms, and what the parent p99 does if those three fan-out children each miss 1% of the time.