Observability
Part 5 of 6 · Resilience PatternsTimeouts, Budgets & Deadline Propagation — End-to-End Latency Caps
Timeouts bound how long one hop may wait; deadlines/budgets bound the whole user request across nested calls. Each hop must shrink the remaining budget (gRPC deadlines, context.WithDeadline, AbortSignal). Anti-pattern: server timeout longer than the client's wait—work continues after the client is gone. Pair with breakers and shedding so you fail fast when budget is nearly exhausted.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
200 ms client SLO, three hops — how do timeouts work?
Prefer
One deadline, shrinking remainder
Client deadline is now+200 ms. Each hop uses min(hop cap, remaining − safety margin). If pricing still needs 120 ms with 100 ms left, fail early.
- gRPC context, Go WithDeadline, AbortSignal, or a deadline header.
- Retries must fit leftover budget or they do not start.
- Hedging races a duplicate after delay — same deadline, same idempotency.
Alternative
Fixed 200 ms on every child, or 30 s server vs 2 s client
Ignoring time already spent blows the user SLO. Server work after the client abandoned is ghost work: pool waste and load amplification.
- Infinite wait + breaker may trip eventually — threads stuck until then.
- Unbounded hedges need [idempotency keys](/studies/api-idempotency-keys) and a hard deadline.
- Prefer breaker + budget over unbounded hedges.
Budget shrinks down the call graph
Short sequence: four participants, short notes, lifelines stay in the card.
- 1
Client deadline 200 ms
Absolute time when the user request must finish. - 2
API → inventory
Call with remaining 150 ms after local work. - 3
Inventory 40 ms OK
Remainder shrinks again before pricing. - 4
Pricing needs 120 ms
Not enough leftover → fail early (DeadlineExceeded). - 5
Partial response
504/503 with partial cart beats ghost work.
Overview
Timeouts bound how long one hop may wait; deadlines/budgets bound the whole user request across nested calls. Each hop must shrink the remaining budget.
Interviewers ask you to design deadline propagation for a 3-service fan-out under a 200 ms client SLO.
Per-hop timeout vs remaining budget
| Concept | Meaning | Example |
|---|---|---|
| Per-hop timeout | Max wait for this call | HTTP client timeout 100 ms to inventory |
| Deadline / budget | Absolute time when the user request must finish | Client deadline = now + 200 ms |
| Remaining budget | Deadline − now (− safety margin) | 200 − 80 elapsed = 120 ms left |
| Propagation | Pass deadline downstream | gRPC deadline, HTTP header, context |
Sequence
- 1
User
Step1 Client deadline 200ms
- 2
User → API
GET /cart budget=200ms
- 3
API → Inventory
Step2 Call remaining 150ms
- 4
Inventory → API
40ms OK
- 5
API → Pricing
Step3 Call remaining 100ms
- 6
Pricing
Step4 Need 120ms: fail early
- 7
Pricing → API
DeadlineExceeded
- 8
API → User
504/503 partial cart
Lesson map
Timeouts, Budgets & Deadline Propagation — End-to-End Latency Caps
Timeouts bound how long one hop may wait; deadlines/budgets bound the whole user request across nested calls. Each hop must shrink the remaining budget (gRPC deadlines, context.WithDeadline, AbortSignal). Anti-pattern: server timeout longer than the client's wait—work continues after the client is gone. Pair with breakers and shedding so you fail fast when budget is nearly exhausted.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB u["User"] a["API"] b["Inventory"] c["Pricing"] u -->|GET /cart| a a -->|Step2 Call| b b -->|40ms OK| a a -->|Step3 Call| c c -->|DeadlineExceeded| a a -->|504/503 partial| u
Rules of thumb
- Nested calls shrink budget: Never pass the full client timeout unchanged to every child.
- Safety margin: Reserve time for serialization, retries (if any), and response write.
- Timeout ≤ remaining budget: Cap each hop at min(configured_hop_timeout, remaining).
- Anti-pattern: Client waits 2 s; server times out at 30 s → ghost work, pool waste.
- Hedging vs retry: Hedging races a second request after delay; still must respect the same deadline and idempotency. Prefer breaker+budget over unbounded hedges.
Propagation mechanisms (concept)
- gRPC: Automatic deadline propagation via context.
- Go:
context.WithDeadline/WithTimeoutpassed through. - Node/TS:
AbortSignal+ absolute deadline timestamp header. - HTTP: Custom
X-Request-Deadline(epoch ms) or OpenTelemetry baggage — team convention matters.
Sandbox: budget that shrinks per hop (Python)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Same helper (TypeScript)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Comparative teaching
| Approach | Pros | Cons |
|---|---|---|
| Fixed per-hop only | Simple config | Ignores upstream wait already spent |
| Propagated deadline | End-to-end SLO compliance | Needs platform support / convention |
| Infinite wait + breaker | Breaker may trip eventually | Threads stuck until then |
| Hedging | Hides tail latency | Extra load; needs idempotency + same budget |
Pitfalls
Client SLO 200 ms. Inventory hop cap 100 ms, pricing 100 ms, tax 80 ms. Sketch remaining budget after a 40 ms inventory hit. When do you refuse to start pricing? Where does a retry fit?
Interview Q&A
Client timeout 200 ms, three sequential deps — how set timeouts?
Answer
Propagate a shared deadline; each hop uses min(hop_cap, remaining − margin). Do not give each hop 200 ms.
Why is server timeout greater than client timeout bad?
Answer
Server keeps working after the client abandoned; wastes bulkhead slots and can amplify load.
How do deadlines interact with retries?
Answer
Retries must fit inside remaining budget; if budget is less than hop needs, do not retry — fail or shed. Cross-link Retry Storms; do not recap jitter.
Hedging vs retry?
Answer
Hedging sends a concurrent duplicate after a delay; retry waits for failure first. Both need idempotency and a hard deadline.
What to metric?
Answer
deadline_exceeded_total, remaining budget histogram at each hop, timeout cause tags (local vs downstream). Pair with triad traces.
gRPC vs HTTP propagation?
Answer
gRPC often propagates deadlines on the context automatically. HTTP needs a team convention (X-Request-Deadline or baggage) plus AbortSignal on Node.
Safety margin — what is it for?
Answer
Serialization, retries if any, and writing the response. Without it, the last hop starts with a lie.
Infinite wait plus a breaker?
Answer
The breaker may trip eventually, but threads stay stuck until then. Budgets fail fast before the pool fills.
Fan-out under one budget?
Answer
Parallel children share the same remaining budget, not a full copy each. The parent still must finish after the slowest child.
504 vs 503 on deadline?
Answer
Either can be correct; be consistent and tag the cause. Partial cart plus 504/503 beats completing work the user will never see.