gRPC Deadlines, Cancellation & Error Model
gRPC calls carry a deadline (absolute time), not only a per-hop timeout. Cancellation propagates so servers stop work the client no longer needs. Errors are status codes + optional details + trailers, not HTTP status alone. Interviewers compare this to HTTP codes and ask which statuses are retryable.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
200 ms client SLO, nested hop
Prefer
One deadline, shrinking remainder
Deadline is now+200 ms on the user RPC. Service A passes now+150 ms to B. If B is slow, A returns DEADLINE_EXCEEDED and cancel stops B's leftover work.
- Timeout is relative on one hop; deadline is absolute for the tree.
- Retries must still be idempotent and fit leftover budget — pointer to idempotency keys.
- Trailers carry grpc-status / grpc-message plus optional google.rpc.Status details.
Alternative
Per-hop timeout only, or server waits longer than the client
Work continues after the client is gone. Streaming handlers keep producing. UNAVAILABLE and INTERNAL get the same retry policy.
- INTERNAL is a bug — page it. UNAVAILABLE is transient — retry with care.
- Do not invent a parallel HTTP code for every RPC; gateways map explicitly.
- Breaker/bulkhead mechanics live in Resilience — pointer only.
Deadline shrinks; cancel stops leftover work
Four beats, three participants. Budget theory in the Resilience timeouts lesson.
- 1
User deadline now+200ms
Absolute instant for the whole RPC tree. - 2
A calls B with now+150ms
Each hop must shrink remaining budget. - 3
B exceeds deadline
A surfaces DEADLINE_EXCEEDED to the user. - 4
Cancel signal on B
Stop work. Streaming must tear down the HTTP/2 stream.
Overview
gRPC calls carry a deadline (absolute time), not only a per-hop timeout. Cancellation propagates so servers stop work the client no longer needs. Errors are status codes + optional details + trailers, not HTTP status alone.
Interviewers compare this to HTTP codes and ask which statuses are retryable.
Deadlines vs timeouts
| Concept | Meaning |
|---|---|
| Timeout | Wait at most T from now on one hop |
| Deadline | Absolute instant for the whole RPC tree |
| Budget | Remaining time after nested calls (see Resilience timeouts) |
Propagate remaining budget: each hop must shrink the deadline. Anti-pattern: server timeout longer than client wait — work continues after the client is gone.
Sequence
- 1
User → Service A
1 Deadline now+200ms
- 2
Service A → Service B
2 Propagate now+150ms
- 3
Service B → Service A
3 Slow, exceeds deadline
- 4
Service A → User
4 DEADLINE_EXCEEDED
- 5
Service B
5 Cancel, stop work
Lesson map
gRPC Deadlines, Cancellation & Error Model
gRPC calls carry a deadline (absolute time), not only a per-hop timeout. Cancellation propagates so servers stop work the client no longer needs. Errors are status codes + optional details + trailers, not HTTP status alone. Interviewers compare this to HTTP codes and ask which statuses are retryable.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB u["User"] a["Service A"] b["Service B"] u -->|1 Deadline| a a -->|2 Propagate| b b -->|3 Slow, exceeds| a a -->|4| u
Cancellation
- Client cancels → server context/errgroup should abort.
- Deadlines imply cancel when exceeded.
- Streaming: cancel must tear down the HTTP/2 stream.
Status codes (know these cold)
| Code | Typical meaning | Retry? |
|---|---|---|
| OK | Success | n/a |
| INVALID_ARGUMENT | Bad request args | No |
| NOT_FOUND | Missing entity | No |
| ALREADY_EXISTS | Conflict create | No |
| FAILED_PRECONDITION | System not ready for call | Usually no |
| ABORTED | Concurrency conflict | Often yes (with backoff) |
| UNAVAILABLE | Transient dependency | Yes |
| DEADLINE_EXCEEDED | Time budget gone | Sometimes (idempotent only) |
| RESOURCE_EXHAUSTED | Quota / overload | Yes with care / shed |
| INTERNAL | Bug / invariant | No (fix) |
| UNAUTHENTICATED | Missing/invalid creds | No |
| PERMISSION_DENIED | Authz failed | No |
| CANCELLED | Caller cancelled | No |
| UNIMPLEMENTED | Method missing | No |
| UNKNOWN | Catch-all | Rarely |
Trailers carry grpc-status / grpc-message (and optional rich google.rpc.Status details).
Comparative vs HTTP
| HTTP | Rough gRPC map | Caution |
|---|---|---|
| 400 | INVALID_ARGUMENT | Not 1:1 |
| 401 | UNAUTHENTICATED | |
| 403 | PERMISSION_DENIED | |
| 404 | NOT_FOUND | |
| 409 | ABORTED / ALREADY_EXISTS | Pick carefully |
| 429 | RESOURCE_EXHAUSTED | |
| 499 client closed | CANCELLED | nginx-ism |
| 503 | UNAVAILABLE | |
| 504 | DEADLINE_EXCEEDED |
Do not invent a parallel HTTP code for every RPC — gateways map explicitly.
Sandbox: deadline budget + status stand-ins (Python)
Conceptual — not a gRPC runtime. UNAVAILABLE is the only retryable code in this toy.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Same ideas (TypeScript)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Deep dive · Resilience cluster — remaining budget only
Per-hop vs end-to-end budget, safety margins, and nested fan-out live in Timeouts, Budgets & Deadline Propagation. This page is how gRPC encodes that: deadline on the call context, cancel, status, trailers. Circuit breakers and bulkheads are not this lesson.
Pitfalls
User → A → B. A already spent 50 ms. What deadline do you pass B? If B needs 180 ms, what status does A return, and what must B do with its leftover work?
Interview Q&A
Deadline vs timeout?
Answer
Deadline is absolute end time for the RPC tree; timeout is often relative per attempt/hop.
Is DEADLINE_EXCEEDED retryable?
Answer
Only if the operation is idempotent and you still have a product-level budget — often you surface failure instead. Pointer: idempotency keys.
UNAVAILABLE vs INTERNAL?
Answer
UNAVAILABLE = transient; INTERNAL = bug. Retry the first, page on the second.
Where is grpc-status returned?
Answer
Typically in HTTP/2 trailers at end of stream (and sometimes headers for errors).
How does cancel interact with streams?
Answer
Cancel closes the stream; handlers must watch context to stop producing.
FAILED_PRECONDITION vs ABORTED?
Answer
Precondition: system state does not allow the call; ABORTED: concurrency/transaction conflict — often retryable.
Why propagate deadlines?
Answer
Prevents orphan work and cascading latency amplification. Ties to Resilience timeouts — do not recap breakers here.
Mapping 429?
Answer
Usually RESOURCE_EXHAUSTED — combine with shedding / Retry-After semantics at the edge.