gRPC Load Balancing, Name Resolution & Channel Health
gRPC clients open a channel to a resolved set of addresses and may balance RPCs client-side (pick_first, round_robin, …) or through a proxy (Envoy/L7). Name resolution (DNS, xDS), keepalive, health checks, and sticky affinity for bidi decide whether streams survive deploys. Cross-link the general Load Balancing cluster — here we focus on gRPC-specific channel behavior.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Who picks the backend?
Prefer
Channel policy you can name (client-side or proxy)
Client-side: stub picks from the resolver (pick_first, round_robin, ring_hash). Proxy/L7: Envoy owns policy, mTLS, retries. xDS when you need dynamic config. Bidi gets affinity so the stream is not reset.
- Client-side is mesh-friendly and low hop latency; every client needs the policy.
- Proxy wins for central policy. Extra hop is the cost.
- Do not dump Maglev/P2C here — algorithms lesson in the Load Balancing cluster.
Alternative
pick_first for every client, DNS TTL measured in minutes
Herd to one pod. Dead addresses linger. Bidi resets on rebalance. Keepalive either hammers battery/LBs or never notices a blackhole.
- Headless Services expose pod IPs for client LB; ClusterIP pushes LB to kube-proxy/Envoy.
- Ignoring TRANSIENT_FAILURE yields spiky errors — wait-for-ready + UNAVAILABLE retries.
- Sidecar often owns LB/mTLS; app dials localhost — Sidecar cluster, pointer only.
Channel from dial to failover
Resolver, policy, READY, multiplex, keepalive, re-resolve. Not Maglev lookup tables.
- 1
Dial the target
dns:///users.svc.cluster.local:443 (or xDS). - 2
Resolver returns addresses
Plus optional service config JSON. - 3
LB policy picks a subchannel
pick_first, round_robin, or sticky for bidi. - 4
HTTP/2 READY
Many RPCs multiplex on one connection per subchannel. - 5
Keepalive + health
On fail: re-resolve / failover. Tune idle timeouts behind L4 LBs.
Overview
gRPC clients open a channel to a resolved set of addresses and may balance RPCs client-side (pick_first, round_robin, …) or through a proxy (Envoy/L7). Name resolution (DNS, xDS), keepalive, health checks, and sticky affinity for bidi decide whether streams survive deploys.
L4 vs L7, Maglev, least-conn: Load Balancing hub — do not re-teach.
Client-side LB vs proxy
| Mode | How | Pros | Cons |
|---|---|---|---|
| Client-side | Stub picks backend from resolver | Low hop latency; mesh-friendly | Every client needs policy |
| Proxy / L7 | Envoy, nginx, cloud LB | Central policy, mTLS, retries | Extra hop; config surface |
| xDS control plane | Lookaside / sidecar config | Dynamic, rich | Ops complexity |
Name resolution
- Target string:
dns:///users.svc.cluster.local:443 - Resolver returns addresses (+ optional service config JSON).
- LB policy chooses subchannel.
- Connectivity states: IDLE → CONNECTING → READY → TRANSIENT_FAILURE → SHUTDOWN.
pick_first vs round_robin
| Policy | Behavior | Use when |
|---|---|---|
| pick_first | One backend until failure | Simple; sticky by accident |
| round_robin | Rotate among READY subchannels | Spread unary load |
| ring_hash / sticky | Consistent hash on key | Session affinity, bidi |
Bidi / streaming often needs affinity so the stream is not reset mid-flight when the LB re-picks.
Flow
- 1
1 Dial dns:///svc:443
- next2 Resolver returns A/B/C
- 2
2 Resolver returns A/B/C
- next3 LB picks subchannel
- 3
3 LB picks subchannel
- next4 HTTP/2 conn READY
- 4
4 HTTP/2 conn READY
- next5 RPCs multiplexed
- 5
5 RPCs multiplexed
- next6 Keepalive + health
- 6
6 Keepalive + health
- next7 Fail: re-resolve
- 7
7 Fail: re-resolve
Lesson map
gRPC Load Balancing, Name Resolution & Channel Health
gRPC clients open a channel to a resolved set of addresses and may balance RPCs client-side (pick_first, round_robin, …) or through a proxy (Envoy/L7). Name resolution (DNS, xDS), keepalive, health checks, and sticky affinity for bidi decide whether streams survive deploys.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1 Dial dns:///svc:443"] b["2 Resolver returns A/B/C"] c["3 LB picks subchannel"] d["4 HTTP/2 conn READY"] a -->|1 Dial dns:///svc:443| b b -->|2 Resolver returns A/B/C| c c -->|3 LB picks subchannel| d
Channel health & keepalive
- Keepalive pings detect dead NATs and blackholed connections.
- Health checking (
grpc.health.v1) removes unready backends fromround_robin. - Connection reuse: many RPCs share one HTTP/2 connection per subchannel — watch head-of-line only at app layer; HTTP/2 multiplexes streams.
- Tune idle timeouts carefully behind L4 load balancers.
Comparative pitfalls
| Pitfall | Symptom | Fix |
|---|---|---|
| DNS TTL too long | Traffic to dead pods | Short TTL + health + pod readiness |
| pick_first + many clients | Herd to one pod | round_robin / proxy |
| Bidi without affinity | Stream resets | Sticky / ring hash / pod IP |
| Keepalive too aggressive | Battery / LB drops | Align with infra idle timeouts |
| Ignoring TRANSIENT_FAILURE | Spiky errors | Wait-for-ready + retries on UNAVAILABLE |
Sandbox: toy resolver + round_robin (Python)
In-memory subchannels. Production uses the gRPC resolver + LB policy stack.
Press Run. Snippets must be self-contained — no network, files, or native modules.
pick_first vs round_robin (TypeScript)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Pitfalls
Unary GetUser and bidi Chat against Kubernetes. Which Service type for client-side round_robin? What affinity does Chat need through Envoy?
Interview Q&A
What does a gRPC channel represent?
Answer
A virtual connection to a target that manages subchannels, LB, and connectivity.
pick_first vs round_robin?
Answer
pick_first uses one backend; round_robin spreads across READY subchannels.
Why is client-side LB common in gRPC?
Answer
HTTP/2 long-lived connections + rich service config; proxies still win for central policy.
How does DNS interact with Kubernetes?
Answer
Headless services expose pod IPs for client LB; ClusterIP pushes LB to kube-proxy/Envoy instead.
Why sticky for bidi?
Answer
The stream is pinned to a process; rebalancing mid-stream breaks it (same idea as WebSocket affinity).
What is wait-for-ready?
Answer
Queue RPCs until a subchannel is READY instead of failing fast on transient disconnects.
Health check protocol?
Answer
Standard grpc.health.v1.Health/Check — integrate with LB policies. Broader health/drain: health checks lesson — pointer only.
Relation to mesh sidecars?
Answer
Sidecar often owns LB/mTLS; app uses a local channel to localhost — see Sidecar & Service Mesh.