Networking
Part 4 of 6 · Load balancingHealth Checks, Slow Start & Connection Draining
Active vs passive health; HTTP vs TCP checks; thresholds; slow start after recover; connection draining / deregistration delay; fail-open vs fail-closed.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
How you learn a backend is dead
Prefer
Active probes + passive outliers
Probes give a predictable signal with no user traffic. Outliers catch 'process up, handler dead' that /healthz missed.
- unhealthy_threshold > 1 so one blip does not eject.
- Shallow liveness + readiness for traffic; deep checks sparingly.
- Drain ≥ p99 (+ WebSocket budget) or deploys abort in-flight work.
Alternative
TCP-only, or kill with deregistration_delay = 0
accept() is cheap and lies. Zero drain is a 502 generator. Deep /healthz that touches the DB cascades when the DB dips.
- Passive-only is silent when there is no traffic.
- Fail-open on a truly dead pool serves garbage.
- All pods unhealthy at once often means a shared deep check.
Deploy runbook
Health is the gate. Drain is the courtesy. Slow start is the hangover prevention.
- 1
Start new targets
Wait healthy_threshold successive probes before they exist. - 2
Attach + slow start
Ramp weight 0 → full over 30–120s if cold cache/JIT/pools matter. - 3
Deregister old
deregistration_delay ≥ p99, plus stream budget. No new assignments. - 4
Inflight → 0 or force-close
Then terminate. Rollback: keep the prior target group warm; flip L7 weights.
Overview
Design health signaling that removes bad backends quickly without flapping, ramp traffic safely after recovery (slow start), and evacuate instances without dropping in-flight work (connection draining). Know fail-open vs fail-closed.
You should be able to:
- Separate Kubernetes liveness (restart me) from readiness (stop sending traffic).
- Size drain to the longest request you refuse to abort — WebSockets included.
- Explain why a region-wide 503 from "all unhealthy" is often a shared deep check.
Active vs passive health
- Active: the LB probes on an interval (HTTP / TCP / gRPC). Predictable. Probe load. May miss real-path failures (
/healthz200, checkout 500). - Passive / outlier: observe real errors and latency. Catches app failures probes miss. Needs traffic and volume thresholds (poison-client patterns).
Best practice: both (Envoy outlier + active checks).
HTTP vs TCP checks
| Check | Cheap? | Catches "process up, app dead"? | Cascade risk |
|---|---|---|---|
TCP accept() | Yes | No | Low |
HTTP GET /healthz 200 | Medium | Only if the handler is honest | Low if shallow |
| Deep dependency ping | Costly | Yes, early | High — DB dip marks every pod bad |
gRPC grpc.health.v1 | Medium | If you implement it | Same shallow/deep split |
Prefer shallow liveness + readiness for traffic. Deep checks belong in alerts, not as the only eject signal.
Thresholds: interval, timeout, healthy_threshold, unhealthy_threshold. Too aggressive → flaps. Too slow → serve 5xx longer.
Slow start after recover
Cold caches, JIT, and connection pools melt under a full RR share. Slow start ramps weight 0 → full over 30–120s (Envoy and some cloud LBs). Pair with the scheduler: RR will otherwise dump 1/N onto a stone-cold box. Depth: algorithms.
Connection draining
Deregister, stop new assignments, let in-flight finish, force-close after timeout, then terminate. Notes on the LB — not labeled self-loops — keep the sequence readable.
WebSockets and gRPC streams need drain times matching max stream length, or send an application-level GOAWAY / close frame first.
Sequence
- 1
Orchestrator → Load Balancer
1 Deregister target
- 2
Load Balancer
2 Stop new assignments
- 3
Clients → Load Balancer
3 New request
- 4
Load Balancer
4 Route to other healthy targets
- 5
Clients → Target
5 In-flight continues
- 6
Target → Clients
6 Complete within drain window
- 7
Load Balancer → Target
7 Force-close after timeout
- 8
Orchestrator → Target
8 Terminate instance
Diagrams - step by step
Three small diagrams. Step numbers in the labels give the animation order. The lesson map under Diagram 1 plays those steps.
Diagram 1 - Happy path: deploy with health, slow start and drain
Flow
- 1
Step 1 New target starts
- nextStep 2 Passes healthy_threshold probes
- 2
Step 2 Passes healthy_threshold probes
- nextStep 3 Attach with slow start - weight ramps over 30-120 s
- 3
Step 3 Attach with slow start - weight ramps over 30-120 s
- nextStep 4 Old target deregistered - no new assignments
- full weight on a cold cacheFailure path - cold target melts and flaps
- 4
Step 4 Old target deregistered - no new assignments
- nextStep 5 In-flight requests finish within the drain window
- 5
Step 5 In-flight requests finish within the drain window
- nextStep 6 Force-close leftovers, terminate the instance
- 6
Step 6 Force-close leftovers, terminate the instance
- 7
Failure path - cold target melts and flaps
A new target passes healthy_threshold probes, then joins with slow start so its weight ramps over 30-120 s. The old target is deregistered and gets no new assignments. In-flight requests finish inside the drain window, leftovers are force-closed, and the instance terminates. Full weight on a cold cache melts the new target and flaps it.
Lesson map
Health Checks, Slow Start & Connection Draining
Diagram 1 walks 6 steps from Step 1 New target starts through Step 6 Force-close leftovers, terminate the instance.
Architecture. Step 1 New target starts Ready. Step 2 Passes healthy_threshold probes Ready. Step 3 Attach with slow start - weight ramps over 30-120 s Ready. Step 4 Old target deregistered - no new assignments Ready. Step 5 In-flight requests finish within the drain window Ready. Step 6 Force-close leftovers, terminate the instance Ready. Failure path - cold target melts and flaps Ready
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB A["Step 1 New target starts Ready"] B["Step 2 Passes healthy_threshold probes Ready"] C["Step 3 Attach with slow start - weight ramps over 30-120 s Ready"] D["Step 4 Old target deregistered - no new assignments Ready"] E["Step 5 In-flight requests finish within the drain window Ready"] F["Step 6 Force-close leftovers, terminate the instance Ready"] X["Failure path - cold target melts and flaps Ready"] A -->|continues| B B -->|continues| C C -->|continues| D D -->|continues| E E -->|continues| F C -->|full weight on a cold cache| X
Diagram 2 - Failure path: deep health check takes down every pod
Sequence
- 1
Shared DB → Shared DB
Step 1 DB has a short latency dip
- 2
Load balancer → All pods
Step 2 deep healthz probe also checks the DB
- 3
All pods → Load balancer
Step 3 every pod reports unhealthy
- 4
Load balancer → Load balancer
Step 4 fail-closed - zero healthy targets, 503 for everyone
- 5
Load balancer
Step 5 a brief DB blip became a full outage
- 6
Load balancer
Fix - shallow readiness, unhealthy_threshold above 1, fail-open when all fail
A short database latency dip makes a deep healthz probe fail on every pod. Fail-closed then has zero healthy targets and returns 503 for everyone. A shallow readiness check, an unhealthy_threshold above 1, and fail-open when every target fails keep a dependency blip from becoming an outage.
Diagram 3 - Decision: probe type and fail mode
Decisions
- ?
Step 1 Traffic always flowing?
- yesActive probes plus passive outlier detection
- no - idle periodsActive probes required
- 2
Active probes plus passive outlier detection
- nextStep 2 What should the probe gate?
- 3
Active probes required
- nextStep 2 What should the probe gate?
- ?
Step 2 What should the probe gate?
- readiness for trafficShallow check plus thresholds
- dependency lossDeep check that alerts only
- 5
Shallow check plus thresholds
- nextStep 3 All targets unhealthy?
- 6
Deep check that alerts only
- Wrong pick as the LB gateShared dependency dip ejects every pod
- ?
Step 3 All targets unhealthy?
- consistency-critical writesFail-closed with 503
- most read APIsFail-open to the last known set
- 8
Fail-closed with 503
- 9
Fail-open to the last known set
- 10
Shared dependency dip ejects every pod
Traffic that is always flowing can add passive outlier detection. Idle periods need active probes. Readiness for traffic should be a shallow check with thresholds. A deep dependency check should alert, not gate the load balancer. Consistency-critical writes fail closed with 503. Most read APIs fail open to the last known set. Using the deep check as the load balancer gate ejects every pod when the shared dependency dips.
Fail-open vs fail-closed
- Fail-closed: 503 if zero healthy — safer for consistency-sensitive writes.
- Fail-open: keep last backends / ignore health — avoids a total SEV1 from probe bugs.
Fail-open danger: you may send to truly dead nodes. Fail-closed danger: a bad /healthz takes the product down.
Comparative health designs
| Design | Fast deploys | Catches logic bugs | Probe cost | Silent death |
|---|---|---|---|---|
Shallow /healthz | Yes | Maybe not | Low | If handler lies |
| Deep dependency check | Slow | Early SEV | High | Cascade when dep dips |
| Passive only | N/A | Yes, with traffic | None | Yes, idle pool |
| Active only | Predictable | Misses probe ≠ user path | Steady | Rare |
Sandbox: health state machine (Python)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Sandbox: drain gate (TypeScript)
Synchronous on purpose — the playground worker does not wait on setTimeout.
Press Run. Snippets must be self-contained — no network, files, or native modules.
p99 request is 8s. WebSocket rooms last 30 minutes. Deregistration delay is 30s. What aborts on deploy? What do you send the client before SIGTERM?
Interview Q&A
Liveness vs readiness?
Answer
Liveness: restart me (deadlock, wedged event loop). Readiness: stop sending traffic (warming, draining, dependency not ready). Mixing them is how you crash-loop a slow starter.
Why is unhealthy_threshold greater than 1?
Answer
One lost probe or a GC pause should not eject a good box. Threshold > 1 trades a little detection delay for no flaps.
What is slow start?
Answer
Ramp weight after a backend becomes healthy so cold caches/JIT/pools are not hit with a full RR share.
Drain vs terminate now?
Answer
Drain preserves in-flight UX. Terminate-now is a 502. Size the window to p99 plus stream budget.
Fail-open danger?
Answer
You may send to truly dead nodes if probes are wrong in the other direction — or if you ignore health when the pool is empty.
All pods unhealthy at once?
Answer
Usually a shared deep check (DB down) or a bad deploy of /healthz. Fail-closed then 503s the region. Pair with GLB failover.
Passive false positive?
Answer
A poison client or a single bad URL pattern. Require volume thresholds and consecutive error rates, not one 500.
ALB deregistration too low?
Answer
Long requests and WebSockets are aborted mid-deploy. Raise deregistration_delay or GOAWAY first.