Networking
Part 1 of 6 · Load balancingLoad Balancing — L4 vs L7, Algorithms & Health Checks
What an LB owns (VIP, backends, health, drain); L4 vs L7 matrix; algorithm zoo; health/drain; global LB; sticky vs stateless — comparative tradeoffs for senior interviews.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Design scenario
Same prompt for every reader.
Requirements
Design the load balancer in front of a stateless API. Clients use one VIP. Backends come and go.
Traffic / scale
About 20k requests/s across 30 instances, with uneven request cost.
Latency
Healthy requests stay under 100 ms at p99. Health checks must not sit on the request path.
Consistency
No session affinity. Any healthy instance can serve any request.
Availability
Lose one instance or one zone and keep serving. Draining must finish in-flight requests.
Failure assumptions
- A backend can accept TCP and still fail the request.
- A deploy can remove instances faster than new ones become healthy.
Constraints
- Do not pin users to one instance unless the API stores session state.
Prompt
Sketch the VIP, the health and drain path, and the algorithm you would start with.
API
What does the client connect to, and what does a backend health endpoint return?
Data
What does the balancer store about each backend, and how does it pick the next one?
Architecture
Where do L4 and L7 sit, and how does drain interact with the health check?
What sits in front of a fleet
Prefer
Health-aware L4/L7 LB with drain
Clients hit a stable VIP. Only healthy backends get new work. Deploys deregister, finish in-flight, then die.
- L7 for HTTP/gRPC path routing and retries; L4 for raw TCP, max PPS, or E2E TLS.
- Algorithm matches traffic shape — RR is not a personality.
- Fail-open vs fail-closed is an explicit SEV tradeoff.
Alternative
DNS round-robin or sticky app servers
Looks simple on a whiteboard. TTL delay, no health, and in-memory sessions fail the follow-up.
- Resolver caches ignore your 60s TTL more often than you want.
- Sticky pins users to dying nodes and fights autoscaling.
- An API gateway is auth/rate-limit — not a substitute for L4 packet scale.
Request path the hub exists to name
Each hop is a later lesson. The failure modes are why interviews start here.
- 1
Resolve a VIP
DNS or Anycast GLB picks a region / PoP. Depth: global LB. - 2
Regional L4 or L7 proxy
Tuple vs Host/path. TLS terminate or passthrough. Depth: L4 vs L7. - 3
Pick a backend
RR, least-conn, Maglev, P2C. Skew and hot keys break naive hash. Depth: algorithms. - 4
Only if healthy
Active + passive checks. Unhealthy_threshold > 1 or you flap. Depth: health. - 5
Drain before kill
No new assignments; finish in-flight; timeout then SIGTERM. WebSockets need a longer window.
Overview
Every production service that scales beyond one process needs a load balancer. Interviewers probe four choices: layer (L4 vs L7), split (algorithms), removal (health + drain), and geography (global LB) — plus when sticky sessions are a smell versus a necessity.
This hub is the map. The five sibling pages are the whiteboard depth.
You should be able to:
- Draw client → DNS/Anycast → regional LB → healthy backends, with one draining.
- Say why DNS-only LB fails interviews (TTL, no health, uneven resolvers).
- Name Maglev vs
hash % Nas ~1/N remap vs almost-all shuffle — without re-teaching vnode rings.
What a load balancer owns
| Piece | Job | Failure if missing |
|---|---|---|
| VIP / frontend | Stable address clients hit | Clients hard-code instance IPs; deploys break them |
| Backend pool | Instances that serve work | Nowhere to send traffic |
| Health awareness | Route only to healthy members | 5xx from dead boxes until a human notices |
| Drain / deregistration | Finish in-flight before kill | Mid-request 502s on every deploy |
| Policy | Algorithm, timeouts, TLS mode, retries | Wrong layer, retry storms, or TLS surprises |
LB vs the alternatives
| Approach | Wins | Loses |
|---|---|---|
| In-line LB (NLB/ALB/HAProxy/Envoy) | Health, drain, one VIP, HA | Extra hop; you operate or pay for it |
| Client-side LB (gRPC / Envoy sidecar) | No extra hop; locality; retries in-process | Every client must implement policy; harder for browsers |
| DNS round-robin only | Cheap | Slow TTL; no health; resolver ≠ user |
| Cloud LB (NLB/ALB/GLB) | Managed HA | Cost; less custom logic |
| API gateway | Auth, rate limits, WAF hooks | Not a substitute for L4 packet scale |
L4 vs L7 overview
- L4: IP/port (optionally TCP/UDP). Lower latency. NLB / HAProxy TCP. Passthrough TLS is common.
- L7: Headers, path, method, gRPC. ALB / NGINX / Envoy HTTP. Terminate TLS. Richer routing.
Rule of thumb: L7 for HTTP APIs; L4 for raw TCP, extreme PPS, or an end-to-end TLS mandate.
Depth: L4 vs L7 proxies.
Algorithm zoo overview
RR / weighted RR, least-conn, least-time, random, P2C, Maglev / consistent hash. Skew and hot keys break naive hashing. Maglev remaps ~1/N on membership change; hash % N reshuffles almost everything.
Do not teach vnode rings here. That lesson is consistent hashing. Hot-key overflow is bounded loads.
Depth: balancing algorithms.
Health / drain overview
Active vs passive checks; thresholds; slow start after recover; connection draining; fail-open (still route if all unhealthy) vs fail-closed (503).
Depth: health checks, slow start, drain.
Global LB overview
DNS geo/latency, Anycast, health-aware DNS, active-active vs active-passive, split-brain, RTO. A regional ALB survives an AZ. It does not survive a region.
Depth: global load balancing.
Sticky vs stateless overview
Prefer JWT + a shared session store. Sticky couples clients to instances and fights scale/failover. Consistent hash at the edge is for cache locality, not in-memory sessions.
Depth: sticky vs stateless.
Architecture (failure paths)
Healthy A/B take new work. C is draining — no new connections. A 5xx spike on C fails over to A. Region loss is a DNS/Anycast problem, not a local RR tweak.
Flow
- 1
1 Client
- next2 DNS / Anycast GLB
- 2
2 DNS / Anycast GLB
- pick region3 Regional L4/L7 LB
- 3
3 Regional L4/L7 LB
- healthy4a Backend A
- healthy4b Backend B
- drain no new4c Backend C
- 4
4a Backend A
- 5
4b Backend B
- 6
4c Backend C
- in-flight then 5xx4a Backend A
Diagrams - step by step
Three small diagrams. Step numbers in the labels give the animation order. The lesson map under Diagram 1 plays those steps.
Diagram 1 - Happy path: from DNS to a healthy backend
Flow
- 1
Step 1 Client resolves the service name
- nextStep 2 Global LB picks a region
- 2
Step 2 Global LB picks a region
- nextStep 3 Regional L4 or L7 LB accepts the connection
- 3
Step 3 Regional L4 or L7 LB accepts the connection
- nextStep 4 Algorithm picks a healthy backend
- target is drainingStep 6 No new connections to the draining target
- 4
Step 4 Algorithm picks a healthy backend
- nextStep 5 Backend serves the request
- 5
Step 5 Backend serves the request
- 5xx spikeFailure path - health checks eject the backend, traffic moves to healthy ones
- 6
Step 6 No new connections to the draining target
- 7
Failure path - health checks eject the backend, traffic moves to healthy ones
The client resolves the service name. The global load balancer picks a region, and the regional L4 or L7 load balancer accepts the connection. The algorithm picks a healthy backend, which serves the request. A draining target gets no new connections. A 5xx spike is the failure path: health checks eject that backend and traffic moves to the ones that are still healthy.
Lesson map
Load Balancing — L4 vs L7, Algorithms & Health Checks
Diagram 1 walks 6 steps from Step 1 Client resolves the service name through Step 6 No new connections to the draining target.
Architecture. Step 1 Client resolves the service name Ready. Step 2 Global LB picks a region Ready. Step 3 Regional L4 or L7 LB accepts the connection Ready. Step 4 Algorithm picks a healthy backend Ready. Step 5 Backend serves the request Ready. Step 6 No new connections to the draining target Ready. Failure path - health checks eject the backend, traffic moves to healthy ones Ready
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB A["Step 1 Client resolves the service name Ready"] B["Step 2 Global LB picks a region Ready"] C["Step 3 Regional L4 or L7 LB accepts the connection Ready"] D["Step 4 Algorithm picks a healthy backend Ready"] E["Step 5 Backend serves the request Ready"] F["Step 6 No new connections to the draining target Ready"] X["Failure path - health checks eject the backend, traffic moves to healthy ones Ready"] A -->|continues| B B -->|continues| C C -->|continues| D D -->|continues| E C -->|target is draining| F E -->|5xx spike| X
Diagram 2 - Failure path: region down but DNS answers are cached
Sequence
- 1
Client → DNS
Step 1 resolve the API name
- 2
DNS → Client
Step 2 Region A address with TTL 300 s
- 3
Region A → Region A
Step 3 Region A goes down
- 4
Client → Region A
Step 4 cached answer keeps sending traffic to A - errors
- 5
DNS → DNS
Step 5 health check marks A down, new answers point to B
- 6
Client → Region B
Step 6 client moves only after its TTL expires
- 7
Client
Fix - short TTLs, client retries, or an anycast VIP
The client caches Region A's address for the TTL. When Region A goes down, that cached answer keeps sending traffic there. DNS can mark A down and answer with B immediately, but this client does not move until its TTL expires. Short TTLs, client retries, or an anycast VIP close the gap.
Diagram 3 - Decision: layer and algorithm
Decisions
- ?
Step 1 Route on host, path or headers?
- yesL7 LB - ALB, Envoy, NGINX
- no - raw TCP or end-to-end TLSL4 LB - NLB, HAProxy TCP
- Tempting shortcutDNS round-robin only - TTL lag and no health checks
- 2
L7 LB - ALB, Envoy, NGINX
- nextStep 2 Request shape?
- 3
L4 LB - NLB, HAProxy TCP
- nextStep 2 Request shape?
- ?
Step 2 Request shape?
- uniform short requestsRound-robin or P2C
- long-lived streamsLeast-conn or least-request
- key affinity for cachesMaglev or ring hash
- 5
Round-robin or P2C
- 6
Least-conn or least-request
- 7
Maglev or ring hash
- 8
DNS round-robin only - TTL lag and no health checks
Routing on host, path, or headers needs an L7 load balancer. Raw TCP or end-to-end TLS stays on L4. Uniform short requests can use round-robin or power of two choices. Long-lived streams want least-conn or least-request. Cache key affinity wants Maglev or a ring hash. DNS round-robin alone has TTL lag and no health checks.
Sandbox: RR vs least-conn (Python)
Same request costs, two backends. RR equalizes count. Least-conn equalizes concurrency.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Same idea (TypeScript)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Pitfalls
Draw three backends behind one regional LB. Mark C draining. A client opens a new TCP connection — where does it go? An in-flight request on C — what happens during a 30s deregistration delay? Now kill the whole region. Which control plane moves traffic, and what still sits in resolver caches?
Interview Q&A
When is L4 enough?
Answer
Homogeneous TCP, TLS passthrough, or max packets-per-second. Path or header routing needs L7.
Why does DNS-only LB fail interviews?
Answer
Resolver TTL delay, no health, and the resolver's location is not the user's. Uneven resolvers also unbalance traffic.
RR vs least-conn?
Answer
RR if work is uniform. Least-conn for long-lived or varying cost (uploads, WebSockets, mixed RPC).
What is connection draining?
Answer
Stop new assignments, finish in-flight within a timeout, then deregister. WebSockets/gRPC streams need a window that matches max stream length — or send GOAWAY first.
Fail-open vs fail-closed?
Answer
Open still routes if every backend looks unhealthy (probe bugs should not become a SEV1). Closed returns 503 (safer for consistency-sensitive writes).
Why avoid sticky sessions?
Answer
They couple clients to instances: autoscaling, drain, and multi-region all get worse. Prefer a shared session store or JWT. Depth: sticky vs stateless.
Maglev vs modulo hashing?
Answer
Maglev remaps about 1/N of lookups when a backend joins or leaves. hash % N reshuffles almost all. Neither fixes a hot key. Rings/vnodes: consistent hashing.
Active-active multi-region gotcha?
Answer
Replication lag, split-brain writes, and sticky users pinned to a sick region. GLB controls traffic RTO; databases control RPO.