Networking
Part 6 of 6 · Load balancingSticky Sessions vs Stateless — Affinity Tradeoffs & Consistent Hashing at the Edge
Cookie affinity and source-IP hash hurt scale/failover; prefer JWT/session store; when sticky is forced (WebSocket); consistent hash for cache locality without sticky sessions.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Where session state lives
Prefer
Stateless app + JWT or shared session store
Any healthy replica can serve the next request. Drain is a timeout, not a user-eviction problem. GLB can move regions.
- Autoscaling actually receives traffic.
- Canaries are header-based, not cookie-pinned to one box.
- Edge hash still pins cache shards for hit rate.
Alternative
Cookie sticky or source-IP hash onto in-memory sessions
Works for a legacy mid-migration. It encodes state into instances. CGNAT melts one backend. Mobile IPs rotate.
- Deploys wait for cookie TTL or drop logins.
- Affinity fights multi-region failover.
- PII concentrates on one VM — a compliance smell.
Overview
Session affinity (sticky sessions) hurts scale and failover. It is forced for some protocols (WebSockets, legacy servers). You can still get cache locality with consistent hashing without pinning user sessions to a single app instance.
You should be able to:
- Draw cookie sticky vs JWT + Redis vs Maglev-to-cache-shard on one whiteboard.
- Say "Maglev is not an in-memory session strategy."
- Migrate off sticky in four steps without a big-bang outage.
What sticky means
After the first request, the LB sends the same client to the same backend using cookie affinity, source-IP hash, or an explicit session header. The backend keeps in-memory session state — that is the root problem.
Why sticky hurts
- Autoscaling: new nodes get little traffic; old nodes stay hot.
- Deploy / drain: users pinned to a dying node until the cookie expires. Depth: drain.
- Failure: all sessions on that box are lost or forced to re-login.
- Imbalance: one heavy user pins a box; RR cannot help.
- Multi-region: affinity fights GLB; users stay trapped in a sick region.
Prefer JWT (stateless auth claims) + Redis/Memcached/DB session store + any healthy backend. JWT vs opaque and revocation: JWT study.
Comparative affinity strategies
| Strategy | Works for | State still on the node? | Interview verdict |
|---|---|---|---|
| Cookie sticky | Browsers | Yes, unless externalized | Transitional only; bound duration |
| Source-IP hash | No cookie clients | Yes | Almost never for public HTTP — CGNAT |
| JWT + shared store | Any HTTP client | No | Default modern APIs |
| Consistent hash by cache key | CDN / edge cache | N/A (shard, not session) | Locality feature |
| WebSocket natural affinity | One TCP/h2 stream | Conn lifetime | Required; drain + externalize presence |
Forced sticky: WebSockets
A WebSocket is one connection to one backend for its lifetime. Operationally: connection draining with a long timeout; externalize room/presence state; signal clients to reconnect on deploy. GraphQL subscriptions and SSE: treat like WebSockets for drain. Layer: L4 vs L7.
Consistent hashing at the edge
Hash user_id or URL to a Maglev / ring-hash cache shard for hit rate — app servers stay stateless and RR/P2C balanced.
Sticky sessions pin app instances. Consistent hash pins cache shards. Only the second is a scalability feature.
Do not re-teach vnode rings here. Placement, ~1/N movement, and RF walks: consistent hashing hub. A viral URL still melts one shard: hot keys / bounded loads. Maglev as an LB table: algorithms.
Flow
- 1
1 Client
- next2 Edge LB
- 2
2 Edge LB
- hash URL3a Cache shard A
- hash URL3b Cache shard B
- 3
3a Cache shard A
- next4 Stateless app pool
- 4
3b Cache shard B
- next4 Stateless app pool
- 5
4 Stateless app pool
- next5 Shared session store
- 6
5 Shared session store
Diagrams - step by step
Three small diagrams. Step numbers in the labels give the animation order. The lesson map under Diagram 1 plays those steps.
Diagram 1 - Happy path: stateless apps with consistent hashing to caches
Decisions
- 1
Step 1 Request with JWT reaches the edge LB
- nextStep 2 Edge hashes URL or key to a cache shard
- 2
Step 2 Edge hashes URL or key to a cache shard
- nextStep 3 Cache hit?
- ?
Step 3 Cache hit?
- yesStep 4a Serve from the cache shard
- noStep 4b Any app node via RR or P2C
- 4
Step 4a Serve from the cache shard
- 5
Step 4b Any app node via RR or P2C
- nextStep 5 App reads the session from the shared store
- session kept in node memoryFailure path - node dies and users are logged out
- 6
Step 5 App reads the session from the shared store
- nextStep 6 Respond - the node keeps no session state
- 7
Step 6 Respond - the node keeps no session state
- 8
Failure path - node dies and users are logged out
A request with a JWT hits the edge load balancer, which hashes the URL or key to a cache shard. A hit is served from that shard. A miss goes to any app node via round-robin or power of two choices. The app reads the session from the shared store and keeps no session state on the node. Keeping the session in node memory logs users out when the node dies.
Lesson map
Sticky Sessions vs Stateless — Affinity Tradeoffs & Consistent Hashing at the
Diagram 1 walks 7 steps from Step 1 Request with JWT reaches the edge LB through Step 6 Respond - the node keeps no session state.
Architecture. Step 1 Request with JWT reaches the edge LB Ready. Step 2 Edge hashes URL or key to a cache shard Ready. Step 3 Cache hit? Ready. Step 4a Serve from the cache shard Ready. Step 4b Any app node via RR or P2C Ready. Step 5 App reads the session from the shared store Ready. Step 6 Respond - the node keeps no session state Ready. Failure path - node dies and users are logged out Ready
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB A["Step 1 Request with JWT reaches the edge LB Ready"] B["Step 2 Edge hashes URL or key to a cache shard Ready"] C["Step 3 Cache hit? Ready"] D["Step 4a Serve from the cache shard Ready"] E["Step 4b Any app node via RR or P2C Ready"] F["Step 5 App reads the session from the shared store Ready"] G["Step 6 Respond - the node keeps no session state Ready"] X["Failure path - node dies and users are logged out Ready"] A -->|continues| B B -->|continues| C C -->|yes| D C -->|no| E E -->|continues| F F -->|continues| G E -->|session kept in node memory| X
Diagram 2 - Failure path: sticky sessions during scale-out and deploy
Sequence
- 1
LB with cookie stickiness → Node 1 old
Step 1 existing users pinned to Node 1 by cookie
- 2
LB with cookie stickiness → Node 2 new
Step 2 autoscaler adds Node 2 - only new users go there
- 3
Node 1 old → Node 1 old
Step 3 Node 1 stays hot while Node 2 idles
- 4
Node 1 old → LB with cookie stickiness
Step 4 Node 1 crashes or is drained for a deploy
- 5
LB with cookie stickiness
Step 5 every in-memory session on Node 1 is lost - forced re-login
- 6
LB with cookie stickiness
Fix - externalize sessions to Redis or a DB, then disable stickiness
Cookie stickiness pins existing users to the old node, so an autoscaled new node only receives new users and the old node stays hot. When that node crashes or is drained for a deploy, every in-memory session on it is lost and users must log in again. Externalize sessions to Redis or a database, then turn stickiness off.
Diagram 3 - Decision: affinity strategy
Decisions
- ?
Step 1 Connection-oriented - WebSocket, SSE, subscriptions?
- yesNatural affinity, long drain, state in Redis
- noStep 2 Need cache locality by key?
- legacy in-memory sessionsCookie sticky - transitional only
- 2
Natural affinity, long drain, state in Redis
- ?
Step 2 Need cache locality by key?
- yesConsistent hash at the edge to cache shards
- noStateless app - JWT plus shared store, RR or P2C
- Tempting shortcutSource-IP hash
- 4
Consistent hash at the edge to cache shards
- 5
Stateless app - JWT plus shared store, RR or P2C
- 6
Cookie sticky - transitional only
- Wrong pick long termBreaks autoscale, deploys and failover
- 7
Breaks autoscale, deploys and failover
- 8
Source-IP hash
- Wrong pickCGNAT piles thousands of users on one node
- 9
CGNAT piles thousands of users on one node
WebSockets, server-sent events, and subscriptions need natural affinity, a long drain, and state in Redis. Cache locality by key is consistent hashing at the edge to cache shards. Otherwise the app stays stateless: a JWT plus a shared store, balanced with round-robin or power of two choices. Cookie stickiness is only a transition off in-memory sessions. Source-IP hash piles CGNAT users onto one node.
Migration off sticky
- Introduce a Redis (or DB) session store; dual-write.
- Read from Redis; keep stickiness as a safety net.
- Disable stickiness; watch 401s / session misses.
- Remove in-memory session code.
Canaries: prefer a header canary on stateless apps. Cookie sticky pins users to the canary instance.
Sandbox: sticky vs failover (Python)
hash() is randomized per process — use a stable digest so the demo is deterministic.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Sandbox: edge shard pick (TypeScript)
This is cache locality, not a user session.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Edge cases interviewers love
- Mobile rotating IPs break source-IP sticky.
- A BFF cookie session should not require sticky if the BFF is stateless.
- GraphQL subscriptions / SSE: drain like WebSockets.
- Compliance: sticky can concentrate PII on one VM.
- ALB stickiness duration: bound it tightly; transitional only.
You ship a canary with cookie affinity. 5% of users stick to the canary box for 12 hours. What happens to that box's CPU? How would a header canary on a JWT API behave instead?
Interview Q&A
Why hate sticky sessions?
Answer
They encode state into instances. Autoscaling, drain, failover, and multi-region all get worse. Externalize state.
Source-IP affinity issue?
Answer
CGNAT and corporate NAT pile thousands of users onto one IP and melt one backend. Mobile IPs also rotate.
Is Maglev sticky?
Answer
It is a stable mapping for keys or flows. Use it for caches and packet affinity, not for in-memory user sessions.
How do you drain sticky users?
Answer
Stop issuing new stickies, wait for cookie TTL, force reconnect. Better: externalize state first so drain is just in-flight HTTP.
WebSocket without losing chat?
Answer
Store presence and history in Redis (or similar). On deploy, GOAWAY / close and let the client reconnect to any node.
ALB stickiness duration?
Answer
Bound it tightly. Treat it as a migration knob, not an architecture.
Canary plus sticky?
Answer
Pins a slice of users to the canary instance, not the canary build. Prefer header-based canaries on stateless apps.