Messaging
Part 5 of 6 · WebSockets & MQTTScaling WebSockets — Sticky Sessions, Fan-out & Backpressure
A WebSocket is pinned to the process that accepted the TCP connection. Sticky affinity reduces reconnect churn, but multi-node rooms need a fan-out bus. Bounded queues, coalesce, and disconnect-slow-clients keep one laggy socket from OOM-ing the node.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Sticky solves routing; the bus solves multi-node rooms
Prefer
Externalize room state + pub/sub (or MQTT broker cluster)
Affinity is optional if every node can reach every socket via a bus. Keep local socket maps tiny. Coalesce ticks so one slow client cannot pin memory.
- Redis pub/sub is ephemeral — gone if nobody is subscribed, gone on restart unless you also persist.
- Kafka (or Redis Streams) when the fan-out must also be a log.
- HPA on connections per pod + event-loop lag, not CPU alone.
Alternative
Sticky in-memory rooms and hope
Simple until failover. User A on node 1 cannot talk to user B on node 2. A deploy that SIGKILLs sockets looks like an outage. One slow ticker OOMs the process.
- Sticky hurts HTTP APIs — the LB lesson still applies; WS is the forced exception for connection lifetime.
- Serverless WS still needs idempotency and backpressure.
- At-least-once buses duplicate; clients dedupe by event id.
From connect to a slow client
Interviews start at 'why can't node 1 push to B?' then at the OOM.
- 1
Connect
Any healthy node. Sticky optional if state is shared. Depth: LB sticky lesson. - 2
Join room
Track socket id locally and register interest on the bus. - 3
Event
PUBLISH to the bus. Nodes with members write frames. - 4
Slow consumer
Queue hits max. Drop oldest or last-value wins. Or disconnect. Depth: bufferedAmount. - 5
Deploy
Stagger close 1001, jitter reconnect, warm capacity first. Depth: handshake + LB drain.
Overview
Scale beyond one WS process: why sticky affinity appears, how Redis (or NATS/Kafka) fan-out decouples connections from publishers, and how backpressure prevents memory melt.
This deepens the sticky-sessions and L4 vs L7 docs. It does not recap Maglev, least-conn, or health-check thresholds — those stay in the load-balancing hub.
Why sticky shows up
A WebSocket is pinned to the process that accepted the TCP connection. An L7 LB with cookie or consistent-hash affinity keeps reconnects warm and simplifies in-memory room maps.
The LB sticky lesson explains why sticky hurts HTTP APIs. For WebSocket it is often forced for connection lifetime, but room/presence state must still be externalized for failover.
L4 can pass TCP; path-based routing and auth cookies need L7. Depth: L4 vs L7.
Do not stop at sticky
If user A is on node 1 and user B on node 2, node 1 cannot push to B without a shared bus.
Pattern: each node SUBSCRIBEs to Redis channels for rooms it hosts; on a local message, PUBLISH to Redis; other nodes forward to their local sockets.
Alternatives: NATS, Redis Streams, or Kafka when you also need retention. Bridge carefully — Kafka is for durability; Redis pub/sub is ephemeral fan-out. Do not replace the bus with a log "because exactly-once" — Kafka delivery is a different job.
Fan-out topology
- Client connects → any healthy node (sticky optional if state is shared).
- Join room → node tracks socket id locally + registers interest on the bus.
- Event → publish to bus → nodes with members write frames.
- Deploy → close 1001, drain, clients reconnect (handshake, connection draining).
Flow
- 1
1 Client A on node 1
- next2 PUBLISH room 42 to Redis
- 2
2 PUBLISH room 42 to Redis
- next3 Redis pub/sub
- 3
3 Redis pub/sub
- next4 WS node 2
- 4
4 WS node 2
- next5 Frame to client B
- 5
5 Frame to client B
- next6 Slow client: shed or coalesce
- 6
6 Slow client: shed or coalesce
Lesson map
Scaling WebSockets — Sticky Sessions, Fan-out & Backpressure
A WebSocket is pinned to the process that accepted the TCP connection. Sticky affinity reduces reconnect churn, but multi-node rooms need a fan-out bus. Bounded queues, coalesce, and disconnect-slow-clients keep one laggy socket from OOM-ing the node.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1 Client A on node 1"] pub["2 PUBLISH room 42 to Redis"] bus["3 Redis pub/sub"] w2["4 WS node 2"] a -->|1 Client A on node 1| pub pub -->|2 PUBLISH room 42 to Redis| bus bus -->|3 Redis pub/sub to 4 WS node 2| w2
Single-column on purpose. A side-by-side "edge | workers | bus" graph shrinks desktop fonts in the ~664px column.
Backpressure
Producers must not outrun socket buffers.
| Tactic | Use when |
|---|---|
| Bounded per-connection queues | Default — never unbounded |
| Drop / coalesce outdated ticks | Prices, presence — last-value wins |
| Disconnect slow consumers | Protect the node |
| Token bucket per user | Fairness across rooms |
bufferedAmount in browsers | Pause sends when the outbound queue is high |
Without backpressure, one slow client OOMs the node.
Comparative options
| Approach | Wins | Loses |
|---|---|---|
| Sticky only | Simple | Bad failover; no multi-node rooms |
| Sticky + Redis pub/sub | Common chat pattern | Ephemeral; lost on Redis restart unless you also persist |
| Stateless WS + external session + bus | Best HA | Slightly more reconnect logic |
| MQTT broker cluster | Broker handles fan-out; app is a publisher | Different ops — brokers |
| Serverless WS (Cloudflare / API GW) | Managed fan-out | Limits, pricing, still design idempotency |
Horizontal pod autoscaler metric: connections per pod + event loop lag — not only CPU.
Sandbox: tiny fan-out + backpressure (Python)
Twelve ticks, max queue 8 — oldest dropped (coalesce).
Press Run. Snippets must be self-contained — no network, files, or native modules.
Same idea (TypeScript)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Pitfalls
Draw client A on node 1, client B on node 2, Redis in the middle. A sends "hello". Which process writes the frame to B? Now kill Redis. What do in-flight ticks do? Now A is slow and bufferedAmount climbs — what do you drop, and what do you never drop (chat vs prices)?
Interview Q&A
Why sticky for WebSockets?
Answer
The connection lives on one process. Affinity reduces reconnect churn. You still externalize room and presence state so failover and multi-node rooms work. Depth: sticky vs stateless.
Is Redis pub/sub durable?
Answer
No. Messages vanish if there is no subscriber. A Redis restart drops in-flight fan-out. Use Streams or Kafka when you need retention. Depth: Kafka delivery for the durable path.
How does this relate to the LB sticky doc?
Answer
Same affinity tradeoffs. WebSocket is the forced exception for connection lifetime. This page adds the bus, backpressure, and drain. Do not recap why sticky hurts REST checkouts.
L4 vs L7 for WebSocket?
Answer
L4 passes TCP. Path-based routing and auth cookies need L7. Idle timeouts are an L7 (or proxy) concern. Depth: L4 vs L7.
What is bufferedAmount?
Answer
The browser's outbound queued bytes on that WebSocket. Pause or coalesce sends when it is high so you do not pile application queues on top of TCP buffers.
Thundering herd on deploy?
Answer
Stagger 1001 closes. Jitter client reconnect. Warm capacity first. Same family as connection draining.
What do you HPA on?
Answer
Connections per pod and event-loop lag (or p99 write time). CPU lags behind a blocked event loop with 50k idle sockets.
At-least-once fan-out?
Answer
The bus may duplicate on retry. Clients dedupe by event id. Do not pretend Redis pub/sub is exactly-once.
Sticky optional?
Answer
Yes, if every node can deliver via the bus and reconnect state lives outside the process. Sticky then becomes a latency optimization, not a correctness requirement.
When is Kafka the bus?
Answer
When fan-out must also be replayable — analytics, late joiners who need history, audit. Do not pay log overhead for ephemeral cursor moves. Cross-link Kafka topics.
Serverless WebSocket catch?
Answer
Managed fan-out has connection quotas and per-message pricing. You still design idempotency, still drain, still backpressure. The cloud vendor is not a protocol choice.
MQTT instead of a WS farm?
Answer
If the product is topic routing for devices, let a broker cluster fan out and keep the app as a publisher. Depth: brokers. Browsers still often speak WS or MQTT-over-WS at the edge.