Operating systems
Part 6 of 6 · Linux I/O Models & Event LoopsThread-per-Connection vs Event Loop vs Async Tasks - Blocking the Loop, C10K & Backpressure
The concurrency model you choose sets where your server breaks. **Thread-per-connection** breaks on memory and context switches as concurrent connections grow, and on pool exhaustion when a dependency slows down. **Event loops** break when anything blocks the loop: a CPU-heavy JSON parse, a sync file read, a blocking DNS lookup, a regex with catastrophic backtracking. All one loop's connections stall together and timers (including health checks) fire late. **Async tasks and green threads** (Tokio, asyncio, goroutines, virtual threads) remove the per-connection thread cost but not the need for **backpressure**: if you accept or read faster than you can process or write, queues and buffers grow until memory runs out. The senior answer is always the same three moves: keep the I/O path non-blocking, offload CPU work to a bounded pool, and propagate backpressure (bounded queues, `write()`-returns-false / `drain`, stop reading, shed load at the edge).
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What resource does each model run out of?
Answer
Threads run out of stacks and pool slots. Event loops run out of time on the loop. Unbounded async runs out of memory.
L2
Why does one slow Node route hurt every route?
Answer
They share one loop. A long callback delays timers and I/O for every connection on it.
L3
What is backpressure?
Answer
A signal that the consumer is slower than the producer. Propagate it. Do not buffer it away.
L4
Where does TCP generate that signal?
Answer
If you stop reading, the receive window fills, the sender's send buffer fills, and send returns EAGAIN or write returns false.
L5
Why is an unbounded queue a bad spike absorber?
Answer
It converts overload into latency and memory, then you do work for clients that already timed out.
L6
What did C10K change, and what did it not change?
Answer
Non-blocking readiness replaced a thread per client. It did not remove the need to bound buffers, fds, and downstream calls.
L7
How do goroutines and virtual threads change the choice?
Answer
Waiting no longer costs an OS thread. You still need semaphores, timeouts, and you still have to avoid pinning the carrier.
Failure modes
Unbounded queue hides overload until OOM
The process accepts faster than it finishes, memory tracks the gap, and the OOM killer is the first hard signal.
Event-loop lag fails every health check
A 300 ms sync callback delays timers. The orchestrator restarts pods and the remaining pods take more load.
One slow dependency owns the thread pool
Every worker blocks in that client. Healthy endpoints queue behind it.
Misconceptions
Slicing work is as good as leaving the loop.
Yielding between 5 ms slices lets timers run. The core is still busy. CPU-heavy work belongs on a worker or another service.
Async automatically means bounded.
gather on 100,000 URLs opens 100,000 sockets unless you add a semaphore.
A thread pool does not need backpressure if it is large.
Little's law grows the in-flight set with latency. The pool still fills, just later.
Interviewer traps
Answering 'add a queue' to overload.
Say bounded queue plus reject or pause, and name the client-visible signal.
Blaming the GIL for loop lag that is a blocking call.
A blocking requests.get or time.sleep inside async def stalls the asyncio loop even before the GIL matters. The GIL page is the CPU-parallelism question.
Design scenario
Same prompt for every reader.
Requirements
Name loop lag, unbounded buffering, and pool exhaustion as three different failures. Say the metric you would add for each.
Traffic / scale
One hot JSON route, a producer faster than its consumer, and a thread pool shared across endpoints.
Latency
Loop lag adds the same delay to every client on the process. Queue wait adds delay only to queued work, until memory runs out.
Consistency
Rejected work must be explicit. Queued work that outlives the client is wasted.
Availability
Restarting a lagged pod concentrates load on the others. An OOM kill drops the queue.
Failure assumptions
- One callback can run for hundreds of milliseconds.
- The downstream can slow down.
- The queue has no max size today.
Constraints
- Do not fix loop lag by growing the queue.
- Do not share one unbounded pool across every dependency.
Prompt
A Node process serves a cheap health route and a route that parses a multi-megabyte JSON body synchronously. p99 rises for both when the heavy route is called. Separately, a worker pulls from a third party into an unbounded in-memory queue, and RSS climbs until the pod is OOMKilled. A thread pool in front of that third party has no timeout.
API
What status do you return when the bounded queue is full?
Data
What does a growing Send-Q mean?
Architecture
Which sibling pages cover bulkheads and admission control?
What absorbs a fast producer
Prefer
A bounded queue, then pause or reject
The producer waits or the edge returns 429 or 503. Admitted work keeps its latency. TCP does the same thing when you stop reading and the window fills.
- Memory stays capped.
- Clients see slowness or a fast refusal, not a timeout after the fact.
- The same idea is drain, isWritable, and max.poll.records.
Alternative
An unbounded buffer in front of a slow downstream
The queue absorbs the spike until it is the incident. Requests outlive their clients, you do work nobody wants, and then the process OOMs.
- Latency grows with the queue.
- A blocked loop makes every endpoint late, not just the slow one.
- Green threads do not remove the need for a semaphore.
The gap has to land somewhere
Unbounded growth is the failure arm. A bound and an early reject are the other two.
- 1
Arrivals outrun the downstream
Little's law says in-flight work rises with latency. Something must hold or refuse that work. - 2
An unbounded queue grows
Memory, GC, and timeouts hit everyone. The crash is late. - 3
A bounded queue pauses the producer
The socket stops reading, the TCP window fills, and the client slows down. - 4
Admission control rejects early
429 or 503 keeps latency stable for the requests you accepted.
The three models under pressure
| Model | Scales to | Breaks when | Symptom | Fix |
|---|---|---|---|---|
| Thread-per-connection / bounded pool | Hundreds to low thousands of concurrent requests | Slow downstream inflates in-flight count (Little's law) | Pool exhausted, queueing, timeouts, thread dumps full of socketRead | Timeouts, bulkheads per dependency, load shedding, async client |
| Event loop (Node, Nginx, Netty) | 100k+ connections per core | Any callback takes long (CPU or blocking call) | Event-loop lag spikes, p99 for all clients jumps, health checks fail | Slice work, worker threads, async drivers, lag monitoring |
| Async tasks / green threads | Millions of tasks | Unbounded spawning; blocking calls on executor threads | Memory growth, scheduler starvation, pinned carriers | Semaphores / bounded queues, spawn_blocking, avoid pinning |
Decisions
- 1
1. Requests arrive faster than downstream drains
- next2. What absorbs the gap?
- ?
2. What absorbs the gap?
- unbounded queue or buffer3a. Memory grows, GC pauses, latency climbs
- bounded queue plus backpressure3b. Producer pauses, socket stops reading
- admission control at edge3c. Reject early with 429 or 503
- 3
3a. Memory grows, GC pauses, latency climbs
- next4a. OOM kill or timeouts for everyone
- 4
4a. OOM kill or timeouts for everyone
- 5
3b. Producer pauses, socket stops reading
- next4b. TCP window fills, client slows down
- 6
4b. TCP window fills, client slows down
- 7
3c. Reject early with 429 or 503
- next4c. Admitted requests keep good latency
- 8
4c. Admitted requests keep good latency
Lesson map
Thread-per-Connection vs Event Loop vs Async Tasks - Blocking the Loop, C10K & Backpressure
>-
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Requests arrive faster than downstream drains"] b["2. What absorbs the gap?"] c["3a. Memory grows, GC pauses, latency climbs"] d["3b. Producer pauses, socket stops reading"] e["3c. Reject early with 429 or 503"] c2["4a. OOM kill or timeouts for everyone"] d2["4b. TCP window fills, client slows down"] e2["4c. Admitted requests keep good latency"] a -->|continues| b b -->|unbounded queue or buffer| c c -->|continues| c2 b -->|bounded queue plus backpressure| d d -->|continues| d2 b -->|admission control at edge| e e -->|continues| e2
Blocking the event loop (runnable)
A "health check" timer should fire after 10 ms. One 300 ms synchronous loop delays it by about 300 ms, and on a real server that's 300 ms of extra latency for every connection on that loop. Slicing the work lets the timer run on time.
// One CPU-heavy callback delays EVERY other client on an event loop.
// Runs with plain node (setTimeout/Date are global). Times are rounded.
function busy(ms: number) { const end = Date.now() + ms; while (Date.now() < end) {} }
function measure(label: string, work: (done: () => void) => void): Promise<void> {
return new Promise((resolve) => {
const scheduled = Date.now();
// A "health check" timer that should fire after 10ms.
setTimeout(() => {
const lag = Date.now() - scheduled - 10;
console.log(`${label.padEnd(26)} timer lag ~${Math.round(lag / 50) * 50}ms (expected ~0)`);
}, 10);
work(() => setTimeout(resolve, 30));
});
}
(async () => {
// BAD: 300ms of synchronous CPU in one tick: the loop can't run timers or I/O.
await measure("blocking 300ms in one tick", (done) => { busy(300); done(); });
// BETTER: slice the work and yield between slices so timers/I/O interleave.
await measure("300ms sliced into 5ms steps", (done) => {
let left = 60;
const step = () => { busy(5); if (--left > 0) setTimeout(step, 0); else done(); };
step();
});
console.log("best: move CPU work to worker_threads / a separate service");
})();Output:
blocking 300ms in one tick timer lag ~300ms (expected ~0)
300ms sliced into 5ms steps timer lag ~0ms (expected ~0)
best: move CPU work to worker_threads / a separate serviceCommon hidden loop-blockers: JSON.parse or JSON.stringify of multi-MB payloads, synchronous fs.*Sync calls, crypto.pbkdf2Sync, big regexes (ReDoS), large array sorts, template rendering, sync logging to a slow disk, and in Python requests or time.sleep inside async def.
Backpressure: bounded vs unbounded (runnable)
# Backpressure: a fast producer and a slow consumer. Without a bound, the queue
# (memory) grows without limit; with a bound, the producer is paused instead.
# asyncio, Python 3.8+.
import asyncio
async def run(maxsize: int) -> int:
q: asyncio.Queue = asyncio.Queue(maxsize=maxsize) # 0 = unbounded
peak = 0
async def producer():
nonlocal peak
for i in range(2000):
await q.put(i) # bounded queue: waits here when full
peak = max(peak, q.qsize())
await q.put(None)
async def consumer():
while (item := await q.get()) is not None:
if item % 50 == 0:
await asyncio.sleep(0.001) # slow downstream (DB, client socket)
await asyncio.gather(producer(), consumer())
return peak
unbounded = asyncio.run(run(0))
bounded = asyncio.run(run(64))
print(f"unbounded queue peak depth: {unbounded} (memory grows with the gap)")
print(f"bounded(64) queue peak depth: {bounded} (producer paused; same as socket write() -> drain)")Output:
unbounded queue peak depth: 2000 (memory grows with the gap)
bounded(64) queue peak depth: 64 (producer paused; same as socket write() -> drain)Expectedunbounded queue peak depth: 2000 (memory grows with the gap) bounded(64) queue peak depth: 64 (producer paused; same as socket write() -> drain)
Press Run. Snippets must be self-contained — no network, files, or native modules.
The same principle shows up at every layer: Node stream.write() returns false when the internal buffer passes highWaterMark, and you wait for 'drain' (or use pipeline()); asyncio StreamWriter.drain(); Netty Channel.isWritable() and write-buffer water marks; gRPC and HTTP/2 flow-control windows; Kafka consumer max.poll.records and pause/resume; and ultimately the TCP receive window, which pushes back on the sender when you stop reading.
C10K to C10M: what changed
- C10K (Dan Kegel, 1999): how do you serve 10,000 concurrent clients on one box? Answer: non-blocking I/O with scalable readiness (epoll, kqueue) instead of a thread or process per client.
- Today: tens or hundreds of thousands of connections per machine are routine for gateways, WebSocket servers and push services. The bottlenecks move to memory per connection (TLS state, buffers), the kernel network stack, and the cost of syscalls.
- C10M thinking: when the kernel itself is the bottleneck, people batch syscalls (io_uring), offload (kTLS, NIC features), or bypass the kernel (DPDK, AF_XDP). Few application teams need this; most need correct backpressure first.
What happens if you choose differently
- Thread pool in front of a slow dependency without timeouts: every thread ends up waiting on the slow call, healthy endpoints are starved too. Use per-dependency bulkheads and timeouts.
- Event loop with one blocking SDK call: everything looks fine at low load; at peak, loop lag spikes, health checks fail, the orchestrator restarts pods, and the remaining pods get more load (cascading failure).
- Async everywhere without limits:
asyncio.gatherover 100k URLs or a goroutine per message with no semaphore opens 100k sockets, hits fd limits or downstream rate limits. Bound concurrency explicitly. - Unbounded in-memory queues "to absorb spikes": they turn a latency problem into a memory problem and hide overload until OOM. Bounded queue + reject or block is honest.
How to measure it in production
- Event-loop lag: Node
perf_hooks.monitorEventLoopDelay()(p99 above a few tens of ms is a red flag), asyncio debug modeslow_callback_duration, Netty pending tasks per loop. - Thread pool saturation: active vs max threads, queue depth, rejected tasks.
- Run-queue and context switches:
vmstat(r,cs),pidstat -w, Goruntime/metricsscheduler latencies. - Socket buffers:
ss -tmifor send-queue backlog per connection; growingSend-Qmeans slow clients. - fd usage:
/proc/PID/fdcount vsulimit -n.
Pros and cons
| Approach | Pros | Cons |
|---|---|---|
| Thread-per-connection | Simple, CPU work parallel by default, readable stacks | Memory per connection, pool exhaustion, context switches |
| Event loop | Huge concurrency, low memory, great for proxies | Fragile to blocking, needs async libraries, single-core per loop |
| Async tasks / green threads | Scale plus readable code | Easy to over-spawn; blocking or pinning pitfalls |
| Explicit backpressure | Stable memory and latency under overload | More code paths (pause, resume, reject), must propagate end to end |
Interview Q&A
Your Node service's p99 jumps for all endpoints when one endpoint is called. What's going on?
Answer
That endpoint is blocking the event loop, likely CPU work (big JSON, sync crypto, regex) or a sync fs call. Every connection on the loop waits. Confirm with event-loop delay metrics and a CPU profile, then move the work to worker_threads, stream or chunk it, or make it a separate service.
What is backpressure and where does it come from in a TCP server?
Answer
It's the signal that a downstream consumer is slower than the producer. In TCP, when the receiver doesn't read, its window fills and the sender's send buffer fills, so send returns EAGAIN or write() returns false. A good server propagates that upstream: stop reading from the source, pause the producer, or reject new work, instead of buffering in memory.
Thread pool or event loop for a service that calls a slow third-party API?
Answer
Either can work if you bound concurrency and set timeouts. A thread pool needs enough threads for throughput x latency and a bulkhead so the slow API can't take all threads. An event loop handles the waiting cheaply but still needs a concurrency limit to avoid flooding the API and growing memory.
How do goroutines or virtual threads change this picture?
Answer
Blocking-style code no longer costs an OS thread per wait, so thread-per-request style scales. You still need limits (semaphores, bounded channels), timeouts and context cancellation, and you need to watch for operations that pin the carrier thread or make real blocking syscalls.
Why is an unbounded queue a bad overload strategy?
Answer
It converts overload into unbounded latency and memory growth. Requests sit in the queue past their client timeout, so you do work nobody is waiting for, then crash. Bounded queues plus rejection keep latency for admitted work predictable.
What does Node stream.write returning false mean?
Answer
The internal buffer passed highWaterMark. Wait for drain, or use pipeline, before writing more. Ignoring false is the unbounded user-space buffer. asyncio StreamWriter.drain and Netty Channel.isWritable are the same signal.
Which signals tell you the loop is late versus the pool is full?
Answer
Loop lag is monitorEventLoopDelay in Node, slow_callback_duration in asyncio debug, and pending tasks per Netty loop. Pool saturation is active versus max threads, queue depth, and rejected tasks. Send-Q growth means a slow peer. fd count versus the rlimit means a leak or unbounded accepts.
What is the C10M step, and who should ignore it?
Answer
When the kernel stack itself dominates, people batch syscalls with io_uring, offload with kTLS, or bypass the kernel. Most services are not there. They still fail from a blocked loop or an unbounded queue first.
Check yourself
Pick a queue in a service you run. Write its max size, what the producer does when it is full, and what the client observes. If the max size is missing, that is the finding.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: bulkheads, load shedding, performance engineering, the Python GIL, the JavaScript event loop.