Operating systems
Part 1 of 6 · Linux I/O Models & Event LoopsLinux I/O Models - Blocking, Non-Blocking, epoll, io_uring & Event Loops
Every server spends most of its life **waiting**: for a client to send bytes, for a disk, for a downstream service. An I/O model is simply the answer to "what does a thread do while it waits?". With **blocking I/O** the thread sleeps inside `read()` and you need one thread per in-flight connection. With **non-blocking I/O + readiness multiplexing** (`select`/`poll`/`epoll`/`kqueue`) one thread asks the kernel "which of my 50,000 sockets are ready?" and only touches those. With **completion-based I/O** (`io_uring` on Linux, IOCP on Windows) you hand the kernel the whole operation and buffer and collect results later. Runtimes such as Node/libuv, Nginx, Netty, Tokio and the Go netpoller are all built from these pieces; knowing which one sits under your framework explains its scaling limits, its failure modes ("someone blocked the event loop") and its tuning knobs.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What does a thread do while it waits in blocking I/O?
Answer
It sleeps inside read or write. You need a thread per in-flight connection.
L2
What does readiness tell you that completion does not?
Answer
Readiness says a call would not block. You still perform the read. Completion says the operation finished and the buffer is filled.
L3
Why is select a poor fit past a thousand sockets?
Answer
FD_SETSIZE is 1024 on glibc, and every wait copies and scans the whole set.
L4
Why does epoll cost O(ready) rather than O(watched)?
Answer
The interest set lives in the kernel. epoll_wait returns only the ready list.
L5
Apply Little's law to 20,000 requests per second at 200 ms.
Answer
In-flight work is 4,000. A thread per request needs 4,000 threads. An event loop needs about one per core, plus a bound.
L6
Why does Node have a threadpool if the loop is single-threaded?
Answer
Regular files always look ready to epoll, and getaddrinfo and some crypto block. libuv runs those off the loop.
L7
When do you still pick a bounded thread pool?
Answer
Moderate concurrency, blocking drivers you cannot replace, or CPU-heavy handlers. Set timeouts and shed load when the pool fills.
Failure modes
One blocking read freezes the loop
A single CPU-heavy or blocking callback stalls every connection on that loop, including health checks.
Thread pool follows a slow dependency
Latency rises, Little's law raises the in-flight count, and new requests queue or time out.
File I/O left on the epoll thread
Files always report ready. The read blocks on a page-cache miss and the loop stops.
Misconceptions
Non-blocking and asynchronous I/O are the same thing.
Non-blocking returns EAGAIN and you still read. Asynchronous completion finishes the operation and fills your buffer.
An event loop removes the need for limits.
Idle sockets are cheap. Unbounded accepts, buffers, and downstream calls are not.
Green threads make blocking syscalls free.
The runtime parks on EAGAIN. A real blocking syscall or a pinned carrier still occupies an OS thread.
Interviewer traps
Describing Nginx as 'just threads, but fewer'.
Say one worker per core sleeps in epoll_wait and only the ready sockets run.
Sizing the pool from average QPS and ignoring latency.
Concurrency is rate times time in system. A slower dependency doubles the threads you need.
Design scenario
Same prompt for every reader.
Requirements
Name the I/O model for each workload. Say what fails if the chat gateway uses a thread per socket, or if the image resize runs on the event-loop thread.
Traffic / scale
200,000 mostly idle sockets, plus a CPU-heavy API, a blocking-driver app, and a modest CRUD service.
Latency
Chat is waiting, not computing. Image latency is CPU. The JDBC app waits inside the driver.
Consistency
Each connection still needs framing, partial writes, and ordered handling on its loop or thread.
Availability
A blocked loop fails health checks for every client on that process. An unbounded thread pool fails when the dependency slows.
Failure assumptions
- One callback may call a blocking driver.
- Latency of a downstream can double.
- Files may be read from the same process as the sockets.
Constraints
- Do not put CPU work on the socket loop.
- Do not give the chat gateway a thread per connection.
Prompt
A chat gateway holds about 200,000 WebSockets with almost no CPU per message. An image API does 80 ms of CPU at a few hundred concurrent requests. A legacy admin app must call a blocking JDBC driver. An internal CRUD service sits around 800 concurrent requests.
API
Which workload stays on a bounded thread pool, and why?
Data
What is in flight at 20,000 requests per second and 200 ms?
Architecture
Where do libuv, Nginx, Tokio, and the Go netpoller sit?
What the waiting thread should do
Prefer
Readiness or completion when connections sit idle
One loop thread sleeps in epoll_wait or reaps a completion ring. Idle sockets cost kernel state, not stacks. CPU work and blocking libraries leave the loop.
- Work per wake-up tracks ready sockets, not open sockets.
- Little's law still sizes in-flight requests. Timeouts and backpressure stay mandatory.
- Files are not 'ready' to epoll. Offload them or use io_uring.
Alternative
A thread per connection with no bound
The code is a straight read and write. It is the right shape at hundreds of concurrent requests, and the wrong shape when a slow dependency multiplies the in-flight count.
- Each idle waiter holds a stack and a scheduler slot.
- Latency doubles, in-flight work doubles, and the pool fills.
- Green threads hide the park, and still need a limit.
Pick the wait, then name the stall
The three branches are the cluster. Sibling pages hold the syscall, the multiplexer, the ring, the runtime, and the overload.
- 1
A request arrives on a socket
The bytes are not here yet. The model is the answer to what the thread does until they are. - 2
Blocking parks an OS thread
read returns only when data arrives. Two hundred idle clients means two hundred threads. - 3
Readiness returns the ready set
epoll or kqueue wakes one loop. The handler reads until EAGAIN and must not block. - 4
Completion returns a filled buffer
io_uring or IOCP runs the operation. The loop reaps the result. The buffer was committed at submit.
The I/O models compared
| Model | What the waiting thread does | Syscalls per request | Concurrency limit | Typical users |
|---|---|---|---|---|
| Blocking, thread-per-connection | Sleeps in read()/write() | ~2 (read, write) | Threads: memory for stacks, context switches, scheduler | Classic Apache prefork/worker, JDBC-era Java, Python WSGI + threads |
| Non-blocking busy polling | Spins calling read() and gets EAGAIN | Unbounded (wasted) | One core burned per loop | Almost never on purpose (only kernel-bypass / DPDK with dedicated cores) |
Readiness multiplexing (select/poll) | Sleeps in select() until any fd is ready | 1 wait + reads, but O(n) scan per wait | ~1k fds (select FD_SETSIZE), slow beyond a few thousand | Old portable code, small tools |
Readiness, scalable (epoll/kqueue) | Sleeps in epoll_wait(); kernel returns only ready fds | 1 wait per batch + reads/writes | Hundreds of thousands of fds per thread | Nginx, Node/libuv, Netty, Redis, Go netpoller, Tokio (mio) |
Completion (io_uring, IOCP) | Submits operations into a ring, reaps completions | Can be ~0 per op (batched or kernel-polled) | Ring sizes, registered buffers | Windows servers (IOCP), high-IOPS storage engines, newer Linux runtimes |
| Green threads / virtual threads over epoll | Code looks blocking; runtime parks the lightweight thread and multiplexes underneath | Same as epoll | Millions of goroutines / virtual threads | Go, Java 21 virtual threads, Erlang BEAM |
Decisions
- 1
1. Request arrives on socket fd
- next2. Who waits for the bytes?
- ?
2. Who waits for the bytes?
- blocking3a. OS thread sleeps in read()
- readiness3b. fd registered with epoll or kqueue
- completion3c. Submit read plus buffer to io_uring or IOCP
- 3
3a. OS thread sleeps in read()
- next4a. Need 1 thread per waiting client
- 4
4a. Need 1 thread per waiting client
- 5
3b. fd registered with epoll or kqueue
- next4b. One loop thread wakes for ready fds only
- 6
4b. One loop thread wakes for ready fds only
- next5b. Handler does non-blocking read until EAGAIN
- 7
5b. Handler does non-blocking read until EAGAIN
- 8
3c. Submit read plus buffer to io_uring or IOCP
- next4c. Kernel fills buffer, posts completion
- 9
4c. Kernel fills buffer, posts completion
- next5c. Loop reaps completion, data already in buffer
- 10
5c. Loop reaps completion, data already in buffer
Lesson map
Linux I/O Models - Blocking, Non-Blocking, epoll, io_uring & Event Loops
>-
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1. Request arrives on socket fd"] b["2. Who waits for the bytes?"] c["3a. OS thread sleeps in read()"] d["3b. fd registered with epoll or kqueue"] e["3c. Submit read plus buffer to io_uring or IOCP"] c2["4a. Need 1 thread per waiting client"] d2["4b. One loop thread wakes for ready fds only"] e2["4c. Kernel fills buffer, posts completion"] d3["5b. Handler does non-blocking read until EAGAIN"] e3["5c. Loop reaps completion, data already in buffer"] a -->|continues| b b -->|blocking| c c -->|continues| c2 b -->|readiness| d d -->|continues| d2 d2 -->|continues| d3 b -->|completion| e e -->|continues| e2 e2 -->|continues| e3
Same workload, two models (runnable)
Two hundred clients, each sending one small message. Thread-per-connection needs two hundred threads to wait; the event loop needs one.
# Same workload, two I/O models: thread-per-connection vs one event-loop thread.
# Runs anywhere with Python 3.8+ (selectors picks epoll on Linux, kqueue on macOS).
import selectors, socket, threading
N = 200 # simulated client connections
def make_pairs(n):
# socketpair() gives us connected sockets without needing a real network.
return [socket.socketpair() for _ in range(n)]
# --- Model A: one OS thread per connection (blocking calls) -----------------
def thread_per_connection():
pairs = make_pairs(N)
done = []
def serve(server_side):
data = server_side.recv(64) # blocks this thread until bytes arrive
server_side.sendall(data.upper()) # blocking write
done.append(1)
threads = [threading.Thread(target=serve, args=(s,)) for _, s in pairs]
for t in threads: t.start()
peak = threading.active_count() # every waiting client pins a thread + stack
for i, (c, _) in enumerate(pairs): c.sendall(f"hi{i}".encode())
for t in threads: t.join()
ok = all(c.recv(64) == f"HI{i}".encode() for i, (c, _) in enumerate(pairs))
for c, s in pairs: c.close(); s.close()
return peak, len(done), ok
# --- Model B: one thread, readiness multiplexing (non-blocking) -------------
def event_loop():
pairs = make_pairs(N)
sel = selectors.DefaultSelector()
for _, s in pairs:
s.setblocking(False) # never let one socket stall the loop
sel.register(s, selectors.EVENT_READ)
for i, (c, _) in enumerate(pairs): c.sendall(f"hi{i}".encode())
handled = 0
while handled < N:
for key, _ in sel.select(timeout=1): # kernel tells us WHICH fds are ready
s = key.fileobj
s.send(s.recv(64).upper()) # ready => this won't block
sel.unregister(s); handled += 1
ok = all(c.recv(64) == f"HI{i}".encode() for i, (c, _) in enumerate(pairs))
for c, s in pairs: c.close(); s.close()
return threading.active_count(), handled, ok, type(sel).__name__
peak, n_a, ok_a = thread_per_connection()
threads_b, n_b, ok_b, sel_name = event_loop()
print(f"thread-per-connection: handled={n_a} correct={ok_a} peak_threads>={min(peak, N)}")
print(f"event loop ({sel_name}): handled={n_b} correct={ok_b} threads={threads_b}")Output:
thread-per-connection: handled=200 correct=True peak_threads>=200
event loop (EpollSelector): handled=200 correct=True threads=1Why the event loop wins for many connections, and what happens if you choose something else
- Memory. A Linux thread reserves an 8 MiB virtual stack by default (
ulimit -s) and commits real pages as it grows, plus kernel structures. Tens of thousands of mostly-idle threads cost gigabytes and make the scheduler work hard. An idle socket in an epoll set costs a few hundred bytes of kernel state plus your per-connection struct. - Context switches. Every wake-up of a blocked thread is a scheduler round trip. An event loop processes many ready fds per wake-up, so the cost is amortized.
- Little's law decides how much concurrency you need. In-flight requests = throughput x latency. At 20,000 req/s with 200 ms of mostly-waiting latency you have 4,000 requests in flight. Thread-per-request needs 4,000 threads; an event loop needs one per core.
- If you choose threads anyway, it works fine at hundreds to low thousands of concurrent requests, the code is simpler, and blocking libraries "just work". The failure mode is a slow downstream: latency rises, in-flight count rises with it (Little's law again), the pool exhausts, and new requests queue or time out. Bound the pool and shed load (see the bulkheads and load-shedding pages).
- If you choose an event loop, the failure mode flips: one CPU-heavy or accidentally blocking callback freezes every connection on that loop. You must keep handlers short, push CPU work to worker pools, and use non-blocking drivers end to end.
- Green threads (Go goroutines, Java virtual threads) give you blocking-style code on top of epoll. The runtime parks the goroutine when a socket isn't ready and resumes it later. You still pay for unbounded concurrency in memory and downstream load, and blocking syscalls or pinned carrier threads (e.g.,
synchronizedblocks in early virtual-thread JDKs) can reduce the benefit.
Choosing a model for a workload (runnable)
// Pick an I/O model from workload shape. Pure TypeScript, runs with tsc + node.
type Workload = {
name: string;
concurrentConns: number; // open connections at once
cpuMsPerRequest: number; // CPU work per request
blockingLibs: boolean; // must call blocking drivers/SDKs?
};
function chooseModel(w: Workload): string {
// Many idle-ish connections + little CPU: readiness event loop wins on memory.
if (w.concurrentConns > 5_000 && w.cpuMsPerRequest < 5 && !w.blockingLibs)
return "event loop (epoll/kqueue) or async runtime";
// Heavy CPU per request: the loop would stall; use a worker pool sized to cores.
if (w.cpuMsPerRequest >= 20)
return "event loop for I/O + worker pool for CPU";
// Blocking libraries you can't replace: threads are simpler and honest.
if (w.blockingLibs)
return "thread pool (bounded) / thread-per-request";
// Modest scale: either works; green threads (Go) give both ergonomics and scale.
return "either; goroutines/virtual threads keep blocking-style code";
}
const workloads: Workload[] = [
{ name: "chat gateway (WebSockets)", concurrentConns: 200_000, cpuMsPerRequest: 0.2, blockingLibs: false },
{ name: "image resize API", concurrentConns: 300, cpuMsPerRequest: 80, blockingLibs: false },
{ name: "legacy JDBC admin app", concurrentConns: 150, cpuMsPerRequest: 3, blockingLibs: true },
{ name: "internal CRUD service", concurrentConns: 800, cpuMsPerRequest: 2, blockingLibs: false },
];
for (const w of workloads) console.log(`${w.name.padEnd(28)} -> ${chooseModel(w)}`);Output:
chat gateway (WebSockets) -> event loop (epoll/kqueue) or async runtime
image resize API -> event loop for I/O + worker pool for CPU
legacy JDBC admin app -> thread pool (bounded) / thread-per-request
internal CRUD service -> either; goroutines/virtual threads keep blocking-style codeExpectedchat gateway (WebSockets) -> event loop (epoll/kqueue) or async runtime image resize API -> event loop for I/O + worker pool for CPU legacy JDBC admin app -> thread pool (bounded) / thread-per-request internal CRUD service -> either; goroutines/virtual threads keep blocking-style code
Press Run. Snippets must be self-contained — no network, files, or native modules.
Where common runtimes sit
| Runtime | Network I/O | File / blocking work | Notes |
|---|---|---|---|
| Node.js (libuv) | epoll / kqueue / IOCP, one loop thread | libuv threadpool (default 4, UV_THREADPOOL_SIZE) for fs, DNS getaddrinfo, some crypto | One loop per process; scale with cluster/workers |
| Nginx | epoll / kqueue, one worker process per core | Optional thread pools for blocking file reads | reuseport spreads accepts across workers |
| Netty (JVM) | NIO selector or native epoll / io_uring transports | Offload to separate executors | EventLoopGroup with one loop per thread |
| Tokio (Rust) | mio over epoll / kqueue / IOCP, work-stealing multi-thread runtime | spawn_blocking pool | tokio-uring is a separate crate |
| Go | netpoller over epoll / kqueue / IOCP inside the runtime | Blocking syscalls hand off the OS thread (M) so other goroutines keep running | You write blocking code; runtime multiplexes |
| Python asyncio | selectors (epoll / kqueue), or Proactor on Windows | run_in_executor thread pool | GIL means CPU work belongs in processes |
The knobs that matter in every model
- Readiness vs completion: readiness tells you "you can read now without blocking"; completion tells you "your read is done". Files are always "ready" to epoll, which is why file I/O goes to threadpools or io_uring.
- Level vs edge triggering: level keeps notifying while data remains; edge notifies once per change, so you must drain to
EAGAIN. - Partial reads and writes: TCP is a byte stream; you always need framing and must keep unsent bytes for the next writable event.
- Backpressure: if you read faster than you can write downstream, buffers grow without bound. Stop reading (or stop accepting) when outbound queues are full.
- Accept distribution: with multiple loops, decide how new connections are spread (
SO_REUSEPORT,EPOLLEXCLUSIVE, a single acceptor thread).
Pros and cons at a glance
| Approach | Pros | Cons |
|---|---|---|
| Thread-per-connection | Simple linear code, works with any blocking library, CPU work is naturally parallel | Memory and context-switch cost per connection, pool exhaustion under slow downstreams |
| Event loop (readiness) | Massive concurrency per core, predictable memory, the standard for proxies and gateways | Callback/async discipline, one blocking call stalls everyone, file I/O needs offload |
| Completion (io_uring / IOCP) | Fewest syscalls, true async file I/O, batching and fixed buffers | Linux-specific API, younger ecosystem, security-hardening concerns (often disabled in containers) |
| Green / virtual threads | Blocking-style code at event-loop scale | Hidden pinning/blocking pitfalls, unbounded concurrency still needs limits |
Interview Q&A
Why can Nginx hold 100k connections on a few cores when a thread-per-connection server can't?
Answer
Nginx registers every socket with epoll and runs one worker per core. A worker sleeps in epoll_wait() and wakes with only the ready sockets, so idle connections cost a few hundred bytes of kernel state, not a thread stack and scheduler slot. Work per wake-up is proportional to ready sockets, not total sockets.
What's the difference between non-blocking and asynchronous I/O?
Answer
Non-blocking means a call returns immediately (EAGAIN) instead of waiting; you still do the read yourself when readiness says so. Asynchronous (completion) I/O means you start the operation and the kernel finishes it and notifies you, with the data already in your buffer. epoll is non-blocking + readiness; io_uring and IOCP are completion-based.
Why does Node.js use a threadpool if it's "single-threaded"?
Answer
The JavaScript runs on one loop thread, but regular file I/O can't be made non-blocking with epoll (files always report ready), and getaddrinfo and some crypto are blocking library calls. libuv runs those on a small threadpool and posts completions back to the loop. Saturating that pool (e.g., many fs calls plus pbkdf2) is a classic latency bug.
When would you still pick thread-per-request?
Answer
Moderate concurrency, blocking-only drivers, CPU-heavy handlers, or a team that benefits from simple stack traces. Bound the pool, set timeouts, and shed load. Java virtual threads and Go goroutines now give that coding style at event-loop scale.
How does Little's law help you size this?
Answer
Concurrency = arrival rate x time in system. If latency doubles because a dependency slows down, your in-flight requests double too. Thread pools exhaust; event loops grow buffers. Either way you need timeouts, limits, and backpressure.
Where does io_uring fit?
Answer
It's Linux's completion API: shared submission and completion rings, batching many ops into one io_uring_enter() (or zero with SQPOLL), real async file I/O, registered buffers. It shines for high-IOPS storage and syscall-bound proxies. Many platforms restrict it for security reasons, so check before relying on it.
Why can a file read block even after you set O_NONBLOCK?
Answer
Regular files always look readable and writable to epoll. If the page is not cached, read blocks on disk anyway. That is why libuv uses a threadpool and why io_uring matters for storage.
What does a full send buffer mean for the model you picked?
Answer
The peer or the network is slower than you. Blocking write sleeps the thread. Non-blocking write returns a short count and then EAGAIN. Buffering the rest in an unbounded user-space queue just moves the stall into memory. Stop reading or shed load.
Check yourself
Take a service you run. Write down concurrent connections, CPU time per request, and whether a driver blocks. Pick the row from the workload snippet and say what the failure looks like if you pick the other row.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: Virtual memory, the JavaScript event loop, Go goroutines, Rust and Tokio, bulkheads.