Linux I/O Models & Event Loops
Studies in this cluster, in series order. Each one keeps its own URL.
Operating systems
Virtual memory, paging, the page cache, and Linux I/O models: blocking calls, epoll, io_uring, and event-loop backpressure.
Linux I/O Models & Event Loops
6 studies- 1.Linux I/O Models - Blocking, Non-Blocking, epoll, io_uring & Event LoopsEvery server spends most of its life **waiting**: for a client to send bytes, for a disk, for a downstream service. An I/O model is simply the answer to "what does a thread do while it waits?". With **blocking I/O** the thread sleeps inside `read()` and you need one thread per in-flight connection. With **non-blocking I/O + readiness multiplexing** (`select`/`poll`/`epoll`/`kqueue`) one thread asks the kernel "which of my 50,000 sockets are ready?" and only touches those. With **completion-based I/O** (`io_uring` on Linux, IOCP on Windows) you hand the kernel the whole operation and buffer and collect results later. Runtimes such as Node/libuv, Nginx, Netty, Tokio and the Go netpoller are all built from these pieces; knowing which one sits under your framework explains its scaling limits, its failure modes ("someone blocked the event loop") and its tuning knobs.
- 2.File Descriptors & Non-Blocking I/O - Syscalls, EAGAIN, Partial Writes & FramingOn Unix, a socket, pipe, file, eventfd or timerfd is a **file descriptor (fd)**: a small integer indexing a per-process table that points at a kernel object. Every byte you move crosses the user/kernel boundary through a **syscall** (`read`, `write`, `recv`, `send`, `accept`). In **blocking** mode, `read()` on an empty socket puts your thread to sleep until data arrives. With `O_NONBLOCK` set (via `fcntl` or `SOCK_NONBLOCK`), the same call returns immediately with `-1` and `errno = EAGAIN` (also spelled `EWOULDBLOCK`). Non-blocking mode alone is useless (you'd spin); it's the building block that readiness APIs like epoll sit on. Two consequences every server must handle: **partial writes** (the kernel accepted only part of your buffer) and **arbitrary read boundaries** (TCP is a byte stream, so you need framing).
- 3.select vs poll vs epoll vs kqueue - Readiness, Level vs Edge Triggering & Thundering HerdsReadiness multiplexing lets one thread wait on many fds. **`select`** (1983, BSD) passes bitmaps of fds into the kernel on every call, is capped at `FD_SETSIZE` (1024 on glibc), and both kernel and app scan all of them. **`poll`** removes the cap with an array of `pollfd`, but still copies and scans the whole set every call: O(watched). **`epoll`** (Linux 2.6) keeps the interest set inside the kernel (`epoll_ctl` once per fd) and `epoll_wait` returns only the ready ones: O(ready). **`kqueue`** (FreeBSD/macOS) is the BSD equivalent and also handles timers, signals, process and file events through one API. On top of that you choose **level-triggered** (keep telling me while data remains, the default and the safest) or **edge-triggered** (`EPOLLET`, tell me once per change, and you must drain to `EAGAIN`). At multi-thread scale you also have to handle the **thundering herd** and accept distribution.
- 4.io_uring & Zero-Copy - Submission/Completion Rings, Batching, sendfile & splice**io_uring** (Linux 5.1, 2019, by Jens Axboe) is a completion-based I/O interface. The app and kernel share two ring buffers in memory: the **submission queue (SQ)** where you write operation descriptors (SQEs: read this fd into this buffer, accept, send, fsync, open...), and the **completion queue (CQ)** where the kernel posts results (CQEs with your `user_data` tag and a result code). One `io_uring_enter()` syscall can submit hundreds of operations; with **SQPOLL** a kernel thread polls the SQ and you may need no syscall at all. Unlike epoll it works for **regular files** as well as sockets, and supports registered files and fixed buffers to cut per-op overhead. The cost: a large, fast-moving kernel attack surface, so many platforms restrict it. Next to it sit the classic **zero-copy** tools: `sendfile`, `splice`, `MSG_ZEROCOPY`, and `mmap`, which reduce copies rather than syscalls.
- 5.Reactor vs Proactor - How libuv, Nginx, Netty, Tokio & the Go Netpoller WorkTwo design patterns turn OS I/O primitives into a programming model. A **reactor** waits for *readiness* (epoll, kqueue, select) and dispatches to a handler that then performs the non-blocking read or write itself: "socket 7 is readable, go read it". A **proactor** starts an *operation* and dispatches the *completion*: "your read on socket 7 finished, here are 4,096 bytes". Linux servers are mostly reactors because epoll is readiness-based; Windows IOCP and Linux io_uring are completion-based, so runtimes built on them are proactors. Every major runtime is some arrangement of these: **libuv** (Node) is a single-threaded reactor plus a threadpool that fakes completions for files; **Nginx** runs one reactor per worker process; **Netty** runs one reactor per event-loop thread; **Tokio** runs reactor-driven futures on a work-stealing pool; **Go's netpoller** hides a reactor under goroutines so your code looks blocking.
- 6.Thread-per-Connection vs Event Loop vs Async Tasks - Blocking the Loop, C10K & BackpressureThe concurrency model you choose sets where your server breaks. **Thread-per-connection** breaks on memory and context switches as concurrent connections grow, and on pool exhaustion when a dependency slows down. **Event loops** break when anything blocks the loop: a CPU-heavy JSON parse, a sync file read, a blocking DNS lookup, a regex with catastrophic backtracking. All one loop's connections stall together and timers (including health checks) fire late. **Async tasks and green threads** (Tokio, asyncio, goroutines, virtual threads) remove the per-connection thread cost but not the need for **backpressure**: if you accept or read faster than you can process or write, queues and buffers grow until memory runs out. The senior answer is always the same three moves: keep the I/O path non-blocking, offload CPU work to a bounded pool, and propagate backpressure (bounded queues, `write()`-returns-false / `drain`, stop reading, shed load at the edge).