Operating systems
Part 4 of 6 · Linux I/O Models & Event Loopsio_uring & Zero-Copy - Submission/Completion Rings, Batching, sendfile & splice
**io_uring** (Linux 5.1, 2019, by Jens Axboe) is a completion-based I/O interface. The app and kernel share two ring buffers in memory: the **submission queue (SQ)** where you write operation descriptors (SQEs: read this fd into this buffer, accept, send, fsync, open...), and the **completion queue (CQ)** where the kernel posts results (CQEs with your `user_data` tag and a result code). One `io_uring_enter()` syscall can submit hundreds of operations; with **SQPOLL** a kernel thread polls the SQ and you may need no syscall at all. Unlike epoll it works for **regular files** as well as sockets, and supports registered files and fixed buffers to cut per-op overhead. The cost: a large, fast-moving kernel attack surface, so many platforms restrict it. Next to it sit the classic **zero-copy** tools: `sendfile`, `splice`, `MSG_ZEROCOPY`, and `mmap`, which reduce copies rather than syscalls.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What do the two rings hold?
Answer
The submission queue holds SQEs, the operations you want. The completion queue holds CQEs, the results.
L2
What is io_uring_enter for?
Answer
It submits the queued SQEs and can also wait for completions. With SQPOLL a kernel thread polls the SQ and you may skip the syscall.
L3
Why are completions out of order?
Answer
Ops run at the same time on different sockets and devices. They finish in device order, not submit order.
L4
What is user_data?
Answer
A 64-bit tag on the SQE, usually a request id, copied onto the CQE so you can find the request.
L5
Why can epoll not replace io_uring for files?
Answer
Files always look ready. A cache miss blocks inside read. io_uring can complete buffered and direct file I/O in the background.
L6
What does sendfile avoid?
Answer
The copy from the page cache into a user buffer and back out to the socket. It does not let you modify the bytes.
L7
When do you refuse to depend on io_uring?
Answer
When the container seccomp profile or kernel.io_uring_disabled can turn it off. Feature-detect and keep an epoll path.
Failure modes
io_uring denied by seccomp or sysctl
Docker and containerd seccomp profiles often block the io_uring syscalls, and kernel.io_uring_disabled can turn the API off. A binary with no epoll fallback fails closed.
user_data ignored, completions applied in submit order
A slow disk op and a fast socket op finish out of order. The wrong request is completed.
sendfile assumed to survive TLS
Without kTLS the bytes still visit user space for encryption, and the zero-copy win disappears.
Misconceptions
io_uring is just a faster epoll.
epoll reports readiness. io_uring performs the operation and posts a completion, including for files.
Zero-copy means zero syscalls.
sendfile and splice skip a user-space copy. They are still syscalls unless something else batches them.
MSG_ZEROCOPY is always a win.
Page pinning and the error-queue notification cost more than a copy for small sends.
Interviewer traps
Promising io_uring for a portable web tier.
Say epoll through a mature runtime, and io_uring only for storage or a measured syscall-bound path, with a fallback.
Forgetting kTLS when asked why Kafka uses sendfile.
The broker sendfile path is about page-cache bytes. TLS stays zero-copy only if the kernel encrypts.
Design scenario
Same prompt for every reader.
Requirements
Assign io_uring, sendfile, or epoll-plus-threadpool to each path. Say what you do when the ring is blocked, and what TLS does to sendfile.
Traffic / scale
High-IOPS reads, multi-megabyte file responses, and a container fleet with a default seccomp profile.
Latency
A threadpool hides file latency until it saturates. A denied ring is a startup or first-request failure.
Consistency
Out-of-order CQEs must still complete the request that submitted them.
Availability
No fallback means a seccomp change is an outage. A user-space copy means CPU, not correctness.
Failure assumptions
- The seccomp profile may deny io_uring.
- Completions are not in submit order.
- TLS may terminate in the process.
Constraints
- Do not require io_uring for ordinary socket I/O.
- Do not claim sendfile transforms bytes.
Prompt
A log service on NVMe wants async file reads. A static-file edge copies 4 MiB responses through user space. A platform team denies io_uring in the default container profile after a kernel advisory. Kafka-style consumers serve segments that are already in the page cache, sometimes over TLS.
API
What tag correlates a CQE with a request?
Data
Which call moves page-cache bytes to a socket without a user buffer?
Architecture
Where does this meet the storage-engine and Kafka pages?
Batch the operation, or skip the user-space copy
Prefer
io_uring for storage, sendfile for unchanged file bytes, epoll as the fallback
High-IOPS file workloads want completion. A static file or a Kafka log segment wants sendfile. Sockets on a mature runtime can stay on epoll until a profile says the syscall is the cost.
- One enter can submit the whole batch.
- user_data matches a CQE back to a request.
- If the ring is blocked, epoll plus a threadpool still serves.
Alternative
io_uring everywhere, or read plus send for static files
A container seccomp profile can reject the ring at deploy time. A user-space copy of a multi-megabyte file spends CPU the kernel could have skipped.
- SQPOLL and registered buffers are version-specific.
- Small MSG_ZEROCOPY sends lose to a normal copy.
- TLS breaks sendfile unless the kernel encrypts (kTLS).
Submit a batch, reap in completion order
The last step matches user_data. The ring does not promise submit order.
- 1
Write SQEs into the shared ring
Each entry names the op, the fd, the buffer, and a user_data tag. That write is not a syscall. - 2
Publish the batch
io_uring_enter submits everything queued, or an SQPOLL kernel thread notices the tail move. - 3
The kernel runs the ops
Sockets, the page cache, and NVMe complete independently. - 4
Reap CQEs and match user_data
Completion order is device order. The buffer is already filled. A missing fallback matters when seccomp denies the ring.
io_uring vs epoll vs Linux AIO
| epoll + read/write | Linux AIO (io_submit) | io_uring | |
|---|---|---|---|
| Model | Readiness: "you may read now" | Completion | Completion (plus poll ops for readiness) |
| Syscalls per op | ~2 (wait + read/write) | 1 submit + 1 reap, batched | Batched submit; ~0 with SQPOLL |
| Regular files | Always "ready", blocks on cache miss | Only with O_DIRECT, often silently blocks | Real async for buffered and direct I/O |
| Operation coverage | Any fd syscall you call yourself | Read/write/fsync only | Read, write, accept, connect, send, recv, open, close, statx, fsync, splice, timeouts, linked chains... |
| Buffers | Your buffer at read time | Pinned at submit | Fixed/registered buffers, provided buffer rings |
| Maturity / safety | Very mature | Legacy, limited | Rapidly evolving; frequent security restrictions |
Sequence
- 1
Application → Submission ring
1. write SQEs for read, send, accept with user_data tags
- 2
Application → Kernel
2. io_uring_enter submits whole batch, or SQPOLL thread picks them up
- 3
Kernel → Kernel
3. run ops on sockets, page cache, NVMe in parallel
- 4
Kernel → Completion ring
4. post CQEs in completion order, not submit order
- 5
Application → Completion ring
5. reap CQEs by advancing head, no syscall needed
- 6
Application → Application
6. match user_data to request, buffer already filled
Lesson map
io_uring & Zero-Copy - Submission/Completion Rings, Batching, sendfile & splice
>-
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB app["Application"] sq["Submission ring"] k["Kernel"] cq["Completion ring"] app -->|1. write SQEs for read, send, accept with user_data tags| sq app -->|2. io_uring_enter submits whole batch, or SQPOLL thread picks them up| k k -->|4. post CQEs in completion order, not submit order| cq app -->|5. reap CQEs by advancing head, no syscall needed| cq
The ring model and batching (simulation, runnable)
Real io_uring needs liburing or a native binding (Rust io-uring, Go and Python wrappers vary in maturity). This simulation models the two properties that matter in design discussions: one syscall submits a whole batch, and completions arrive out of order, so you correlate by user_data.
// io_uring's shape, simulated: a submission ring (SQ) and completion ring (CQ)
// shared with the kernel. NOT real io_uring (needs liburing / a native binding);
// this models why batching cuts syscalls and why completions arrive out of order.
type Sqe = { op: "read" | "write" | "accept"; fd: number; userData: number };
type Cqe = { userData: number; res: number };
class Ring {
sq: Sqe[] = []; cq: Cqe[] = [];
enterCalls = 0; // io_uring_enter(2) syscalls
submit(sqe: Sqe) { this.sq.push(sqe); } // just a memory write: no syscall
enter(latency: (s: Sqe) => number) {
this.enterCalls++; // ONE syscall submits the whole batch
const batch = this.sq.splice(0);
// Kernel finishes ops in whatever order the devices complete them.
batch.sort((x, y) => latency(x) - latency(y))
.forEach(s => this.cq.push({ userData: s.userData, res: s.op === "write" ? 512 : 4096 }));
}
reap(): Cqe[] { return this.cq.splice(0); } // reading the CQ is also syscall-free
}
const ring = new Ring();
const ops: Sqe[] = Array.from({ length: 64 }, (_, i) => ({
op: (["read", "write", "accept"] as const)[i % 3], fd: 10 + i, userData: i,
}));
ops.forEach(o => ring.submit(o));
ring.enter(s => (s.fd * 37) % 11); // fake per-op device latency
const done = ring.reap();
console.log(`ops=${ops.length} io_uring_enter calls=${ring.enterCalls} (epoll+read/write would be ~${ops.length + 1})`);
console.log("first 5 completions by user_data:", done.slice(0, 5).map(c => c.userData).join(","));
console.log("completions out of submit order:", done.some((c, i) => c.userData !== i));Output:
ops=64 io_uring_enter calls=1 (epoll+read/write would be ~65)
first 5 completions by user_data: 1,12,23,34,45
completions out of submit order: trueExpectedops=64 io_uring_enter calls=1 (epoll+read/write would be ~65) first 5 completions by user_data: 1,12,23,34,45 completions out of submit order: true
Press Run. Snippets must be self-contained — no network, files, or native modules.
When io_uring pays off, and what happens if you choose something else
- High-IOPS storage engines (databases, log stores, object stores on NVMe): async buffered and direct file I/O without a threadpool. Without it you either block threads on disk or run a big pool and pay context switches. ScyllaDB/Seastar, TigerBeetle, and newer Postgres versions (asynchronous I/O work landing in Postgres 18) moved in this direction.
- Syscall-bound network proxies: batching accept/recv/send can cut CPU per request when syscall overhead dominates (post-Spectre mitigations made syscalls more expensive). Gains are workload-specific; epoll is already good for sockets, and many benchmarks show modest network wins.
- If you choose epoll instead for sockets: mature, debuggable, portable through libuv/mio/Netty, and fast enough for almost everyone. You keep a threadpool for files.
- If you choose io_uring without a fallback: you may discover in production that the container runtime blocks it. Recent Docker and containerd default seccomp profiles deny the io_uring syscalls, Linux 6.6 added the
kernel.io_uring_disabledsysctl, and Google restricted io_uring on ChromeOS, Android apps and its production servers after it accounted for a large share of kernel exploit submissions to its kCTF program. Always feature-detect and fall back to epoll + threadpool.
Zero-copy: fewer copies, not fewer syscalls
A naive static-file server does read(file) (kernel page cache to user buffer) then send(socket) (user buffer to socket buffer): two copies plus two syscalls per chunk. Zero-copy primitives keep the bytes in the kernel.
| Technique | What it avoids | Use when | Catch |
|---|---|---|---|
sendfile(out, in) | User-space copy: page cache straight to socket | Serving static files (Nginx sendfile on, Kafka consumers reading log segments) | No transformation of bytes; TLS needs kTLS to keep it zero-copy |
splice / tee | User copy between any fd and a pipe | Proxies moving bytes socket to socket | Needs a pipe in the middle; Linux-specific |
MSG_ZEROCOPY (Linux 4.14) | Copy into socket buffer for large sends | Large (~10 KB+) sends | Completion notifications via error queue; worse for small sends |
mmap + write | One copy (file mapped into address space) | Random access over large files | Page faults, SIGBUS on truncate, TLB pressure |
| Kernel bypass (DPDK, AF_XDP) | Kernel network stack entirely | Packet processing at line rate | Dedicated cores busy-polling, you re-implement TCP or use a userspace stack |
sendfile on a real kernel (runnable)
# Zero-copy-ish transfer: read()+send() copies through user space;
# os.sendfile() lets the kernel move page-cache pages straight to the socket.
# Linux (and macOS for sockets). Python 3.8+.
import os, socket, tempfile, threading
payload = os.urandom(4 * 1024 * 1024) # 4 MiB "static asset"
with tempfile.NamedTemporaryFile(delete=False) as f:
f.write(payload); path = f.name
def drain(sock, out):
got = bytearray()
while chunk := sock.recv(1 << 16):
got += chunk
out.append(bytes(got))
def transfer(use_sendfile: bool):
a, b = socket.socketpair()
out = []
t = threading.Thread(target=drain, args=(b, out)); t.start()
user_copies = 0
with open(path, "rb") as f:
if use_sendfile:
offset, size = 0, os.path.getsize(path)
while offset < size: # kernel copies file -> socket; no user buffer
offset += os.sendfile(a.fileno(), f.fileno(), offset, size - offset)
else:
while chunk := f.read(1 << 16): # kernel -> user buffer (copy 1)
user_copies += len(chunk)
a.sendall(chunk) # user buffer -> kernel socket (copy 2)
a.shutdown(socket.SHUT_WR); t.join(); a.close(); b.close()
return out[0] == payload, user_copies
ok1, c1 = transfer(use_sendfile=False)
ok2, c2 = transfer(use_sendfile=True)
os.unlink(path)
print(f"read+send : intact={ok1} bytes copied through user space={c1}")
print(f"sendfile : intact={ok2} bytes copied through user space={c2}")Output:
read+send : intact=True bytes copied through user space=4194304
sendfile : intact=True bytes copied through user space=0Pros and cons
| Approach | Pros | Cons |
|---|---|---|
| io_uring | Batched/zero syscalls, true async files, broad op coverage, linked ops | Linux 5.1+ (features vary by version), security restrictions, harder debugging |
| epoll + threadpool | Mature, portable via libraries, well understood | Files need threads; ~2 syscalls per op |
| sendfile / splice | Big CPU savings for byte shoveling | Only when you don't transform bytes (or with kTLS) |
| Kernel bypass | Lowest latency, highest packet rate | Burns dedicated cores, huge operational complexity |
Interview Q&A
How is io_uring different from epoll?
Answer
epoll is readiness: it tells you a socket is ready and you still do read/write syscalls. io_uring is completion: you queue the operation with its buffer into a shared ring, the kernel performs it and posts a completion, and many operations share one io_uring_enter (or none with SQPOLL). It also works for regular files, which epoll can't.
Why do completions come back out of order and how do you handle it?
Answer
Operations run concurrently on different devices and sockets, so they finish in device order. Each SQE carries a 64-bit user_data, usually a pointer or request id, which you use to find the request when the CQE arrives. Linked SQEs (IOSQE_IO_LINK) enforce ordering when you need it.
Would you adopt io_uring for a new web service?
Answer
Usually not directly. For sockets, epoll via a mature runtime is fast and well understood, and many hosted container environments block io_uring. I'd consider it for storage-heavy components, behind a runtime that feature-detects and falls back.
What does sendfile save, and why does Kafka care?
Answer
It avoids copying file bytes into user space and back, so the kernel moves page-cache pages to the socket directly. Kafka consumers read log segments that are already in page cache, so sendfile lets brokers serve them with little CPU. With TLS, that only stays zero-copy if the kernel does the encryption (kTLS).
When is MSG_ZEROCOPY a bad idea?
Answer
For small sends: page pinning and completion notification overhead exceed the copy cost. The kernel docs suggest it pays off for larger writes.
What are fixed buffers and registered files for?
Answer
They let the kernel pin buffers and file references once, instead of mapping them on every operation. The win shows up when per-op overhead dominates, which is the same regime as batching.
What does IOSQE_IO_LINK do?
Answer
It chains SQEs so the next operation starts only after the previous one succeeds. Use it when order matters. Independent ops should stay unlinked so devices can complete in parallel.
How is Linux AIO different from io_uring?
Answer
io_submit is an older completion API that really only covers read, write, and fsync, and it often blocks unless you use O_DIRECT. io_uring covers accept, connect, send, recv, open, close, statx, splice, and more, including buffered file I/O.
Check yourself
List 64 mixed reads and writes. Mark which ones a single io_uring_enter submits, and which user_data values can legally arrive first. Then say what your process does if that syscall returns EPERM.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: Virtual memory, storage engines, Kafka, HTTP, TLS, and QUIC.
Go Deeper
- io_uring(7) Linux manual page
- io_uring_setup(2) Linux manual page (SQPOLL, ring sizes)
- Lord of the io_uring (unixism.net guide)
- LWN: Ringing in a new asynchronous I/O API
- Google Security Blog: Learnings from kCTF VRP's 42 Linux kernel exploits submissions
- sendfile(2) Linux manual page
- splice(2) Linux manual page
- Linux kernel docs: MSG_ZEROCOPY