Operating systems
Part 2 of 6 · Linux I/O Models & Event LoopsFile Descriptors & Non-Blocking I/O - Syscalls, EAGAIN, Partial Writes & Framing
On Unix, a socket, pipe, file, eventfd or timerfd is a **file descriptor (fd)**: a small integer indexing a per-process table that points at a kernel object. Every byte you move crosses the user/kernel boundary through a **syscall** (`read`, `write`, `recv`, `send`, `accept`). In **blocking** mode, `read()` on an empty socket puts your thread to sleep until data arrives. With `O_NONBLOCK` set (via `fcntl` or `SOCK_NONBLOCK`), the same call returns immediately with `-1` and `errno = EAGAIN` (also spelled `EWOULDBLOCK`). Non-blocking mode alone is useless (you'd spin); it's the building block that readiness APIs like epoll sit on. Two consequences every server must handle: **partial writes** (the kernel accepted only part of your buffer) and **arbitrary read boundaries** (TCP is a byte stream, so you need framing).
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What is a file descriptor?
Answer
A per-process integer index into the kernel fd table, pointing at an open file description and then at a socket, pipe, file, or eventfd.
L2
What does EAGAIN mean?
Answer
The call would block. Wait for readable, writable, or the next accept. It is normal control flow.
L3
What is a short write?
Answer
send copied fewer bytes than you passed because the send buffer filled. The rest stays in your outbound buffer.
L4
Why does TCP need framing?
Answer
It is a byte stream. One send can arrive as many reads, and several sends can arrive as one read.
L5
Why does O_NONBLOCK not make file reads asynchronous?
Answer
Regular files always look ready. A page-cache miss blocks inside read. Use a threadpool or io_uring.
L6
When do you arm EPOLLOUT?
Answer
Only while you still have bytes to send. A level-triggered socket is almost always writable, so a permanent EPOLLOUT spins.
L7
What do dup and fork share?
Answer
The open file description. Offset and O_NONBLOCK can be shared across processes, which surprises epoll interest lists.
Failure modes
Short write dropped after EAGAIN
send returned a short count, the rest was discarded, and the client waits forever for bytes that will never arrive.
Busy loop on EAGAIN
The server retries recv without waiting for readiness and burns a core.
Fd leak on the error path
Connections are not closed, the process hits EMFILE, and new accepts fail.
Misconceptions
EAGAIN is a connection error.
It means try again when the fd is ready. Close the socket only on a real error or EOF.
One recv equals one request.
That happens to work on localhost. Segmentation and Nagle coalescing split and merge messages.
Non-blocking without a multiplexer is a strategy.
It is a spin. Readiness is what makes O_NONBLOCK useful.
Interviewer traps
Saying you will 'just use sendall' on a non-blocking socket.
sendall loops until the buffer drains. On a non-blocking fd that loop returns as soon as the kernel returns EAGAIN.
Treating a file fd like a socket fd.
O_NONBLOCK does not make disk reads return EAGAIN on a cache miss.
Design scenario
Same prompt for every reader.
Requirements
Separate the short-write bug, the EAGAIN spin, and the fd leak. Say what the framer must do when a read ends inside the length header.
Traffic / scale
Many connections, large responses, and occasional idle periods.
Latency
Hung clients wait on bytes that were dropped. The spin shows up as a hot core with no useful I/O.
Consistency
Frames must reassemble across arbitrary read cuts. Writes must eventually send every accepted byte.
Availability
EMFILE stops new connections even while old ones are still open.
Failure assumptions
- The send buffer can fill.
- Reads can split a frame.
- Error paths can skip close.
Constraints
- Do not spin on EAGAIN.
- Do not assume one read is one message.
Prompt
A proxy reads a length-prefixed request and writes a large response. Under load, clients hang after the first few kilobytes. CPU on the loop is pegged even when traffic is idle, and overnight the process starts failing accept with EMFILE.
API
What do you store per connection after a short send?
Data
What do you do with a 4-byte length that arrives as 1 byte and then 3?
Architecture
Why is this still true on the HTTP and WebSocket pages?
Who keeps the unsent bytes
Prefer
Non-blocking fd, readiness wait, explicit outbound buffer
The call returns EAGAIN or a short count. You keep the tail, arm EPOLLOUT only while bytes remain, and you stop reading when that buffer hits its limit.
- EAGAIN is not an error.
- One read is not one message.
- A full send buffer is backpressure from the peer.
Alternative
Blocking read and a forgotten short write
The thread sleeps inside the syscall, which is fine on its own stack. Inside a loop, that sleep freezes everyone. Dropping the unsent tail hangs the client.
- sendall hides the loop only in blocking mode.
- Retrying recv in a tight loop burns a core.
- A leaked fd becomes EMFILE.
From an empty socket to the rest of the write
The sequence on this page is the whole path, including the short write.
- 1
recv on an empty non-blocking socket
The kernel returns -1 and EAGAIN. The thread does not sleep. - 2
Wait for readiness
Bytes land in the receive buffer. epoll or kqueue wakes the loop. - 3
Copy what is there
recv copies into user space. The chunk is not a message until the framer says so. - 4
Keep the short write
send accepts part of the buffer, then EAGAIN. The remainder waits for writable.
Blocking vs non-blocking, call by call
| Call | Blocking fd, no data / buffer full | Non-blocking fd, no data / buffer full | What you must do |
|---|---|---|---|
read/recv | Thread sleeps until >= 1 byte or EOF | Returns -1, EAGAIN | Wait for readable, then read until EAGAIN |
write/send | Thread sleeps until all bytes fit in the send buffer | Returns a short count, then EAGAIN | Keep leftover bytes, wait for writable (EPOLLOUT) |
accept | Sleeps until a connection is queued | EAGAIN when the backlog is empty | Loop accept until EAGAIN on each wake |
connect | Sleeps through the TCP handshake | Returns EINPROGRESS | Wait for writable, then check SO_ERROR |
read on a regular file | Blocks on disk if not in page cache | Still blocks: files ignore O_NONBLOCK for reads | Threadpool or io_uring for real async file I/O |
Sequence
- 1
App thread → Kernel socket buffer
1. recv() on empty non-blocking socket
- 2
Kernel socket buffer → App thread
2. -1 EAGAIN, returns immediately
- 3
Remote peer → Kernel socket buffer
3. segment arrives, bytes land in receive buffer
- 4
Kernel socket buffer → App thread
4. readiness event via epoll or kqueue
- 5
App thread → Kernel socket buffer
5. recv() copies bytes to user space
- 6
App thread → Kernel socket buffer
6. send() large reply
- 7
Kernel socket buffer → App thread
7. short count, send buffer full
- 8
App thread
keep leftover bytes, wait for writable
- 9
Kernel socket buffer → App thread
8. writable event after peer ACKs drain buffer
- 10
App thread → Kernel socket buffer
9. send() the remainder
Lesson map
File Descriptors & Non-Blocking I/O - Syscalls, EAGAIN, Partial Writes & Framing
>-
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB app["App thread"] k["Kernel socket buffer"] peer["Remote peer"] app -->|1. recv() on empty non-blocking socket| k k -->|2. -1 EAGAIN, returns immediately| app peer -->|3. segment arrives, bytes land in receive buffer| k k -->|4. readiness event via epoll or kqueue| app app -->|5. recv() copies bytes to user space| k app -->|6. send() large reply| k k -->|7. short count, send buffer full| app k -->|8. writable event after peer ACKs drain buffer| app app -->|9. send() the remainder| k
EAGAIN and partial writes on a real kernel (runnable)
# Blocking vs non-blocking fds: EAGAIN on read, and partial writes when the
# kernel socket buffer fills. Linux/macOS, Python 3.8+.
import errno, socket
a, b = socket.socketpair()
b.setblocking(False) # O_NONBLOCK: calls return instead of waiting
# 1) Read with no data: blocking would sleep; non-blocking fails fast with EAGAIN.
try:
b.recv(1024)
except BlockingIOError as e:
print("empty read ->", errno.errorcode[e.errno]) # EAGAIN (aka EWOULDBLOCK)
# 2) Write until the kernel send buffer is full: send() returns a SHORT count,
# then EAGAIN. Real servers must keep the leftover bytes and retry on writable.
b.setsockopt(socket.SOL_SOCKET, socket.SO_SNDBUF, 4096)
chunk = b"x" * 65536
total, partial_writes = 0, 0
while True:
try:
n = b.send(chunk)
total += n
if n < len(chunk):
partial_writes += 1 # kernel accepted only part of our buffer
except BlockingIOError:
break # buffer full: wait for EPOLLOUT/writable
print("buffer full ->", f"accepted {total > 0}, partial_writes>=1: {partial_writes >= 1}")
# 3) The reader drains; now the writer could make progress again.
drained = 0
a.setblocking(False)
while True:
try:
drained += len(a.recv(65536))
except BlockingIOError:
break
print("drained == accepted:", drained == total)Output:
empty read -> EAGAIN
buffer full -> accepted True, partial_writes>=1: True
drained == accepted: TrueWhat the kernel actually does on each call
- Syscall entry: the CPU switches to kernel mode. Costs on the order of 100 ns to a few microseconds depending on CPU and mitigations (Spectre/Meltdown made syscalls noticeably more expensive). This is why batching matters.
- fd lookup: the integer indexes your process's fd table to find the
struct fileand socket. - Copy:
readcopies from the kernel receive buffer into your buffer;writecopies your bytes into the send buffer. TCP then segments, sends and retransmits on its own. - Wait or return: if nothing can be done, blocking mode puts the thread on a wait queue for that socket; non-blocking mode returns
EAGAIN.
Buffer sizes (SO_RCVBUF, SO_SNDBUF, autotuned by net.ipv4.tcp_rmem/tcp_wmem) decide when "full" happens. A full send buffer on your side is the kernel telling you the client (or network) is slower than you: that's backpressure, and ignoring it means buffering unboundedly in user space.
Framing: TCP doesn't preserve message boundaries (runnable)
One send("GET /a") can arrive as three reads, and two sends can arrive as one. Length-prefix, delimiter (\r\n), or self-describing formats (HTTP/2 frames, protobuf with a length header) solve it. The decoder below receives the stream sliced at awkward offsets and still produces whole messages.
// TCP is a byte stream: one write can arrive as many reads, or several writes as one.
// A length-prefixed frame decoder must handle both. Pure TypeScript.
class FrameDecoder {
private buf = new Uint8Array(0);
// Feed any chunk; returns every complete frame found so far.
push(chunk: Uint8Array): string[] {
const merged = new Uint8Array(this.buf.length + chunk.length);
merged.set(this.buf); merged.set(chunk, this.buf.length);
this.buf = merged;
const frames: string[] = [];
// Frame = 4-byte big-endian length + payload.
while (this.buf.length >= 4) {
const len = new DataView(this.buf.buffer, this.buf.byteOffset).getUint32(0);
if (this.buf.length < 4 + len) break; // partial frame: wait for more bytes
frames.push(new TextDecoder().decode(this.buf.subarray(4, 4 + len)));
this.buf = this.buf.slice(4 + len);
}
return frames;
}
}
function encode(msg: string): Uint8Array {
const body = new TextEncoder().encode(msg);
const out = new Uint8Array(4 + body.length);
new DataView(out.buffer).setUint32(0, body.length);
out.set(body, 4);
return out;
}
// Two messages concatenated, then sliced at awkward offsets (like real reads).
const wire = new Uint8Array([...encode("GET /a"), ...encode("GET /bb")]);
const cuts = [3, 7, 12, 15]; // split inside the header and inside the payload
const dec = new FrameDecoder();
let start = 0;
for (const end of [...cuts, wire.length]) {
const got = dec.push(wire.subarray(start, end));
console.log(`read bytes [${start},${end}) -> frames: ${JSON.stringify(got)}`);
start = end;
}Output:
read bytes [0,3) -> frames: []
read bytes [3,7) -> frames: []
read bytes [7,12) -> frames: ["GET /a"]
read bytes [12,15) -> frames: []
read bytes [15,21) -> frames: ["GET /bb"]Expectedread bytes [0,3) -> frames: [] read bytes [3,7) -> frames: [] read bytes [7,12) -> frames: ["GET /a"] read bytes [12,15) -> frames: [] read bytes [15,21) -> frames: ["GET /bb"]
Press Run. Snippets must be self-contained — no network, files, or native modules.
What happens if you get it wrong
- Assuming one read = one message: works on localhost tests, breaks under real network segmentation or Nagle coalescing. Symptoms: truncated JSON, "unexpected end of input", requests merged.
- Ignoring short writes:
send()returned fewer bytes than you passed and you dropped the rest. The client hangs waiting for bytes that never come. In blocking modesendall()loops for you; in non-blocking mode you must queue the remainder. - Busy looping on EAGAIN: retrying
recvin a tight loop burns 100% CPU. Always wait on readiness. - Blocking fd inside an event loop: one
readon a blocking socket (or a blocking DNS lookup) freezes every connection on that loop. - Leaking fds: each connection holds an fd; forget to close on error paths and you hit
EMFILE("too many open files"). Watchulimit -nand/proc/PID/fd. - Unhandled EINTR: a signal can interrupt a blocking call; well-written code retries. (Python 3.5+ retries automatically per PEP 475.)
Pros and cons
| Mode | Pros | Cons |
|---|---|---|
| Blocking fds | Simple, sequential code; kernel does the waiting | One thread per in-flight operation; no way to wait on many fds at once without threads |
| Non-blocking fds + readiness | One thread multiplexes thousands; explicit control of buffers and backpressure | You own state machines, partial I/O, framing and EAGAIN handling |
| Non-blocking without readiness | None in practice | Burns CPU spinning |
Interview Q&A
What is a file descriptor, really?
Answer
A per-process integer index into the kernel's fd table, which points to an open file description (offset, flags) that points to the underlying object: a socket, pipe, file, eventfd. dup and fork share the open file description, which is why offsets and O_NONBLOCK can be shared across processes.
What does EAGAIN mean and what should a server do with it?
Answer
"The operation would block." On read: no data now, wait for readable. On write: the send buffer is full, keep the remaining bytes and wait for writable. On accept: the backlog is empty. It's normal control flow, not an error.
Why doesn't O_NONBLOCK make file reads asynchronous?
Answer
Regular files are always considered ready; if the page isn't cached, read blocks on the disk regardless. That's why libuv uses a threadpool for fs operations and why io_uring was such a big deal for storage.
How do you handle a `send()` that returns fewer bytes than requested?
Answer
Keep an outbound buffer per connection, register interest in writability, and flush the remainder when the fd becomes writable. If that buffer grows past a limit, stop reading from the source (backpressure) or drop the slow client.
Why is framing needed on TCP but not UDP?
Answer
TCP is a reliable byte stream with no message boundaries; UDP delivers datagrams whole or not at all. On TCP you add length prefixes or delimiters and parse incrementally.
What should a server do with EINTR?
Answer
A signal interrupted a blocking call. Retry it. Python 3.5 and later retries automatically for most syscalls (PEP 475). A non-blocking loop that treats every -1 as fatal will drop connections on a timer signal.
How does a full SO_SNDBUF become backpressure?
Answer
The kernel is telling you the peer is slower than the sender. If you copy the refused bytes into an unbounded user-space queue, memory grows until you stall harder. Stop reading the source, or drop the slow client, once the outbound buffer crosses its limit.
Why does connect on a non-blocking socket return EINPROGRESS?
Answer
The handshake has started and has not finished. Wait until the fd is writable, then read SO_ERROR. Zero means the connection is up. A non-zero code is the handshake failure.
Check yourself
Sketch a per-connection struct with an inbound frame buffer and an outbound byte queue. Write the transitions for readable, short write, and outbound-full.
Elsewhere in the library
These pages stay as they are. This lesson only points at them: Virtual memory, HTTP, TLS, and QUIC, WebSocket fan-out.