HTTP/1.1 — Keep-Alive, Pipelining & Head-of-Line Blocking
HTTP/1.1 made persistent connections the default so TCP and TLS are amortized. Pipelining failed in practice because responses stay in order, which is application head-of-line blocking.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
What did keep-alive change from HTTP/1.0?
Answer
HTTP/1.1 reuses the connection unless the peer sends Connection close. HTTP/1.0 needed an explicit keep-alive token and implementations disagreed.
L2
Why is pipelining unused?
Answer
Responses are ordered, so a slow first response blocks later ones. Intermediaries mishandled the pipeline. Browsers disabled it.
L3
How did browsers work around the concurrency limit?
Answer
About six connections per origin, plus domain sharding onto extra hostnames.
L4
What is the safe pipelining setting?
Answer
One in-flight request per connection unless you own both peers and have measured the proxies in between.
L5
How do you see a keep-alive bug in production?
Answer
Connect time spikes, resets after idle, and pool wait. Compare the client idle timer with the proxy idle timer.
L6
Is head-of-line blocking unique to HTTP?
Answer
No. Any ordered multiplex on one FIFO can block, including TCP. QUIC moves recovery onto each stream.
L7
Should a new API still accept HTTP/1.1?
Answer
Yes as a fallback and for debug. Performance paths should negotiate HTTP/2 or HTTP/3.
Failure modes
Idle timeout mismatch
The proxy closes an idle socket and the client still thinks it is live. The next write gets a reset.
Too few connections
Chatty pages queue behind one large response because pipelining is off and the pool is tiny.
Too many connections
A deploy or a retry opens a burst of sockets and exhausts file descriptors.
Slowloris-class stalls
Clients drip request headers and pin a worker per socket. Header timeouts and a per-IP cap are the H1-shaped mitigations.
Misconceptions
Connection keep-alive means responses may arrive out of order.
Keep-alive only reuses the socket. Order is still one response after the previous one.
Pipelining is how browsers got parallel downloads.
Browsers used multiple connections. Pipelining was specified and then abandoned.
Disabling keep-alive makes timeouts simpler.
It makes every call pay TCP and TLS. Fix the idle timers instead.
Interviewer traps
Explaining HTTP/2 frames when the question was the HTTP/1.1 FIFO.
One sentence that streams replace pipelining, then stay on order, sharding, and idle resets.
Redesigning the load balancer while naming the idle timer.
State the timer inequality and point at the L4 versus L7 page for proxy design.
Design scenario
Same prompt for every reader.
Requirements
Small JSON calls must not wait behind the image. A pause must not surface as a connection reset. The server must not hold idle sockets for many minutes.
Traffic / scale
A few hundred browsers, each opening the historical per-origin connection cap.
Latency
A small JSON call starts without waiting for the image body.
Consistency
A reset is retried only for the idempotent GETs.
Availability
Header timeouts drop slowloris-style sockets without dropping healthy keep-alive clients.
Failure assumptions
- The proxy idle timeout is shorter than the client idle timeout.
- Someone sets pipelining above 1 in front of an old proxy.
Constraints
- Do not shard the API hostname to raise the connection cap.
- Do not turn keep-alive off as the fix.
Prompt
A page loads one large image and twenty small JSON calls through an HTTP/1.1 client. Users on the office network see resets after they pause for a minute.
API
Which requests may share a socket, and which get their own connection?
Data
What timers do you record on the client and on the proxy?
Architecture
When do you stop tuning HTTP/1.1 and negotiate HTTP/2?
How HTTP/1.1 tries to do two things at once
Prefer
Several keep-alive connections, one request in flight on each
This is what browsers actually shipped. A slow image occupies one socket. A JSON call can use another.
- Each socket still pays one handshake, then reuses it.
- HOL is per connection, not across the pool.
- The cap is small, so a chatty page still queues.
Alternative
Pipelining on one socket
The client writes request two before response one finishes. The server must answer in order. A slow first response blocks the second, and many proxies got this wrong.
- Specification allowed it. Browsers turned it off.
- A single stalled response stalls the pipeline.
- HTTP/2 streams exist because this approach failed.
Reuse the socket, do not reorder the responses
One handshake, then a FIFO. Concurrency comes from more sockets, until HTTP/2.
- 1
Pay TCP and TLS once
The first request on a fresh socket pays the handshake. HTTP/1.1 then keeps the connection unless a peer sends Connection close. - 2
Send the next request after the response
With pipelining left at 1, the client waits. That is slower than streams and safer than a broken pipeline. - 3
Open a few more sockets
Browsers used about six per origin so one large body did not block every call. Do not raise that with extra hostnames on HTTP/2. - 4
Close before the proxy does
Client idle shorter than proxy idle. Otherwise the pool writes onto a socket the proxy already reset.
Overview
HTTP/1.1 made persistent connections the default. Without them every request pays a TCP handshake and, on HTTPS, a TLS handshake. With them the socket stays up until an idle timeout or an explicit close.
That win is not concurrency. A persistent HTTP/1.1 connection is still one response at a time. The next lesson is the version that interleaves streams.
Keep-alive
HTTP/1.0 needed Connection: Keep-Alive, and stacks disagreed. HTTP/1.1 persistence is the default. A peer that wants the old behavior sends Connection: close.
The response is delimited by Content-Length or by Transfer-Encoding: chunked. If the client cannot find the end of the body, it cannot safely start the next request on that socket. Trailers are uncommon on HTTP/1.1 APIs. They show up later, on HTTP/2, which is why gRPC is not an HTTP/1.1 story.
Flow
- 1
Dial TCP and TLS
- nextGET /a
- 2
GET /a
- next200 for /a
- 3
200 for /a
- nextReuse the socket
- 4
Reuse the socket
- nextGET /b
- 5
GET /b
- next200 for /b
- 6
200 for /b
- nextIdle timeout closes
- 7
Idle timeout closes
Lesson map
HTTP/1.1 — Keep-Alive, Pipelining & Head-of-Line Blocking
HTTP/1.1 made persistent connections the default so TCP and TLS are amortized. Pipelining failed in practice because responses stay in order, which is application head-of-line blocking.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["Dial TCP and TLS"] b["GET /a"] c["200 for /a"] d["Reuse the socket"] a -->|Dial TCP and TLS to GET /a| b b -->|GET /a to 200 for /a| c c -->|200 for /a to Reuse the socket| d
Pros: fewer handshakes, fewer ephemeral ports, less CPU on a TLS site.
Cons: idle sockets hold memory and file descriptors. A NAT or proxy idle timeout can kill a socket the client still lists as live. A slow client can pin a worker, which is the slowloris shape.
| Approach | In flight per connection | HOL | What shipped |
|---|---|---|---|
| Sequential keep-alive | 1 | Each call waits | Safe HTTP/1.1 default |
| Pipelining | Many writes, ordered reads | High | Disabled in browsers |
| Several connections | 1 each, about six origins | Per socket | The browser workaround |
| HTTP/2 streams | Many | TCP still, not the app | Next lesson |
Pipelining
Pipelining writes request 2 before response 1 is done. Responses must come back in the same order. A large or slow first response blocks every later response on that connection. That is application head-of-line blocking: an earlier message on a FIFO stops an independent later message.
Proxies that buffered badly, reordered, or failed the connection made pipelining unsafe on the open internet. Chrome never turned it on for general browsing. Firefox removed it. Node's undici defaults pipelining to 1, which means "do not pipeline."
Leave it at 1 unless you control the client, the server, and every proxy between them.
Head-of-line blocking on one connection
Flow
- 1
One H1 connection
- nextLarge response 1
- 2
Large response 1
- nextJSON request waits
- 3
JSON request waits
- nextResponse 1 finishes
- 4
Response 1 finishes
- nextResponse 2 starts
- 5
Response 2 starts
Interview line: HOL is when an earlier message prevents a later independent message from making progress on a shared FIFO. Multiple connections shrink the blast radius to one socket. They do not remove the FIFO. TCP HOL under loss is the HTTP/2 problem, covered next. Stream-level recovery is the QUIC problem after that.
Domain sharding
Sites split assets across a.example.com and b.example.com so the browser's per-origin cap applied twice. The cost was extra DNS, extra TLS, and worse cookie and cache locality. HTTP/2 and HTTP/3 made that split harmful. Consolidate origins. Cache-key and hierarchy questions belong on CDN cache hierarchy, not here.
Timers that interact with keep-alive
| Timer | Who sets it | Too short | Too long |
|---|---|---|---|
| Client idle | HTTP client | Needless reconnect and TLS | Stale sockets after NAT |
| Server idle | App or reverse proxy | Clients flap | File descriptors pile up |
| Proxy idle | The load balancer | Reset after a pause | Slowloris has more time |
| Request | Client | False timeouts under load | A hung call holds a worker |
Prefer client idle shorter than proxy idle. The client closes first and opens a fresh socket for the next call, instead of writing into a reset. Document the numbers either way. The mismatch is the usual "curl works, the pool does not" bug. How that proxy is built is L4 versus L7. This page only owns the timer relationship.
Slowloris, the HTTP/1.1 shape
An attacker opens many connections and drips headers slowly. Each socket pins a worker. Mitigations that belong on this page: a request-header timeout, a cap on concurrent connections per client address, and an L7 front door that can drop the drip. HTTP/2 SETTINGS change the shape of the attack. They do not replace application rate limits.
Expect: 100-continue and a proxy that buffers chunked bodies forever are the rarer cousins. Name them if the interviewer has seen a large upload stall.
When HTTP/1.1 is still the right answer
Prefer it when you are debugging with a raw socket, talking to an appliance that only understands text, or sitting behind a middlebox that cannot speak HTTP/2. Prefer a newer version for browsers, mobile APIs, and any fan-out where many responses must overlap.
A minimal request that asks for persistence:
GET /v1/items HTTP/1.1
Host: api.example.com
Connection: keep-alive
Accept: application/jsonThe Connection token is optional on HTTP/1.1. Sending close is the explicit opt-out.
Sandbox: plan reuse without a network
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Production clients set the same idea as numbers: undici pipelining: 1, a small connections cap, and keepAliveTimeout shorter than the proxy idle timeout. The sandbox only shows the plan.
The client idle timeout is 60 seconds. The proxy idle timeout is 30 seconds. A user pauses for 45 seconds and then the client sends the next GET on the pooled socket. What does the client observe, and which timer do you change?
Interview Q&A
What did keep-alive change versus HTTP/1.0?
Answer
HTTP/1.1 keeps the TCP connection open for the next request unless a peer sends Connection: close. HTTP/1.0 required an explicit keep-alive token, and stacks did not agree. The practical win is amortizing TCP and TLS across calls.
Why is pipelining rarely used?
Answer
Responses have to return in order, so a slow first response blocks later ones on that connection. Intermediaries mishandled pipelines. Browsers disabled the feature. Independent streams, which HTTP/2 added, are the replacement.
How did browsers work around HTTP/1.1 concurrency limits?
Answer
They opened about six connections per origin, and sites sharded assets onto more hostnames to multiply that cap. Sharding costs DNS and TLS. On HTTP/2 and HTTP/3 it is usually the wrong move.
How do you detect keep-alive trouble in production?
Answer
Watch connect time, resets after an idle gap, and time spent waiting for a pool socket. Compare the client idle timeout with the proxy idle timeout. Curl hides the bug because it often uses one fresh connection and exits.
Is head-of-line blocking unique to HTTP?
Answer
No. Any protocol that multiplexes independent messages onto one ordered FIFO can block this way. TCP itself does it under loss. HTTP/3 moves retransmission onto QUIC streams so one loss does not freeze the siblings.
Should new APIs still speak HTTP/1.1?
Answer
Yes, as a fallback and so a human can debug with a raw client. Negotiate HTTP/2 or HTTP/3 on the performance path. Do not make text framing the steady state for a chatty browser.
What does a chunked response change?
Answer
The client learns the end of the body from the chunk framing instead of Content-Length. Until that end is known, the next request on the same socket is not safe. A proxy that buffers chunks forever looks like a hung keep-alive connection.