Connections at Scale — Pooling, Timeouts, Retries & Failure Modes
At scale the HTTP version matters less than the pool, the deadlines, and the retry policy. Wrong defaults cause reset storms, double writes, and tail latency that no cipher suite will fix.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Question ladder
L1
Why not dial a new TCP connection per request?
Answer
The handshake, especially TLS, costs more than a short handler. A pool spreads that cost across many calls.
L2
How does an HTTP/2 pool differ from an HTTP/1.1 pool?
Answer
HTTP/1.1 pools connections and runs one response at a time on each. HTTP/2 and HTTP/3 pool sessions and multiplex streams up to the peer cap.
L3
What is a retry budget?
Answer
A cap on the fraction of calls that may be retried, so a brownout cannot multiply traffic without bound.
L4
How should timeouts be ordered?
Answer
The client total deadline should be shorter than the upstream request timeout, with connect, TLS, and read as slices of that deadline. Client idle should be shorter than proxy idle.
L5
What breaks on a rolling deploy?
Answer
Pools pinned to dying tasks, GOAWAY ignored, DNS caches holding the old address, and readiness flipping before the pool has drained.
L6
When is hedging acceptable?
Answer
For an idempotent read whose duplicate cost is cheap. It is a bad default for writes.
L7
What do you graph?
Answer
Pool wait, in-use and idle counts, connect errors, TLS errors, version mix, RST rate, and refused streams.
Failure modes
Pool exhaustion
Every socket or stream is busy. New calls queue until the deadline. The upstream may be healthy and still look down from the client.
Stale keep-alive
The proxy closed an idle socket. The client's next write gets a reset. Count it as a connect failure and dial again.
Retry storm
Every client retries on the same cadence after a blip. Full jitter and a budget break the lockstep.
GOAWAY race
The server is draining. Clients that open new streams on that session get refused. Move them to a new session.
DNS pin
A long-lived pool keeps a task that scale-in already removed. Recycle connections on a max age so DNS is consulted again.
Misconceptions
A 30 second timeout per phase is conservative and therefore safe.
Connect plus read plus three retries can hold a worker for minutes. The user already left.
HTTP status 200 means a retry would have been harmless.
The dangerous retry is the one you send when you do not know whether the first attempt ran.
Creating a client object per request still pools, because the library is smart.
The pool lives on the client or the agent. A new client each call is a new pool with nothing in it.
Interviewer traps
Designing a circuit-breaker state machine instead of the pool.
Pools and deadlines are the first line. Breakers are a later policy. Stay on caps, timeouts, and idempotence.
Redrawing the load balancer to explain an idle reset.
State the timer inequality and point at the L4 versus L7 page for the proxy itself.
Design scenario
Same prompt for every reader.
Requirements
Catalog reads may retry with jitter. Charges run at most once. A brownout must not multiply traffic. Deploys drain sessions instead of resetting them.
Traffic / scale
A few thousand requests per second, fan-out two upstreams, HTTP/2 to each.
Latency
The caller deadline is 800 ms. Each upstream gets a slice, not a fresh 30 seconds.
Consistency
A timed-out charge is reconciled with an idempotency key, not with a blind second POST.
Availability
When the pool is exhausted the caller fails fast instead of queueing without bound.
Failure assumptions
- The upstream sends GOAWAY during a rollout.
- One percent of calls time out and every client retries immediately.
- DNS updates while the pool still holds the old addresses.
Constraints
- Do not hedge the charge.
- Do not build a new client per request.
Prompt
A service calls payments and a catalog. Payments charges a card. Catalog is a read. After a one-second brownout, both upstreams fall over from retry traffic. Deploys drop in-flight reads.
API
Which calls are retried, and what key makes a charge safe to repeat?
Data
Which pool gauges page you before the error rate does?
Architecture
How does a draining upstream tell you to open a new session?
What the pool is counting
Prefer
Sessions and streams, with a deadline on the call
HTTP/2 and HTTP/3 want few connections and many streams. HTTP/1.1 wants a small socket cap and one response at a time. Both want one budget for the whole RPC.
- Reuse beats a handshake on every short call.
- A cap turns overload into a queue or a fast failure.
- Retries spend the same budget. They do not get a new 30 seconds.
Alternative
One connection per request, or an infinite retry
A client constructed inside the handler never fills a pool. A retry loop without jitter and without an idempotence check multiplies the outage and the writes.
- Handshake CPU tracks QPS.
- A brownout becomes a synchronized reconnect.
- A timeout plus a POST becomes two charges.
Borrow, bound, then maybe retry
The HTTP version changes the thing you borrow. It does not change the safety rule.
- 1
Reuse a live stream or socket
HTTP/1.1 borrows a connection. HTTP/2 and HTTP/3 borrow a stream on a session that is under the peer cap. - 2
Dial only under the cap
If nothing is idle and you are under the max, pay DNS, TCP or QUIC, and TLS. If you are at the max, queue briefly or fail. - 3
Arm one deadline
Connect, handshake, and body read are slices. The caller cancels the whole call when the budget is gone. - 4
Retry with jitter if a copy is safe
GET and a keyed write can back off. A bare POST that charges cannot. A budget caps how much extra traffic retries may add.
Overview
Keep-alive explains why a socket is worth reusing. HTTP/2 explains streams and GOAWAY. TLS explains why the first dial is expensive. This page is what you configure so those facts survive production.
The usual outage is not "we picked the wrong cipher." It is a pool that waits forever, a retry that lines up, or a client that keeps writing to a task that already sent GOAWAY.
Pool units
| Protocol | What you pool | Knobs that show up in incidents |
|---|---|---|
| HTTP/1.1 | TCP connections | Max per host, idle TTL, pipelining left at 1 |
| HTTP/2 and HTTP/3 | Sessions, then streams | Max concurrent streams, session age, GOAWAY handling |
| TLS on those sessions | Resumed tickets inside the pool | Warm sessions, ticket reuse, handshake errors |
Decisions
- 1
Outbound request
- nextIdle stream free?
- ?
Idle stream free?
- YesReuse session
- NoUnder the cap?
- 3
Reuse session
- nextApply the deadline
- ?
Under the cap?
- YesDial and TLS
- NoQueue or shed
- 5
Dial and TLS
- nextApply the deadline
- 6
Queue or shed
- 7
Apply the deadline
- nextSafe to retry?
- ?
Safe to retry?
- YesJitter then retry
- NoReturn the error
- 9
Jitter then retry
- 10
Return the error
Lesson map
Connections at Scale — Pooling, Timeouts, Retries & Failure Modes
At scale the HTTP version matters less than the pool, the deadlines, and the retry policy. Wrong defaults cause reset storms, double writes, and tail latency that no cipher suite will fix.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["Outbound request"] b["Idle stream free?"] c["Reuse session"] d["Under the cap?"] a -->|Outbound request to Idle stream free?| b b -->|Yes| c b -->|No| d
Do not create the client, the agent, or the http.Client inside the request function. The pool is on that object. A new object per call is an empty pool.
Browsers pool for you and prefer HTTP/2 or HTTP/3. You control less there. On a server, you control the numbers, and the defaults are often too patient.
One deadline, several slices
Think in layers, and make the sum fit the caller's budget.
- Connect. DNS plus TCP or QUIC.
- Handshake. Sometimes folded into connect. Still a slice of the same budget. The TLS lesson is why this slice exists.
- Time to first byte and body. A slow server and a slow body are different graphs, same deadline.
- Total deadline. Absolute cancel. When it fires, in-flight reads stop counting as "just a bit longer."
- Idle pool timeout. When an unused socket is closed. Keep it shorter than the proxy idle timeout so you do not write onto a reset. The proxy product is L4 versus L7.
A 30 second connect timeout, a 30 second read timeout, and three retries can hold a worker for minutes. That is not conservative. That is how a slow dependency eats the process. Give each hop a slice of the caller's budget and leave margin for the hops after this one.
Flow
- 1
Connect slice
- nextHandshake slice
- 2
Handshake slice
- nextFirst byte slice
- 3
First byte slice
- nextBody read slice
- 4
Body read slice
- nextCaller deadline fires
- 5
Caller deadline fires
Retries
| Call | Retry by default? | Why |
|---|---|---|
| GET or HEAD | Yes, with backoff and jitter | A second read should not change state. A GET that charges is a broken API. Do not paper over it here. |
| PUT or POST with an idempotency key | Yes, if the server stores the key | The server recognizes the duplicate. The key is part of the API contract. |
| POST that charges or creates | No | A timeout means you do not know if the first attempt ran. |
| 408, 429, 503 | Often | Honor Retry-After when it is present. |
| 400, 401, 403, 404 | No | The next attempt will fail the same way. |
| TCP reset in the middle of a POST | Only with a key | You cannot see whether the server committed. |
Full jitter sleeps a random value between zero and min(cap, base * 2^attempt). Equal retries without randomness become a second traffic spike. A retry budget caps the share of calls allowed to try again, so a 20 percent failure rate cannot become a 60 percent load increase.
Hedging, meaning a second attempt in parallel before the first finishes, is acceptable for a cheap idempotent read. It is a good way to double-charge if you point it at a write.
Failure catalog
- Pool exhaustion. In-use equals the cap, wait time climbs, and the handler is fine. Shed or fail the wait. Do not add an unbounded queue in front of a saturated upstream.
- Stale keep-alive. The load balancer closed the idle socket. The next request gets a reset. Treat that reset as "dial again," once, for an idempotent call. Fix the timer inequality so it is rare.
- Retry storm. Synchronized backoff after a blip. Jitter plus a budget.
- GOAWAY race. The upstream is draining. Stop opening streams on that session and move work to a new one. Ignoring
GOAWAYis how a rolling deploy looks like a random 500. Frame details are the HTTP/2 lesson. - DNS TTL ignored. The pool keeps the old address after scale-in. Set a max connection age so you re-resolve. Readiness on the new task is not enough if old pools never leave.
- Max streams. HTTP/2 answers
REFUSED_STREAM. Open another session only if you mean to, or back off. A tight cap is a signal, not a bug by itself. - Partial body, then a timeout, then a retry. The server may have committed. This is the idempotency-key case.
- Workers blocked on a sync client. The pool has spare sockets and the process has no threads left. Async or a larger worker cap does not fix an infinite read timeout.
Server side, briefly
- Align the upstream idle timeout with the clients. The client should give up the socket first.
- Recycle connections by request count or age so DNS and load shift.
- On shutdown: stop accepting, send
GOAWAYorConnection: close, drain, then exit. A sidecar that kills the process first wastes that sequence. Drain order next to a proxy is mesh architecture. - Health checks should be cheap and separate from the heavy route. How those checks are specified on a load balancer stays on the L4 versus L7 page.
Libraries, the one-line version
| Client | What interviews want |
|---|---|
Go net/http | The Transport pools. Pass a context deadline. Honor Close. |
| Java HttpClient, OkHttp | Name the pool and the dispatcher. Do not allocate either per call. |
| Python httpx or aiohttp | One Client or session for the process. Limits and Timeout are objects, not an afterthought. |
| Node fetch / undici | Keep-alive is on the agent or dispatcher. The global fetch defaults surprise people. |
| Browser | You inherit the browser pool and HTTP version. You do not get a max-sockets knob. |
Sandbox: caps and jitter without sleeping
Press Run. Snippets must be self-contained — no network, files, or native modules.
Press Run. Snippets must be self-contained — no network, files, or native modules.
The real client uses AbortController or a context so the cancel is a signal, not a boolean after the fact. The sandbox only shows the inequality.
Client idle is 90 seconds. The proxy idle is 60 seconds. During a rollout the old task sends GOAWAY and the pool still has its address cached for five minutes. Which symptom do you see on the next request, and which two settings do you change first?
Interview Q&A
Why not open a new TCP connection per request?
Answer
A short handler is often cheaper than DNS, TCP, and TLS. The pool pays that cost once and reuses the socket or the QUIC session. Per-request clients throw the pool away.
How do HTTP/2 pools differ from HTTP/1.1 pools?
Answer
HTTP/1.1 concurrency is more connections, each with one response in flight. HTTP/2 and HTTP/3 concurrency is more streams on fewer sessions, bounded by MAX_CONCURRENT_STREAMS and by the session lifetime. Tuning only "max connections" under HTTP/2 misses the stream cap.
What is a retry budget?
Answer
A limit on how many calls, or what fraction of calls, may be attempted again. Without it, a dependency that fails 30 percent of the time can receive a full extra copy of that traffic from well-meaning clients.
How do you align timeouts?
Answer
Give the caller one deadline. Spend connect, handshake, and read from it. Keep that deadline inside whatever the upstream proxy will wait, with room for the next hop. Separately, keep the client idle timeout shorter than the proxy idle timeout so stale sockets are closed by the client.
What breaks after a rolling deploy?
Answer
Pools stick to tasks that are exiting. Clients that ignore GOAWAY open streams the server will refuse. DNS and the pool keep the old address. Readiness can go green on the new tasks while the old connections are still the ones in use.
When is hedging acceptable?
Answer
When the call is an idempotent read and doing it twice is cheap. Send the second attempt only if the first is slow, and still count both against the deadline and the budget. Do not hedge a write.
How do you observe pool health?
Answer
Wait time for a socket or stream, in-use and idle gauges, dial errors, TLS handshake errors, HTTP version mix, RST rate, and refused streams. A rising wait time with a flat upstream CPU means the cap, not the handler.
Where do circuit breakers fit?
Answer
After pools and deadlines. The pool limits concurrency and the deadline limits waiting. A breaker stops new calls when the dependency is already failing. This page does not design the breaker. It keeps the first line correct so the breaker is not hiding an unbounded retry.