Observability
Part 3 of 6 · Resilience PatternsBulkheads — Thread Pools, Connection Pools, Queues & Failure Domains
Bulkheads isolate failure domains so a slow or chatty dependency cannot exhaust shared threads, connections, or CPU and starve unrelated work. Separate pools, queue depth limits, and coarse Kubernetes resource limits are all bulkheads at different layers. Unlike circuit breakers (which trip on health), bulkheads always cap capacity per dependency.
- 1Gist
- 2Maps
- 3Q&A
- 4Sandbox
Voice readout needs Web Speech Synthesis in this browser.
Recommendations spike — does checkout survive?
Prefer
Per-dependency pools plus a bounded queue
Recs get their own semaphore/executor and a short queue. Checkout keeps its slots even while recs saturates. Reject locally when the recs queue is full.
- Always on — unlike a breaker that only trips when unhealthy.
- Queue cap (e.g. 2× pool size) beats unbounded wait.
- Critical-path latency stays stable while the non-critical pool saturates.
Alternative
One big shared pool, or only Kubernetes limits
Shared maximizes utilization until one dep fails — then site-wide p99. Pod CPU/memory limits stop node noisy neighbors, not one HTTP client eating all threads inside the pod.
- Tail latency holds workers; queueing delay becomes everyone's p99.
- Retries amplify; the breaker may not trip if calls eventually succeed slowly.
- K8s limits are necessary and still too coarse for outbound I/O isolation.
From shared blast radius to isolated pools
Single column: the anti-pattern, then the split. No side-by-side subgraphs.
- 1
Shared pool
Checkout, recs, and search share one thread/connection pool. - 2
Recs latency spikes
Workers block on the slow client. Unrelated endpoints wait. - 3
Checkout times out
Revenue path misses SLO even though payments are healthy. - 4
Split pools
Per-dep executors / semaphores / connection pools. - 5
Bound the queue
Beyond 2× pool size, fail fast locally — then shed at admission if you are the bottleneck.
Overview
A classic outage: recommendations HTTP client shares a pool with checkout; recommendations latency spikes; all threads block; checkout timeouts; revenue drops. Interviews expect you to name isolation boundaries.
Unlike circuit breakers (which trip on health), bulkheads always cap capacity per dependency.
Isolation mechanisms
| Layer | Mechanism | Isolates | Limitation |
|---|---|---|---|
| Threads / async workers | Per-dep executor / semaphore | CPU scheduling and waiters | Does not limit remote connections alone |
| Connection pools | Per-host max connections | Sockets to one dep | Shared CPU still contended |
| Queues | Bounded queue + reject | Memory and wait time | Need reject/shed policy |
| Processes / containers | K8s CPU/memory limits | Noisy neighbor on node | Coarse; cold-start cost |
| Tenants | Quotas per customer | Multi-tenant blast radius | Product policy needed |
Flow
- 1
1 Anti-pattern: shared pool
- next2 Recs + checkout + search
- 2
2 Recs + checkout + search
- next3 One slow dep starves all
- 3
3 One slow dep starves all
- next4 Split into per-dep pools
- 4
4 Split into per-dep pools
- next5 Checkout pool
- 5
5 Checkout pool
- next6 Recs pool
- 6
6 Recs pool
- next7 Search pool
- 7
7 Search pool
Lesson map
Bulkheads — Thread Pools, Connection Pools, Queues & Failure Domains
Bulkheads isolate failure domains so a slow or chatty dependency cannot exhaust shared threads, connections, or CPU and starve unrelated work. Separate pools, queue depth limits, and coarse Kubernetes resource limits are all bulkheads at different layers. Unlike circuit breakers (which trip on health), bulkheads always cap capacity per dependency.
Architecture. Architecture
Select a node to see why it exists, or an edge to see the protocol, direction, effect, and consequence.
Mermaid export
flowchart TB a["1 Anti-pattern: shared pool"] b["2 Recs + checkout + search"] c["3 One slow dep starves all"] d["4 Split into per-dep pools"] a -->|1 Anti-pattern: shared pool| b b -->|2 Recs + checkout + search| c c -->|3 One slow dep starves all| d
Bulkhead vs circuit breaker
| Bulkhead | Circuit breaker | |
|---|---|---|
| Question | How much capacity may this dep use? | Is this dep healthy enough to call? |
| Always on | Yes — hard caps | Only trips when unhealthy |
| Failure mode | Queue/reject when full | Fail-fast when Open |
| Together | Best practice: pool + breaker per dep | Breaker without bulkhead still shares threads |
Why one shared pool is dangerous: Tail latency from one dep holds workers; queueing delay becomes site-wide p99; retries amplify; circuit breaker may not trip if calls eventually succeed slowly.
Queue depth and rejection
- Cap queue length (e.g. 2× pool size). Beyond that: fail fast locally (better than unbounded wait).
- Prefer load shedding at admission when your queues are deep.
- Expose metrics:
pool_active,pool_queued,pool_rejected.
Size from Little's Law: concurrency ≈ arrival_rate × latency, plus headroom. Set the queue bound so max wait ≤ remaining deadline budget.
Sandbox: semaphore bulkhead (Python)
Per-dependency semaphore with bounded wait. Slow recs cannot block checkout.
Press Run. Snippets must be self-contained — no network, files, or native modules.
Same idea (TypeScript)
Press Run. Snippets must be self-contained — no network, files, or native modules.
Comparative teaching
- Vs larger shared pool: More throughput until one dep fails — then worse blast radius. Prefer isolation over raw size.
- Vs only K8s limits: Pod limits stop node-level noisy neighbors but not "one HTTP client eats all threads inside the pod."
- Vs circuit breaker alone: Breaker trips after damage starts; bulkhead prevents the damage from spreading even while Closed.
Pitfalls
Arrival 100 QPS checkout at 50 ms, 400 QPS recs at 200 ms. Using Little's Law, sketch limits and a queue bound so max wait fits a 200 ms remaining budget on checkout. What happens if recs stays on the shared pool?
Interview Q&A
What is a bulkhead in distributed systems?
Answer
A capacity partition (threads, connections, queues, processes) that confines failure impact to one dependency or tenant.
Shared pool vs per-dependency pools?
Answer
Shared maximizes utilization; per-dep protects latency SLOs of critical paths. Prefer per-dep for outbound I/O to unreliable services.
How do you size a bulkhead?
Answer
From concurrency ≈ arrival_rate × latency (Little's Law), plus headroom; set queue bound so max wait ≤ remaining deadline budget.
Relationship to Kubernetes resource limits?
Answer
Limits are coarse bulkheads between pods/nodes. Still need in-process isolation for multiple outbound deps in one service.
What metrics prove bulkheads work?
Answer
Per-pool active/queued/rejected; critical-path latency stable while non-critical pool saturates during an incident. Dashboards sit next to SLOs.
Why is a shared pool dangerous even with a breaker?
Answer
Slow-but-successful calls may not trip the breaker. Workers stay held; site-wide p99 moves. Bulkheads cap that blast radius while Closed.
Threads vs connection pools?
Answer
Per-dep executors isolate waiters and CPU scheduling. Per-host max connections isolate sockets. You often need both — a thread cap does not limit remote connections alone.
What do you do when the bulkhead queue is full?
Answer
Fail fast locally. If you are the bottleneck, shed at admission rather than queueing work that will miss the deadline.
Tenant quotas as bulkheads?
Answer
Per-customer quotas confine multi-tenant blast radius. That is product policy plus capacity, not a substitute for per-dep I/O pools.
How do retries interact with a shared pool?
Answer
Retries amplify hold time on the shared workers. Isolation plus a breaker beats "just retry." Jitter lives on Retry Storms.