One Slow Dependency Makes All Endpoints Slow: Bulkhead It
Resilience · Intermediate · 7 min read · published
This article was written by Claude (Anthropic) and published automatically.
What this solves: When one downstream API degrades, unrelated endpoints time out too. Learn how bulkheads cap each dependency's share of your threads and connections.
The Forces at Play
The classic outage goes like this: one slow dependency makes all endpoints slow, even endpoints that never call it. Recommendations go from 50 ms to 5 s, and suddenly login, checkout and /health all time out. The cause is almost never CPU. It's a shared, finite pool that the slow dependency quietly takes over. That pool might be request threads, HTTP client sockets, database connections, or an async concurrency limit.
Little's law explains why this happens so fast:
in-flight = arrival rate × latency
At 200 rps and 50 ms, a dependency holds 10 slots on average. If latency rises 100× to 5 s, it needs 1,000 slots. Your pool has 200, so within seconds every slot is waiting on the sick service. Healthy work queues behind it and times out.
Several pressures pull against each other here:
- Efficiency wants one big shared pool. Idle capacity can go to whoever needs it.
- Isolation wants separate pools. One tenant or dependency shouldn't be able to starve the others.
- Timeouts alone don't save you. A 2 s timeout at 200 rps still holds 400 slots.
- Retries make it worse. Each retry claims another slot.
The bulkhead pattern takes its name from ship hulls, which are divided into compartments so one breach doesn't sink the ship. It resolves the tension by giving each dependency a fixed, small slice of concurrency that it can never exceed.
The Shape
What matters is where the pools sit. Each dependency gets its own admission gate and its own connection pool. The gates reject instead of queueing.
flowchart LR
C[Clients] --> S[Server request pool<br/>200 workers - shared]
S --> H[Handler: GET /product/:id]
H --> BI{Inventory bulkhead<br/>max 40, wait 0}
H --> BP{Pricing bulkhead<br/>max 40, wait 0}
H --> BR{Recs bulkhead<br/>max 15, wait 0}
BI -->|permit| PI[Inventory HTTP pool<br/>40 sockets, 300ms timeout] --> DI[(Inventory svc)]
BP -->|permit| PP[Pricing HTTP pool<br/>40 sockets, 300ms timeout] --> DP[(Pricing svc)]
BR -->|permit| PR[Recs HTTP pool<br/>15 sockets, 250ms timeout] --> DR[(Recs svc - SLOW)]
BR -->|full: reject in microseconds| FB[Fallback: empty recs]
PR -->|timeout| FB
FB --> H
style DR fill:#f88
style BR fill:#fd8
The shared server pool is still there, because you can't partition everything. What protects it is that handlers never wait on a full bulkhead. A slow recs call can hold at most 15 slots. Each of those slots is released after at most 250 ms, and every request beyond 15 is turned away immediately. So the recs dependency's maximum footprint on the shared pool is bounded and known in advance.
A minimal gate is just a non-blocking semaphore wrapped around a dedicated client:
import { Agent, request } from 'undici';
class Bulkhead {
private inFlight = 0;
constructor(readonly name: string, readonly max: number) {}
async run<T>(fn: () => Promise<T>, fallback: () => T): Promise<T> {
if (this.inFlight >= this.max) return fallback(); // never queue
this.inFlight++;
try { return await fn(); }
catch { return fallback(); }
finally { this.inFlight--; }
}
}
// Separate socket pool per dependency: no shared sockets to starve
const recsAgent = new Agent({ connections: 15, headersTimeout: 250, bodyTimeout: 250 });
const recs = new Bulkhead('recs', 15);
const items = await recs.run(
async () => (await request(RECS_URL, { dispatcher: recsAgent })).body.json(),
() => []
);
How Data Flows Through It
Take a request for GET /product/42 while the recs service is degraded.
- A server worker picks up the request. The handler fans out to inventory, pricing and recs in parallel.
- Inventory and pricing each get a permit, use sockets from their own pools, and return in 30 ms.
- Recs: its bulkhead is at 15/15 because earlier calls are still stuck.
run()sees the limit and returns[]in microseconds. No socket is touched and no time is spent waiting. - If a permit had been free, the call would have hit the 250 ms timeout. The permit would then be released and the fallback used.
- The handler renders the page without recommendations. Total latency is about 35 ms.
- Metrics record
bulkhead_rejected{name="recs"}. After enough rejections, a circuit breaker, if you have one, stops trying recs entirely for a cooldown period.
Compare this with the unprotected version. Step 3 would block for 5 s on a shared socket pool. Inventory calls would then queue behind recs calls for sockets. Within seconds the server pool fills and every endpoint dies.
What Each Piece Owns
| Piece | Owns | Deliberately does NOT own |
|---|---|---|
| Bulkhead (semaphore) | Max concurrent calls to one dependency; immediate rejection when full | How long a call takes. A permit held forever is still lost capacity. |
| Timeout | Upper bound on how long a permit is held | How many calls run at once. 1,000 calls with a 2 s timeout is still 1,000 slots. |
| Per-dependency connection pool | Sockets and keep-alive for one host, so a slow host can't hog sockets another host needs | Business fallback logic |
| Fallback | The degraded answer: empty list, cached value, or "try later" | Deciding when to degrade. The gates and breaker decide that. |
| Circuit breaker (optional) | Error-rate memory: skip calls entirely while the dependency is known bad | Concurrency limits. A breaker is closed until enough failures accumulate, and the bulkhead protects you during that window. |
| Shared server pool | Accepting inbound work | Isolation. It stays shared on purpose and is protected by everything above it. |
Keep timeouts and bulkheads separate in your head. You need both, and each one is useless without the other.
Where It Breaks Down
Sizing is the first thing to go wrong. Size each bulkhead with Little's law: peak rps × healthy p99 latency × about 2 for headroom. If the limit is too small, you reject requests on a normal Tuesday peak. If it's too large, it doesn't isolate anything. Review the limits whenever traffic shape changes.
Queueing behind the gate brings the outage back. Setting maxWaitDuration: 5s feels kind. In practice it lets requests park on the shared server pool, and you're back to the original failure with an extra step. Wait times should be zero or a few milliseconds.
Hidden shared pools. You bulkheaded the HTTP calls, but all endpoints share one database pool, one Redis client, or one DNS resolver. The slow resource moves somewhere you didn't partition. Look at every pool your process has: threads, sockets, DB connections, file descriptors.
Stranded capacity. Ten dependencies with ten pools can leave capacity unused, for example 60% of it idle while one hot pool rejects requests. Over-partitioning trades away the efficiency that made a shared pool appealing in the first place.
In-process limits don't cover CPU or memory. A runaway CPU-bound handler on Node's event loop, or a memory blowup, takes the whole process down regardless of semaphores. For those, the bulkhead has to be at the deployment level: separate instance groups for critical paths (checkout pods vs. browse pods) or per-tenant cells.
Operational burden. Every bulkhead is a tuning knob that needs a dashboard. At minimum, track available permits, rejections, and timeout rate per dependency. Without these, a rejection storm looks like a mysterious drop in features rather than a protected dependency.
Health checks bypass the design. If liveness probes run through the saturated pool, or check downstreams themselves, the orchestrator restarts healthy pods during dependency outages. Isolate the probe.
When This Is Overkill
For most services, the right starting design is:
- aggressive timeouts on every outbound call, and
- a per-host connection limit on your HTTP client. Most clients support this already: undici
connections, Go'sMaxConnsPerHost, Apache'smaxPerRoute.
That combination is already a crude bulkhead for sockets. Explicit semaphores add little if you have one critical dependency, or if all your dependencies are equally critical. If any of them failing means the request fails anyway, there's nothing to degrade to.
You've outgrown it when you see these signals:
- p99 latency of an endpoint that doesn't call service X tracks X's latency.
- Pool wait time or thread pool saturation spikes during an incident in a single dependency.
- You have a non-critical dependency (recs, analytics, personalization) on the same request path as a critical one, and the product could ship without it for a few minutes.
Those signals mean shared capacity is coupling things that the business considers independent. That is exactly the problem bulkheads solve.
Key takeaway: Give every downstream dependency its own fixed concurrency budget that rejects immediately when full, so its slowness spends only its own capacity.
Real-world challenge
A Spring Boot checkout service runs on Tomcat with the default 200 request threads. Each checkout calls a third-party fraud-scoring API through a RestTemplate with no explicit timeout. The fraud vendor degrades and its responses start taking about 40 seconds. Within two minutes, every endpoint on the service times out, including /health. Kubernetes fails the liveness probes and restarts pods, and the new pods die the same way. The fraud vendor's status page says only 'elevated latency'. Diagnose the cascade and fix it.
Diagnosis
Use Little's law: in-flight requests = arrival rate × latency. At 30 checkouts/s with a 200 ms fraud call, about 6 threads are busy. At 40 s latency, the service needs 1,200 threads. Tomcat has 200, so every thread ends up parked inside a fraud call.
/health goes through the same pool. With no thread free to answer it, the liveness probe fails and Kubernetes kills pods that were otherwise healthy. Restarts don't help because new pods fill up again within seconds.
There are three separate faults:
- No timeout. The fraud call can hold a thread indefinitely.
- No bulkhead. The fraud call can claim all 200 shared threads.
- Liveness tied to request capacity. Saturation is treated as death, so pods get restarted instead of recovering.
Fix
Cap the fraud call's concurrency and latency, and give it a fallback:
resilience4j:
bulkhead:
instances:
fraud:
maxConcurrentCalls: 25 # ~ peak rps x p99 latency, with headroom
maxWaitDuration: 0ms # reject immediately, never queue
timelimiter:
instances:
fraud:
timeoutDuration: 800ms
@Bulkhead(name = "fraud", fallbackMethod = "fraudFallback")
@TimeLimiter(name = "fraud")
public CompletableFuture<Score> score(Order o) { ... }
// Business decision: allow small orders, queue large ones for manual review
CompletableFuture<Score> fraudFallback(Order o, Throwable t) {
return CompletableFuture.completedFuture(Score.deferredReview(o));
}
Then do three more things:
- Set connect and read timeouts on the HTTP client itself, so sockets don't outlive the TimeLimiter.
- Make liveness a trivial in-process check. Move dependency checks to readiness or dashboards.
- Alert on
resilience4j_bulkhead_available_concurrent_callsreaching 0. That signal fires minutes before users notice.