Retries Overload a Service Right After It Recovers: Retry Budgets That Stop the Metastable Loop
Resilience · Advanced · 7 min read · published
This article was written by Claude (Anthropic) and published automatically.
What this solves: Your downstream comes back healthy, then instantly collapses again under a wall of retries. Here's the retry-budget and jitter architecture that breaks that loop.
The Forces at Play
When retries overload a service right after it recovers, the instinct is to blame capacity. It isn't capacity. A dependency comes back healthy, serves traffic beautifully for ten seconds, and then dies again — and the loop repeats until someone manually drains traffic. This is a metastable failure: the system has a stable healthy state and a stable dead state, and once it's knocked into the dead one, the load it generates itself is enough to hold it there.
The forces pulling against each other:
- Retries make individual requests more reliable. A single dropped packet or a pod rotating out shouldn't surface as a user-visible error. Per-request, retries are unambiguously good.
- Retries make the system less reliable in aggregate. They are a positive feedback loop: failures cause retries, retries cause load, load causes failures. The amplification factor is exactly the thing you configured to help you.
- Backoff spreads retries in time but doesn't bound their volume. Exponential backoff with three attempts is still 3x load when the failure rate is 100%.
- Amplification compounds across hops. Gateway retries × BFF retries × service retries = 27x for three layers of "just 3 attempts".
The architecture below separates whether this request may retry from whether the system as a whole can afford more retries right now — and puts the second decision in a component that sees aggregate traffic.
The Shape
The centrepiece is a client-side retry controller that sits between your call site and the transport, plus a server-side shedding policy that can tell first attempts apart from retries.
flowchart TB
subgraph Caller["Order Service (caller)"]
CS["Call site\nattempt 1"]
RC{"Retry controller"}
BKT["Retry budget\ntoken bucket\nrefilled by successes"]
CB["Circuit breaker\nclosed / open / half-open"]
JIT["Jitter scheduler\nsleep = rand(0, base*2^n)"]
DL["Deadline\npropagated from inbound req"]
end
subgraph Callee["Inventory Service"]
LB["Load balancer"]
Q["Admission queue\ndepth-aware"]
SHED{"Shed policy\nx-retry-attempt > 0?"}
WRK["Worker pool"]
end
CS --> RC
RC -->|"is breaker open?"| CB
RC -->|"tokens available?"| BKT
RC -->|"time left?"| DL
RC -->|"all yes"| JIT
JIT -->|"retry attempt n"| LB
CS -->|"first attempt always allowed"| LB
RC -.->|"any no: fail fast"| CS
LB --> Q
Q --> SHED
SHED -->|"retry + queue deep"| DROP["429 shed early"]
SHED -->|"first attempt"| WRK
WRK -->|"2xx"| BKT
WRK -->|"5xx / timeout"| CB
DROP -.->|"Retry-After"| RC
The critical edge is WRK -->|2xx| BKT: retry capacity is minted by successes. No successes, no retries. That single feedback direction is what converts a positive feedback loop into a negative one.
How Data Flows Through It
Follow one POST /orders during a recovery.
- Inbound request arrives at the order service with a 2s deadline. The deadline is stored in the request context; every downstream call inherits the remaining time, not a fresh 500ms timeout.
- First attempt to inventory goes straight through. First attempts are never budgeted or breaker-gated for latency reasons — the breaker gates them only when fully open.
- Inventory is still saturated. Its admission queue depth is 800; the request carries no
x-retry-attemptheader, so the shed policy admits it. It times out anyway at 400ms. - Back in the retry controller. It asks three questions in order:
- Is the circuit breaker open? If yes → fail fast, no network call.
- Does the retry budget have a token? During the outage, successes were near zero, so the bucket sits at its floor (say 5 retries/sec across the whole process). 95% of callers get refused here.
- Is there deadline left? 400ms spent of 2000ms, yes.
- This request wins a token. The jitter scheduler sleeps
rand(0, 100ms * 2^1)— a uniform draw, not a fixed 200ms, so the surviving retries don't arrive as a synchronised pulse. - Second attempt goes out with
x-retry-attempt: 1. Inventory's queue has drained to 120 because 95% of the retry traffic never left the callers. It's admitted and succeeds in 30ms. - The success refills the budget. The bucket now permits slightly more retry traffic. As inventory's success rate climbs, retry capacity climbs with it — a self-throttling ramp instead of a cliff.
The whole recovery is gradual by construction. Nobody wrote a ramp-up schedule.
What Each Piece Owns
Retry budget (token bucket). Owns the aggregate question: can this process afford another retry this second? Refilled by observed successes at a fixed ratio (10% is the common default), with a small absolute floor so low-traffic endpoints aren't permanently locked out. It does not know why a call failed, does not decide attempt counts, and deliberately does not coordinate across processes — per-process budgets are enough because the ratio is what bounds amplification, not the absolute number.
Circuit breaker. Owns the latency problem, not the load problem. When a dependency is fully dead, the breaker stops you from burning your caller's deadline on calls that can't succeed. It does not replace the budget: breakers are all-or-nothing and have hysteresis, so they're bad at the partial-degradation case where 40% of calls work.
Jitter scheduler. Owns de-synchronisation. Full jitter (rand(0, base * 2^n)) rather than equal jitter, because the outage itself synchronises every client's clock. It owns nothing about whether to retry.
Deadline propagation. Owns the truth that a retry after the caller gave up is pure waste. It does not own retry policy — it's a veto, not a decision.
Server-side shed policy. Owns prioritisation under overload. It's the only component that can distinguish a first attempt from a retry and preserve the former. It does not own client behaviour, though Retry-After is a hint clients should honour.
type RetryBudget interface {
// Deposit is called on every successful call; it mints retry capacity.
Deposit()
// Withdraw returns false when retries would exceed the configured
// ratio of recent successes. Callers MUST fail fast on false.
Withdraw() bool
}
type BudgetConfig struct {
TTL time.Duration // window over which successes count, e.g. 10s
MinPerSec float64 // floor, e.g. 5 retries/sec regardless of traffic
RetryRatio float64 // e.g. 0.1 -> at most 10% amplification
}
Note what the interface cannot express: a per-request "please, this one is important". That's intentional. Priority belongs in the shed policy on the server, where it can be evaluated against actual queue depth.
Where It Breaks Down
The floor becomes the attack surface. MinPerSec exists so a service making 2 calls/minute can still retry. With 500 caller pods, a floor of 5/sec/pod is 2,500 retries/sec hitting a dead dependency. The floor must be sized against fleet size, not per-process intuition. This is the first thing that bites at scale.
Budgets are per-process and per-dependency. A shared connection pool or a shared thread pool means one unbudgeted dependency can still starve everything. Budgets bound retry load; they don't bound resource coupling.
Nested retries defeat everything. If your gateway retries and your service retries, the effective ratio is 1.1 × 1.1 on the load but the amplification during total failure is multiplicative on attempt counts. The fix is a policy decision, not a library one: retry at exactly one layer, usually the one closest to the dependency, and mark requests as non-retryable for everyone above.
Partial failure is the hard case. One bad shard returns errors while nine are healthy. Successes keep refilling the budget, so retries keep flowing to the bad shard — and because it's the one failing, it absorbs a disproportionate share. Per-endpoint or per-shard budgets fix this and multiply your operational surface.
Operationally, the budget is the burden. It fails silently and correctly: requests get refused locally with no downstream error to trace. Without a dedicated retries_refused_by_budget counter and an alert on it, your on-call sees elevated error rates and no cause. Instrument the refusal path before you ship the budget.
Half-open breaker thundering herd. When 500 pods' breakers all transition to half-open at the same time, you get 500 probes at once. Jitter the half-open transition too.
When This Is Overkill
For most services, the correct design is much smaller:
- Retry only idempotent reads, once, with full jitter. No budget, no breaker. This is genuinely fine below a few hundred RPS against a dependency with spare headroom.
- Or: don't retry at all, and propagate the error. If your caller is a browser with a user behind it, a fast error plus a retry button is better architecture than a hidden 3-attempt loop that burns 2 seconds.
- A timeout plus a bulkhead (bounded concurrency per dependency) prevents most cascading failure with a fraction of the machinery. Concurrency limits bound load directly; they just don't preserve availability during transient blips the way retries do.
The specific signal you've outgrown the simple version: your dependency's recovery is not monotonic. Plot its success rate after an incident. If it climbs and then drops — a sawtooth — you have a metastable loop and you need a budget. If it climbs and stays up, your retry load is small relative to capacity and adding a budget is complexity you'll have to debug later.
A second signal: you can't answer "what is my worst-case amplification factor across all hops?" in under thirty seconds. If the answer is unknown, it's larger than you think.
Key takeaway: Cap retries as a fraction of your success traffic (a budget), not as a per-request count — per-request limits let load amplify exactly when the dependency is weakest.
Real-world challenge
An order service calls an inventory service. After a 4-minute inventory outage, inventory pods come up healthy, serve traffic for ~15 seconds, then CPU pegs and liveness probes start killing them. This repeats in a loop for 20 minutes. Order-service dashboards show outbound RPS at 6x normal. Inventory's own request latency histogram shows most requests completing in 30ms before the pod dies. What's happening and how do you break the cycle?
Diagnosis. This is a metastable failure loop, not a capacity problem. Inventory can serve 30ms requests fine — it just can't serve 6x normal volume. The 6x comes from two sources:
- Retry amplification. Every caller has 3 attempts; during the outage all of them exhausted and are now re-queued or being re-driven by upstream retries too. If order-service is itself retried by an API gateway, amplification multiplies (3 × 3 = 9x).
- Queued work released at once. Clients that backed off with a fixed schedule all wake in the same window, synchronised by the outage's start time.
The clue that separates this from genuine overload: per-request latency is healthy right up until death. The system is being killed by arrival rate, not by slow work.
Fix, in order of impact:
retry:
budget:
ttl: 10s
min_retries_per_second: 5 # floor so low-traffic paths still retry
retry_ratio: 0.1 # retries <= 10% of successes
backoff: full_jitter # sleep = rand(0, base * 2^n)
base: 100ms
max_attempts: 3 # still needed, but no longer the main control
budget_exhausted: fail_fast
- Add the retry budget at the order-service client. While inventory is down there are no successes, so the budget drains to the floor and retries essentially stop.
- Switch to full jitter so the surviving retries de-synchronise.
- Mark retried calls with a header (
x-retry-attempt) and have inventory shed retried requests first when its queue depth crosses a threshold — protect first attempts. - Turn off retries at the gateway for this path, or make retries non-nestable, to kill multiplicative amplification.
- Widen the liveness probe threshold so overloaded-but-alive pods aren't restarted, which removes capacity exactly when it's needed.
After the change, expect the recovery curve to be slow and boring rather than a sawtooth — that's the point.