Retries Overload a Service Right After It Recovers: Retry Budgets That Stop the Metastable Loop

Resilience · Advanced · 7 min read · published

This article was written by Claude (Anthropic) and published automatically.

What this solves: Your downstream comes back healthy, then instantly collapses again under a wall of retries. Here's the retry-budget and jitter architecture that breaks that loop.

The Forces at Play

When retries overload a service right after it recovers, the instinct is to blame capacity. It isn't capacity. A dependency comes back healthy, serves traffic beautifully for ten seconds, and then dies again — and the loop repeats until someone manually drains traffic. This is a metastable failure: the system has a stable healthy state and a stable dead state, and once it's knocked into the dead one, the load it generates itself is enough to hold it there.

The forces pulling against each other:

The architecture below separates whether this request may retry from whether the system as a whole can afford more retries right now — and puts the second decision in a component that sees aggregate traffic.

The Shape

The centrepiece is a client-side retry controller that sits between your call site and the transport, plus a server-side shedding policy that can tell first attempts apart from retries.

flowchart TB
    subgraph Caller["Order Service (caller)"]
        CS["Call site\nattempt 1"]
        RC{"Retry controller"}
        BKT["Retry budget\ntoken bucket\nrefilled by successes"]
        CB["Circuit breaker\nclosed / open / half-open"]
        JIT["Jitter scheduler\nsleep = rand(0, base*2^n)"]
        DL["Deadline\npropagated from inbound req"]
    end

    subgraph Callee["Inventory Service"]
        LB["Load balancer"]
        Q["Admission queue\ndepth-aware"]
        SHED{"Shed policy\nx-retry-attempt > 0?"}
        WRK["Worker pool"]
    end

    CS --> RC
    RC -->|"is breaker open?"| CB
    RC -->|"tokens available?"| BKT
    RC -->|"time left?"| DL
    RC -->|"all yes"| JIT
    JIT -->|"retry attempt n"| LB
    CS -->|"first attempt always allowed"| LB
    RC -.->|"any no: fail fast"| CS

    LB --> Q
    Q --> SHED
    SHED -->|"retry + queue deep"| DROP["429 shed early"]
    SHED -->|"first attempt"| WRK
    WRK -->|"2xx"| BKT
    WRK -->|"5xx / timeout"| CB
    DROP -.->|"Retry-After"| RC

The critical edge is WRK -->|2xx| BKT: retry capacity is minted by successes. No successes, no retries. That single feedback direction is what converts a positive feedback loop into a negative one.

How Data Flows Through It

Follow one POST /orders during a recovery.

  1. Inbound request arrives at the order service with a 2s deadline. The deadline is stored in the request context; every downstream call inherits the remaining time, not a fresh 500ms timeout.
  2. First attempt to inventory goes straight through. First attempts are never budgeted or breaker-gated for latency reasons — the breaker gates them only when fully open.
  3. Inventory is still saturated. Its admission queue depth is 800; the request carries no x-retry-attempt header, so the shed policy admits it. It times out anyway at 400ms.
  4. Back in the retry controller. It asks three questions in order:
    • Is the circuit breaker open? If yes → fail fast, no network call.
    • Does the retry budget have a token? During the outage, successes were near zero, so the bucket sits at its floor (say 5 retries/sec across the whole process). 95% of callers get refused here.
    • Is there deadline left? 400ms spent of 2000ms, yes.
  5. This request wins a token. The jitter scheduler sleeps rand(0, 100ms * 2^1) — a uniform draw, not a fixed 200ms, so the surviving retries don't arrive as a synchronised pulse.
  6. Second attempt goes out with x-retry-attempt: 1. Inventory's queue has drained to 120 because 95% of the retry traffic never left the callers. It's admitted and succeeds in 30ms.
  7. The success refills the budget. The bucket now permits slightly more retry traffic. As inventory's success rate climbs, retry capacity climbs with it — a self-throttling ramp instead of a cliff.

The whole recovery is gradual by construction. Nobody wrote a ramp-up schedule.

What Each Piece Owns

Retry budget (token bucket). Owns the aggregate question: can this process afford another retry this second? Refilled by observed successes at a fixed ratio (10% is the common default), with a small absolute floor so low-traffic endpoints aren't permanently locked out. It does not know why a call failed, does not decide attempt counts, and deliberately does not coordinate across processes — per-process budgets are enough because the ratio is what bounds amplification, not the absolute number.

Circuit breaker. Owns the latency problem, not the load problem. When a dependency is fully dead, the breaker stops you from burning your caller's deadline on calls that can't succeed. It does not replace the budget: breakers are all-or-nothing and have hysteresis, so they're bad at the partial-degradation case where 40% of calls work.

Jitter scheduler. Owns de-synchronisation. Full jitter (rand(0, base * 2^n)) rather than equal jitter, because the outage itself synchronises every client's clock. It owns nothing about whether to retry.

Deadline propagation. Owns the truth that a retry after the caller gave up is pure waste. It does not own retry policy — it's a veto, not a decision.

Server-side shed policy. Owns prioritisation under overload. It's the only component that can distinguish a first attempt from a retry and preserve the former. It does not own client behaviour, though Retry-After is a hint clients should honour.

type RetryBudget interface {
    // Deposit is called on every successful call; it mints retry capacity.
    Deposit()
    // Withdraw returns false when retries would exceed the configured
    // ratio of recent successes. Callers MUST fail fast on false.
    Withdraw() bool
}

type BudgetConfig struct {
    TTL         time.Duration // window over which successes count, e.g. 10s
    MinPerSec   float64       // floor, e.g. 5 retries/sec regardless of traffic
    RetryRatio  float64       // e.g. 0.1 -> at most 10% amplification
}

Note what the interface cannot express: a per-request "please, this one is important". That's intentional. Priority belongs in the shed policy on the server, where it can be evaluated against actual queue depth.

Where It Breaks Down

The floor becomes the attack surface. MinPerSec exists so a service making 2 calls/minute can still retry. With 500 caller pods, a floor of 5/sec/pod is 2,500 retries/sec hitting a dead dependency. The floor must be sized against fleet size, not per-process intuition. This is the first thing that bites at scale.

Budgets are per-process and per-dependency. A shared connection pool or a shared thread pool means one unbudgeted dependency can still starve everything. Budgets bound retry load; they don't bound resource coupling.

Nested retries defeat everything. If your gateway retries and your service retries, the effective ratio is 1.1 × 1.1 on the load but the amplification during total failure is multiplicative on attempt counts. The fix is a policy decision, not a library one: retry at exactly one layer, usually the one closest to the dependency, and mark requests as non-retryable for everyone above.

Partial failure is the hard case. One bad shard returns errors while nine are healthy. Successes keep refilling the budget, so retries keep flowing to the bad shard — and because it's the one failing, it absorbs a disproportionate share. Per-endpoint or per-shard budgets fix this and multiply your operational surface.

Operationally, the budget is the burden. It fails silently and correctly: requests get refused locally with no downstream error to trace. Without a dedicated retries_refused_by_budget counter and an alert on it, your on-call sees elevated error rates and no cause. Instrument the refusal path before you ship the budget.

Half-open breaker thundering herd. When 500 pods' breakers all transition to half-open at the same time, you get 500 probes at once. Jitter the half-open transition too.

When This Is Overkill

For most services, the correct design is much smaller:

The specific signal you've outgrown the simple version: your dependency's recovery is not monotonic. Plot its success rate after an incident. If it climbs and then drops — a sawtooth — you have a metastable loop and you need a budget. If it climbs and stays up, your retry load is small relative to capacity and adding a budget is complexity you'll have to debug later.

A second signal: you can't answer "what is my worst-case amplification factor across all hops?" in under thirty seconds. If the answer is unknown, it's larger than you think.

Key takeaway: Cap retries as a fraction of your success traffic (a budget), not as a per-request count — per-request limits let load amplify exactly when the dependency is weakest.

Real-world challenge

An order service calls an inventory service. After a 4-minute inventory outage, inventory pods come up healthy, serve traffic for ~15 seconds, then CPU pegs and liveness probes start killing them. This repeats in a loop for 20 minutes. Order-service dashboards show outbound RPS at 6x normal. Inventory's own request latency histogram shows most requests completing in 30ms before the pod dies. What's happening and how do you break the cycle?

Diagnosis. This is a metastable failure loop, not a capacity problem. Inventory can serve 30ms requests fine — it just can't serve 6x normal volume. The 6x comes from two sources:

  1. Retry amplification. Every caller has 3 attempts; during the outage all of them exhausted and are now re-queued or being re-driven by upstream retries too. If order-service is itself retried by an API gateway, amplification multiplies (3 × 3 = 9x).
  2. Queued work released at once. Clients that backed off with a fixed schedule all wake in the same window, synchronised by the outage's start time.

The clue that separates this from genuine overload: per-request latency is healthy right up until death. The system is being killed by arrival rate, not by slow work.

Fix, in order of impact:

retry:
  budget:
    ttl: 10s
    min_retries_per_second: 5      # floor so low-traffic paths still retry
    retry_ratio: 0.1               # retries <= 10% of successes
  backoff: full_jitter             # sleep = rand(0, base * 2^n)
  base: 100ms
  max_attempts: 3                  # still needed, but no longer the main control
  budget_exhausted: fail_fast

After the change, expect the recovery curve to be slow and boring rather than a sawtooth — that's the point.