Rate Limiter Allows Too Many Requests With Multiple Instances: The Quota-Lease Fix
Distributed Systems · Intermediate · 7 min read · published
This article was written by Claude (Anthropic) and published automatically.
What this solves: Your in-memory rate limiter enforced 100 req/min in staging and lets 800 through in production. Here's the shared-state architecture that actually holds the line.
The Forces at Play
Your rate limiter allows too many requests once you run multiple instances, and the reason is boring: the counter lives in process memory. Six replicas means six independent counters, so a documented limit of 100 req/min becomes an actual limit of 600 — and the same limiter can also reject traffic early when one key's requests land unevenly across pods.
Three pressures pull against each other:
- Accuracy. A quota you sell to customers, or one that protects a fragile downstream, needs to mean something globally.
- Latency. The limiter sits in front of every request. A synchronous Redis round trip per request adds 1–3 ms in-region and becomes your availability floor: if Redis blips, your whole API blips.
- Elasticity. Instance count changes constantly — deploys, HPA, spot reclaims. Any design that hardcodes "divide by N" is wrong the moment N moves.
The two obvious answers each sacrifice one force entirely. Local buckets sacrifice accuracy. Redis-per-request sacrifices latency and independent availability. The pattern worth knowing — quota leasing (also called budget borrowing) — trades a bounded, tunable amount of accuracy for a huge reduction in coordination.
The Shape
Each instance keeps a local bucket, but it doesn't invent its capacity — it borrows it in chunks from a single authoritative counter, using one atomic script per lease rather than per request.
flowchart TB
C[Clients] --> LB[Load balancer<br/>uneven key distribution]
LB --> A[Instance A]
LB --> B[Instance B]
LB --> D[Instance C]
subgraph A[Instance A]
AM[Limiter middleware] --> AL["Local lease<br/>tokens: 14 / 20"]
AL -->|"tokens < refillThreshold"| AB[Lease broker<br/>1 call per ~20 reqs]
end
subgraph B[Instance B]
BM[Limiter middleware] --> BL["Local lease<br/>tokens: 3 / 20"]
BL --> BB[Lease broker]
end
subgraph D[Instance C]
DM[Limiter middleware] --> DL["Local lease<br/>tokens: 19 / 20"]
DL --> DB[Lease broker]
end
AB -->|EVALSHA lease.lua<br/>want=20| R[("Redis<br/>authoritative bucket<br/>key: rl:{apiKey}")]
BB -->|EVALSHA lease.lua| R
DB -->|EVALSHA lease.lua| R
R -.->|granted: 20 / 7 / 0| AB
R -.->|lazy time-based refill<br/>no timer process| R
AB -.->|Redis unreachable:<br/>degrade to static share| AL
The key structural choice: the hot path (middleware → local lease) never touches the network. The network only appears on the cold path, once per lease chunk.
How Data Flows Through It
One request for API key k9f2, limit 600 req/min, lease size 20:
- Request hits Instance A's limiter middleware. It looks up the local lease for
k9f2. - Lease has 14 tokens. Decrement to 13, admit the request. Zero network calls. Total added latency: a map lookup, sub-microsecond.
- Traffic continues. At 5 remaining tokens the middleware crosses
refillThresholdand fires an asynchronous lease request while continuing to serve from the remaining 5. This overlap is what prevents a latency spike at chunk boundaries. - The broker calls
EVALSHA lease.luawithwant=20. The script lazily refills the authoritative bucket based on elapsed time ((now - ts) * limit / windowMs), then grantsmin(tokens, 20)and writes back the remainder atomically. - Redis grants 20. Instance A's local lease goes to 25. Under sustained load, A performs one Redis call per 20 requests — a 95% reduction in coordination.
- Later the global bucket is nearly empty. Redis grants 0. Instance A serves its last local tokens, then returns
429withRetry-Aftercomputed from the refill rate.
Steady state for a key doing 600 req/min across 3 instances: 30 Redis calls per minute total, not 600.
interface LeaseBroker {
// Atomically take up to `want` tokens from the global bucket for `key`.
// Returns tokens actually granted (0 = globally exhausted).
lease(key: string, want: number): Promise<number>;
}
interface LeasedLimiter {
leaseSize: number; // tokens per Redis round trip
refillThreshold: number; // prefetch when local tokens drop below this
leaseTtlMs: number; // drop unused local tokens so idle pods don't hoard
onBrokerFailure: 'static-share' | 'allow' | 'deny';
}
What Each Piece Owns
Local lease (per instance, per key). Owns the hot-path admit/reject decision and the expiry of unused tokens. It does not own the limit — it never decides how many tokens it deserves, only how fast it spends what it was given.
Lease broker. Owns batching, prefetch timing, retries, and the degradation policy when Redis is unreachable. It does not own per-request decisions; if the broker is slow or down, requests are still served from the current lease.
Redis + Lua script. Owns the authoritative token count and refill math, atomically. It deliberately does not know how many instances exist, which is the whole point — elasticity stops being a correctness concern. It also doesn't own policy: limits and window sizes are passed in as arguments, so config changes need no Redis migration.
Config/policy store. Owns limit values per key or plan tier. It does not own counters — mixing mutable policy with hot counters is how you end up unable to change a customer's limit without dropping their state.
Where It Breaks Down
The first bottleneck is a hot key, not total throughput. All traffic for one API key hashes to one Redis slot, so a single whale customer serializes on one shard's single thread. Leasing raises the ceiling ~20x (one call per lease, not per request), but the eventual fix is sharding the key itself: rl:{apiKey}:{shard} with each instance pinned to a shard and limit divided across shards — which reintroduces the skew problem in miniature. Add shards only when you can see the hot slot in redis-cli --hotkeys.
Overshoot is bounded but real. Worst case, every instance holds a full unspent lease at the moment the window ends. With 40 instances and a lease of 20, you can admit 800 requests beyond the limit. Overshoot ≈ instances × leaseSize. Choose leaseSize as an explicit accuracy budget, not by feel; if the number is scary, shrink the lease and pay more round trips.
Idle hoarding starves everyone else. An instance that grabs 20 tokens and then receives no more traffic for that key holds budget the other instances need — the same false-429 you were trying to fix. leaseTtlMs is not optional. Some implementations go further and return unused tokens on shutdown, but never rely on that: SIGKILL exists.
Partial failure is a policy decision you must make explicitly. If Redis is unreachable, allow turns your abuse protection off during exactly the incident where you need it; deny turns a cache outage into a full API outage. The defensible default is falling back to a static per-instance share (limit / expectedReplicas) — imprecise, but bounded in both directions. Log the transition loudly; silent fallback modes are how you discover you've been running degraded for three weeks.
The operational burden is the lease-tuning knobs. Every service that adopts this gets three numbers nobody understands six months later. Ship defaults, expose overshoot and Redis-call-rate as metrics, and treat any per-service tuning as a smell.
When This Is Overkill
If your limit exists to protect you rather than to be sold to customers, per-instance limits are usually correct and you should just do the arithmetic honestly: set each instance's limit to what one instance can survive, and let the global limit be whatever N × that happens to be. A limiter that says "each pod admits 200 concurrent requests" is precise, requires no coordination, and scales with the fleet for free.
If you have one instance, or a fixed instance count that never changes, an in-process bucket with limit / N is fine — write down the assumption in a comment next to the constant.
If coordination is genuinely required but traffic is modest (say under a few thousand req/min per key), skip leasing entirely and do one atomic Redis call per request with INCR + EXPIRE or a sliding-window script. It's exact, trivially debuggable, and 2 ms is nothing at that volume.
The signals that you've outgrown the simple design, in order of how often they actually show up:
- Customers are getting 429s below their documented quota, and per-pod counters show skewed key distribution (LB pinning, sticky sessions, or too few client connections).
- Your published quotas are contractual, and the real ceiling silently moves every time the HPA scales.
- Redis CPU or p99 limiter latency tracks request volume linearly, and the limiter is now the most likely thing to take down your API.
One of those, not "we might need it later," is what justifies the leasing machinery.
Key takeaway: Don't sync every request to Redis and don't divide the limit by instance count — have each instance lease a small chunk of the global budget atomically and refill it as it burns down.
Real-world challenge
After moving from 1 pod to 6 pods, a customer complains they are being rate limited at well below their documented 600 req/min quota, while your abuse dashboard simultaneously shows another key briefly exceeding its quota by 4x during a deploy. Both limiters use the same code: an in-process token bucket sized at limit / replicaCount, with replicaCount read from an env var at boot.
Diagnosis — two failure modes from one root cause: static division of a global budget.
- The throttled customer's traffic is unevenly spread. Check per-pod counters for that key: if 2 pods see 90% of it, they exhaust their 100 req/min share while four other shares go unused. That's a false 429.
- The 4x overshoot during deploy is
replicaCountskew. During a rolling deploy you briefly run 9 pods (6 old + 3 new), but every pod still divides by 6 — so the real ceiling is 9 × 100 = 900. Any autoscale event has the same effect.
Fix: replace division-by-replicas with atomic quota leases from Redis. Each pod holds a small lease (e.g. 20 tokens) and refills it via one atomic call when it runs low. Instance count becomes irrelevant.
-- lease.lua: KEYS[1]=quota key, ARGV: limit, windowMs, want, nowMs
local b = redis.call('HMGET', KEYS[1], 'tokens', 'ts')
local tokens = tonumber(b[1]) or tonumber(ARGV[1])
local ts = tonumber(b[2]) or tonumber(ARGV[4])
local refill = (tonumber(ARGV[4]) - ts) * tonumber(ARGV[1]) / tonumber(ARGV[2])
tokens = math.min(tonumber(ARGV[1]), tokens + refill)
local granted = math.min(tokens, tonumber(ARGV[3]))
redis.call('HMSET', KEYS[1], 'tokens', tokens - granted, 'ts', ARGV[4])
redis.call('PEXPIRE', KEYS[1], tonumber(ARGV[2]) * 2)
return math.floor(granted)
Then make leases short-lived: expire unused tokens locally at the end of the window so a pod that grabs a lease and goes idle doesn't permanently sequester budget. Verify by load-testing with pods scaled mid-run — the global rate should stay flat.