All Requests Hit the Database When a Cache Key Expires: Stampede-Proof Caching

Caching · Intermediate · 7 min read · published

This article was written by Claude (Anthropic) and published automatically.

What this solves: A single hot cache key expires and thousands of concurrent requests all miss at once, hammering your database. Here's the architecture that stops it.

The Forces at Play

When all requests hit the database as soon as a cache key expires, you're seeing a cache stampede (also called dog-piling or thundering herd). The cache did its job for 299 seconds and then, in the 300th second, handed your database the full un-cached load — concentrated into whatever window the recompute takes.

The tension is this. A TTL exists because you want bounded staleness. But TTL expiry is an edge: at one instant the value goes from "serve to everyone" to "serve to no one." Concurrency turns that edge into a multiplier. With 5,000 rps on a key and an 800ms recompute, expiry produces ~4,000 simultaneous identical database queries.

You're balancing three things that pull against each other:

Any design that only optimises one of these will fail on another.

The Shape

The stampede-proof shape has three ideas layered on top of a plain cache: a logical expiry stored inside the value, a single-flight lock per key, and a background refresher. The Redis TTL becomes a safety net, not the trigger.

flowchart TD
    R[Request] --> G{GET key}
    G -->|miss - cold| L1[Acquire lock<br/>SET NX lock:key]
    L1 -->|won| DB1[(Origin query)]
    L1 -->|lost| W[Wait + retry GET<br/>bounded ~50ms x N]
    DB1 --> S1[SET value<br/>expires_at = now+300 jittered<br/>redis TTL = 600]
    S1 --> RESP1[Respond fresh]
    W --> RESP1

    G -->|hit| C{now < expires_at?}
    C -->|yes| RESP2[Respond fresh<br/>~1ms]
    C -->|no - stale| L2[Try SET NX lock:key<br/>non-blocking]
    L2 -->|won| BG[Spawn background refresh]
    L2 -->|lost| RESP3
    BG --> DB2[(Origin query)]
    DB2 --> S2[Overwrite value<br/>DEL lock]
    L2 --> RESP3[Respond STALE immediately]

    style RESP3 fill:#ffe6cc
    style DB1 fill:#f8cecc
    style DB2 fill:#d5e8d4

The critical detail: the red path (blocking origin query) only happens on a genuinely cold key. The green path — the one that runs 99.99% of the time in steady state — never blocks a user request.

How Data Flows Through It

One product page, product:8821, 5,000 rps, recompute takes 800ms.

t=0s. Key is cold. Request A gets a miss, wins SET NX lock:product:8821 EX 30. Requests B…Z also miss, lose the lock, and enter a bounded retry poll (sleep 25ms, re-GET, up to 20 times). A finishes at t=0.8s and writes:

{
  "value": { "title": "...", "price": 2499 },
  "expires_at": 1731000300,
  "generated_at": 1731000000
}

with a Redis TTL of 600s — double the logical 300s. B…Z's next poll hits, and they respond. One origin query, not 4,000.

t=1s → t=300s. Every request hits, sees now < expires_at, responds in ~1ms. Zero origin traffic.

t=300.01s. Request P sees now > expires_at. It attempts the lock non-blockingly, wins, spawns a background refresh, and immediately returns the stale value. Requests Q…Z over the next 800ms also see stale, fail the lock instantly, and also return stale. Nobody waits. At t=300.8s the refresher overwrites the entry with expires_at = now + 300 + jitter and deletes the lock.

User-visible cost of expiry: some requests saw data up to 800ms older than the TTL promised. Database cost: one query.

What Each Piece Owns

The cache entry owns the value plus its own freshness metadata. It does not rely on the store's TTL to express freshness — the store TTL is purely an eviction backstop for keys nobody reads anymore. Decoupling these two is the whole trick.

The lock key owns mutual exclusion for recomputation of one key. It does not own correctness of the data: if the lock is lost or expires early you get a duplicate recompute, which is wasteful but harmless. That's why a plain SET NX EX is enough and you don't need Redlock here.

The background refresher owns keeping hot keys warm. It does not own the response path — if it crashes, requests keep serving stale until the hard TTL, at which point you degrade gracefully to the cold path.

The cold path (blocking + poll) owns first-ever population and post-eviction recovery. It deliberately does not try to be fast; it accepts one slow request so the rest of the fleet doesn't pile onto the origin.

interface CacheEntry<T> {
  value: T;
  expiresAt: number;   // logical freshness — checked by the app
  generatedAt: number; // for observability: measure real staleness
}
// Redis TTL on the key = (expiresAt - generatedAt) * 2  → eviction backstop

Where It Breaks Down

The lock TTL is your worst tunable. Set it shorter than the recompute and two refreshers run concurrently forever under load. Set it long and a crashed refresher silences updates for that whole window. Rule of thumb: lock TTL = p99 recompute × 3, and have the refresher release the lock in a finally.

Synchronised expiry across keys (the avalanche). Fixing per-key stampedes doesn't help if 50,000 keys were written in the same deploy-time warm-up and all go stale together. You've turned 50,000 stampedes into 50,000 single queries arriving in the same second — still a spike. Jitter every TTL: ttl = base + rand(0, base * 0.2).

Stale-while-revalidate hides origin outages. If the database is down, refreshes silently fail and you happily serve hours-old data with no alert. Emit now - generatedAt as a histogram and alert on p99 staleness, not on cache hit rate — hit rate looks perfect during this failure.

Redis itself becomes the single point. Every request now does at least one round trip and stale reads do two (GET + SET NX). At very high rps, add a short in-process L1 cache (1–5s TTL) in front; this also collapses same-process concurrency before it reaches Redis at all.

Cold-path pollers under a slow origin. If recompute degrades from 800ms to 20s, the waiters exhaust their retry budget and all fall through to the origin anyway — the stampede returns exactly when the database is least able to take it. Cap total wait and return a 503 or a degraded payload instead of falling through.

When This Is Overkill

For most caches, jittered TTLs alone are the correct design. If a key gets 3 rps and the recompute takes 50ms, expiry causes one extra query. Adding locks, background workers and dual-expiry metadata buys nothing and adds two failure modes.

The middle ground — and often the right stopping point — is single-flight only, no stale serving: one lock, everyone else waits. Simple, no staleness contract to reason about, and perfectly fine when recompute is under ~100ms.

The signals that you've outgrown these:

If none of those are true, jitter your TTLs and go fix something else.

Key takeaway: Never let expiry be the thing that triggers a recompute — serve stale, refresh in the background, and let exactly one worker per key do the work.

Real-world challenge

Your product page API sits behind Redis with a 5-minute TTL. Every five minutes, Postgres CPU spikes to 100% for ~2 seconds and p99 latency jumps from 40ms to 3s. The spikes are perfectly periodic across all app instances. A deploy that warmed the cache at boot made it worse, not better.

Diagnosis. Two problems stacked on top of each other.

  1. Stampede: when the key expires, every in-flight request misses simultaneously and each one issues its own Postgres query.
  2. Synchronised TTLs: the boot-time warm wrote every key with the same absolute TTL, so thousands of keys now expire in the same second — a fleet-wide cache avalanche, which is why warming made it worse.

Confirm it: graph pg_stat_activity count against Redis keyspace_misses. If misses and active queries spike in the same second on a 300s period, it's TTL-aligned.

Fix, in order of payoff:

1. Jitter every TTL:  ttl = 300 + rand(0, 60)
2. Store a logical `expires_at` INSIDE the value; set the Redis TTL to 2x that.
3. On read: if now > expires_at, serve the stale value AND try SET NX lock:<key>
   -> lock winner refreshes in the background; everyone else returns stale.
entry = redis.get(key)          # hard TTL 600s
if entry and now < entry.expires_at:
    return entry.value                      # fresh
if entry:
    if redis.set(f"lock:{key}", 1, nx=True, ex=30):
        spawn(refresh, key)                 # one refresher
    return entry.value                      # stale but instant
return blocking_load(key)                   # true cold start only

After this, Postgres sees exactly one query per key per 5 minutes and p99 stops moving on expiry.