Skip to content
José da Cruz

Enterprise Architect, and Life Thinker

José da Cruz

Enterprise Architect, and Life Thinker

The Cache Stampede: How One Expired Key Can Take Down Your Database

josedacruz, September 18, 2026September 18, 2026

TL;DR: A cache stampede happens when a popular cached value expires and a flood of requests all miss the cache at the same instant, sending them straight to your database all at once. It is one of the few bugs that gets worse the more successful your caching strategy is. The fix is not “cache more” — it is controlling who is allowed to refresh a stale value while everyone else waits or gets served something reasonable in the meantime.

Myth: Caching only makes things faster — it can’t make things worse

A cache (a fast, temporary storage layer that sits in front of a slower system, usually a database) is supposed to be a pure win. Read the value once, store a copy somewhere fast, and hand that copy out to everyone who asks for it next. As long as the cache is working, your database barely notices the traffic.

That story is true right up until the cached value expires. Every cache entry has a TTL (time to live — how long a stored value is considered valid before it must be refreshed). When the TTL runs out, the next request for that key is a cache miss: nobody has a fresh copy, so the request has to go fetch one from the real source. Normally that is fine, a single request pays the cost.

The problem is that popular keys do not get requested once. If a product page, a homepage banner, or a “trending” list is cached and read by thousands of users per second, then the instant that key expires, every one of those thousands of requests misses the cache in the same fraction of a second. They all discover the cache is empty, and they all go run the same expensive database query at the same time. Your fast cache just turned into a synchronized trigger for a self-inflicted denial-of-service attack on your own database.

Flowchart showing a cache key expiring and many simultaneous requests all missing the cache and hitting the database
When a hot key expires, every concurrent request discovers the miss independently and repeats the same expensive work.

Myth: A high cache hit rate means you’re safe

Imagine a ticket-booking site running a flash sale. The page for the hottest event is cached with a 99.9% hit rate — an excellent number by any dashboard’s standard. Everything looks calm. Then, once a minute, that cache entry expires, and for the next 200 milliseconds, every single request is a miss, because they all arrive while the value is being recomputed.

A hit rate is an average measured over time. It tells you nothing about how misses are distributed. A cache that misses one request every second, spread evenly, is invisible to your database. A cache that misses two thousand requests in the same 200 milliseconds, once a minute, can knock your database over even though the hit rate for that hour still reads as 99.9%. The average hides the spike, and the spike is the part that takes the site down.

This is why “cache hit rate” alone is a dangerous metric to sleep well on. You also want visibility into concurrent misses per key — how many requests are missing the same key at the same moment — because that number is what actually predicts an outage.

Myth: A shorter TTL is always the safer choice

It feels intuitive that shorter TTLs mean fresher data and lower risk — the cache never gets too stale, so nothing bad can build up. In reality, a shorter TTL means the expiry event, and the stampede risk that comes with it, happens more often. If your hot key refreshes every 10 seconds instead of every 5 minutes, you are giving your database 30 chances an hour to get hit by a synchronized wave of traffic instead of one.

The reverse mistake is just as common: teams panic after an incident and set the TTL to something very long, like an hour, to “reduce risk.” That does make stampedes rarer, but now stale data sticks around much longer, and when that TTL finally does expire, the request volume that has built up behind it is even bigger, so the stampede that does happen is worse.

State diagram showing a cache key moving between fresh, stale-served, and refreshing states
The safer pattern is not picking the “right” TTL, but adding a state between fresh and expired where stale data can still be served while a refresh happens in the background.

TTL length is a knob worth tuning, but it does not solve the underlying problem on its own. What actually matters is what happens in the moment a key expires, not how long it took to get there.

Myth: This only happens at massive scale, so my app is too small to worry about it

Stampedes are usually explained with big-company numbers — millions of requests per second — which makes it easy to assume the problem does not apply to a smaller app. But the trigger is not raw scale, it is concurrency on a single key at the moment it expires. A small internal dashboard queried by fifty people at 9am, all loading the same cached report, can produce the exact same pile-up on a database that normally handles that load fine when it is spread out.

The mechanism is identical whether it is fifty requests or fifty thousand: many callers discover a miss independently, and none of them know that the others are about to ask for the same thing. Scale changes how bad the outage looks. It does not change whether the bug exists.

Myth: Adding more cache servers fixes a stampede

When a cache-related outage happens, the instinctive fix is often to scale the cache layer — add more nodes, increase memory, spread keys across more shards. That helps with a cache that is out of capacity. It does nothing for a stampede, because the problem was never that the cache ran out of room. The problem is that many requests independently decided the cache was empty and all went to fetch the same answer from the same slow place at the same time. More cache servers just means more places for the same miss to happen simultaneously.

The actual fix is to make sure that when a key expires, only one request is allowed to go recompute it. A common pattern for this is called single-flight (or request coalescing): the first request that notices the miss takes a lock on that key, goes and does the expensive work, and every other request that arrives while the lock is held either waits for that first request to finish and reuses its result, or is served the old, slightly stale value in the meantime. Either way, your database sees one query instead of two thousand.

Sequence diagram showing one request acquiring a lock to refresh a cache key while other requests wait or get stale data
With single-flight, one request pays the cost of the refresh and every other concurrent request rides along, instead of repeating the work.

Two other cheap habits pair well with this: adding a small amount of random jitter to TTLs (so a batch of keys cached at the same time do not all expire in the exact same millisecond), and serving stale data for a short grace period after expiry while a background refresh runs, instead of treating “expired” as “must block until we have a fresh value.” Neither requires new infrastructure, just a different rule for what happens the moment a key goes stale.

The takeaway

A cache stampede is not a sign that caching was the wrong idea. It is a sign that the plan only covered the easy 99% of the time — when the cache is warm — and never asked what happens in the split second it is not. The fix is small: make sure only one request refreshes a hot key at a time, and let everyone else either wait briefly or use the slightly stale answer. Your database will never know the traffic spike happened at all.

Related

architecture anti-patternsarchitecture patternscachingdatabasescalability

Post navigation

Previous post
Next post

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

  • architecture (154)
  • artificial-intelligence (6)
  • books (2)
  • courses (4)
  • decision-frameworks (4)
  • finances (1)
  • java (8)
  • life-improvement (20)
  • metrics (3)
  • observability (2)
  • puzzles and challenges (3)
  • reviews (1)
  • security (8)
  • spring-boot (36)
  • Systems Thinking (3)
  • GitHub Weekly: MCP Tools Trending Right Now (September 28 – October 4, 2026)
  • The Bulkhead Pattern: Why One Slow Dependency Shouldn’t Sink the Whole Ship
  • Essential Reading: A Philosophy of Software Design — Fighting Complexity One Module at a Time
  • GitHub Weekly: Top 10 Trending Repos (September 21 – September 27, 2026)
  • The Volunteer Tax: Why Raising Your Hand in a Meeting Makes You the Permanent Owner
©2026 José da Cruz | WordPress Theme by SuperbThemes