Designing a Rate Limiter That Survives a Flash Sale: Token Bucket vs Sliding Window josedacruz, August 13, 2026August 13, 2026 TL;DR: A simple “count requests per minute” rate limiter looks fine in testing but can let through double the traffic right at the boundary between two time windows — exactly when a flash sale or a bot swarm is most likely to hit you. Token bucket and sliding window algorithms fix this in different ways, and picking between them is really a trade-off between memory cost and precision, not a search for one “correct” answer. Picture a small e-commerce team launching a flash sale at 10:00 AM. Their checkout API allows 100 requests per second per customer, enforced by a rate limiter one engineer built in an afternoon: a counter that resets every second. It’s been running fine for months. At 9:59:59, right before the sale starts, a wave of shoppers refreshes the page. The counter allows 100 requests in that last second. Then, at the stroke of 10:00:00, the counter resets to zero — and lets through another 100 requests immediately. Two allowed windows, one right after the other, and suddenly the checkout service is handling 200 requests in about a second, not the 100 it was designed for. The database connection pool maxes out. Real customers start seeing timeouts during the two minutes that matter most all quarter. Nothing was “broken.” The rate limiter did exactly what it was built to do. It just wasn’t built to handle the one moment it actually needed to. What a rate limiter is actually protecting you from A rate limiter’s job is simple to say and easy to get wrong in practice: make sure no client sends more requests than your system can safely handle in a given period of time. “Client” might mean a single user, an API key, or an IP address. The point is the same — somewhere downstream, you have a database, a payment processor, or a queue that has a real limit on how much work it can do per second, and the rate limiter is the guard standing in front of it. Most teams reach for the simplest possible version first: a counter tied to a clock. Count how many requests came in during the current window (say, one second or one minute). If the count is under the limit, allow the request. If not, reject it with an HTTP 429 (“Too Many Requests”). When the window ends, reset the counter to zero and start again. This is called a fixed window counter, and it’s a completely reasonable place to start. The problem only shows up at the edges. The fixed window trap The bug in the flash sale story is a well-known failure mode of fixed windows, sometimes called the “boundary burst” problem. A fixed window only remembers how many requests happened in the current window. It has no memory of what happened in the window right before it. So if a burst of traffic lands right before a window resets, and another burst lands right after, the limiter sees two separate windows that each look compliant — even though, measured over any one-second slice that straddles the reset, twice the intended traffic got through. The fixed window resets to a fresh quota at exactly 10:00:00 — so a burst just before the reset and a burst just after both get waved through. This isn’t a hypothetical corner case. It’s most likely to bite you at the exact moments you care about most: a marketing email that goes out on the hour, a scheduled job that kicks off at midnight, or a flash sale timed to the second. Traffic tends to cluster around round numbers, and round numbers are exactly where a fixed window is weakest. The fix isn’t to abandon the idea of a “budget per period.” It’s to change how that budget is tracked. Token bucket: letting bursts through on purpose A token bucket flips the mental model. Instead of counting requests within a rigid clock window, imagine a bucket that holds a fixed number of tokens — say, 100. Every request has to take one token to proceed. If the bucket is empty, the request is rejected. Separately, on a steady clock (say, 10 tokens every 100 milliseconds), the bucket refills, up to its maximum capacity. The key difference from a fixed window is that there’s no hard reset. Tokens trickle in continuously, and the bucket has a cap on how many can build up at once. That cap is what lets you allow short, legitimate bursts (a user’s browser firing off a handful of requests at once when a page loads) without ever letting sustained traffic exceed the intended average rate. Two independent processes: requests spend tokens, a clock refills them. No window ever resets to a clean slate. For the flash-sale team, a token bucket sized around 100 tokens with a steady refill rate would have meant no double-counted burst at 10:00:00. A customer could still burst briefly if they had unused tokens saved up, but the system could never be pushed to sustain 200 requests per second, because tokens simply wouldn’t exist fast enough to allow it. Token buckets are cheap to run — each client needs just two numbers in memory: a token count and a last-refill timestamp — which is a big part of why they’re the default choice in tools like AWS API Gateway and many load balancers. Sliding window: closer to exact, at a higher cost There’s a third option worth knowing about, especially if you need tighter accuracy than a token bucket gives you: the sliding window. Instead of a fixed clock boundary, a sliding window looks back exactly N seconds from right now, continuously, and counts the requests inside that rolling range. The most precise version, a sliding window log, stores a timestamp for every request and simply counts how many timestamps fall within the last N seconds. This is exact — there’s no boundary trick to exploit — but it can get expensive in memory if a client is sending a lot of traffic, since you’re storing one entry per request. A cheaper compromise, the sliding window counter, blends the current fixed window’s count with a weighted portion of the previous window’s count, based on how far into the current window you are. It’s not perfectly exact, but it’s very close, and it costs about the same as a plain fixed window counter to run. None of these algorithms wins on every axis — the right pick depends on which one your system can least afford to run out of: memory or accuracy. So which do you actually reach for? If you’re protecting a single hot endpoint and can afford a little extra memory per client, a sliding window log gives you the most accurate guarantee. If you’re rate-limiting millions of API keys and memory efficiency matters, a token bucket or sliding window counter gets you 95% of the accuracy for a fraction of the cost. Fixed windows are fine for cases where the limit is a loose guideline rather than a hard capacity constraint — like a generous “1000 requests per day” quota where a brief 2x burst genuinely doesn’t matter. The takeaway The flash-sale team didn’t need a more complicated system. They needed to notice that their rate limiter’s simplicity was hiding an assumption: that traffic arrives smoothly, never clustering right at a clock boundary. That assumption breaks precisely when a rate limiter matters most — during a spike. The general lesson travels well beyond rate limiting. Any time you’re protecting a system with a “budget per time period,” ask what happens at the edges of that period, not just in the steady state. A fixed window is easy to reason about and easy to implement, which is exactly why it’s a common first choice — and exactly why it’s worth replacing before the first real burst finds the seam. Related architecture API Designdistributed systemsrate limitingsystem designtrade-offs