"Design a rate limiter."
It sounds small. It is not.
This question shows up at Stripe, Google, Amazon, Cloudflare, and almost every company that runs a public API. It is short enough to fit in 45 minutes. But it is deep enough to separate a junior answer from a senior one.
In this post, I will walk through the problem the way we teach it in Grokking the System Design Interview. Step by step. Requirements first, then estimates, then algorithms, then architecture, then the hard parts.
Let's go.
What Is a Rate Limiter?
A rate limiter controls how many requests a client can send in a given time window.
If a user is allowed 100 requests per minute, request number 101 gets rejected. Usually with an HTTP 429 Too Many Requests response.
Think of it like a nightclub bouncer with a clicker. The club holds 200 people. The bouncer counts everyone who walks in. When the counter hits 200, the next person waits outside. Nobody is being rude. The building just has a limit.
Rate limiters exist for three main reasons:
- Protect your servers. One buggy script or one bad actor should not take down the whole service.
- Control costs. If every request hits a paid third-party API, unlimited traffic means an unlimited bill.
- Keep things fair. One heavy user should not slow the system down for everyone else.
Step 1: Clarify the Requirements
Never start drawing boxes right away. Ask questions first. In a real interview, this is where you earn your first points.
Here are the questions I would ask:
- Client-side or server-side? Server-side. Client-side limits are easy to bypass.
- What do we limit by? User ID, IP address, API key, or a mix. Let's support all of them through configurable rules.
- What scale? Think large. Millions of users, a few hundred thousand requests per second at peak.
- Distributed? Yes. Many API servers sit behind a load balancer. The limit must hold across all of them.
-
What happens when a request is throttled? Return
429with headers that tell the client when to retry. - Hard or soft limit? Mostly hard. A tiny overshoot during a race is acceptable if it keeps latency low.
From those answers, we can write the requirements down.
Functional requirements
- Limit requests per client based on configurable rules (for example, 100 requests per minute per user).
- Return
429when a client exceeds the limit. - Tell the client its remaining quota through response headers.
Non-functional requirements
- Low latency. The check sits on every request. It must add only a millisecond or two.
- Highly available. If the rate limiter breaks, the API should not break with it.
- Accurate enough. Small errors are fine. Letting 10x the limit through is not.
- Scalable. Works across many servers and regions.
Step 2: Back-of-the-Envelope Estimates
Let's keep the math simple.
- 10 million daily active users.
- Peak traffic of 200,000 requests per second.
- Each counter needs a key (about 50 bytes) plus a count and a timestamp (about 16 bytes). Call it roughly 100 bytes with overhead.
Memory for 10 million active counters:
10,000,000 x 100 bytes = 1 GB
That fits in memory on a single Redis node. But 200,000 checks per second, with a read and a write each, is better spread across a small Redis cluster. That also gives us replication for availability.
The takeaway: storage is cheap here. Latency and coordination are the real problems.
Step 3: Pick a Rate Limiting Algorithm
This is the heart of the interview. There are five classic algorithms. You do not need all of them in detail, but you should know the trade-offs.
1. Token Bucket
A bucket holds tokens. Tokens refill at a fixed rate. Every request takes one token. No token, no entry.
- Bucket size = how big a burst you allow.
- Refill rate = your long-term average limit.
Pros: Simple. Memory efficient. Allows short bursts, which matches real user behavior.
Cons: Two parameters to tune.
This is what Amazon API Gateway and Stripe use. It is the safest default in an interview.
2. Leaky Bucket
Requests enter a queue. The queue drains at a fixed rate, like water leaking from a hole in a bucket. If the queue is full, new requests are dropped.
Pros: Produces a perfectly smooth output rate.
Cons: Bursts of traffic fill the queue with old requests. Recent requests get starved.
Good for systems that need steady processing, like a payment processor that calls a bank API.
3. Fixed Window Counter
Split time into fixed windows, like one minute each. Keep one counter per window. Reject when the counter passes the limit.
Pros: Very simple. One counter per user.
Cons: The edge problem. A user can send 100 requests at 12:00:59 and another 100 at 12:01:00. That is 200 requests in two seconds, double the limit.
4. Sliding Window Log
Store the timestamp of every request. On each new request, drop timestamps older than the window and count what is left.
Pros: Perfectly accurate.
Cons: Memory heavy. A limit of 1,000 requests per hour means storing up to 1,000 timestamps per user.
5. Sliding Window Counter
A hybrid. Keep counters for the current and previous window. Estimate the count with a weighted average.
estimate = current_count + previous_count x (overlap percentage of previous window)
If we are 30% into the current minute, the previous minute still overlaps by 70%. So with 80 requests last minute and 20 so far this minute:
20 + 80 x 0.7 = 76
Pros: Smooths out the edge problem. Memory efficient.
Cons: It is an approximation. Cloudflare measured it at about 0.003% error across 400 million requests, which is more than good enough.
Quick Comparison
| Algorithm | Memory | Accuracy | Allows bursts | Best for |
|---|---|---|---|---|
| Token bucket | Low | Good | Yes | General API limits |
| Leaky bucket | Low | Good | No (smooths) | Steady downstream calls |
| Fixed window | Very low | Weak at edges | Yes, at edges | Simple, low-stakes limits |
| Sliding window log | High | Exact | No | Strict, low-volume limits |
| Sliding window counter | Low | Very good | Limited | High-scale APIs |
What I would say in the interview: "I will go with the token bucket as the default because it is cheap, handles bursts, and is used in production by large API providers. If the business needs strict smoothing, I would switch to sliding window counter."
Step 4: High-Level Design
Now we draw boxes.
Here is the flow:
- Client sends a request.
- API Gateway / Rate Limiter middleware receives it. This is where the check happens.
- The middleware builds a key, like
rate:user:42, and checks the counter in Redis. - Rules Service stores the limits (for example, free tier gets 100 per minute, paid tier gets 1,000). The middleware caches these rules locally and refreshes them every few seconds.
- If the client is under the limit, the request goes to the API servers.
- If not, the middleware returns
429right away. The backend never sees it.
Where should the rate limiter live?
Three options:
- Inside each service. Simple, but every team rebuilds the same logic.
- In the API gateway. Central and consistent. This is the common choice.
- As a separate service. Flexible, but adds a network hop to every request.
I would pick the API gateway. If the company already runs a gateway like Kong, Envoy, or AWS API Gateway, rate limiting is a plugin away.
Why Redis?
- In-memory, so reads and writes take under a millisecond.
- Built-in
INCRandEXPIREcommands. - Supports Lua scripts, which run atomically. That matters a lot. More on that next.
Response headers
Be kind to clients. Tell them where they stand.
HTTP/1.1 429 Too Many Requests
X-RateLimit-Limit: 100
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 1696608000
Retry-After: 30
Good clients read these headers and back off on their own. That reduces load before the limiter even has to say no.
Step 5: Deep Dives (Where Senior Candidates Shine)
A basic design gets you to "hire" at the mid level. These deep dives push you toward senior and staff.
Deep Dive 1: Race Conditions
Picture two API servers handling requests from the same user at the same moment.
- Server A reads the counter: 99.
- Server B reads the counter: 99.
- Both see "under 100" and allow the request.
- Both write 100.
Two requests went through. The counter says one. Do this at scale and limits leak badly.
Fix: Make read-check-write a single atomic step. In Redis, use a Lua script. Redis runs each script start to finish without interruption.
Here is a token bucket in Lua:
-- KEYS[1] = bucket key
-- ARGV[1] = capacity, ARGV[2] = refill rate per second, ARGV[3] = now (ms)
local capacity = tonumber(ARGV[1])
local rate = tonumber(ARGV[2])
local now = tonumber(ARGV[3])
local bucket = redis.call("HMGET", KEYS[1], "tokens", "ts")
local tokens = tonumber(bucket[1]) or capacity
local ts = tonumber(bucket[2]) or now
-- refill based on elapsed time
local elapsed = math.max(0, now - ts) / 1000
tokens = math.min(capacity, tokens + elapsed * rate)
local allowed = 0
if tokens >= 1 then
tokens = tokens - 1
allowed = 1
end
redis.call("HMSET", KEYS[1], "tokens", tokens, "ts", now)
redis.call("EXPIRE", KEYS[1], math.ceil(capacity / rate) * 2)
return { allowed, math.floor(tokens) }
One network round trip. No races. No locks.
Deep Dive 2: Scaling Redis
One Redis node will not last forever. Shard the keys across a Redis cluster using consistent hashing on the client key. All counters for user:42 land on the same shard, so the Lua script still works.
Add a replica per shard for failover.
Deep Dive 3: What If Redis Goes Down?
This is a favorite follow-up. You have two choices:
- Fail open. Let all requests through. The API stays up, but you lose protection for a while.
- Fail closed. Reject all requests. You stay protected, but your API is effectively down.
For most products, fail open is the right call. A short window without limits hurts less than a full outage. Pair it with a local in-memory fallback limiter on each gateway node, so you still get rough protection.
For security-critical limits, like login attempts, fail closed may be the safer choice. Say this trade-off out loud. Interviewers love it.
Deep Dive 4: Multi-Region
If you run in three regions, a global limit gets tricky. Syncing every counter across regions on every request adds too much latency.
Common approaches:
- Local limits per region. Split the global limit (for example, 1,000 per minute becomes about 333 per region). Simple, slightly inaccurate.
- Async sync. Each region counts locally and syncs totals every second or so. Accepts a small overshoot for much lower latency.
Most real systems choose eventual consistency here. Perfect global accuracy is rarely worth the cost.
Deep Dive 5: Reducing Load on Redis
At very high traffic, even Redis can become a hot spot. Two tricks help:
- Local pre-check. Each gateway keeps a small in-memory counter and only syncs with Redis in batches.
-
Hot key handling. A single huge customer can overload one shard. Split their key into sub-keys (
rate:user:42:0torate:user:42:3) and divide their limit across them.
What Interviewers Actually Look For
Here is how I would grade this question, based on years of running these interviews.
| Level | What a strong answer includes |
|---|---|
| Junior / Mid | Clear requirements, one algorithm explained well, gateway plus Redis design |
| Senior | Compares algorithms with trade-offs, solves race conditions, discusses fail open vs fail closed |
| Staff | Multi-region strategy, hot keys, rule management, and how clients should react to 429
|
You do not need every deep dive. Pick two and go deep. Depth beats breadth.
Common Mistakes to Avoid
- Jumping to Redis before asking a single question.
- Explaining five algorithms and never picking one.
- Forgetting the race condition.
- Ignoring what happens when the rate limiter itself fails.
- Storing counters in a relational database. Too slow for a per-request check.
Key Takeaways
- A rate limiter caps how many requests a client can make in a time window.
- Always clarify: what to limit by, scale, distributed or not, hard or soft limits.
- Token bucket is the best default. Sliding window counter is great for high-scale accuracy.
- Put the limiter in the API gateway and store counters in Redis.
- Use Lua scripts to make the check atomic and avoid race conditions.
- Fail open for most APIs. Fail closed for security-critical limits.
- For multiple regions, accept small inaccuracy in exchange for low latency.
FAQ
What HTTP status code should a rate limiter return?
429 Too Many Requests, along with a Retry-After header that tells the client when to try again.
Which rate limiting algorithm is best?
There is no single best one. Token bucket is the most common default because it is simple, cheap, and allows short bursts. Sliding window counter is better when you need smoother limits at scale.
Why use Redis instead of a database?
The check runs on every single request. Redis is in-memory and answers in under a millisecond. A disk-based database would add too much latency.
What is the difference between token bucket and leaky bucket?
Token bucket allows bursts up to the bucket size. Leaky bucket smooths everything into a steady output rate, even if requests arrive in bursts.
Should the rate limiter fail open or fail closed?
Usually fail open, so a limiter outage does not become an API outage. For login or payment endpoints, failing closed can be safer.
How long should I spend on this in a 45-minute interview?
About 5 minutes on requirements, 5 on estimates, 10 on algorithms, 10 on the high-level design, and the rest on one or two deep dives.
Wrapping Up
The rate limiter is a great interview question because it looks easy. The candidates who stand out are the ones who slow down, ask the right questions, and talk openly about trade-offs.
If you want more walkthroughs like this one, including URL shorteners, news feeds, and chat systems, they are all covered in Grokking the System Design Interview.
Which system design question should I break down next? Let me know in the comments.



Top comments (0)