Every serious API needs rate limiting. Most people reach for a library the moment they hear "rate limiter" — express-rate-limit, bottleneck, some Redis-backed package with a dozen config options they'll never touch.
But the algorithm underneath almost all of them — token bucket — is genuinely simple. You can write a correct, production-usable in-memory version in about 50 lines of JavaScript, no dependencies. Understanding it from scratch also means you'll actually know what those library config options are doing instead of copy-pasting them.
Let's build it.
The idea, in one paragraph
Imagine a bucket that holds tokens. Every request needs one token to go through. The bucket refills at a steady rate (say, 10 tokens per second) up to some maximum capacity. If the bucket is empty when a request arrives, the request gets rejected (or delayed). That's it. That's the whole algorithm — the rest is just implementation detail.
The reason token bucket is so popular over simpler approaches ("max N requests per fixed minute") is that it naturally allows bursts while still enforcing an average rate. A user who's been idle can burst up to the bucket's capacity, then has to slow down to the refill rate — which matches how real traffic actually behaves.
The core implementation
class TokenBucket {
constructor({ capacity, refillRatePerSec }) {
this.capacity = capacity; // max tokens the bucket can hold
this.tokens = capacity; // start full
this.refillRatePerSec = refillRatePerSec;
this.lastRefillTimestamp = Date.now();
}
_refill() {
const now = Date.now();
const elapsedSeconds = (now - this.lastRefillTimestamp) / 1000;
const tokensToAdd = elapsedSeconds * this.refillRatePerSec;
if (tokensToAdd > 0) {
this.tokens = Math.min(this.capacity, this.tokens + tokensToAdd);
this.lastRefillTimestamp = now;
}
}
tryConsume(tokensRequested = 1) {
this._refill();
if (this.tokens >= tokensRequested) {
this.tokens -= tokensRequested;
return true;
}
return false;
}
}
That's the entire algorithm — about 25 lines. No timers, no setInterval, no background refill loop. The trick is lazy refill: instead of a ticking clock adding tokens every N milliseconds, we just calculate "how many tokens should have been added since the last check" whenever a request actually comes in. This is simpler, cheaper, and has no drift.
Wrapping it for per-user rate limiting
A single bucket rate-limits your entire server, which isn't usually what you want. In practice you need one bucket per user (or per IP, per API key, etc.). Here's a manager class that handles that:
class RateLimiter {
constructor({ capacity, refillRatePerSec }) {
this.capacity = capacity;
this.refillRatePerSec = refillRatePerSec;
this.buckets = new Map(); // key -> TokenBucket
}
_getBucket(key) {
if (!this.buckets.has(key)) {
this.buckets.set(
key,
new TokenBucket({
capacity: this.capacity,
refillRatePerSec: this.refillRatePerSec,
})
);
}
return this.buckets.get(key);
}
allow(key, cost = 1) {
return this._getBucket(key).tryConsume(cost);
}
}
And that's the full implementation — right around 50 lines total, zero dependencies.
Using it as Express middleware
Here's where it becomes actually useful:
const limiter = new RateLimiter({ capacity: 10, refillRatePerSec: 2 });
function rateLimitMiddleware(req, res, next) {
const key = req.ip; // or req.user.id, or an API key from headers
if (limiter.allow(key)) {
next();
} else {
res.status(429).json({ error: "Too many requests. Slow down." });
}
}
app.use("/api/", rateLimitMiddleware);
With capacity: 10 and refillRatePerSec: 2, each client can burst up to 10 requests instantly, then is limited to a sustained 2 requests/second after that. Tune both numbers to match your actual traffic patterns — capacity controls burst tolerance, refill rate controls the long-run average.
Testing it actually works
A quick sanity check, no test framework needed:
const bucket = new TokenBucket({ capacity: 5, refillRatePerSec: 1 });
// Burst: first 5 requests succeed immediately
for (let i = 0; i < 5; i++) {
console.log(bucket.tryConsume()); // true x5
}
// 6th request fails — bucket is empty
console.log(bucket.tryConsume()); // false
// Wait ~1 second, one token refills
setTimeout(() => {
console.log(bucket.tryConsume()); // true
console.log(bucket.tryConsume()); // false again
}, 1000);
Run that and you'll see exactly the burst-then-throttle behavior we designed for.
The catch: this only works on a single instance
This in-memory version resets if your process restarts, and — more importantly — each server instance has its own separate buckets. If you're running 4 instances behind a load balancer, a client could get 4x their intended limit by hitting different instances.
For a single-server app, side project, or internal tool, this is completely fine as-is. For a distributed production system, you need the buckets to live somewhere shared — which is exactly what Redis-backed rate limiters (like rate-limiter-flexible) are doing under the hood: the same token bucket math, just with INCR and EXPIRE commands against a shared store instead of a local Map. The algorithm you just wrote doesn't change — only where the state lives does.
A quick comparison to other approaches
You'll sometimes see fixed window ("max 100 requests per clock-minute") or sliding window log (timestamp array per user) used instead. Fixed window is simpler but has a nasty edge case: a client can send 100 requests at 11:00:59 and another 100 at 11:01:00 — 200 requests in two seconds, technically within limits. Token bucket doesn't have this problem because tokens refill continuously, not in discrete resets.
Sliding window log is more precise than token bucket but costs more memory (you're storing every timestamp, not just a number). For most APIs, token bucket is the sweet spot: simple, cheap, and burst-tolerant without being exploitable.
Wrapping up
The full rate limiter — bucket logic, per-user manager, and middleware wiring — comes in under 50 lines and has zero external dependencies. You now know exactly what's happening every time you set windowMs and max in some rate-limiting library, because you just built the thing those options configure.
If you want to take it further: swap the in-memory Map for Redis calls to make it distributed, or add a Retry-After header calculated from (tokensRequested - this.tokens) / this.refillRatePerSec so clients know exactly how long to back off.
What rate-limiting edge case has bitten you in production? Drop it in the comments — I'll bet it's a fixed-window bug.
Top comments (0)