DEV Community

Timevolt
Timevolt

Posted on

Building a Rate Limiter: Lessons from The Matrix

The Quest Begins (The "Why")

A few months ago I was on-call for a small SaaS product that suddenly started getting hammered by a burst of traffic. Our API began returning 500s, the database groaned, and the Slack channel filled with frantic “Is it down?” messages. I felt like Neo dodging bullets in the lobby scene—except the bullets were request IDs and I had no idea how to stop them.

The problem was clear: we needed a way to protect our backend from abusive clients without penalizing legitimate users. The first thing I reached for was a simple fixed‑window counter: increment a request count per IP, reset it every minute, and reject when the count exceeded a threshold. It seemed easy, but it turned out to be a trap.

When the traffic spiked at the very start of a window, we allowed a whole minute’s worth of requests in a flash, then slammed the door shut for the rest of the minute. Legitimate users got blocked unfairly, and clever attackers learned to time their bursts just after the reset to squeeze through. I spent three hours debugging this, and when I finally saw the pattern, I felt like I’d just uncovered a hidden cheat code.

The Revelation (The Insight)

The breakthrough came when I stopped thinking about “counting requests in a slice of time” and started thinking about budgeting tokens over time. The token bucket algorithm flips the script:

  1. Bucket capacity = maximum burst you’re willing to allow.
  2. Refill rate = how many tokens are added per second (the sustainable rate).
  3. Each incoming request consumes one token; if the bucket is empty, the request is rejected.

Visually, it looks like this:

+-------------------+
|   Token Bucket    |
|  (capacity = C)   |
|   ^   refill r    |
|   |  tokens/sec   |
|   |               |
|   v               |
+-------------------+
   |   <-- request consumes 1 token
   v
+-------------------+
|   Allowed?        |
+-------------------+
Enter fullscreen mode Exit fullscreen mode

The magic is two‑fold:

  • Burst handling – If the bucket is full, you can serve up to C requests instantly, exactly what you want for legitimate spikes.
  • Smooth throttling – After the burst, tokens only refill at rate r, so the long‑term average traffic can’t exceed r, no matter how cleverly an attacker times their hits.

Compared to the fixed‑window counter, the token bucket eliminates the “window edge” problem and gives you a single knob (refill rate) to control average load, while the capacity knob lets you tune burst tolerance.

Wielding the Power (Code & Examples)

The Struggle: Fixed‑Window Counter (the trap)

# naive fixed-window rate limiter (Python‑like pseudocode)
from time import time
from collections import defaultdict

WINDOW_SECONDS = 60
LIMIT = 100

request_counts = defaultdict(int)
window_start = defaultdict(lambda: time())

def allow_fixed(ip):
    now = time()
    # reset if we've moved past the window
    if now - window_start[ip] > WINDOW_SECONDS:
        window_start[ip] = now
        request_counts[ip] = 0

    request_counts[ip] += 1
    return request_counts[ip] <= LIMIT
Enter fullscreen mode Exit fullscreen mode

What went wrong?

  • At t = 0:01 we could blast 100 requests, then at t = 0:59 another 100, effectively doubling the allowed rate.
  • Resetting the counter erases history, so a client that made 99 requests at t = 0:58 gets a fresh slate at t = 1:00 and can immediately send another 100.

The Victory: Token Bucket (the insight)

import time
import threading

class TokenBucket:
    def __init__(self, rate_per_sec: int, capacity: int):
        self.rate = rate_per_sec          # tokens added each second
        self.capacity = capacity          # max tokens in bucket
        self.tokens = capacity            # start full
        self.timestamp = time.time()
        self._lock = threading.Lock()

    def _refill(self):
        now = time.time()
        elapsed = now - self.timestamp
        # add tokens based on elapsed time, but never exceed capacity
        self.tokens = min(self.capacity,
                          self.tokens + elapsed * self.rate)
        self.timestamp = now

    def allow(self) -> bool:
        with self._lock:
            self._refill()
            if self.tokens >= 1:
                self.tokens -= 1
                return True
            return False

# Usage
limiter = TokenBucket(rate_per_sec=10, capacity=20)  # 10 req/s steady, bursts up to 20

def handler(ip):
    if limiter.allow():
        # process request
        pass
    else:
        # 429 Too Many Requests
        pass
Enter fullscreen mode Exit fullscreen mode

Why this works better:

  • The _refill method adds tokens continuously, so the effective rate is exactly rate_per_sec over any interval, eliminating the edge‑case spikes.
  • The capacity lets you absorb short bursts without rejecting legitimate traffic.
  • A single Lock keeps the structure thread‑safe for a typical web‑service worker pool.

Common Traps to Avoid

  1. Forgetting to lock – If multiple threads call allow simultaneously, two threads could read the same token count and both decrement, letting the bucket go negative. Always protect the critical section.
  2. Using integer math for refill – elapsed * self.rate can produce a fractional token; truncating too aggressively starves the bucket. Keep tokens as a float (or use a fixed‑point representation) and only consume whole tokens when checking >= 1.

Why This New Power Matters

With a token bucket in place, I could finally sleep through on‑call rotations. Our API stayed healthy under both legitimate flash‑sales traffic and the occasional mischievous bot. The insight didn’t just save us from downtime; it gave us a portable primitive we now reuse for:

  • API throttling (different tiers per plan)
  • Webhook delivery (prevent overwhelming subscriber endpoints)
  • Job queue workers (limit how many jobs a worker can pull per second)

The best part? The algorithm is tiny—under 30 lines of code—but its impact is huge. It’s the kind of tool that feels like finding a hidden upgrade in a game: suddenly you can tackle bigger bosses with confidence.

Your Turn

Grab a language of your choice and implement a token bucket limiter for a public endpoint you control. Try tweaking the rate and capacity values and watch how the behavior changes under a simple load‑test script (hey, even curl in a loop works).

Challenge: Build a small dashboard that shows the current token count in real‑time. Seeing the bucket refill live is oddly satisfying—like watching a mana bar recharge after a tough fight.

Once you’ve got it running, share your snippet or a link to a gist in the comments. I’d love to see how you adapt the pattern to your own stack, and maybe we’ll discover another hidden level together.

Happy rate‑limiting! 🚀

Top comments (0)