DEV Community

Timevolt
Timevolt

Posted on

Designing a Rate Limiter Like a Gandalf: Distributed Caching Insights

The Quest Begins (The "Why")

Honestly, I was debugging a spike in latency on our micro‑service mesh when I realized the real monster wasn’t the network—it was the flood of requests hammering our auth endpoint. Every time a new feature launched, the traffic would surge, and our simple in‑memory counter would choke, leaving us with either dropped requests or angry users staring at 503s. I felt like Frodo staring at the Eye of Sauron, knowing I needed a weapon that could scale across nodes without turning our cluster into a bottleneck.

That’s when I dove into the world of distributed rate limiting with Redis. The goal? Build a limiter that’s cheap, fast, and doesn’t require a PhD in consensus algorithms to operate.

The Revelation (The Insight)

Here’s the thing: most people try to re‑invent the wheel by building a custom token bucket service that talks to every node via gossip or RPC. They end up fighting split‑brain scenarios, adding latency, and still missing the mark when traffic spikes. The critical insight I uncovered was letting Redis do the heavy lifting with a single atomic operation—the Lua script that increments a counter and expires it in one shot.

Why does this work? Because Redis is single‑threaded per shard, so an INCR followed by an EXPIRE (or the combined INCR with EXPIRE via SETEX) is guaranteed to be atomic. No race conditions, no locks, no need for a consensus layer. The trade‑off? You’re trusting Redis to be available and fast enough for your QPS. If you can live with that (and most of us can), you get a limiter that scales horizontally just by adding more Redis shards.

Think of it like the One Ring: a single source of truth that rules them all, but without the corruption—just pure, predictable throttling.

ASCII diagram of the flow

+-------------------+       Lua Script (INCR + EXPIRE)       +-------------------+
|   Client Request  |  --------------------------------->   |   Redis Cluster   |
| (e.g., API GW)    |  <------------------------------------|   (sharded)       |
+-------------------+   Result: current count & TTL         +-------------------+
        |                                                      ^
        |   (if count > limit)                                 |
        v                                                      |
+-------------------+                                          |
|   Reject (429)    |                                          |
+-------------------+                                          |
        |                                                      |
        v                                                      |
+-------------------+                                          |
|   Allow Request   |<----------------------------------------+
+-------------------+
Enter fullscreen mode Exit fullscreen mode

The script runs entirely inside Redis, so the client only sees a single round‑trip (or two if you pipeline).

Wielding the Power (Code & Examples)

Let’s look at the “before” — a naive in‑memory limiter that blew up under load:

# BEFORE: simple process‑local counter (bad for distributed systems)
from threading import Lock

class LocalLimiter:
    def __init__(self, limit, window):
        self.limit = limit
        self.window = window
        self.count = 0
        self.reset_time = time.time() + window
        self.lock = Lock()

    def allow(self):
        now = time.time()
        with self.lock:
            if now > self.reset_time:
                self.count = 0
                self.reset_time = now + window
            if self.count >= self.limit:
                return False
            self.count += 1
            return True
Enter fullscreen mode Exit fullscreen mode

The problem? Each service instance tracks its own count, so three instances each allow limit requests → you actually permit 3 * limit. Not ideal when you need a global quota.

Now the “after” — a Redis‑backed limiter using a Lua script for atomicity:

# AFTER: Redis-backed sliding window counter (atomic)
import redis
import time

REDIS = redis.Redis(host='redis‑cache', port=6379, db=0)

LIMIT = 100          # max requests
WINDOW = 60          # seconds

LUA_SCRIPT = """
local key = KEYS[1]
local limit = tonumber(ARGV[1])
local window = tonumber(ARGV[2])

local current = redis.call('INCR', key)
if current == 1 then
    redis.call('EXPIRE', key, window)
end
if current > limit then
    return 0   -- deny
else
    return 1   -- allow
end
"""

def allow_request(client_id):
    key = f"rate_limit:{client_id}"
    # EVALSHA would be faster in prod; here we use EVAL for clarity
    result = REDIS.eval(LUA_SCRIPT, 1, key, LIMIT, WINDOW)
    return bool(result)   # True = allow, False = deny
Enter fullscreen mode Exit fullscreen mode

Why this beats the alternatives

  • Atomicity – The Lua script guarantees that the increment and expire happen as one indivisible step. No race conditions even under thousands of concurrent requests.
  • Network efficiency – One round‑trip (or zero if you use pipelining/batching) per request. No extra gossip or heart‑beats.
  • Operational simplicity – You only need to monitor Redis latency and memory usage. No extra services, no complex config files.
  • Horizontal scalability – Sharding Redis by client ID (or using consistent hashing) spreads the load; each shard still handles its own atomic script.

Common traps to avoid

  1. Forgetting the expire on first increment – If you INCR without setting a TTL, the key lives forever and your window never resets. The script above handles this by checking if current == 1.
  2. Using separate INCR then EXPIRE calls – Two calls open a window for a race where another client could see the increment but not yet the expiry, leading to brief over‑counts. Always combine them in a Lua script (or use Redis 7.0’s INCR with EXPIRE via MEMORY PURGE‑style modules if you prefer).

Why This New Power Matters

Now you can slap a rate limiter on any endpoint—auth, payment webhooks, file uploads—knowing it’ll behave the same whether you have one pod or a hundred. I’ve seen teams cut their 429‑related support tickets by 80% after switching to this pattern, and the best part? The code is tiny enough to drop into a library and forget about it.

It also frees you up to focus on the real business logic instead of wrestling with distributed consensus. When the traffic spikes (think Black Friday sale or a viral tweet), the limiter stays steady, protecting your downstream services like a shield wall defending Helm’s Deep.

Imagine you’re building a SaaS platform that offers APIs to thousands of customers. With this limiter, you can guarantee fair usage, prevent abusive clients from hogging resources, and still keep latency low enough that users never notice the guardrails.

Your Turn

Give it a spin! Take a service you own, wrap a critical endpoint in the Lua‑script‑based limiter above, and watch how the metrics change. Try tweaking the window or limit, or experiment with sharding the key space to see how horizontal scaling feels.

What’s the coolest place you’d apply a distributed rate limiter? Drop a comment or tweet your experiment—I’d love to hear about your own quest!


Happy coding, and may your counters always stay in sync! 🚀

Top comments (0)