DEV Community

Usman Khan
Usman Khan

Posted on Originally published at ctousman.com

Distributed API Rate Limiting & Idempotency at Scale: Redis Keyspace Architecture & Lock Patterns

Distributed API Rate Limiting & Idempotency at Scale: Redis Keyspace Architecture & Lock Patterns

Protecting high-throughput public APIs requires two non-negotiable guarantees: shielding downstream services from traffic spikes and guaranteeing that duplicate client payloads execute exactly once. Here is how to engineer scalable rate-limiting and idempotency subsystems using Redis.

1. Distributed Rate Limiting: Sliding Window Log vs. Token Bucket

When operating a multi-tenant platform, naïve fixed-window rate limiters suffer from boundary burst vulnerabilities—where a client can exhaust 2x their limit during the reset window boundary. To enforce smooth, predictable rate limits, modern systems rely on atomic Redis operations.

The Sliding Window Log Pattern in Redis

Using Redis sorted sets (ZSET), we can track timestamps for every request in a sliding time range, prune expired tokens atomically, and evaluate current volume in O(log N + M) time complexity.

// Atomic Sliding Window Rate Limiter (Lua Script executed in Redis)
const SLIDING_WINDOW_LUA = `
local key = KEYS[1]
local now = tonumber(ARGV[1])
local window = tonumber(ARGV[2])
local limit = tonumber(ARGV[3])
local clearBefore = now - window

-- 1. Remove timestamps older than the sliding window frame
redis.call('ZREMRANGEBYSCORE', key, '-inf', clearBefore)

-- 2. Count current valid requests within window
local currentRequests = redis.call('ZCARD', key)

if currentRequests < limit then
    -- 3. Add current request timestamp with microsecond unique payload
    redis.call('ZADD', key, now, now .. ':' .. redis.call('INCR', 'global:nonce'))
    redis.call('EXPIRE', key, math.ceil(window / 1000))
    return {1, limit - currentRequests - 1} -- Allowed, remaining limit
else
    return {0, 0} -- Rate limit exceeded
end
`;
Enter fullscreen mode Exit fullscreen mode

2. API Idempotency: Eliminating Duplicate Side Effects

Network timeouts, client retries, and webhook delivery loops cause duplicate payload deliveries. An Idempotency Layer intercepts requests containing an Idempotency-Key HTTP header to ensure underlying business logic (e.g., billing charge, database insert) runs exactly once.

State Machine of an Idempotent Request

  • State 1: IN_PROGRESS (Mutex Lock): A distributed lock is acquired on idempotency:{tenant_id}:{key} with a short TTL (e.g., 30 seconds) to prevent concurrent race condition executions.
  • State 2: COMPLETED (Cached Response): Once the handler finishes successfully, the status and serialized HTTP response body are cached in Redis with a long TTL (e.g., 24 to 72 hours).
  • State 3: FAILED: If an unhandled application error occurs during processing, the lock is evicted immediately so subsequent retries can re-execute cleanly.

Implementation Pattern: Middleware Interceptor

// Fastify/Express Middleware Pattern for Idempotent APIs
async function idempotencyMiddleware(req, res, next) {
  const idempotencyKey = req.headers['idempotency-key'];
  if (!idempotencyKey) return next();

  const lockKey = `idempotency:${req.tenant.id}:${idempotencyKey}`;

  // Attempt atomic status set or lock acquisition
  const acquired = await redis.set(
    lockKey,
    JSON.stringify({ status: 'IN_PROGRESS' }),
    'NX',
    'EX',
    30 // 30 second execution lock
  );

  if (!acquired) {
    // Key exists: retrieve cached state or signal lock contention
    const rawState = await redis.get(lockKey);
    if (!rawState) return res.status(429).send({ error: 'Concurrent request lock conflict' });

    const state = JSON.parse(rawState);
    if (state.status === 'IN_PROGRESS') {
      return res.status(409).send({ error: 'Request with this Idempotency-Key is currently processing' });
    }

    if (state.status === 'COMPLETED') {
      // Replay identical response directly from cache
      res.header('X-Cache-Lookup', 'HIT-IDEMPOTENT');
      return res.status(state.statusCode).send(state.body);
    }
  }

  // Intercept res.send to capture downstream output and save state
  const originalSend = res.send.bind(res);
  res.send = (body) => {
    if (res.statusCode >= 200 && res.statusCode < 300) {
      redis.set(
        lockKey,
        JSON.stringify({ status: 'COMPLETED', statusCode: res.statusCode, body }),
        'EX',
        86400 // Cache payload for 24 hours
      );
    } else {
      // Evict key on application errors to allow retries
      redis.del(lockKey);
    }
    return originalSend(body);
  };

  next();
}
Enter fullscreen mode Exit fullscreen mode

3. Operational Pitfalls to Avoid

  • Ignoring Memory Overhead in Redis: Storing full HTTP response bodies in Redis idempotency records can trigger memory pressure. Mitigation: For heavy payloads, store large response blobs in S3/Object Storage and cache only signed retrieval URLs in Redis.
  • Unbounded Redis Cluster Memory Growth: Always explicitly define key expiration (EXPIRE) policies on rate limiters and idempotency records to prevent eviction failures during high-traffic surges.
  • Clock Drift in Distributed Redis Clusters: Ensure NTP synchronization across node clusters when using timestamp-based sorted set scores for rate limiting.

Top comments (0)