DEV Community

Cover image for Production failures: the things that page you · 2. Thundering herd: when the cache expires all at once
Amit chakraborty
Amit chakraborty

Posted on Originally published at amitchakraborty.dev

Production failures: the things that page you · 2. Thundering herd: when the cache expires all at once

Production failures: the things that page you · Chapter 2 of 24 · Reliability · new chapter every Friday morning

By the end of this chapter: Recognise correlated cache expiry and apply jitter and request coalescing.

The problem

Say you are an engineer on call for a content API. A celebrity mentions your site, and traffic spikes to ten times its normal volume. You open your dashboards, applying the structured reading method from Chapter 1, and the symptom is clear: the primary Postgres database is pegged at 100% CPU. Connections are maxed out, and the API is returning 503s.

You check the cache hit rate. It is normally 95%, but right now it is oscillating wildly between 99% and 0%. You assume the cache time-to-live (TTL) is too short, so you deploy a change increasing it from five minutes to thirty minutes. The database recovers instantly. You close the incident. Thirty minutes later, the pager goes off again with the exact same database CPU spike.

You have built a thundering herd. When a cache key expires under heavy load, hundreds of concurrent requests all miss the cache at the exact same millisecond. They all query the database simultaneously. The database stalls, connections queue up, and the API drops requests. Increasing the TTL did not fix the problem; it just delayed the next synchronised failure by thirty minutes.

Before you start

You need a local Redis server and Node.js.

  • Node.js 20.x or higher.
  • Redis 7.2 or higher.
  • The redis npm package (v4.6.x).

Verify your environment by running this in your terminal:

node -v && redis-cli ping
Enter fullscreen mode Exit fullscreen mode

You should see v20.x.x (or higher) and PONG. If redis-cli is not installed, ensure your Redis server is running on the default port (6379).

Why does a fixed TTL synchronise traffic?

A cache with a fixed TTL acts as a traffic synchroniser. If a popular item is not in the cache, the first request fetches it from the database and writes it to Redis with a TTL of exactly 300 seconds.

For the next 299 seconds, the database does no work for that item. The API serves thousands of requests directly from Redis. But exactly 300 seconds later, the key evaporates. If your API is handling 500 requests per second for that item, all 500 requests in that specific second will check Redis, find nothing, and query the database.

The database does not see a smooth average of traffic. It sees zero requests for five minutes, followed by a wall of 500 concurrent queries, followed by zero requests for another five minutes. If those 500 queries require complex joins, the database CPU spikes, queries time out, and the API fails.

How do we decouple the expiry?

You prevent synchronised expiry by adding jitter. Jitter is randomness applied to a fixed interval.

Instead of caching an item for exactly 300 seconds, you cache it for 300 seconds minus a random percentage—usually up to 20%. One request might set a TTL of 281 seconds, the next 245 seconds, the next 299 seconds.

When you apply jitter to a large number of cached items, they no longer expire simultaneously. The database load smooths out. However, jitter only prevents different keys from expiring at the same time. It does not stop 500 concurrent requests from hitting the database when a single, highly popular key expires.

How do we handle concurrent misses?

To protect the database from a single popular key expiring, you must implement request coalescing, also known as single-flight.

When 500 requests ask for the same missing key, only the first request should query the database. The other 499 requests should wait for the first request to finish, and then share its result.

In an asynchronous language like JavaScript, you do this by caching the Promise of the database call in memory, rather than just the final result. When a request comes in, it checks the in-memory map. If a Promise exists for that key, the request awaits it. If not, it creates the Promise, puts it in the map, and queries the database.

Build it

We will build a Node.js script that simulates 50 concurrent requests for a single item that is not in the cache. We will implement both request coalescing and TTL jitter so the database is only queried once.

Step 1: Initialize the project
Create a new directory and install the Redis client. This client handles the connection to your local Redis server.

npm init -y
npm install redis@4.6.10
Enter fullscreen mode Exit fullscreen mode

Step 2: Write the coalescing cache implementation
Create a file named index.js and paste the following code.

const { createClient } = require('redis');

const redis = createClient();
redis.on('error', err => console.error('Redis Client Error', err));

// The in-memory map that holds pending database calls
const pendingRequests = new Map();

// A simulation of a slow database query
async function queryDatabase(id) {
    console.log(`[DB] Executing expensive query for ${id}...`);
    return new Promise(resolve => 
        setTimeout(() => resolve(`Data for ${id}`), 1000)
    );
}

async function getWithCoalescingAndJitter(id, baseTtlSeconds) {
    const cacheKey = `item:${id}`;

    // 1. Check the distributed cache (Redis)
    const cached = await redis.get(cacheKey);
    if (cached) {
        return cached;
    }

    // 2. Check for an in-flight request (Coalescing)
    if (pendingRequests.has(cacheKey)) {
        console.log(`[Cache] Coalescing request for ${id}`);
        return pendingRequests.get(cacheKey);
    }

    // 3. Fetch from DB, cache it, and clean up the pending map
    const promise = (async () => {
        try {
            const data = await queryDatabase(id);

            // Calculate jitter: subtract up to 20% from the base TTL
            const jitter = Math.floor(Math.random() * (baseTtlSeconds * 0.2));
            const finalTtl = baseTtlSeconds - jitter;

            await redis.set(cacheKey, data, { EX: finalTtl });
            console.log(`[Cache] Wrote to Redis with TTL ${finalTtl}s`);

            return data;
        } finally {
            // Always remove the promise, whether it resolved or rejected
            pendingRequests.delete(cacheKey);
        }
    })();

    // Store the promise so subsequent requests can await it
    pendingRequests.set(cacheKey, promise);
    return promise;
}

async function run() {
    await redis.connect();
    await redis.flushAll();

    console.log("Firing 50 concurrent requests...");

    // Fire 50 requests at the exact same time
    const promises = Array.from({ length: 50 }).map(() => 
        getWithCoalescingAndJitter('user-1', 300)
    );

    const results = await Promise.all(promises);

    console.log(`Completed ${results.length} requests.`);
    console.log(`First result: ${results[0]}`);

    await redis.disconnect();
}

run();
Enter fullscreen mode Exit fullscreen mode

Step 3: Run the code and verify the output
Execute the script.

node index.js
Enter fullscreen mode Exit fullscreen mode

You should see output similar to this:

Firing 50 concurrent requests...
[DB] Executing expensive query for user-1...
[Cache] Coalescing request for user-1
[Cache] Coalescing request for user-1
... (repeated 49 times)
[Cache] Wrote to Redis with TTL 284s
Completed 50 requests.
First result: Data for user-1
Enter fullscreen mode Exit fullscreen mode

Notice that despite 50 requests firing simultaneously, the [DB] log only appears once. The other 49 requests found the pending Promise in the map and waited for it. The TTL was set to 284 seconds, demonstrating the jitter subtracting from the 300-second base.

When this breaks

Symptom: Memory usage on the application servers grows linearly until the process crashes with FATAL ERROR: Reached heap limit Allocation failed - JavaScript heap out of memory.
Cause: The pendingRequests map is leaking Promises. This happens if a database call hangs indefinitely without a timeout, or if the code that removes the Promise from the map is inside a .then() block rather than a finally block. If the database call rejects, the .then() block is skipped, the Promise stays in the map forever, and all future requests for that key will await a rejected Promise.
Fix: Always use a finally block to delete the key from the coalescing map, as shown in the implementation above. Enforce strict timeouts on the database queries themselves.

Symptom: Redis throws OOM command not allowed when used memory > 'maxmemory'.
Cause: You implemented jitter by adding a random amount of time to the base TTL, rather than subtracting it. If your base TTL was tuned to exactly fit your available Redis memory, adding 20% to the TTL increases the total number of keys stored at any given moment by 20%, breaching your memory limit.
Fix: Always calculate jitter by subtracting from the maximum acceptable TTL. The base TTL should represent the absolute longest you are willing to serve stale data.

Symptom: The database still sees a thundering herd, just a smaller one.
Cause: In-memory request coalescing only protects the database from concurrent requests hitting the same application instance. If you run 100 Node.js containers, and a popular key expires, each container will allow one request through to the database. You will still see 100 concurrent database queries.
Fix: For moderate scale, 100 queries is acceptable and in-memory coalescing is sufficient. For massive scale, you must implement distributed coalescing using Redis locks, ensuring only one container globally is allowed to query the database for a missing key.

What it costs

Request coalescing trades database CPU for application memory and open connections.

When you coalesce 500 requests, 499 of them are sitting in your application server doing nothing but holding open HTTP connections to your clients, waiting for the single database query to resolve. If the database query takes five seconds, you are holding 500 connections open for five seconds. Under heavy load, this can easily exhaust your application server's connection limits or memory, causing it to drop traffic before it even reaches the coalescing logic.

This commits you to keeping your database queries fast. Coalescing is a shield against concurrent load, not a band-aid for slow queries. If the underlying query is inherently slow, you should not be coalescing requests; you should be pre-computing the cache asynchronously in a background worker so the API never has to wait for the database at all.

In the interview

When an interviewer asks, "How do you protect a database from a sudden traffic spike?", they are probing your understanding of cache failure modes.

A weak answer is, "I would put a Redis cache in front of it." This is weak because it assumes caches never expire and never start empty. It demonstrates that the candidate has read about caching but has not supported a high-traffic cache in production.

A strong answer names the failure mode: "I would use a cache, but I'd need to protect against a cache stampede or thundering herd when keys expire." A senior candidate will explain how they use jitter to prevent correlated expiry windows, and request coalescing to handle concurrent misses.

If you are interviewing for a Staff or Director role, the interviewer is probing your grasp of trade-offs. The follow-up will be: "What happens to the application servers while they wait for the coalesced request?" A successful answer acknowledges the cost: holding those requests consumes memory and connections. The candidate should then pivot to discussing stale-while-revalidate patterns or background cache warming as superior architectural alternatives that avoid holding client connections open entirely.

Your tasks

  1. Observe the herd: Modify the index.js script. Comment out the coalescing check (step 2 in the function). Run the script again. You should see 50 [DB] logs, simulating 50 concurrent database queries. This is the failure mode you are preventing.
  2. Break the cleanup: Restore the coalescing check. Remove the finally block entirely. Modify the queryDatabase function to throw an error for user-2. Fire a request for user-2, catch the error, and then fire a second request for user-2. Verify that the second request hangs forever because the rejected Promise is permanently stuck in the pendingRequests map.
  3. Implement stale-while-revalidate: Modify the implementation so that if the cache has expired, but the data is still in Redis (you will need to store the insertion timestamp inside the cached JSON), the function immediately returns the stale data to the user, but fires the database query in the background to update Redis for the next user.

Your tasks this week

Do the exercises above before the next chapter. Reading a tutorial and doing
one are different activities and only one of them changes what you can build.

Stuck on any of them? Say so — describe what you tried and what happened:
tell me where you got stuck. I read every one, and the questions
that come back more than twice get answered in the next chapter.

Production failures: the things that page you

Chapter 2 of 24. New chapter every Friday morning.
Next: Retry storms.

· The full syllabus and every chapter so far
· Subscribers also get the condensed notes for this chapter, the running
recap of everything the series has covered, and the extended guidance:
subscribe


Written by Amit Chakraborty — founding engineer and senior architect: React Native, AI and RAG systems, production architecture. Portfolio · LinkedIn · GitHub.

Need this built, reviewed or taught to your team? Get in touch or email amit@devamit.co.in. Available for senior and founding engineering roles, consulting and training, remote worldwide.

Top comments (0)