DEV Community

Cover image for 3 Redis Design Failures You Should Avoid Before They Become Production Incidents
Gaurav Talesara
Gaurav Talesara

Posted on

3 Redis Design Failures You Should Avoid Before They Become Production Incidents

Caching is usually introduced for one simple reason:

Take pressure off the database and make reads faster.

The architecture often looks straightforward:

Client
  |
  v
Application
  |
  +----> Redis
  |
  +----> Database
Enter fullscreen mode Exit fullscreen mode

A request checks Redis first.

If the value exists, return it.

If it doesn't, query the database, put the result into Redis, and return it.

Simple.

Until the system meets real production traffic.

That's where caching stops being just a performance optimization and becomes a system-design problem.

Let's walk through how a seemingly simple cache evolves through three stages—and the failure modes that appear at each stage.


1. Stale Data

The first version of a cache is usually cache-aside.

The application owns the caching logic.

The flow is:

Request
   |
   v
Check Redis
   |
   +---- HIT ----> Return cached value
   |
   +---- MISS ---> Query database
                     |
                     v
                 Store in Redis
                     |
                     v
                  Return
Enter fullscreen mode Exit fullscreen mode

For read-heavy workloads, this can dramatically reduce database traffic.

Suppose we have a user profile:

{
  "id": 42,
  "name": "Gaurav",
  "email": "gaurav@example.com"
}
Enter fullscreen mode Exit fullscreen mode

The application stores it under:

user:42
Enter fullscreen mode Exit fullscreen mode

Everything works.

Until the user changes their name.

The database is updated:

Gaurav
    ↓
Gaurav Talesara
Enter fullscreen mode Exit fullscreen mode

But Redis still contains:

Gaurav
Enter fullscreen mode Exit fullscreen mode

Now you have two versions of reality.

Database → Gaurav Talesara
Redis    → Gaurav
Enter fullscreen mode Exit fullscreen mode

The database is correct.

The application can still return the wrong answer.

This is stale data.

Why this matters

Caching introduces another copy of your data.

The moment you do that, you have to answer:

When does the cached copy stop being valid?

That's the beginning of cache consistency design.


2. TTL Helps—but Doesn't Solve Everything

A common next step is adding a TTL (Time To Live).

For example:

user:42
TTL = 5 minutes
Enter fullscreen mode Exit fullscreen mode

After five minutes, Redis considers the entry expired.

The next request becomes a cache miss:

Redis MISS
    |
    v
Database
    |
    v
Fresh value
    |
    v
Redis
Enter fullscreen mode Exit fullscreen mode

This gives you eventual freshness without requiring every application write to explicitly update the cache.

But TTL introduces a trade-off.

Longer TTL

You get:

  • fewer database reads
  • better cache hit rate

But potentially:

  • older data
  • longer periods of inconsistency

Shorter TTL

You get:

  • fresher data

But potentially:

  • more cache misses
  • more database traffic
  • lower cache efficiency

So TTL isn't really a "freshness setting."

It's a consistency vs. performance trade-off.

And some data simply can't wait five minutes.

If a user changes their email address, waiting for TTL expiration may be unacceptable.

That's where explicit invalidation comes in.


3. Invalidation on Write

Instead of waiting for the cache to expire, the application can invalidate the relevant key when the database changes.

For example:

UPDATE users
SET name = 'Gaurav Talesara'
WHERE id = 42;
Enter fullscreen mode Exit fullscreen mode

Then:

DELETE user:42 FROM Redis
Enter fullscreen mode Exit fullscreen mode

The next read becomes:

Redis MISS
    |
    v
Database
    |
    v
Fresh data
    |
    v
Redis
Enter fullscreen mode Exit fullscreen mode

This gives you a much fresher cache.

But now the application has to maintain the relationship between:

database writes ↔ cache invalidation

And that creates its own failure scenarios.

What if the database update succeeds but cache invalidation fails?

You can end up with:

Database → NEW value
Redis    → OLD value
Enter fullscreen mode Exit fullscreen mode

What if the invalidation happens first but the database transaction fails?

Now the cache is empty even though the database still contains the old value.

There is no universal answer here.

The right approach depends on the consistency requirements of the application.

But there is another problem that has nothing to do with stale data.

And this one can take down the database.


4. Cache Stampede

Imagine you have a popular product page:

product:123
Enter fullscreen mode Exit fullscreen mode

Normally it receives thousands of requests.

Because the data is cached, most of those requests never reach the database.

That's exactly what we wanted.

Now the cache entry expires.

At almost the same moment, thousands of requests arrive.

They all see:

CACHE MISS
Enter fullscreen mode Exit fullscreen mode

And they all do this:

Request 1 → DB
Request 2 → DB
Request 3 → DB
Request 4 → DB
...
Request 5000 → DB
Enter fullscreen mode Exit fullscreen mode

The database was protected by Redis a moment ago.

Now thousands of requests are hitting it simultaneously.

This is a cache stampede.

Another name you'll often see for this pattern is thundering herd.

The important point is that the problem isn't simply that the cache missed.

The problem is that many requests independently react to the same cache miss.


5. Defense #1: Single Flight

One way to solve this is to allow only one request to refresh a missing key.

Suppose 5,000 requests miss:

             Redis MISS
                  |
        +---------+---------+
        |                   |
    Request 1            Requests 2-5000
        |                   |
      DB query              WAIT
        |                   |
        +--------+----------+
                 |
          Fresh result
                 |
             Redis SET
                 |
        Everyone gets result
Enter fullscreen mode Exit fullscreen mode

Request #1 becomes responsible for fetching the data.

The other requests wait for the result instead of independently hitting the database.

This pattern is often called single-flight or request coalescing.

The key idea is:

One cache miss should not become 5,000 database queries.

This is particularly useful when a small number of keys receive a very large amount of traffic.


6. Defense #2: TTL Jitter

There is another failure pattern that can happen even when you aren't dealing with one single hot key.

Imagine millions of cache entries are created at roughly the same time.

If every entry gets exactly:

TTL = 300 seconds
Enter fullscreen mode Exit fullscreen mode

then many of them can expire around the same time.

That can create a synchronized wave of cache misses.

Instead, you can introduce jitter.

Rather than:

TTL = 300 seconds
Enter fullscreen mode Exit fullscreen mode

you might use something conceptually like:

TTL = 300 + random(0, 60)
Enter fullscreen mode Exit fullscreen mode

Now expiration is spread across time.

Instead of:

                    ↓
              EVERYTHING EXPIRES
                    ↓
             DATABASE SPIKE
Enter fullscreen mode Exit fullscreen mode

you get something more like:

      ↓      ↓          ↓    ↓       ↓
   expire  expire     expire expire  expire
Enter fullscreen mode Exit fullscreen mode

The database sees a more distributed load pattern.

TTL jitter doesn't eliminate cache misses.

It reduces the probability of synchronized expiration becoming a traffic spike.


7. Defense #3: Hot-Key Replication

Now consider a different situation.

One key becomes extremely popular.

For example:

product:123
Enter fullscreen mode Exit fullscreen mode

might represent a product that suddenly goes viral.

You can have thousands or millions of requests for essentially the same piece of data.

Even if Redis is distributed, concentrating a huge amount of traffic around one key can create a bottleneck.

This is a hot key.

One approach is to replicate frequently accessed data across multiple cache nodes so requests don't all depend on the same location.

Conceptually:

                  product:123
                       |
            +----------+----------+
            |          |          |
            v          v          v
         Redis-1    Redis-2    Redis-3
Enter fullscreen mode Exit fullscreen mode

Now requests can be distributed rather than concentrating entirely on one cache location.

But this introduces another design question:

How do you keep replicated hot data consistent?

Again, the optimization introduces another system-design trade-off.


8. The Pattern Behind All Three Failures

This is the part that's easy to miss.

You start with:

Database
Enter fullscreen mode Exit fullscreen mode

You add Redis:

Application
     |
   Redis
     |
 Database
Enter fullscreen mode Exit fullscreen mode

Now you have better read performance.

Then production reveals stale data.

So you add:

TTL + Invalidation
Enter fullscreen mode Exit fullscreen mode

Then production traffic reveals cache stampedes.

So you add:

Single Flight
TTL Jitter
Request Coalescing
Enter fullscreen mode Exit fullscreen mode

Then traffic concentration reveals hot keys.

So you consider:

Replication
Load Distribution
Failure Isolation
Enter fullscreen mode Exit fullscreen mode

The architecture keeps evolving.

That's because caching isn't simply:

"Put Redis in front of the database."

You're introducing another distributed component with its own:

  • consistency behavior
  • expiration behavior
  • concurrency
  • failure modes
  • traffic patterns
  • recovery requirements

9. What to Design Before Adding Redis

Before introducing a cache, I'd want the team to answer at least these questions.

1. What can be stale?

Not every piece of data has the same consistency requirement.

Ask:

Can this value be 1 second old?
10 seconds?
5 minutes?
Never?
Enter fullscreen mode Exit fullscreen mode

That answer should influence the caching strategy.

2. What invalidates the data?

Define the relationship between writes and cache entries.

For example:

User updated
    ↓
Database committed
    ↓
Invalidate user:42
Enter fullscreen mode Exit fullscreen mode

Don't leave invalidation as an afterthought.

3. What happens when the key expires?

Ask:

1 request misses?
100?
10,000?
1,000,000?
Enter fullscreen mode Exit fullscreen mode

If the answer is "they all query the database," you may have a stampede waiting to happen.

4. What happens when one key becomes extremely popular?

Measure key access patterns.

A distributed cache doesn't automatically mean distributed traffic.

You can still have:

99% of requests → 1 key
Enter fullscreen mode Exit fullscreen mode

5. What happens when Redis is unavailable?

This is one of the most important questions.

If Redis disappears, does the application:

Fail closed?
      or
Fall back to database?
Enter fullscreen mode Exit fullscreen mode

If everything falls back to the database simultaneously, Redis failure can become a database incident.

Your fallback path needs capacity planning too.


10. The Real Lesson

Caching is one of the easiest performance optimizations to add to an architecture.

It's also easy to underestimate what you're introducing.

The progression often looks like this:

Cache
  ↓
Stale Data
  ↓
TTL / Invalidation
  ↓
Cache Stampede
  ↓
Single Flight / Jitter
  ↓
Hot Keys
  ↓
Replication / Distribution
Enter fullscreen mode Exit fullscreen mode

Each step solves a real problem.

And each step introduces another design decision.

That's the real lesson:

A cache doesn't remove complexity. It moves complexity somewhere else in the system.

So before adding Redis, don't only ask:

"How much faster will this make our application?"

Ask:

"What happens when the cache is stale?"

"What happens when the key expires?"

"What happens when 10,000 requests miss simultaneously?"

"What happens when one key receives 100× normal traffic?"

"What happens when Redis itself fails?"

Those questions are where caching stops being a feature and becomes system design.

Top comments (0)