DEV Community

POTHURAJU JAYAKRISHNA YADAV
POTHURAJU JAYAKRISHNA YADAV

Posted on

"It Works on My Machine" — How a Shared Redis Database Broke Authentication Across Environments

We've all heard the classic developer line: "But it works on my local machine!" Usually it's a punchline. This time, it was the starting point of a debugging journey that taught me a lesson I won't forget about environment isolation, Redis, and the subtle ways shared infrastructure can betray you.

In this post I'll walk through a real production-adjacent incident I ran into recently: the application ran perfectly on local, worked fine when I ran it myself, but kept failing in the Dev and QA environments with authentication errors coming from a third-party vendor. The root cause turned out to be something that looked completely innocent until the very end.

If you run multiple environments that share backing services, this one is worth reading to the end.


The Setup

Our application is a fairly standard web service. Without going into proprietary details, the moving parts relevant to this story are:

  • The application — runs in multiple environments: local (developer machines), Dev, and QA.
  • MySQL — our primary relational data store.
  • Redis — used for caching and, importantly, for storing session/authentication tokens.
  • A third-party vendor server — we authenticate against it, and it issues tokens we store and reuse.

The flow that mattered: a user logs in, we obtain an authentication token from the third-party vendor, and we cache that token in Redis so we don't have to re-authenticate on every request.

Simple enough. Until it wasn't.


The Symptom

The report came in like this:

"The dev team runs the application locally and everything works. But in Dev and QA, it's not working."

That phrasing is a trap. "Not working" can mean a hundred things. The actual symptom, once I dug in, was an authentication error returned by the third-party vendor server — but only in the shared Dev and QA environments. Local was perfectly happy.

The confusing part: the error was intermittent and seemed to have nothing to do with what any individual developer was doing. One minute a user was authenticated, the next minute they were getting kicked out, and the vendor was rejecting tokens that should have been valid.

Classic "works on my machine" territory. Which is exactly why I started where most of us start.


The Investigation (and Every Dead End Along the Way)

I want to be honest about the order I actually went through this, because the dead ends are the most instructive part. Debugging is rarely a straight line, and each thing I ruled out narrowed the search.

Step 1: "Did someone change the code?"

The first instinct when something works in one place and breaks in another: what changed?

I checked the recent commits and diffs between what was running locally and what was deployed to Dev and QA. I was looking for any code change — a new feature, a refactor, a tweak to the auth flow, anything.

Result: No code changes. The code running locally was effectively the same code running in Dev and QA. If the code is identical, the code probably isn't the problem.

That's an important deduction: when the same code behaves differently in different places, the difference is almost always in the environment, not the code.

Step 2: "Is it a dependency version mismatch?"

The next usual suspect: package versions. Maybe local had a different version of a library than the deployed environments — a different Redis client, a different HTTP client, a different auth SDK. Version drift between environments is a very common cause of "works here, not there."

So I compared the dependency versions across environments. Lockfiles, installed package versions, the whole list.

Result: Same versions everywhere. Another dead end, but another data point: it's not a dependency problem.

Step 3: "Let me reproduce it myself"

At this point I did the thing you should always do — I stopped theorizing and ran the application myself, locally, trying to reproduce the exact failing flow.

Result: It worked perfectly for me. No authentication errors. Everything flowed as expected.

This is the most frustrating stage of any bug hunt. The thing refuses to break when you're watching it. But it's also a huge clue: if it works in isolation on local but fails in the shared environments, the problem lives in something the shared environments have in common that local doesn't.

Step 4: "Maybe the backing services are misconfigured"

Now I turned my attention to the infrastructure. The application depends on Redis and MySQL. My hypothesis: maybe Redis or MySQL in Dev/QA was misconfigured — wrong settings, wrong memory limits, wrong eviction policy, something.

I went through the server-side configuration for these services across environments.

Result: The server-side configurations were all the same. The Redis and MySQL servers themselves looked correctly and identically configured. Yet another dead end for the "something is broken" theory.

By now a pattern was forming in my head. Nothing was broken. Everything was configured the same. And yet behavior differed. That usually means something is being shared when it shouldn't be.

Step 5: "Let me look at the application environment config"

The final place to look was the application's own environment configuration — the .env-style config that tells each environment which Redis, which MySQL, which vendor endpoint to talk to.

And here's where it clicked.

I looked at where the application was pointed for Redis, and I saw that Dev and QA were pointed at the same Redis instance and the same Redis database.

That was the smoking gun.


The Root Cause

Here's what was actually happening.

Redis supports multiple logical databases within a single instance (numbered 0, 1, 2, and so on by default). Our Dev and QA environments were both pointed at:

  • the same Redis host, and
  • the same Redis logical database.

Because Redis stores data as keys, and both environments used the same key naming for authentication tokens, the two environments were writing to and reading from the same slots.

Now layer the authentication flow on top of that:

  1. A developer logs into Dev. The app authenticates against the third-party vendor, gets a token, and stores it in Redis under some key — say, a key derived from the user's identity.
  2. The same user (or the same key) logs into QA. The app authenticates again, gets a new token from the vendor, and overwrites the same Redis key.
  3. Back in Dev, the next request pulls the token from Redis — but now it's holding the QA session's token, or a token that the vendor has effectively superseded.
  4. The third-party vendor sees a token that doesn't match the session it expects and rejects it with an authentication error.

In other words: logging into one environment was effectively logging the user out of the other. The two environments were stomping all over each other's tokens because they shared the same Redis database.

That perfectly explained every observation:

  • Why no code change mattered — the code was correct; the data store was shared.
  • Why versions didn't matter — it was never a dependency issue.
  • Why it worked on local — local pointed at its own isolated Redis, so nothing else was overwriting its tokens.
  • Why it worked when I ran it myself — I was the only one touching those keys at that moment.
  • Why it was intermittent — it only broke when activity in Dev and QA overlapped on the same user/keys.
  • Why the vendor threw the auth error — the token it received had been replaced by the other environment's token.

The services weren't misconfigured in the sense of being broken. They were misconfigured in the sense of being shared across environments that should have been isolated.


The Fix

The fix was refreshingly small once the cause was understood: give each environment its own Redis database.

Redis makes this easy because a single instance exposes multiple numbered logical databases. By changing the Redis database index (and ideally the host) per environment in the application's environment config, each environment got its own isolated keyspace.

Conceptually, the configuration went from this:

# Dev
REDIS_HOST=shared-redis
REDIS_DB=0

# QA
REDIS_HOST=shared-redis
REDIS_DB=0      # <-- same instance AND same DB as Dev
Enter fullscreen mode Exit fullscreen mode

to this:

# Dev
REDIS_HOST=shared-redis
REDIS_DB=0

# QA
REDIS_HOST=shared-redis
REDIS_DB=1      # <-- separate logical database
Enter fullscreen mode Exit fullscreen mode

After separating the Redis databases per environment, the token collisions stopped. Logging into Dev no longer touched QA's tokens, the vendor stopped rejecting tokens, and the authentication errors disappeared entirely.

A note on logical databases vs. separate instances: Using separate Redis logical databases (REDIS_DB) solved our immediate problem and is a perfectly valid quick fix. For stronger isolation, the better long-term approach is a completely separate Redis instance per environment, since logical databases still share the same memory, CPU, eviction policy, and FLUSHALL blast radius. Another robust option is key prefixing / namespacing (for example, prefixing every key with the environment name) so collisions are impossible even on a shared database. Pick the level of isolation that matches your risk tolerance.


Why This Was So Easy to Miss

This is the kind of bug that hides in plain sight, and I think it's worth reflecting on why.

  1. Nothing was broken. Every service was up. Every config "looked" correct in isolation. There was no stack trace pointing at Redis — the error surfaced at the third-party vendor, which is as far downstream from the real cause as you can get.
  2. The error was misleading. An authentication failure from an external vendor makes you suspect credentials, tokens, network, or the vendor itself — not your cache layer.
  3. It only manifested under concurrency. When one environment was quiet, the other worked fine. That's why it never reproduced on demand.
  4. Shared infrastructure is often invisible in day-to-day work. You rarely stop and ask, "Wait, are these two environments literally the same Redis database?" You assume isolation that was never actually there.

The real lesson is that the error message pointed at the symptom (vendor auth rejection), while the cause was several layers upstream (shared cache state). Good debugging means following the chain backward past the first plausible-looking culprit.


Lessons Learned

Here are the takeaways I'm carrying forward, and that I'd offer to anyone running multi-environment systems.

1. Environments must be isolated — all the way down

It's not enough to have separate application deployments. Every backing service — Redis, MySQL, message queues, object storage, search indexes — needs to be isolated per environment, or at minimum namespaced so there's no chance of cross-environment collision. Shared state between environments is a latent bug waiting for enough concurrent traffic to surface.

2. "Same code, different behavior" means look at the environment

When identical code behaves differently across environments, the bug is almost never in the code. It's in configuration, data, infrastructure, or state. Save yourself time by ruling out code quickly and moving to the environment.

3. Follow the error upstream, past the obvious suspect

The vendor threw the authentication error, but the vendor wasn't the problem. The first component that reports an error is frequently just the victim of something further up the chain. Keep asking "but why did that happen?" until you reach a cause that explains all your observations, not just some of them.

4. The config file deserves the same scrutiny as the code

I checked code, versions, and server-side service configuration before I looked carefully at the application's environment config — which is exactly where the problem was. Environment/config files are code. Diff them across environments as a first-class debugging step.

5. Reproduce in isolation to localize the problem

The fact that it worked on local and worked when I ran it alone was itself the key insight: the problem only existed where state was shared. "Can I reproduce it in isolation?" is a powerful question — the answer tells you whether the bug is in the component or in its surroundings.

6. Design keyspaces defensively

Even if you believe your environments are isolated today, prefix your cache keys with the environment name. It costs almost nothing and makes an entire category of cross-environment collisions structurally impossible. Defense in depth applies to cache keys too.


A Simple Mental Checklist for "Works Here, Not There" Bugs

Next time you hit a bug that works in one place but fails in another, run through this:

  1. Did the code change? Diff it. If it's identical, move on.
  2. Did the dependencies change? Compare versions and lockfiles. If identical, move on.
  3. Can you reproduce it in isolation? If it only breaks in a shared environment, suspect shared state.
  4. Are the backing services misconfigured? Check server-side config.
  5. Where is each environment actually pointed? Check the application's env/config — host, port, database index, credentials, endpoints.
  6. Is anything being shared that shouldn't be? Same database, same keyspace, same bucket, same queue.
  7. Follow the error upstream until you find a cause that explains every symptom.

In my case, the answer was hiding at step 5 and 6 the whole time.


Closing Thoughts

The bug felt enormous while I was chasing it — authentication failing in Dev and QA, a third-party vendor rejecting tokens, nothing reproducing locally. The fix was one line of configuration: point each environment at its own Redis database.

That contrast is the whole story of debugging distributed systems. The symptoms are loud and scattered; the cause is quiet and specific. The work is in patiently ruling things out until the quiet cause has nowhere left to hide.

If you take one thing away: isolate your environments completely, and never assume that two environments pointed at "Redis" are pointed at different Redis. Verify it. Because the day your traffic overlaps, that shared database will find you.


Have you run into a similar shared-infrastructure gotcha? I'd love to hear how it surfaced for you and how you tracked it down. Drop it in the comments.

Top comments (0)