DEV Community

WebSurf Digital
WebSurf Digital

Posted on

It Worked On My Machine. It Took Me 4 Days To Find Out Why It Didn't On The Server.

If you have deployed anything more than once, you already know the sentence.

"It works on my machine."

I said that sentence last month. Then I spent four days proving that my machine is a liar.

Before you ask, no, this is not a post telling you to "just use Docker". Everyone on the internet already told me that, and yes, they are right. But the fix took me ten seconds. The hunt took me four days. And honestly? The hunt is the part I actually learned from.

So let me tell you about it.

The Setup

Small team. Small app. Python backend, React frontend, deployed to a Linux box somewhere in a data center I will never visit.

Everything was green. Tests passed. Lint passed. I pushed on a Tuesday night, closed my laptop, and went to sleep like a person with no enemies.

I woke up to 47 unread messages.

The container was crash-looping. Not slow. Not flaky. Not "sometimes returns a 500". Dead. Restarting every few seconds like it was being paid to suffer.

I opened the logs and found this:

ModuleNotFoundError: No module named 'Utils'
Cool. Weird, but cool. I know for a fact that module exists. I wrote it. I imported it in three different files. It works. I literally ran it four hours ago.

So I did what any reasonable developer does. I pulled the latest main, ran the app locally, and watched it start up perfectly.

Works fine.

Day 1: The Blame Game

Here is a list of things I blamed before I blamed myself:

· The deployment pipeline · requirements.txt · The venv · Python 3.11 (I was on 3.10 locally, obviously that's the issue) · Our DevOps guy (sorry, Marcus) · The Docker image cache · A stale .pyc file · Mercury retrograde

I rebuilt the container locally. It worked. I rebuilt it in CI. It worked. I ran the exact same command the server runs. It worked.

I want you to sit with how frustrating that is. Every single reproduction attempt succeeded. The bug was only reproducible in the one place I couldn't poke at it.

Day 2 and 3: Looking In All The Wrong Places

I checked git status. Clean.

I checked the file was actually committed:

git ls-files | grep -i utils
And there it was. utils.py. Committed. Pushed. Present.

So the file exists, the import exists, and the error says the module doesn't exist. Those three things cannot all be true unless something is lying to me.

I tried adding the module to PYTHONPATH manually in the Dockerfile. It worked. Which meant I had "solved" it, but I had no idea why.

And that's the worst feeling in this job, right? When you fix something without understanding it. You just know it's coming back. Like a horror movie villain.

I reverted the hack. I refused to ship a fix I couldn't explain.

Day 4: The Stupidest Answer Possible

I was explaining the problem to a friend over coffee. I said, out loud, the sentence I had been avoiding:

"Okay so the file is utils.py and the import is from Utils import parse."

He looked at me. I looked at him.

I went home. I renamed the file. I pushed. It deployed.

It worked.

The issue was a capital U.

Here's What Actually Happened

My laptop is a Mac. MacOS uses APFS, which is case-insensitive by default. So when I write from Utils import parse, my Mac shrugs and says "sure, you mean utils.py, no problem buddy."

The production server runs Linux. Linux, like a reasonable operating system, is case-sensitive. When it sees from Utils import parse, it looks for a file literally called Utils.py. Which does not exist. So it throws exactly the error it should throw, and I ignored it for four days.

I was not debugging a bug. I was debugging my operating system lying to me about the rules of reality.

Why This Hurts More Than It Should

Here's the thing that actually bothers me about this. The error message was correct the entire time.

ModuleNotFoundError: No module named 'Utils'

I read that and thought "yes, that's the problem I'm trying to solve." I never read it as a statement of fact. I treated it as a symptom, when it was actually a diagnosis.

And this is such a Mac-developer-on-Linux-prod thing to get bitten by. If you develop on Linux, you have never once had this problem in your life. You're playing on hard mode and you don't even know it.

The Stuff I Do Now

Here are the actual changes I made afterwards, not the Docker one everyone already told me about.

Read the error like it's not a hint, it's an answer.
The error said No module named 'Utils'. Not "no module named utils", not "import failed", not "check your PYTHONPATH". It named the exact thing that didn't exist. If I had taken that literally on day one I would have found this in an hour.

I now have a rule: before I start debugging, I write out what the error message is claiming in plain English, and I check whether that claim is true or false. Nine times out of ten it's true and I've been assuming it's lying.

Check your case sensitivity, it takes 5 seconds.

On your machine

touch testfile && ls TestFile
If that prints something, you're on a case-insensitive filesystem and you are developing with training wheels on. Know it. It changes how much you should trust your imports.

git ls-files does not tell you the truth about case.
This one got me. On my Mac, git ls-files | grep -i utils happily matched utils.py. The -i flag is what did it. I was so used to ignoring case that I built it into my own debugging command.

If you actually need to check the real filename:

git ls-files | grep "utils"
No -i. Let it be case sensitive. Let it hurt. That's the point.

Renaming a file case-only on Mac is a nightmare.
Because git on a case-insensitive filesystem doesn't see utils.py → Utils.py as a change. You have to force it:

git mv utils.py temp_name.py
git mv temp_name.py Utils.py
git commit -m "fix: actually case the filename correctly"
Two steps. Ugly. Works. Write it down somewhere, because you will need it again.

The Part That Actually Matters

The fix was one character. I'm not going to pretend the fix was interesting.

What's interesting is that I had a correct error message sitting in front of me for four days and I treated it as noise instead of information. I was so sure the problem was somewhere interesting — the container, the pipeline, the environment — that I couldn't accept it was somewhere boring.

Bugs are usually boring. That's why they're so hard to find. We're all out here looking for a villain when the answer is a capital letter.

If you're currently stuck on something and it's been more than a day, take the error message, read it out loud, and ask yourself if it is simply, literally, boringly true.

It usually is.

What's the dumbest bug that took you the longest to find? I'll go first, clearly. Drop yours in the comments so I feel less alone.

Top comments (1)

Collapse
 
shieldxbot profile image
shieldx •

The gap between a local dev environment and a production server is usually where the most expensive debugging hours go. I once spent an entire afternoon chasing a ghost bug that turned out to be a subtle difference in how the local Docker daemon handled filesystem permissions compared to the production Linux kernel. It is rarely the logic itself that fails, but rather the environmental assumptions we take for granted, like specific library versions, environment variables, or even time zone offsets. Moving toward a strict containerization strategy or using tools like Nix to ensure bit-for-bit reproducibility across machines is usually the only way to stop playing this guessing game.