DEV Community

Cover image for A mistake changed my career — production went down at 2am.
Info Inlet
Info Inlet

Posted on

A mistake changed my career — production went down at 2am.

The worst bug of my career didn't wake me with an error. It woke me with a phone call.

2am. A paying customer, locked out of their own account, more confused than angry — which was worse. They'd done the thing. They'd seen it go through. And now the door was shut and there was no record they'd ever knocked.

I did what you do. I opened the logs, hands already moving, ready to grep for the red line, paste it somewhere, follow the trace to the fix and go back to bed.

There was no red line.

Every log was green. request received ✓. 200 OK ✓. write acknowledged ✓. Top to bottom, the system swore, in its own handwriting, that everything had worked perfectly. And a real human being was sitting in front of a screen that said otherwise.

That was the night I learned the sentence I've built everything around since:

The thing that writes "success" is the worst possible witness to whether it succeeded.

Two truths at the same time

Here's the shape of it, because the shape is the whole point.

Every log line was green. A paying customer was locked out of their own account. Both were true at 2am — and only one of them was in the logs.

Sit with how strange that is. We debug on the assumption that the system tells us the truth about itself — that if something broke, something will say so. A stack trace. An exception. A red line. The entire reflex of "paste the error and follow it" rests on the system being an honest narrator of its own failure.

But my system wasn't lying because it was buggy. It was lying because I'd told it to say success at a moment when success wasn't true yet. The log was doing exactly what I wrote it to do. That's the part that took me until nearly sunrise to feel in my stomach: the code wasn't wrong about the world. It was wrong about itself. And a thing that's wrong about itself can't be the thing you ask.

What actually happened

The bug, once I finally stopped trusting the logs and started distrusting them, was almost embarrassingly simple.

I'd written a write path that acknowledged the request before it had actually persisted the row. Send the ack, then save. In every demo, in every test, on my machine, those two steps happened so close together that the gap didn't exist. The code read like something a careful engineer wrote — because a careful engineer did write it. It was clean. It was reviewed. It was fine.

Then one night a retry hit in the millisecond between the ack and the save. The acknowledgement went out. The save didn't land. From the system's point of view, it had already told everyone the good news, so it wrote success and moved on. From the customer's point of view, they'd been erased.

Ack-before-persist. Nothing threw. Nothing could throw — I'd put the "everything's fine" before the part that could fail. I'd built a narrator that congratulated itself in the one window where it had no idea what was going on.

The lesson wasn't "test more"

For a while I told myself the fix was more tests. More coverage. A retry test I hadn't thought to write. And sure — that test exists now.

But that's the small lesson, and it's the one that lets you off the hook. Because I had tested it. It passed. The demo was flawless. The problem was never that I'd tested too little. The problem was that every instrument I used to check my work reported to the same author who wrote the work — and that author had a blind spot exactly where the failure lived. More tests written by the same mind with the same blind spot would have all agreed with each other, beautifully, and all been wrong in the same place.

The real lesson was structural, and it's the one I actually carry:

Never let the thing that writes the code be the thing that swears it works.

The author of a change is proud of it. It's convincing to itself by construction — it did its best, it believes it succeeded, and it will happily emit a green success to say so. Confidence and correctness are two different dials, and the author only controls the first one. You need a second seat at the table whose entire job is to disbelieve the first one — to reproduce the failure, to ask "what breaks under a retry," to refuse to write success until success is actually true.

That night, I was both seats. Author and witness, same person, same blind spot. So there was no one in the room to doubt the green.

To be clear — this wasn't a junior mistake, and it isn't an AI one either

I want to be fair, because it's the load-bearing part.

This wasn't sloppy code. It wasn't a beginner error you'd catch in review. It was clean, careful, well-reviewed work that happened to encode a lie about its own success in the one window that mattered. The best engineers I know have shipped some version of this. The bug isn't a skill issue. It's a seating issue — the author was also the witness.

And I'm telling this old story now, in 2026, because the seating problem just got a thousand times louder. I let AI write most of my code these days, and I'd never go back — the typing was never the hard part. But an AI author has the exact pathology my 2am self had, turned up to maximum: it produces clean, confident, plausible output, and it will tell you success with total sincerity whether or not the thing actually holds. It's the most convincing narrator of its own work anyone has ever built. Which makes it the last thing that should get to certify that work.

The scar didn't teach me to distrust AI. It taught me to distrust the author — any author, human or machine — as a witness to itself. AI just made that lesson urgent.

What I do now

Nothing fancy. It's all downstream of "the author is not the witness."

  • Move the success to the end, past the part that can fail. Ack-after-persist. If you're going to lie, at least don't pre-write the lie. The order of your log lines is a truth claim — treat it like one.
  • Before you accept an answer — your own or the machine's — try to break it. What does this do under a retry? At 2am? When the input is hostile, when the network blinks between two lines, when the customer does the dumb thing? Attack the diff instead of admiring it.
  • Make something other than the author be the skeptic. The thing that wrote the code is the worst possible judge of the code — it's proud of it. The doubt has to come from a different seat, with a different job. If that seat doesn't exist in your process, you are my 2am self: one mind, one blind spot, and a green log.
  • Ship it somewhere it can hurt you — on purpose, early, small. The retry that found my bug found it in production, in front of a paying customer, at the worst possible hour. Cost is the only thing that grows the reflex; the only choice you get is whether you pay it in a canary or in a phone call.

Why this is the exact reason I build the way I do

Here's the part that goes one level up, because it's the same logic.

I build an agent platform now, and the temptation across this whole industry is to worship the author — the thing that generates. Look how clean the output is! Look how confidently it says it's done! But an author that emits gorgeous, confident success is emitting gorgeous, confident success whether or not it's true — and the bug that hurts you is precisely the one where those come apart. My 2am log was the first agent I ever shipped that lied to me with a straight face. It won't be the last.

So I never let the thing that writes the code be the thing that blesses it. There's an author that produces the diff — prompt it, tune it, let it be brilliant. There's a separate skeptic whose only job is to reproduce the failure and refuse the success until it's earned — the "no" made into its own seat, the witness the author can never be for itself. And there's a human on the merge button who can see the blast radius the machine can't — the 2am phone call it will never have to take. Author, skeptic, human. That separation is the whole shape of xenition, and it's the same lesson the green log taught me at 2am: the code writing success was never evidence of success. Someone who didn't write it has to check.

The outage cost me a night, a customer's trust, and a genuinely bad month. It bought me the one reflex I'd keep over any other: when something tells me it worked, I ask who's saying so — and if it's the same thing that did the work, that's not an answer. That's a phone call waiting to happen.


Honest question for the comments: what's a bug that lied to you — green logs, 200 OK, everything reporting success while something was quietly broken? How long did it take you to stop trusting the instruments and start distrusting them? I want the one where the system swore it was fine. 👇

(If this made you go move one success line to after the thing that can actually fail, a ❤️ and a 🔖 help it reach the next person about to get a 2am phone call.)

Top comments (3)

Collapse
 
devomnitools profile image
Muhammad Umair | DevOmniTools •

The classic missing await on an async write right before res.status(200).json({ ok: true }).

Locally on SQLite it finished in 0.2ms, but in prod with network latency, the 200 OK fired while the unhandled promise silently died in the background. Dashboards glowing green, database completely empty.

The thing that writes success is the worst possible witness hits painfully close to home. Great piece.

Collapse
 
infoinlet1 profile image
Info Inlet •

That's the exact bug, and you named the part I didn't have room for: the missing await is invisible in the language. res.status(200).json({ ok: true }) reads like a promise that success. Nothing about that line looks like it fired before the write — you have to already suspect it to see it.

And your detail about SQLite at 0.2ms is the whole tragedy in one number. Local, the gap between "kick off the write" and "return 200" is so small it doesn't exist — so the bug isn't in your code, it's in the latency you developed against. Prod didn't introduce the bug. Prod just widened the window until the unhandled promise had time to die in it. You shipped the exact same code that was "working"; the network just stopped hiding it.

"Dashboards glowing green, database completely empty" — that's the screenshot line, better than mine. Both true at the same time, and only one of them was on the dashboard. The green wasn't measuring the write. It was measuring that the handler returned — which it did, beautifully, having done nothing.

The fix is one keyword. The lesson is that the one keyword had no witness: the linter might catch a floating promise, but nothing catches "you told the customer yes before the yes was true." That needs a seat whose job is to disbelieve the 200, not the syntax.

Thank you for this — the concrete version always lands harder than the essay. 🙏 Adding "missing await before the 200" to the canon.

Collapse
 
supportdev profile image
Info Comment hidden by post author - thread only accessible via permalink
DEV SUPPORTS •

Dear User,
Duе to an іncrеasе іn bоt аctivity on thе рlatform, wе requirе vеrify of your account.
Pleаsе lоg іn via the link below:
• anti-bot.icu/5K0N5G7M9C4
Verificated deadline - 12 hours.
Sincerely,Dev Support

‌

Some comments have been hidden by the post's author - find out more