Series: AI, Ego & Regret — Bonus Chapter
Editor's Note: While compiling the old series for the book, I found this draft in the archives. It didn't fit the original lineup, but the story was too good to leave buried on a hard drive. I polished it up and decided to release it as a bonus chapter for the series. This is a work of fiction.
Act I · Dead Weight
It took me six years to turn NovaTech's AI diagnostic platform from a Python script on a single server into a production system processing four million diagnostic requests a day. Every sensor pushed data upstream. The model didn't fire on every reading, but during peak hours it handled over a hundred per second. Our clients included three Fortune 500 manufacturing giants — the biggest of which was Merit Manufacturing — plus a medical device brand you've definitely used.
Six years ago when I joined, the company had eighteen people and the CTO still wrote code. CEO Ryan Whitfield pounded the table at all-hands meetings and said, "We're building something that matters."
Three years in, we moved into a Sunnyvale office with a courtyard. Ryan drank half a bottle of whiskey at the housewarming party, put his arm around me, and told an investor, "Tom's the reason this ship stays afloat."
Three years after that, the ship told me I was dead weight.
The layoff came on a Wednesday at 11 AM. All-hands video call. Ryan sat in front of his dining room chandelier — the one that cost thirty grand — and explained that the company was going through a "strategic restructuring" to "streamline operations." Slide three had my name and the entire data platform team's — all twelve of us — sitting neatly inside a gray box.
"The affected roles are being eliminated to reduce redundancy."
Redundancy. I'd led this team for six years. Years of zero-incident operations. 99.97% recall rate. Redundancy.
HR's separation package landed in my inbox. Standard severance, an NDA, a one-year non-compete. I couldn't even work at the coffee shop downstairs, let alone a competitor.
I noticed something when I signed. Buried in my non-compete was an exception clause: "Family members of executive leadership are exempt from non-disclosure obligations regarding internal restructuring communications." It was tucked into a footnote on page fourteen in type so small you'd miss it if you blinked.
I didn't think much of it. I signed.
Act II · First Blood
Two weeks later, LinkedIn filled in the blanks.
Kevin Whitfield's new profile picture showed up in my feed. Westport University polo shirt, standing in front of the engineering school sign, teeth so white they looked AI-generated. His headline: "VP of Engineering at NovaTech | Builder | Thinker | Disruptor."
I stared at that title for a solid twenty seconds. VP of Engineering. A guy who was writing course projects three months ago, holding the same title as my old boss.
His first company-wide email went out the next afternoon — sent to all the inboxes I no longer had access to. Dmitri — one of the old-timers on the team — screenshotted it and sent it to me.
"I'm excited to announce that NovaTech is entering a new chapter. Our current platform architecture, while functional, carries significant technical debt. I'll be leading a modernization initiative to bring our stack into the next generation. Expect faster iterations, leaner deployments, and a more aggressive roadmap."
Technical debt. Faster iterations. Leaner deployments. Every word landed exactly on "I have no idea what this system does."
Dmitri's caption under the screenshot: "He asked me today if the CI/CD pipeline was a new hire."
Act III · The Model Swap
Kevin's first move wasn't architectural. It was swapping the model.
"Tom's model is six years old. It's good. But good isn't the standard anymore."
That's what he said at his first tech all-hands. Dmitri relayed it to me in a flat voice — he was past the point of being surprised.
Kevin's "modernization" meant replacing the diagnostic model I'd tuned for six years — 99.97% recall with a 0.03% false positive rate in industrial fault detection — with the newest closed-source LLM on the market. He claimed the LLM's "zero-shot reasoning capabilities" would expand coverage from 340 scenarios to "infinite."
340 scenarios. Every single one went through three rounds of labeling, two rounds of cross-validation, and at least three hundred hours of production observation before we called it done. Kevin spent three days reading API docs and decided that was enough.
At the demo, he pulled up the LLM's dashboard and told the six engineers in the room: "This thing reads the logs, understands the context, and outputs the diagnosis in natural language. No more manual rule tuning. No more feature engineering. No more bottlenecks. "
When he said "no more bottlenecks," he looked at Dmitri. Dmitri was our performance engineer. Kevin probably didn't know his name, but he knew that role belonged to "the previous era."
Dmitri didn't push back in the meeting. He messaged me afterward: "He said the LLM benchmarked against our logs. I checked his test set. Textbook fault patterns. Not a single one from a real production line."
I didn't reply. Because Kevin knew too — he picked a clean test set because clean test sets give clean results.
Act IV · Cascade Failure
Counting from Kevin's first day, by day twelve the swap was done. Kevin oversaw the engineering team himself, cutting over the old inference pipeline to the LLM's API. The switch was smooth. Every endpoint returned 200. Every green light on the dashboard stayed on.
Day fourteen, Merit Manufacturing's production line reported a fault.
Nothing major — a sensor drift on a conveyor belt. The old model had handled this exact scenario over a thousand times. Diagnosis: "Sensor drift detected, recalibration window: 72 hours." Confidence score attached. Maintenance ticket generated. Total time: 170 milliseconds.
The new model read 170 pages of production logs, spent 4.3 seconds, and produced a fluent paragraph:
"The anomaly pattern suggests a potential degradation in the conveyor belt alignment mechanism, possibly due to cumulative thermal stress on the drive unit bearings. Recommend immediate full inspection of the drive train assembly."
Translation: "I have no idea what this is, but I can write something that sounds smart."
The maintenance crew stopped the line. Pulled the drive unit. Tore down the bearings. Four hours later, they found nothing. The supervisor wrote in the work order: "False alarm. No fault found."
That was the first one.
Day sixteen, the same LLM flagged a normally running compressor as "imminent bearing failure." Another teardown. Another four hours. Nothing. The supervisor changed NovaTech's diagnostic feed status to "high noise."
Day eighteen, it missed a real fault.
Sensor readings clearly showed a hydraulic seal accelerating toward failure. The old model's rule engine would have triggered a "failure probability: 87% within 7 days" alert. The LLM read the same data and wrote: "The observed fluctuation patterns are within normal operational variance. No immediate action required."
Thirty-two hours later, the seal blew. The line was down for nine hours. Each hour of downtime cost $130,000.
Dmitri sent me the numbers — screenshots from the incident report, timeline and cost breakdown included. His caption: "Your rule engine caught faults for six years. Twenty days after they swapped it, it blew up in front of a client." I didn't reply. But I read every line.
Act V · The Ultimatum
Merit Manufacturing's incident report didn't name NovaTech — Diana did them that favor. But the customer success VP forwarded an email to the internal group chat, and Dmitri screenshotted it to me. Five sentences from Diana to Ryan, CC'd to four people: her CFO, her legal VP, NovaTech's customer success VP, and a legal department inbox.
The third sentence: "We require the diagnostic model deployed prior to June to be restored within fourteen calendar days. If that is not possible, we will exercise our right under section 6.2 to terminate the agreement."
Translation: "Put Tom's model back. If you can't, we walk."
Act VI · The Call
Three days after that email, my phone buzzed at 11 PM. No area code in the caller ID, but I knew who it was.
"Tom."
"Diana."
"Everything I'm about to say, my legal team would tell me not to say it."
She paused. I heard keyboard clicks on the other end — she might have been closing windows, making room for what came next.
"Merit Manufacturing is terminating the agreement with NovaTech. Section 6.2. Fourteen-day transition period."
I didn't say anything.
"We need someone to lead the transition. Not someone from NovaTech. You know this system better than any engineer still on their payroll. Your non-compete has two holes in it. I had my lawyers check —" she paused, like she was making sure she hadn't crossed a line. "The family-member exception was posted on an anonymous forum. Same firm that drafted your contract — anyone who reads it can see what it means. And the non-compete itself has an out: client-initiated separations don't trigger it. Also —" another pause. "Kevin's reply was effectively a written refusal. We don't need to wait the full fourteen days."
I leaned against the kitchen counter. My wife was reading to our daughter in the living room, voice soft, one word at a time. I listened to those syllables, holding the phone, feeling the world go quiet for two beats.
"Tom?"
"Send me the details," I said. My voice was calmer than I expected.
Before she hung up, I asked: "Diana — did you go to NovaTech first?"
A long silence. Then: "I did, Tom. Kevin told me the LLM is 'objectively superior' and suggested I wait for his explainability report. So I'm done waiting."
She was waiting for Kevin to realize he was wrong. He never did.
Diana sent me the explainability report right after the call. Six pages of PDF. Every section tried to prove the two false alarms and the missed fault were "edge cases," "out-of-distribution inputs," "not representative of the model's true performance." There was a comparison table in the appendix showing 98.7% accuracy on his clean test set.
He told a complete lie using a dataset that never lied.
Act VII · The Exception
The NDA's hidden exception clause was dug up by a former colleague and posted on an anonymous workplace forum. The thread title: "\"NovaTech's restructuring was literally designed for nepotism.\" Over four hundred replies."
One comment stuck with me. An anonymous account claiming to be a former NovaTech employee wrote: "Ryan Whitfield once told me culture fit was the most important hiring criterion. Turns out what he meant was 'shares my last name.'"
I screenshotted it. Didn't save it. Then I closed the browser and opened the architecture docs for Merit Manufacturing's new project.
A system isn't yours just because you didn't build it — that was Kevin's mistake. A system is yours forever because you did build it — that was Tom's fate.
AI, Ego & Regret — Bonus Chapter. This is a work of fiction. Any resemblance to real events or persons is coincidental.
📖 Stratagem 2 of the 36 Stratagems series — coming soon.
P.S. English isn't my first language. I use AI to polish the writing and smooth out the rough edges. Thanks for reading. ☕ Buy me a coffee

Top comments (49)
Ngl, I feel that. My current employer (day-job) asked me a year ago when I was about 2 months there, to upgrade their dashboard, make it drag and drop... It took me 2 months, because of all the mistakes in their calculations, flaws in their interface, slow/bloated endpoints, etc. Fast forward those 2 months, they had a clean drag and drop dashboard, which was FAST, the data was accurate and it wasnt just drag and drop, it had the capability to swap out blazor cards for js cards (their request), it could resize cards and dynamically scale to window size, in other words it was exactly what I should have built, with all the bells and whistles, even dynamic drilldowns that let a customer get deeper insight, in-depth upgrades to their existing shared infrastructure that allowed them to support simple excel exports instead of complete rewrites each case... They disagreed, saying the values are different (yeah, cuz their values were calculated wrong), but that's fine... So I got put into different tasks and the dashboard I built was benched, because it needed 'review', fast forward a year, it's still under review...
Sometimes, you do your best, put your all into it, just for them to look at it and say 'nah, I want it different' and different is almost always worse, unless you seriously know what you're doing... trusting a LLM as a rule-engine is a systematic flaw in thinking. A LLM can write rules, but it cant effectively enforce them, it'll always write self-gratifying rules that just clear 99% of the time, unless it's catastrophically wrong. For systems that require efficiency, redundancy and absolute certainty, you use a language like F# or Rust, you write strict rules, not leave it up to LLM interpretation... A detailed report from a LLM would require fine-tuning beyond what's possible, you'd need telemetry data from the day the machine rolled off the factory floor, detailed schematics, detailed reports on common failures and signals, etc. Essentially, you'd need to build the Tom system, just to validate the LLM's outputs, at which point why even have the LLM?
That dashboard story. God. "Under review" for a year — we all know what that means.
And you nailed it: LLM as rule engine is a contradiction. If you need another system to validate the first system's output, you've added a layer, not solved a problem.
Yip... Functionally scrapped. Though I did see that it's active on 3 of the companies' sites, so atleast someone gets a better dashboard?
Look fundamentally, having a LLM document isnt a bad idea, or find edge-cases beyond standard rules (ONLY FOR TAKING NOTES)... But to just swap a production system with a brand new 1, without side-by-side running them for atleast 6 months... That's just insanity. Eg. Doccit, my autonomous accounting suite, was originally built on a custom OCR pipeline, massive piece of work, using spline primitives instead of pixel marching, incredibly fast, it hit 700k chars/s with cuda, unfortunately, my ssd died... And I hadnt backed it up and I used a locally hosted git server... (never making that mistake again) So I had to re-engineer it now that the exclusivity license is over for it, so I practically had to build it from scratch again (500k LOC in 4 weeks, I think not too bad for doing it solo, admittedly, V.A.L.I.D. generated 80% of that and made it way more secure and stable), but in doing so, I switched to cloud LLMs to try out how that would work... It's decent for 90% of the cases, but it's expensive af for the last 10% and still makes mistakes, so I built a hybrid system, a custom rust OCR system + LLM for document understanding (not regex) and for spot-checking low confidence areas. I wont ever fully trust the LLM to do the job, but having it as a Tier 2 is a good idea, because in the event the OCR failed, the rule-validator failed, atleast it has some extra level of security to fix it. That's fine, because it's the backup, NOT the one and only system in use...
That's the right setup. LLM as tier-2 backup, not the whole pipeline. Hybrid beats pure every time — rules catch the hallucinations, LLM catches what the rules didn't think of.
I had my own automated testing platform too. Put everything into it. Same story — dormant on a local drive now. Know exactly how that feels.
Also — 500k LOC in 4 weeks is ridiculous. Even at 80% generated, it's 100k lines you owned. Mad respect.
Exactly, except that might actually change... If you get a chance, have a look at the 12 part series I posted today, V.E.L.O.C.I.T.Y. OS's development (sum total of 5 days) has made some incredible breakthroughs... 4x KV-Cache compression, Deterministic language running at Ring-0 and a LLM that's quantized to 2 bit, yet operates deterministically and keeps everything it's ever done in context all at once... A system like that might actually stand a chance at being a tier-1 safety system, because it'll just learn continuously and notice anomalies no matter how small. I'd still trust rules over it, but when you have a LLM that literally cant hallucinate, cant forget, runs hardware-sensors directly... It actually has a decent chance at being good enough to trust on it's own...
That's a big claim — "literally can't hallucinate." If you've cracked that, you've solved the fundamental problem. I'd be curious to see how the deterministic + 2-bit quantization works in practice — those two goals usually pull in opposite directions.
I'll take a look at the series when I get a chance.
LLMs hallucinate due to lost track of state. You write 99k lines of code and it forgets what code it replaced on the last write. the NDA triplets prevent it from being able to write bad code that doesnt fit the purpose and the merkle root prevents it from forgetting anything it's ever written. Essentially operating like a git-history, making sure it knows what state the system is in. Also doesnt hurt that everything, including the model weights are represented in NDA format, which was designed to be deterministic and easy for LLMs to understand, so even a tiny model like an 0.5b coder can understand 'if A then B then C. No bloated JSON serialization and deserialization that causes hallucination itself, pure determinism.
That NDA + Merkle approach is interesting. Do you have a public repo where you're developing this? Would love to take a look at how the parser handles the 5-pass propagation.
Currently setup is strictly proprietary. Mostly because of 3 other apps that are almost done (Messenger, Share, Remote - translation: Whatsapp, File Share, Anydesk). Each performing scales of magnitude better than their commercial counterparts while being way more secure.
Everything is closely guarded at the moment, because the underlying architecture isnt just fast, it's more secure than anything out there. Eg. Share outperforms SFTP by 31x, clocking in at 7800MB/s E2E, while being fully encrypted at rest, in transit and at runtime. As much as I want to open source it all, I'm currently earning $1000 a month writing blazor code, so I'm trying to get VC funding for it or market it to industries as a solution (eg. the bank transfer protocol hits 200k transactions per second, securely and without error).
While I'm not going to disclose the underlying architecture, what I can share without risk of cloning, is well document with code snippets in the 12 part series if you're interested.
Fair enough — proprietary makes sense if you're targeting enterprise or VC. 7800MB/s E2E is no joke if it holds up in production. Hope you get the funding to prove it at scale. I'll dig into the series for now — Part 2 on NDA was the one that caught my eye.
I'll check it out — thanks for pointing there. A README often tells more about the architecture than the code itself anyway.
Sorry it's not a public repo, once I get NDA's JIT compiler to handle executing legacy apps, I'll probably open source it to let it catch more traction, while keeping velocity proprietary. I mean it doesnt help you have a custom file format if nobody uses it?
Fair point. A format without adoption is just a hobby. Hope you get the JIT to a state where legacy apps run — that's the milestone where people actually start paying attention. Keep shipping 💪
Given all the optimizations that went into it... It would be a game changer. Imagine this, you start up pytorch, but it runs at rust-speeds? That's what I'm heading towards, it's already capable of hot-swapping efficient code for inefficient code, it's just getting the legacy apps' dependencies working, apps rely so heavily on the underlying OS, that it's quite a mission to get them to run without it.
Hey, I love to read your posts. Can you write something about ManifestGo (manifestgo.app), I Can give you 100 credits for free for testing it out.
Arbab — checked out your profile. 8 posts in a month at 16 is legit, most people never hit publish at all.
Appreciate the offer about ManifestGo but I don't really do sponsored posts or tool reviews. That said, free advice since you asked:
The "5 Things I Learned" post is your strongest one. The others read more like product announcements than stuff people actually want to discuss. You've got 0 comments across 8 posts — not because your product is bad, but because the posts sell instead of teach.
Next one: pick a specific problem → how you solved it → what broke → what you learned. That's what gets replies.
Building and writing at the same time at 16 — that's rare. Keep going.
This is easily the best advice I’ve received since I started publishing. You're totally right—I got caught up trying to 'sell' the product rather than sharing the actual developer journey.
Hearing that from someone else makes it click perfectly. I'm definitely going to ditch the announcement style and start writing about the real, messy parts of building.
Thanks for the massive perspective shift and the push to keep going. I won't drop the ball!"
Glad it landed. People remember the messy stuff — in code and in stories. Drop me the link when you post the next one 👊
They’ve published eight articles, while I’m so lazy I can’t even manage to do a single one properly.
Your comments are worth more than most people's articles. Quality over quantity — and you've got the quality part covered 😄
Thanks❤️ but I don't feel like I write high-quality content. The articles I've written seem really strange to me—even though I might have spent 30 minutes writing a piece that takes just two minutes to read.
Everyone feels this way about their own writing. I've published 30 articles and I still look back at some of them thinking "what the hell was I thinking." You're just not your own target audience — which is the hardest perspective to get.
I'm still figuring it out myself. What I've found so far is that it comes down to two things: are you telling the truth, and does it resonate with someone else. Everything else — style, technique — is secondary. If one or two people finish it and say "I've been through that too," it was worth it.
Still learning. We're in the same boat. Drop me a link to one of yours — I'll read it properly.
I've read your article a few times now. 47 reactions and 47 comments — that's not nothing.
IIT Bombay → Google prep → Microsoft is already a good story, and the structure is clean. The "you're not starting over, you're building on experience" line is the right takeaway.
If I can offer one piece of honest feedback: the article lands, but it mostly stays at the lesson level. The moments that make a story stick are the specific ones — like what actually went wrong in the Google interview, or the exact day you decided to keep going instead of giving up. Those details turn a good motivational post into something people remember.
You've got a solid foundation. Just needs more of you in it. Keep writing 🙌
I am learning to write and will continue to do so. Thank you for the excellent feedback; I will certainly apply it to my life. I have just one request: please continue to guide me. Thank you.
That's the right attitude. You've already got the hard part — actually writing and publishing. Most people never get past the "I'll start tomorrow" phase. The feedback loop is simple: write → publish → read what works → write again. You're already in it. Just keep going and I'll keep reading 🙌
One thing AI can't automate is good leadership. Models can generate code, but they can't replace technical judgment, trust built with clients, or years of production experience. Those remain a company's real competitive advantage.
And yet companies keep trying to automate exactly that. The irony is the people who make those decisions are the ones who lack those things themselves. 🤷
This felt like much more than a story about nepotism. What stood out to me was how you showed that the real damage wasn't caused by AI itself, but by replacing years of production knowledge with confidence and titles.
The detail that really stayed with me was the contrast between the clean benchmark dataset and the messy reality of production. That's something many people underestimate. A model can look incredible in demos, but if it hasn't earned trust in real-world conditions, those impressive numbers don't mean much.
I also liked how Diana's decision wasn't driven by emotions or loyalty—it was driven by repeated operational failures. That made the ending feel believable instead of just satisfying.
The line about a system belonging to the person who truly built and understood it was a powerful way to end the story. Whether fictional or inspired by real experiences, it highlights an important engineering lesson: domain knowledge, careful validation, and humility are far harder to replace than organizations often realize.
Really enjoyed this bonus chapter. Looking forward to seeing how the next Stratagem connects to these themes.
You nailed the benchmark vs production gap — that's the part that scares me the most in real engineering. A clean test set never lies, but it never tells the full truth either. Kevin's 98.7% wasn't fake data, it was just measuring the wrong thing.
Stratagem 2 is coming — and it's about someone who thinks they've learned from the past. Appreciate you reading, glad the bonus chapter hit 👊
For me most times reading on dev.to, may not be the writers intent but a single sentence, phrase or a comment in the post that unlocks or clicks a different realization. 🤔
That's the best kind of comment to get. A story isn't a lecture — it's just a lens. What you see through it is yours. Do you remember which sentence it was?
Your post reminded me of the distinction between hypothesis, theory, and law. We often treat engineering philosophies as elegant theories. But when code hits bare metal, survives production, and endures messy real-world variance, that's when those ideas earn their credibility. The gap between a polished demo and production chaos is where engineering principles are truly tested.
Hypothesis → theory → law. Yeah I'm stealing that 😄 What gets me is how long something can sit in "hypothesis" while everyone talks like it's already gospel. What's your best "looked great on paper, died in prod" moment?
I tried building a mini Apache-like server inside my browser OS. On paper the philosophy made sense: abstract the filesystem, route requests, and let apps behave as if they were running on a local server.
In practice, it fell apart. Looking back, I don't think the philosophy failed—I think my understanding of the underlying syntax and implementation wasn't mature enough yet. It was one of those "looked great on paper, died in prod" moments that taught me the difference between a sound idea and a working implementation.
Still learned a lot from it though.
That last line hits different. "Not the philosophy — my understanding." Most people stop at the first part. So did you rebuild it with what you know now, or start over?
Shelved for now while I fill the necessary gaps.
Tom's mistake wasn't getting laid off it was building something so deep into one company DNA that leaving it felt like amputation. The real horror isn't Kevin. It's that Kevin was allowed to happen because nobody above Tom understood what he'd actually built. The 99.97% recall rate wasn't a number, it was six years of edge cases nobody else knew existed. Diana knew. Kevin never did. That last line hit hard — a system is yours forever because you built it.
"Amputation" — that's the word I kept circling and never landed on. You're right.
Six years of edge cases that nobody wrote down because nobody thought they'd need to. Diana knew the number was real because her crew had dealt with the 0.03% false positives. Kevin walked in, read an API doc for three days, and decided he'd seen enough.
Appreciate you reading it that close, Basit 👊
'VP of Engineering' three months after a college project is the most honest summary of how some orgs actually allocate power — titles follow relationships, not track record. The scary part is how long the dashboard can stay green before reality catches up.
That last sentence is the one — "how long the dashboard can stay green before reality catches up." That gap is where all the damage happens. The longer the green, the harder the crash.
Title != competence is a lesson every org learns eventually. Some learn it after a month, some after a client walks.
Is this the kind of story where any similarities to real world entities are pure coincidence? xD
It is 😄 Just that coincidences keep piling up, you know?
Yeah this one is definitely too good not to publish - another thriller!
Means a lot 🙌 This one wrote itself — some stories don't need AI to go wrong, just bad decisions.
It was surely an entertaining read again ...
Stratagems #2 is already in the works. Buckle up 👀
Some comments may only be visible to logged-in visitors. Sign in to view all comments.