Exhaust the enemy's strength without fighting. Weaken the strong by nurturing the soft.
— The 36 Stratagems, "Wait at Leisure While the Enemy Labo...
For further actions, you may consider blocking this person and/or reporting abuse
three weeks of watching an agent in prod before touching alert config. baselines from pure observation are different - synthetic tests miss the real distribution. the POC that validates by watching first is underrated.
You'd like him. That's almost exactly what the real P told me. "If I touch the config in week one, I'm testing my own assumptions, not the system." Took me a solid minute to realize waiting is the test.
that framing hit different. waiting IS the test - and the hardest part is it looks identical to just not getting around to it
Really enjoyed this one, and the mechanism underneath the story is the part I keep running into for real. The whole thing turns on the gap between detection rate and coverage, and those aren't even the same axis. Detection rate is a number on a dashboard that moves every week, coverage is how much of your actual failure distribution the thing can even see, and P's point is that the second number was frozen from day one while everyone fought over the first. That's Goodhart in a hard hat, the moment detection rate becomes the thing vendors optimize it stops measuring anything, which is also exactly why P refuses to reveal which metrics he tracks. A held-out eval the vendor can't see is the only kind that stays honest.
But the sharpest bit for me is buried in the last table. The 12 modes that caused 9 of the P0 incidents were covered by neither vendor, and the 49 both covered hadn't fired in two years. So the real failure wasn't the vendors gaming their numbers, it was that the entire POC measured against a 61-mode standard taxonomy that barely overlapped the actual incident history. Detection rate on the wrong test set is worse than no metric, because it manufactures confidence in the exact places you're blind. The fix is almost boring next to the drama: build your eval set from your own two years of incidents, not the vendor's taxonomy, and the whole illusion collapses in week one instead of month three. P basically did that quietly and then waited. Great series.
You got it. That last paragraph especially — "build your eval set from your own two years of incidents" — is the whole thing in one sentence. Boring answer, nobody wants to hear it, because it means doing your homework instead of buying a shiny dashboard.
What stuck with me after writing: the VP never asked about FirmCore's own coverage for three months. Not lazy — just didn't want to find out he didn't know what he was measuring. Once you ask that question, you can't un-ask it. P knew.
Thanks for reading this deep. Comments like this are why I keep writing. 🔥
Yeah, and that reframe is the darker version of it. It's not that the VP couldn't ask, it's that asking makes you the owner of the answer. The coverage question stays unasked because whoever raises it inherits the bad news and the cleanup, so the quiet incentive is to keep it fuzzy. P had no skin in that game, which is the whole reason an outside evaluator can even see straight. Good series, curious where #5 goes.
"Whoever raises it inherits the bad news" — that's the line I wish I'd written. The whole system is set up to reward not knowing. P could see because P had nothing to lose.
First six stories are solo runs for each of the six protagonists from the teaser. #5's lead is the exact opposite of P — still sitting on my hard drive though. Being a QA in real life means I keep auditing my own stories for bugs before I ship them 😂 When it feels right, I'll hit publish. No rush. Slow and steady.
Also — once you've read a few of these, I'd love to know which protagonist hits closest to home for you. Let me know when you've met them all 👀
Ha, the auditing-your-own-stories-for-bugs-before-shipping thing is painfully relatable, that's basically my whole job description with a different noun. I'll take you up on it, going to read back through the series and tell you which one hits closest once I've met them all. The "P had nothing to lose, so P could see straight" line is going to live in my head for a while though. Keep them coming.
Haha it's basically an occupational hazard at this point. Welcome to the AI rabbit hole series — comments like yours are what keep the keyboard smoking 🔥⌨️
Really enjoying this series so far. It's been fun seeing how the different characters and stories start connecting together.
Looking forward to the next one 😄
Glad you're spotting the connections — they only get tighter from here 😄 Next one's a different beast though. Thanks for sticking with the series!
The Stratagems format is a good frame for AI evaluation failures. "Wait at leisure while the enemy labors" captures exactly why vendor demonstrations and POC narratives are so unreliable â the vendor is optimizing for showing you success, not for revealing where their coverage is thin. P's move was to eliminate the performance entirely and just observe.
The interesting thing from the monitoring side: the coverage numbers weren't wrong because MonitorAI lied. They were wrong because coverage was measured against a vendor-curated list of failure modes rather than against actual production incidents. The 47.5% real coverage vs 97.8% claimed coverage gap is a data labeling problem more than a product defect â if you'd asked MonitorAI to define coverage before the engagement started, they'd have given you the same 97.8%. The gap came from a mismatch between whose fault taxonomy was used.
This shows up in AI monitoring broadly: "we cover X% of failure modes" is only meaningful if you've agreed on what the failure modes are and whose incidence data you're counting against. Production data drift and edge-case sensor behavior are outside the vendor's training distribution, so they're not covered â not because the system is bad, but because the problem definition excluded them.
The lesson I'd take for agent evaluation more broadly: define your failure modes from production observation before you evaluate any tool against them. The POC should confirm that a tool handles failures you've already seen, not define failures for the tool to optimize against.
Spot on about the taxonomy mismatch — P's 47.5% wasn't a gotcha. MonitorAI genuinely thought they had 97.8%. The gap isn't lying, it's framing: they measured against what they already knew to look for. P measured against what actually broke. Those are different lists.
The part I enjoyed most was P figuring this out week one and just... waiting. Say it on day one and everyone argues. Let the data say it three months later and people just nod.
Your last line hits it — define failures from what's already burned you, not from what a vendor optimizes for.
Fifth stratagem's almost ready — just polishing a few things. Would love to hear what you think when it drops.
Mic drop with 4 tables 😂 It's reminding me of anime, when the "I was secretly 10 steps ahead" plot-twist hits.
Haha right? Like dude just flexed right in your face and walked away. 😂
The gap between "99% coverage" and 29 of 61 real failure modes is the whole story, and it's why a coverage number with no denominator attached is close to meaningless. What made P's read-only pipeline work is that it counted the incidents the monitor never fired on, not the ones it caught. That inverted metric, misses over hits, is the one I'd want on the wall before trusting any monitoring claim.
"Misses over hits" — that's the whole story in fewer words than I managed. The hard part about inverted metrics isn't defining them. It's that nobody signs a budget for "the things you didn't see." P pulled it off because he never asked who was paying for it.
"If you know too early, no one believes you." Story of my life.
The worst part is when you're proven right later and nobody remembers you said it first 😂
ahhahah this... or they play dumb
Stratagem #5 is still being thoroughly reviewed. Stay tuned 😄
Ofc I will, omg, this is an absolute godsend. Every story turns into a full-on cinematic thriller, I swear to God xD. Your stuff and orbithealien.com 's IG are the best find, always make my day.
OK, I finally got some time and I have to praise you (it's long, no TL;DR today—people can unleash their bots on it if they want xD).
Anyways: your style is soo immersive, funny, and just awesome. The Video Game/Terminator summaries absolutely killed me. That is exactly how my brain works when I visualize tech disasters, see people screwing up, or finally get angry after putting up with shit. I have seen no one write like that in ages. Your imagination and the way your brain jumps between ideas is insane. I keep spamming all my cool tech friends with your series, and they all said you are dope. An idea: hink about a podcast or audio stories, please. It would be hilarious to give those corp bros voices—the ones who love to "open the kimono" while proudly showing off their low-hanging... dried fruits khm 😂 Throw in all the buzzword jargon and have the narrators absolutely nail that intonation.
You're also damn good with tech. You can tell this isn't researched or faked, but you actually live this stuff. The depth, curiosity, and joy of pulling systems apart is all over your protagonists. I think that's why they feel so believable.
My depth, breadth, and passion for tech align completely with your protagonists’ vibe. The only difference is I don’t play the games, so Stratag(A)mes(sic!) also hit hard hahah. My Slavic pride and stupid moral compass mean I always end up stepping back and giving people everything I know regardless of how they treat me. I just can’t seem to override it. Honestly, it feels great seeing your guys win. Somewhere in the back of my head I’m always like "...yeah, that could’ve been me too." (But goody two-shoes, alas)
P.S. ^^ Hmm, hnow I think about it, I actually think these stories helped a little. Yesterday I finally did it for once (almost). Things are about to break because I’m leaving some sc*m place. Instead of "crying wolf" in the channels, I just dropped notes in the docs, left all the links, and encouraged them to read. (They probably won’t. They always hated me for documenting stuff. Maybe their bots will look. 😂)
Ofc after 24h I felt bad and had to message one colleague and tell him, if it breaks, call me. I’ll fix it. (For free ofc, because there are still a few normal people there, and I don’t want clients going through hell.)
Still... the fact I managed to stay mad for 24 hours and fought the urge to tell them everything feels like a huge win, omg. The sentence "If you know too early, no one believes you" actually did the job.
Thank you so much. These rock. Just laughed at " SentryWave followed right behind: "99.7% coverage, 7-day deployment" — bigger numbers, bolder font with my friend, he is also dead. Damn bolder font nailed it xD
Oh no — this might be my favorite comment on the whole series. You went and wrote something this heartfelt — feels like my brain just got a shot of pure focus, thinking at 10x speed. Pretty sure my keyboard evolved on the spot and is about to start writing stories on its own any minute now. 😂
"Stay mad for 24h, offer to fix it for free, feel guilty about it anyway" — congrats, you're officially a protagonist in training. Just need a stratagem name and a habit of never telling anyone your full plan. 😂
You and your friend laughing at "bolder font" together — that's exactly why I keep sneaking those little details in. 😂
About the podcast idea — not ready for that yet hahaha. I'm actually super shy in real life. But who knows, maybe someday. 😅
Also — you're the first person who noticed that robot AI summary at the end. Congrats. Why's it there? Gotta keep that one to myself for now. 🤫
Tell your tech friends I owe them a beer 🍺
Honestly though, this series is still alive only because of every like and comment from people like you. So before anything else — thank you. From the bottom of my keyboard. 🙏
A monitoring POC without tests is basically theater. The useful question is what failure would change the team's behavior, and whether the system catches that failure before users do.
Exactly. Whole POC was built to make pretty dashboards, not to tell anyone anything useful. That one question — "what failure would actually change what you do" — would've wrapped that meeting in 5 minutes. Which is exactly why nobody in that room wanted to ask it.
That question is uncomfortable in the best way.
If a monitoring POC cannot name the failure that would change a decision, it is really a dashboard POC. The useful test is whether the system can make someone stop, rollback, page, or change priority with confidence.
Only had a few spare minutes, but I couldn't skip your post. Worth every minute! Great read. ❤️🔥
Haha 32 more in the backlog. All-you-can-eat buffet, no rush 😂
"All-you-can-eat" - yes, chef! ;-)
🤣hahaha
Coming right up, chef 🫡😂
Okay, honest question — should P ever get an actual name, or are we just rolling with the letter forever? And since we're here, what is P's gender? My brain cells are running on fumes here. Drop your thoughts, I'm lost 🤣
@unitbuilds this cursed red flag catch you?🤣
Yip... 😅
Nice! Across all these stories, I like the protagonists' cool and collected demeanor ... :-)
Wait til you meet in #5 — he's the one who breaks the calm 😂 Thanks leob.