DEV Community

Cover image for Stratagems #7: P Watched an AI That Only Looked One Way. The 99.97% Was Real. It Just Missed Everything That Mattered.

Stratagems #7: P Watched an AI That Only Looked One Way. The 99.97% Was Real. It Just Missed Everything That Mattered.

xulingfeng on July 07, 2026

"Show nothing, hold everything." — The Thirty-Six Stratagems, Create Something Out of Nothing Previously on this series: #4: P Walked Into an AI...
Collapse
 
unitbuilds profile image
UnitBuilds

Ooooh good one, third cup, I was half expecting 'third drumstick' 😭

Loving the series, especially cuz it points out real attack vectors that go unnoticed. You train an AI, but if you corrupt your own training data with test data, how can you not expect it to fail? Pattern matching is what AI does best, if you teach it to not flag pen-tests, then it's literally useless, because the whole point of pen-testing is that it actually catches it... The read-only strategy is a good one, especially when you consider ports tend to be open for egest and block ingest, with only secured APIs going through to trigger known events, eg. reading a database... Smart tactic and one that isnt often used, yet that's how millions of passwords leak.

Collapse
 
xulingfeng profile image
xulingfeng

Third cup spotted 😄 "third drumstick" would've been great actually, now I'm second-guessing myself.
The training data thing though — that's the one that keeps me up. It's so obvious in hindsight. You teach a pattern matcher to ignore the exact thing it's supposed to catch. But nobody thinks "green = blind" in the moment. They see the pass and move on. I've done it too.
Read-only strategy is one of those "too simple to be real" things until you actually trace a breach and realize yeah, a basic filter would've stopped the whole thing cold. Most of the nasty leaks I've seen weren't fancy — just a port doing exactly what it was told, and nobody asking if it should be listening.
Anyway, appreciate you reading this deep. Means more than the drive-by "great post" — you're actually kicking the tires.
So... does this mean you're getting KFC today? 🍗

Collapse
 
unitbuilds profile image
UnitBuilds

You know... Just maybe, because you are such a generous person, you really made me crave it yesterday, but I controlled myself... today is a different story (it's 2:43am my side) and I'm seriously contemplating lunch break drive-thru (it's like 2 min away from my apartment, it'll be 1h away (hopefully) in 2 months time 😭 though I really hope everything goes well and we get the house, it would be awesome! (it comes with the pool table)

Thread Thread
 
xulingfeng profile image
xulingfeng

Man, you're gonna be writing the next series updates from next to a pool table. That's an upgrade I didn't see coming.
And trust me — the moment you sink a clean shot, the 1-hour KFC drive becomes a distant memory. Worth it. (Mostly. Get the drumsticks.)🤣

Thread Thread
 
unitbuilds profile image
UnitBuilds

It's gunna be worth it, given that rent is only 30% more and still 20% less than renting... I'll just have to convince them that for the app, they need to deliver to the next town (I think worth it, for a complete solution?) 😂

Thread Thread
 
unitbuilds profile image
UnitBuilds

Today (8th's) game is gunna be a fun 1, you ever play Witcher 3's gwent?

Collapse
 
hemapriya_kanagala profile image
Hemapriya Kanagala

Another interesting one! The idea that a model can achieve an impressive metric while still missing what actually matters is something we see surprisingly often. Really enjoying how this series blends technical ideas with storytelling.

Collapse
 
xulingfeng profile image
xulingfeng

Thanks Hemapriya — "surprisingly often" is the scariest part, right? The metric says green, the room's on fire. Glad you're still along for the ride

Collapse
 
jugeni profile image
Mike Czerwinski

The whitelist detail is the tell I would keep reading for. Whitelisting your pentester's certificate tier is not laziness; it is measured-against-owned-spec. The 99.97% is real, and it is measured against a definition of intrusion the same team wrote. When the vendor owns both the classifier and the taxonomy the classifier grades against, the classifier's accuracy is a claim about the taxonomy, not about the world.

The medical-data leak sat outside the taxonomy. That is not a detection failure. That is a scope statement dressed as a metric. The report says "99.97% of the things we agreed to look for." Which is a different sentence from "99.97% of what mattered to the client."

What I would want P to bring back is not a hole in the classifier. It is the write-time definition of the alert set, dated and signed, next to the report's percentage. A percentage without a co-timestamped taxonomy is a display metric, not evidence. The whitelist too complete is the earliest visible artifact of the same choice.

Collapse
 
xulingfeng profile image
xulingfeng

You're right — and I think that's the exact tension P operates in. The #7 engagement shipped what was in scope, and the taxonomy itself wasn't in scope.
Let's just say #9 has a different kind of delivery. Different engagement, different approach.
Curious what you'll make of it.

Collapse
 
yune120 profile image
Yunetzi

Nice angle - the 99.97% is real until it isn't. AI can ace clean tests but miss messy humans, ethics, and edge cases. Metrics hype meets real-world cost. Let's demand accountability and test for nuance, not just shiny numbers.

Collapse
 
xulingfeng profile image
xulingfeng

Appreciate you reading it that close. "Metrics hype meets real-world cost" — that's exactly the gap the whole series keeps circling back to. The 99.97% isn't wrong, it's just... looking the wrong way. What's your take — at what point does a metric stop being useful and start being a liability?

Collapse
 
xulingfeng profile image
xulingfeng

Quick note for series readers — MedTech isn't a random name. If you want to know whose AI monitoring missed this leak too — go back and find it. 👀
Also lost count of how many threads I've buried across these stories already. You've probably caught a few. The rest won't surface until way later. That's the whole 36-part thing. 🤣

Collapse
 
leob profile image
leob • Edited

Some nice good old enjoyable "Sherlock Holmes" style detective work - I'm seeing a recurring pattern though:

"a company/org choosing to ignore 'stuff' because 'we know it doesn't matter'" - can we call it "complacency", or "laziness"? ;-)

P.S. I think at some point I'll need to revisit some previous episodes to figure out the various 'strands'

Collapse
 
xulingfeng profile image
xulingfeng

Glad you spotted the pattern — it's not accidental.
I can finally say it out loud: when you stare into the abyss, the abyss stares back.
And yeah, revisiting previous episodes is a good call. You've already passed a few clues without knowing it. The answer's been right in front of you.😄

Collapse
 
technogamerz profile image
𝐓𝐡𝐞 𝐋𝐚𝐳𝐲 𝐆𝐢𝐫𝐥

Good Article 👍🏻

Collapse
 
xulingfeng profile image
xulingfeng

😁🎉