DEV Community

Cover image for Stratagems #9: Lena and P Watched Two AI Suppliers Fight. The Logs Said Neither Was Clean.

Stratagems #9: Lena and P Watched Two AI Suppliers Fight. The Logs Said Neither Was Clean.

xulingfeng on July 09, 2026

Watch the fires burning on the far shore. Don't cross until they've burned themselves out. — The 36 Stratagems, Watch the Fires Burning Across the...
Collapse
 
hemapriya_kanagala profile image
Hemapriya Kanagala

I like how this one flipped the focus from choosing between two vendors to questioning whether either was the right choice in the first place. Sometimes the best move really is to step back, gather your own evidence, and avoid getting pulled into someone else's battle. Looking forward to seeing where Lena's in-house approach goes.

Collapse
 
xulingfeng profile image
xulingfeng

That's exactly the pivot I wanted #9 to land on — the question isn't "which one do we pick," it's "why are we picking from this menu at all?" Glad that came through.
Lena's in-house approach is going to take some interesting turns in the coming stories. Let's just say she didn't walk away from that audit empty-handed 😄

Collapse
 
leob profile image
leob

Yeah suspense built up well, and keeping the reader curious about what the outcome will be (and with the exact final outcome being pretty unexpected) ...

I have a feeling that these stories would make for a great movie script - perfect for these contemporary suspense/thriller movies, with their dark atmosphere and silent/mysterious characters, and a layered plot with multiple strands ... makes sense, or totally not?

Collapse
 
xulingfeng profile image
xulingfeng

Haha🤣, you're making me float up here. I'm already casting the movie in my head — Cameron directing, Nolan writing the screenplay, Keanu Reeves and Leo DiCaprio as the leads, and the rest of the cast from The Matrix. 😂

Collapse
 
leob profile image
leob

Yeah that's it, The Matrix vibes - could be me, but I'm getting "images" in my head when I'm reading these stories, it's definitely got a visual quality ... nice experiment - use AI to turn these stories into a low budget movie (YouTube) ? :-) :-) :-)

Thread Thread
 
xulingfeng profile image
xulingfeng

That's a really interesting idea. AI-generated video is pretty mature now — I've noted it down. Might give it a go once the series is done. The main issue is the time cost though — writing these stories while sneaking time at work is already pushing it, haha 😂

Thread Thread
 
leob profile image
leob

Your future second career - AI Movie Director!

Collapse
 
technogamerz profile image
𝐓𝐡𝐞 𝐋𝐚𝐳𝐲 𝐆𝐢𝐫𝐥

Make a movie in Blender, but don't forget blender is not free the price is your soul

Collapse
 
alexshev profile image
Alex Shev

The supplier-fight framing is painfully real. Logs are useful because they move the argument away from who sounds more credible. But even logs need provenance: who can write them, what gets omitted, and whether the customer can see the raw events instead of a polished dispute summary.

Collapse
 
xulingfeng profile image
xulingfeng

Appreciate this — and "polished dispute summary" cuts exactly where it should. That tension between raw events and a curated view is basically the spine of the series at this point.

8 (Alex's hidden dashboard) was built precisely because the AI dashboard was serving a polished summary that hid 1,530 low-confidence anomalies. The suppliers in #9 were fighting over whose dashboard looked greener, but the raw logs showed both were hiding something. Same pattern, different layer.

So yeah — provenance is the next question after "trust the logs." And it's the harder one, because answering it usually means auditing the audit layer, which almost no one does until after the incident.

Collapse
 
alexshev profile image
Alex Shev

Auditing the audit layer is the phrase that matters. Once dashboards become part of the dispute, the dashboard itself needs provenance: raw event access, write boundaries, retention rules, and who can rewrite the story after the fact. Otherwise the audit log is just another narrative surface with better typography.

Collapse
 
technogamerz profile image
𝐓𝐡𝐞 𝐋𝐚𝐳𝐲 𝐆𝐢𝐫𝐥

Great story! I really enjoyed how the plot unfolded and how it kept me guessing until the end.

I have a question out of curiosity. If the story hadn't gone in this direction, how else do you think it could have evolved? For example, if Lena hadn't found those logs, or if both AI suppliers had actually been completely clean, what path would you have taken? Would it have become more of a psychological battle, a technical investigation, or even a story about trust and human decision-making?

I'm also curious about your writing process. Did you have this ending planned from the beginning, or did the story naturally evolve as you wrote it? Were there any alternative plot twists or endings that you considered but ultimately decided not to use?

I'd love to hear your thoughts. It's always fascinating to learn how authors imagine different possibilities before choosing the final version. Looking forward to the next chapter!

Collapse
 
xulingfeng profile image
xulingfeng

Q1:
You have to put yourself in Lena's shoes here. She's the kind of person who sees everything as chess pieces — the two suppliers are pieces, P is a piece (a more useful one, but one she'll eventually pay for treating that way). I honestly think Lena would squeeze both suppliers dry and end up building her own testing platform out of the rubble. So yeah — it always comes back to Lena building her own AI platform, something she can actually trust.
Q2:
Great question. First, I already nailed down each character's personality in the series announcement post. Second, the 36 stratagems themselves are a natural constraint — kind of like setting boundaries when you're vibe coding with AI. The plot can go anywhere as long as it stays inside the stratagem's frame. That makes writing easier in some ways, but it also limits creative freedom. And on top of that, there are multiple interconnected threads running between the stratagems, so that's another layer of constraints.
Overall I have a pretty complete outline for the series — every character arc and every stratagem's story exists in some form. As for how many times I've revised it? Hard to say. Every time I publish a new story I go back and audit the whole outline, fill in gaps. I've definitely considered a lot of alternate plots — started writing, felt something was off, scrapped it and pivoted. Only to circle back to the first version sometimes 😂 Just like work — the boss always ends up wanting the first draft.

Collapse
 
technogamerz profile image
𝐓𝐡𝐞 𝐋𝐚𝐳𝐲 𝐆𝐢𝐫𝐥

I really enjoyed reading this story; it felt like I was watching a movie. You truly write amazingly.

Thread Thread
 
xulingfeng profile image
xulingfeng

Haha, that's such high praise it almost makes me feel like I've successfully switched careers — from QA to professional writer. I'm gonna be waking up multiple times tonight just to smile about it 😂

Thread Thread
 
technogamerz profile image
𝐓𝐡𝐞 𝐋𝐚𝐳𝐲 𝐆𝐢𝐫𝐥

Haha,good! Just remember -- you deserve the compliments too.
Don't smile too much,or people will start asking why you are so happy tommorow!

Thread Thread
 
xulingfeng profile image
xulingfeng

I'm already trying my best to hold back the smile. Just 10 more minutes until I clock out and I'm free. 🤣

Collapse
 
motedb profile image
mote

The detail I can't stop thinking about is TestShield's one-liner: 'your AI test coverage is lower than you think.' It's not a metric, it's a frame flip. Both vendors showed real numbers, and both were arguably right — which is the whole trap. A benchmark is only as honest as the thing it defines as 'normal,' and if your definition of normal is the loophole, 99.97% just becomes a very confident way to miss everything that matters.

This is why I'm suspicious of any single vendor-supplied eval. The only number I'd trust is one I can re-run against my own distribution, not theirs. What's the most underrated evaluation dimension you've seen a team actually get burned by — the one that never shows up in the sales deck?

Collapse
 
xulingfeng profile image
xulingfeng

The dimension that never makes the deck: the cost of building evaluation dependency into your delivery pipeline before you've validated the eval itself.
In Stratagems #10, Leo's experience might be the most common version of this: not that the vendor lied, but that the team committed to the eval format before they understood what it was optimizing for.
Looking forward to your take on Stratagem #10.

Collapse
 
unitbuilds profile image
UnitBuilds

Oh snap, 3 months ago... We all know whose model was chucked 3 months ago...

Collapse
 
xulingfeng profile image
xulingfeng

😏 Someone read between the lines. Let's just say the fine-tune looked great on paper.

Collapse
 
xulingfeng profile image
xulingfeng

First crossover. Two observers, one table. Neither says much — but the espresso and the hash say everything. 🃏