DEV Community

Pragmatic AI Tools Series' Articles

Back to Debashish Ghosal's Series
Your Flaky Tests Aren't Flaky
Cover image for Your Flaky Tests Aren't Flaky

Your Flaky Tests Aren't Flaky

Comments
6 min read
I Planned 10 LLM Evaluation Experiments And Only Ran 1. It Was Enough.
Cover image for I Planned 10 LLM Evaluation Experiments And Only Ran 1. It Was Enough.

I Planned 10 LLM Evaluation Experiments And Only Ran 1. It Was Enough.

5
Comments 2
13 min read
I Thought Building Agent Observability Was a Detector Problem. I Was Wrong.
Cover image for I Thought Building Agent Observability Was a Detector Problem. I Was Wrong.

Real production data broke 35 detectors instantly

I Thought Building Agent Observability Was a Detector Problem. I Was Wrong.

27
Comments 19
8 min read
I Built Scenario Packs for Agent Regression Testing. The Integration, Not the Judge, Broke Me.
Cover image for I Built Scenario Packs for Agent Regression Testing. The Integration, Not the Judge, Broke Me.

Adapters need conformance suites too

I Built Scenario Packs for Agent Regression Testing. The Integration, Not the Judge, Broke Me.

18
Comments 21
15 min read
I Stopped Trusting AI Agents With Tools. So I Built a Gatekeeper.
Cover image for I Stopped Trusting AI Agents With Tools. So I Built a Gatekeeper.

83 real agents tested zero mocks

I Stopped Trusting AI Agents With Tools. So I Built a Gatekeeper.

54
Comments 61
14 min read
I Shipped an Agent Gatekeeper (v0.1). 14 Developers Showed Me What I Missed. Here's v0.2 — a Control Plane.
Cover image for I Shipped an Agent Gatekeeper (v0.1). 14 Developers Showed Me What I Missed. Here's v0.2 — a Control Plane.

I Shipped an Agent Gatekeeper (v0.1). 14 Developers Showed Me What I Missed. Here's v0.2 — a Control Plane.

8
Comments 2
13 min read
9 Bugs That All Looked Like a Working System
Cover image for 9 Bugs That All Looked Like a Working System

Silent statistical traps in self-improving loops

9 Bugs That All Looked Like a Working System

20
Comments 14
10 min read
I Built an AI That Rewrites Its Own Prompts — Its Safety Gate Rejected Every Single Edit
Cover image for I Built an AI That Rewrites Its Own Prompts — Its Safety Gate Rejected Every Single Edit

I Built an AI That Rewrites Its Own Prompts — Its Safety Gate Rejected Every Single Edit

24
Comments 7
8 min read
I Let an LLM Rewrite Its Own Prompt. The Real Win Was the Gate That Rejected It.
Cover image for I Let an LLM Rewrite Its Own Prompt. The Real Win Was the Gate That Rejected It.

I Let an LLM Rewrite Its Own Prompt. The Real Win Was the Gate That Rejected It.

11
Comments 1
6 min read
My Agent Found Real Improvements. The Statistics Still Killed the Promotion.
Cover image for My Agent Found Real Improvements. The Statistics Still Killed the Promotion.

My Agent Found Real Improvements. The Statistics Still Killed the Promotion.

11
Comments 3
6 min read
I Tried 4 Models to Save My Self-Improving Agent. All 4 Failed.
Cover image for I Tried 4 Models to Save My Self-Improving Agent. All 4 Failed.

I Tried 4 Models to Save My Self-Improving Agent. All 4 Failed.

24
Comments 2
7 min read
When Your Benchmark Finally Tells the Truth
Cover image for When Your Benchmark Finally Tells the Truth

When Your Benchmark Finally Tells the Truth

13
Comments 6
8 min read
We Could Have Shipped on Local Models Alone
Cover image for We Could Have Shipped on Local Models Alone

We Could Have Shipped on Local Models Alone

6
Comments 2
12 min read
A Better Model Improved the Numbers. It Didn't Fix the Product.
Cover image for A Better Model Improved the Numbers. It Didn't Fix the Product.

A Better Model Improved the Numbers. It Didn't Fix the Product.

15
Comments 4
10 min read
The 6-Line Fix That Outperformed My Entire Matcher Week
Cover image for The 6-Line Fix That Outperformed My Entire Matcher Week

The 6-Line Fix That Outperformed My Entire Matcher Week

18
Comments 5
11 min read
My 3B Model Found a Shortcut. It Took Me Three Fixes to Close It.
Cover image for My 3B Model Found a Shortcut. It Took Me Three Fixes to Close It.

My 3B Model Found a Shortcut. It Took Me Three Fixes to Close It.

14
Comments 2
9 min read
I Tried to Poison My Agent's Rule Store. It Produced 20 Triggers. Zero Got In.
Cover image for I Tried to Poison My Agent's Rule Store. It Produced 20 Triggers. Zero Got In.

I Tried to Poison My Agent's Rule Store. It Produced 20 Triggers. Zero Got In.

13
Comments 1
9 min read
A Rule Can Be Specific and Still Be Too Broad
Cover image for A Rule Can Be Specific and Still Be Too Broad

A Rule Can Be Specific and Still Be Too Broad

12
Comments 3
10 min read
I Shipped a Fix That Fixed Nothing. Here's Why I Kept It.
Cover image for I Shipped a Fix That Fixed Nothing. Here's Why I Kept It.

I Shipped a Fix That Fixed Nothing. Here's Why I Kept It.

20
Comments 3
10 min read
The Model Wrote the Right Rule and My Replay Rejected It: The Extraction-vs-Replay Split
Cover image for The Model Wrote the Right Rule and My Replay Rejected It: The Extraction-vs-Replay Split

The Model Wrote the Right Rule and My Replay Rejected It: The Extraction-vs-Replay Split

17
Comments 7
6 min read
My Extraction Score Was 0.08 and the Model Was Innocent: Rebuilding the Ruler
Cover image for My Extraction Score Was 0.08 and the Model Was Innocent: Rebuilding the Ruler

My Extraction Score Was 0.08 and the Model Was Innocent: Rebuilding the Ruler

15
Comments 3
8 min read
0/60 Wasn't the Model: The Empty Haystack Behind My Two Worst Corpora
Cover image for 0/60 Wasn't the Model: The Empty Haystack Behind My Two Worst Corpora

0/60 Wasn't the Model: The Empty Haystack Behind My Two Worst Corpora

19
Comments 7
7 min read
Killed by the Word 'git': One Token of Coincidence, 40 Points of Pass Rate
Cover image for Killed by the Word 'git': One Token of Coincidence, 40 Points of Pass Rate

Killed by the Word 'git': One Token of Coincidence, 40 Points of Pass Rate

18
Comments 9
8 min read
A Floor of 0.80 and a Ceiling of 0.63: The Semantic Channel That Never Fired
Cover image for A Floor of 0.80 and a Ceiling of 0.63: The Semantic Channel That Never Fired

A Floor of 0.80 and a Ceiling of 0.63: The Semantic Channel That Never Fired

17
Comments 2
7 min read