DEV Community

ai Series' Articles

Back to Debashish Ghosal's Series
i was using a frontier model to update jira tickets. here's how i fixed it.
Cover image for i was using a frontier model to update jira tickets. here's how i fixed it.

i was using a frontier model to update jira tickets. here's how i fixed it.

2
Comments
6 min read
Why Agent Evaluation Is Harder Than Model Evaluation
Cover image for Why Agent Evaluation Is Harder Than Model Evaluation

Builder scars reveal untrustworthy agent paths

Why Agent Evaluation Is Harder Than Model Evaluation

19
Comments 25
8 min read
I Built an Agent Eval Harness. Real Agents Broke the Clean Version of the Story
Cover image for I Built an Agent Eval Harness. Real Agents Broke the Clean Version of the Story

Real-world agents break clean evals

I Built an Agent Eval Harness. Real Agents Broke the Clean Version of the Story

29
Comments 41
9 min read
AI Made Prototyping Free. That Is Exactly Why Your Portfolio Strategy Matters Now.
Cover image for AI Made Prototyping Free. That Is Exactly Why Your Portfolio Strategy Matters Now.

AI Made Prototyping Free. That Is Exactly Why Your Portfolio Strategy Matters Now.

14
Comments 3
13 min read
I Ran 157 Agent Plans Against a Real LLM. The Problem Wasn't Execution. It Was Planning.
Cover image for I Ran 157 Agent Plans Against a Real LLM. The Problem Wasn't Execution. It Was Planning.

Explores why self-review fails agents

I Ran 157 Agent Plans Against a Real LLM. The Problem Wasn't Execution. It Was Planning.

33
Comments 32
8 min read
I Told My LLM Critic to Be Adversarial. It Started Blocking Plans for Being 'Not Thorough Enough.'
Cover image for I Told My LLM Critic to Be Adversarial. It Started Blocking Plans for Being 'Not Thorough Enough.'

Code-enforced allowlists fix prompt drift

I Told My LLM Critic to Be Adversarial. It Started Blocking Plans for Being 'Not Thorough Enough.'

16
Comments 26
4 min read
The Planner Made the Same 3 Mistakes Every Time. A Bigger Model Didn't Fix It.
Cover image for The Planner Made the Same 3 Mistakes Every Time. A Bigger Model Didn't Fix It.

Exposes the three recurring failure patterns

The Planner Made the Same 3 Mistakes Every Time. A Bigger Model Didn't Fix It.

21
Comments 28
5 min read
I Ran 170 Agent Goals for $0.49. The Field Test Found 10 Issues That Unit Tests Never Would.
Cover image for I Ran 170 Agent Goals for $0.49. The Field Test Found 10 Issues That Unit Tests Never Would.

I Ran 170 Agent Goals for $0.49. The Field Test Found 10 Issues That Unit Tests Never Would.

17
Comments 5
13 min read
I Tried to Prompt-Inject My Own Agent Engine. It Didn't Work. Here's Why.
Cover image for I Tried to Prompt-Inject My Own Agent Engine. It Didn't Work. Here's Why.

Structural gates beat prompt safety

I Tried to Prompt-Inject My Own Agent Engine. It Didn't Work. Here's Why.

41
Comments 11
12 min read
A Reader Audited My OSS Release in Public. He Found the Contradictions I Missed.
Cover image for A Reader Audited My OSS Release in Public. He Found the Contradictions I Missed.

A Reader Audited My OSS Release in Public. He Found the Contradictions I Missed.

18
Comments 8
7 min read
Most AI Second Opinions Are Fake. I Built a Two-LLM Review Engine to Prove It.
Cover image for Most AI Second Opinions Are Fake. I Built a Two-LLM Review Engine to Prove It.

Tested on 70 real PRs for just $0.53 total cost

Most AI Second Opinions Are Fake. I Built a Two-LLM Review Engine to Prove It.

17
Comments 12
11 min read
My LLM Critic Disagreed With Itself on Every Trial. The Safe Part Was the Code I Didn’t Trust It to Touch.
Cover image for My LLM Critic Disagreed With Itself on Every Trial. The Safe Part Was the Code I Didn’t Trust It to Touch.

My LLM Critic Disagreed With Itself on Every Trial. The Safe Part Was the Code I Didn’t Trust It to Touch.

21
Comments 8
5 min read
Most AI Second Opinions Are Theater. I Built a System That Actually Fights Back.
Cover image for Most AI Second Opinions Are Theater. I Built a System That Actually Fights Back.

Reveals why blind reviews fail

Most AI Second Opinions Are Theater. I Built a System That Actually Fights Back.

15
Comments 10
14 min read
I Thought My Multi-Agent Debate Engine Was Broken. The Real Bug Was the Prompt.
Cover image for I Thought My Multi-Agent Debate Engine Was Broken. The Real Bug Was the Prompt.

I Thought My Multi-Agent Debate Engine Was Broken. The Real Bug Was the Prompt.

15
Comments
31 min read
The Best Model Pair in My Field Test Was Also the Least Trustworthy
Cover image for The Best Model Pair in My Field Test Was Also the Least Trustworthy

The Best Model Pair in My Field Test Was Also the Least Trustworthy

22
Comments 7
12 min read
Two Projects, One Problem — What PlannerCritic and AdversarialDebate Each Got Wrong
Cover image for Two Projects, One Problem — What PlannerCritic and AdversarialDebate Each Got Wrong

Two Projects, One Problem — What PlannerCritic and AdversarialDebate Each Got Wrong

17
Comments 3
10 min read
The Same Model Debating Itself Was More Self-Critical Than Two Different Models
Cover image for The Same Model Debating Itself Was More Self-Critical Than Two Different Models

The Same Model Debating Itself Was More Self-Critical Than Two Different Models

14
Comments 1
13 min read
I Added a Fourth Model Mid-Run. It Changed What My Field Test Could Prove.
Cover image for I Added a Fourth Model Mid-Run. It Changed What My Field Test Could Prove.

I Added a Fourth Model Mid-Run. It Changed What My Field Test Could Prove.

11
Comments
7 min read
The Gate That Stayed Silent — When a Blocker Count That Drops Reads as Improvement
Cover image for The Gate That Stayed Silent — When a Blocker Count That Drops Reads as Improvement

The Gate That Stayed Silent — When a Blocker Count That Drops Reads as Improvement

12
Comments 5
5 min read
My LLM Critic Flip-Flops on Every Run. That's Fine — Because a Frozenset Decides What's Fatal.
Cover image for My LLM Critic Flip-Flops on Every Run. That's Fine — Because a Frozenset Decides What's Fatal.

My LLM Critic Flip-Flops on Every Run. That's Fine — Because a Frozenset Decides What's Fatal.

13
Comments 8
7 min read
I Published Every Flaw My Safety Tool Can't Catch. It Made It More Credible, Not Less.
Cover image for I Published Every Flaw My Safety Tool Can't Catch. It Made It More Credible, Not Less.

I Published Every Flaw My Safety Tool Can't Catch. It Made It More Credible, Not Less.

10
Comments 4
8 min read
The Edit That Fixed 4 Tasks and Broke 1
Cover image for The Edit That Fixed 4 Tasks and Broke 1

The Edit That Fixed 4 Tasks and Broke 1

15
Comments
8 min read
My Self-Improving Agent Still Couldn't Improve. That Was the Breakthrough.
Cover image for My Self-Improving Agent Still Couldn't Improve. That Was the Breakthrough.

My Self-Improving Agent Still Couldn't Improve. That Was the Breakthrough.

9
Comments 2
7 min read
I Compared a Local 4B, a Better Cloud Model, and Role Separation. The Results Were Weird.
Cover image for I Compared a Local 4B, a Better Cloud Model, and Role Separation. The Results Were Weird.

I Compared a Local 4B, a Better Cloud Model, and Role Separation. The Results Were Weird.

11
Comments 2
6 min read
I Thought This Was a Classification Problem. It Wasn't.
Cover image for I Thought This Was a Classification Problem. It Wasn't.

I Thought This Was a Classification Problem. It Wasn't.

10
Comments
6 min read
I Thought the Optimizer Was the Product. I Was Wrong. The Gate Was.
Cover image for I Thought the Optimizer Was the Product. I Was Wrong. The Gate Was.

I Thought the Optimizer Was the Product. I Was Wrong. The Gate Was.

8
Comments 2
5 min read
I Thought Role Separation Would Fix the Optimizer. It Didn't.
Cover image for I Thought Role Separation Would Fix the Optimizer. It Didn't.

I Thought Role Separation Would Fix the Optimizer. It Didn't.

7
Comments 2
6 min read