DEV Community

Testing

Find those bugs before your users do! 🐛

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
My AI Agent Found a 5-Sigma Result on Day One. I Deleted It.

My AI Agent Found a 5-Sigma Result on Day One. I Deleted It.

Comments
4 min read
Don't Trust the Transcript: A Pytest Harness That Audits What Your AI Coding Agent Actually Did

Don't Trust the Transcript: A Pytest Harness That Audits What Your AI Coding Agent Actually Did

Comments
6 min read
You don't retrain For You when the product changes. You change a weight. I simulated that loop.

You don't retrain For You when the product changes. You change a weight. I simulated that loop.

1
Comments
4 min read
Stop Picking LLMs by Vibes: A Reproducible Evaluation Harness You Can Run for Free

Stop Picking LLMs by Vibes: A Reproducible Evaluation Harness You Can Run for Free

Comments
5 min read
I Stopped Trusting My Gut on AI Coding Models. Here's the 30-Minute Test Rig I Use Instead

I Stopped Trusting My Gut on AI Coding Models. Here's the 30-Minute Test Rig I Use Instead

Comments
6 min read
Your Agent's Permission Slip Belongs in Version Control, Not in a Prompt

Your Agent's Permission Slip Belongs in Version Control, Not in a Prompt

Comments
7 min read
Which parts of this are real? Should be a question with an answer

Which parts of this are real? Should be a question with an answer

1
Comments
7 min read
I Stopped Trusting My Agent's Boundaries Until I Could Break Them in a Throwaway Sandbox

I Stopped Trusting My Agent's Boundaries Until I Could Break Them in a Throwaway Sandbox

Comments
5 min read
Shadow-Gate Your LLM-Generated SQL: A Replay Test Against a Frozen Fixture Database

Shadow-Gate Your LLM-Generated SQL: A Replay Test Against a Frozen Fixture Database

Comments
6 min read
A Boundary-Failure Test Plan for Coding Agents You Can Run on Free Model Tiers

A Boundary-Failure Test Plan for Coding Agents You Can Run on Free Model Tiers

Comments
5 min read
The Question Nobody Asks About Free Coding Models: How Many of Their Patches Break Something Else?

The Question Nobody Asks About Free Coding Models: How Many of Their Patches Break Something Else?

Comments
5 min read
Gate the Toolbelt: A Filesystem-First Smoke Test for Agentic Coding Models

Gate the Toolbelt: A Filesystem-First Smoke Test for Agentic Coding Models

Comments
5 min read
Model Swaps Are Boundary Events: Gate Agent Tool Changes With a Deterministic Replay Lane

Model Swaps Are Boundary Events: Gate Agent Tool Changes With a Deterministic Replay Lane

Comments
6 min read
Silent Model Drift Will Break Your Prompts: A Weekly Drift Detector You Can Run for Free

Silent Model Drift Will Break Your Prompts: A Weekly Drift Detector You Can Run for Free

Comments
4 min read
A Reproducible Sandbox Loop for AI-Generated Code: Generate, Isolate, Assert

A Reproducible Sandbox Loop for AI-Generated Code: Generate, Isolate, Assert

Comments
4 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.