DEV Community

Testing

Find those bugs before your users do! 🐛

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
My quality gate wasn't strict. It was dead.

My quality gate wasn't strict. It was dead.

Comments
4 min read
Hype Cycles Don't Ship My Code: How I Gate New LLMs with a Self-Written Eval Deck

Hype Cycles Don't Ship My Code: How I Gate New LLMs with a Self-Written Eval Deck

Comments
5 min read
Your Feed Says the New Model Is Great. Mine Says Prove It in Under an Hour

Your Feed Says the New Model Is Great. Mine Says Prove It in Under an Hour

Comments
4 min read
I Broke My AI Decision System 20 Ways to See Which Layer Would Catch What

I Broke My AI Decision System 20 Ways to See Which Layer Would Catch What

Comments
8 min read
The Receipt Was Cryptographically Valid. The Action Didn't Match the Judgment.

The Receipt Was Cryptographically Valid. The Action Didn't Match the Judgment.

1
Comments 1
11 min read
Your Agent Reads Untrusted Text All Day. Here's How I Grade What It Does With It.

Your Agent Reads Untrusted Text All Day. Here's How I Grade What It Does With It.

Comments
6 min read
Pick Your LLM With a Scoreboard, Not a Hunch: A Two-File Eval That Costs Nothing to Re-Run

Pick Your LLM With a Scoreboard, Not a Hunch: A Two-File Eval That Costs Nothing to Re-Run

Comments
6 min read
A Two-Hour Fit Test for AI Coding Models on Your Own Codebase

A Two-Hour Fit Test for AI Coding Models on Your Own Codebase

Comments
4 min read
My AI Agent Found a 5-Sigma Result on Day One. I Deleted It.

My AI Agent Found a 5-Sigma Result on Day One. I Deleted It.

Comments
4 min read
I Installed 300 MCP Servers From PyPI. About 43% Don't Start.

I Installed 300 MCP Servers From PyPI. About 43% Don't Start.

Comments
4 min read
Stop Picking LLMs by Vibes: A Reproducible Evaluation Harness You Can Run for Free

Stop Picking LLMs by Vibes: A Reproducible Evaluation Harness You Can Run for Free

Comments
5 min read
Your Agent's Permission Slip Belongs in Version Control, Not in a Prompt

Your Agent's Permission Slip Belongs in Version Control, Not in a Prompt

Comments
7 min read
I Stopped Trusting My Gut on AI Coding Models. Here's the 30-Minute Test Rig I Use Instead

I Stopped Trusting My Gut on AI Coding Models. Here's the 30-Minute Test Rig I Use Instead

Comments
6 min read
Don't Trust the Transcript: A Pytest Harness That Audits What Your AI Coding Agent Actually Did

Don't Trust the Transcript: A Pytest Harness That Audits What Your AI Coding Agent Actually Did

Comments
6 min read
A Boundary-Failure Test Plan for Coding Agents You Can Run on Free Model Tiers

A Boundary-Failure Test Plan for Coding Agents You Can Run on Free Model Tiers

Comments
5 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.