DEV Community

Testing

Find those bugs before your users do! 🐛

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Don't Wire a Coding Model Into Your Workflow Until It Passes Your Own Harness

Don't Wire a Coding Model Into Your Workflow Until It Passes Your Own Harness

Comments
5 min read
A Race You Can't Reproduce Is Still a Race: Testing Model-Generated C++ Concurrency Fixes

A Race You Can't Reproduce Is Still a Race: Testing Model-Generated C++ Concurrency Fixes

Comments
6 min read
What Breaks First When You Swap a Local Coding Model for a Free Hosted One? A Failure-Mode Probe Suite

What Breaks First When You Swap a Local Coding Model for a Free Hosted One? A Failure-Mode Probe Suite

Comments
5 min read
Fuzz Your Agent's Tool Calls Before You Ship: A Reproducible Boundary Test Harness

Fuzz Your Agent's Tool Calls Before You Ship: A Reproducible Boundary Test Harness

Comments
6 min read
Red-Teaming AI Coding Agents Without a Budget: A Boundary Test Suite on Free Models

Red-Teaming AI Coding Agents Without a Budget: A Boundary Test Suite on Free Models

Comments
6 min read
Testing HL7-to-FHIR Pipelines Without a Hospital: Mocking HAPI FHIR with respx

Testing HL7-to-FHIR Pipelines Without a Hospital: Mocking HAPI FHIR with respx

1
Comments
15 min read
My quality gate wasn't strict. It was dead.

My quality gate wasn't strict. It was dead.

Comments
4 min read
Hype Cycles Don't Ship My Code: How I Gate New LLMs with a Self-Written Eval Deck

Hype Cycles Don't Ship My Code: How I Gate New LLMs with a Self-Written Eval Deck

Comments
5 min read
Your Feed Says the New Model Is Great. Mine Says Prove It in Under an Hour

Your Feed Says the New Model Is Great. Mine Says Prove It in Under an Hour

Comments
4 min read
I Broke My AI Decision System 20 Ways to See Which Layer Would Catch What

I Broke My AI Decision System 20 Ways to See Which Layer Would Catch What

Comments
8 min read
The Receipt Was Cryptographically Valid. The Action Didn't Match the Judgment.

The Receipt Was Cryptographically Valid. The Action Didn't Match the Judgment.

1
Comments 1
11 min read
Your Agent Reads Untrusted Text All Day. Here's How I Grade What It Does With It.

Your Agent Reads Untrusted Text All Day. Here's How I Grade What It Does With It.

Comments
6 min read
Pick Your LLM With a Scoreboard, Not a Hunch: A Two-File Eval That Costs Nothing to Re-Run

Pick Your LLM With a Scoreboard, Not a Hunch: A Two-File Eval That Costs Nothing to Re-Run

Comments
6 min read
A Two-Hour Fit Test for AI Coding Models on Your Own Codebase

A Two-Hour Fit Test for AI Coding Models on Your Own Codebase

Comments
4 min read
I Installed 300 MCP Servers From PyPI. About 43% Don't Start.

I Installed 300 MCP Servers From PyPI. About 43% Don't Start.

Comments
4 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.