DEV Community

Testing

Find those bugs before your users do! 🐛

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Une IA peut-elle prouver qu’elle a raison ?

Une IA peut-elle prouver qu’elle a raison ?

Comments
4 min read
A Two-Model Regression Harness for Evaluating a New Low-Cost Model Release

A Two-Model Regression Harness for Evaluating a New Low-Cost Model Release

Comments
5 min read
The Free-Model Agreement Test for AI Code Generation

The Free-Model Agreement Test for AI Code Generation

Comments
6 min read
A No-Cost Harness for Comparing Free Coding-Agent Models and Runtimes

A No-Cost Harness for Comparing Free Coding-Agent Models and Runtimes

Comments
5 min read
The MiniMax H3 Hype Arrived Before the Numbers. A Small Team Built a Gate Instead.

The MiniMax H3 Hype Arrived Before the Numbers. A Small Team Built a Gate Instead.

Comments
5 min read
I Tried to Verify an AI Agent Benchmark. Here's the Bundle I Wish Everyone Shipped

I Tried to Verify an AI Agent Benchmark. Here's the Bundle I Wish Everyone Shipped

Comments
6 min read
Testing AI C++ Race Fixes: TSan Harness

Testing AI C++ Race Fixes: TSan Harness

Comments
6 min read
I Thought Role Separation Would Fix the Optimizer. It Didn't.

I Thought Role Separation Would Fix the Optimizer. It Didn't.

12
Comments 6
6 min read
The model said it read the report. It didn't.

The model said it read the report. It didn't.

2
Comments 1
8 min read
I tested whether a Bedrock guardrail blocks the right things. It blocked a math question.

I tested whether a Bedrock guardrail blocks the right things. It blocked a math question.

Comments
7 min read
How Developers Think About Software Testing in the AI Era

How Developers Think About Software Testing in the AI Era

Comments
20 min read
Gate the Patch, Not the Model: A Three-Check Loop for Untrusted Code Outputs

Gate the Patch, Not the Model: A Three-Check Loop for Untrusted Code Outputs

Comments 1
5 min read
We asked ten agents for a test that must go red. Five wrote one that could not.

We asked ten agents for a test that must go red. Five wrote one that could not.

Comments
8 min read
When a New Model Drops, Hype Is Not a Benchmark

When a New Model Drops, Hype Is Not a Benchmark

Comments
2 min read
From Bug Log to Free Server Regression Gate: 142 Bugs, 51 Tests

From Bug Log to Free Server Regression Gate: 142 Bugs, 51 Tests

Comments
5 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.