DEV Community

Testing

Find those bugs before your users do! 🐛

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
I Tried to Verify an AI Agent Benchmark. Here's the Bundle I Wish Everyone Shipped

I Tried to Verify an AI Agent Benchmark. Here's the Bundle I Wish Everyone Shipped

Comments
6 min read
How Do You Actually Test an AI System? A Layered Strategy From Five Tools I Built

How Do You Actually Test an AI System? A Layered Strategy From Five Tools I Built

Comments 1
5 min read
AI Agents Are Great at Exploratory Testing. Regression Needs Repeatable Assets.

AI Agents Are Great at Exploratory Testing. Regression Needs Repeatable Assets.

Comments 2
6 min read
Testing AI C++ Race Fixes: TSan Harness

Testing AI C++ Race Fixes: TSan Harness

Comments
6 min read
Dozens of Resumes, One Call From an Old Colleague

Cold applications fail while networks win

Dozens of Resumes, One Call From an Old Colleague

94
Comments 78
6 min read
App Intents, AppFunctions, and the Interface Your Test Suite Has Never Driven

App Intents, AppFunctions, and the Interface Your Test Suite Has Never Driven

6
Comments 1
6 min read
I tested whether a Bedrock guardrail blocks the right things. It blocked a math question.

I tested whether a Bedrock guardrail blocks the right things. It blocked a math question.

Comments
7 min read
The model said it read the report. It didn't.

The model said it read the report. It didn't.

2
Comments 1
8 min read
How Developers Think About Software Testing in the AI Era

How Developers Think About Software Testing in the AI Era

Comments
20 min read
trelix v3.2.2 to v3.2.5: The Source Tree Was Fine. The Published Package Wasn't.

trelix v3.2.2 to v3.2.5: The Source Tree Was Fine. The Published Package Wasn't.

Comments 1
11 min read
Gate the Patch, Not the Model: A Three-Check Loop for Untrusted Code Outputs

Gate the Patch, Not the Model: A Three-Check Loop for Untrusted Code Outputs

Comments 1
5 min read
We asked ten agents for a test that must go red. Five wrote one that could not.

We asked ten agents for a test that must go red. Five wrote one that could not.

Comments
8 min read
When a New Model Drops, Hype Is Not a Benchmark

When a New Model Drops, Hype Is Not a Benchmark

Comments
2 min read
From Bug Log to Free Server Regression Gate: 142 Bugs, 51 Tests

From Bug Log to Free Server Regression Gate: 142 Bugs, 51 Tests

Comments
5 min read
A Free Model's Review Had No Evidence. The Maintainer Added a Claim Ledger.

A Free Model's Review Had No Evidence. The Maintainer Added a Claim Ledger.

Comments
4 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.