DEV Community

Testing

Find those bugs before your users do! 🐛

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Parity Suite 3.0.0: measuring the wrong thing

Parity Suite 3.0.0: measuring the wrong thing

Comments
6 min read
Delegating to AI Means Governing the Environment

Delegating to AI Means Governing the Environment

Comments
12 min read
Free Model Endpoints Make Good Fuzz Generators and Bad Oracles

Free Model Endpoints Make Good Fuzz Generators and Bad Oracles

Comments
4 min read
A Three-Pass Hygiene Check for Free Coding Model Diffs

A Three-Pass Hygiene Check for Free Coding Model Diffs

Comments
3 min read
A Free Model Answer Is a Hypothesis. Give It an Error Budget Before It Reaches Main.

A Free Model Answer Is a Hypothesis. Give It an Error Budget Before It Reaches Main.

Comments
4 min read
Replay AI-Generated Migrations Against a Throwaway Database Before You Merge Them

Replay AI-Generated Migrations Against a Throwaway Database Before You Merge Them

Comments
5 min read
MiniMax H3 Is Making Noise. My First Check Isn’t the Leaderboard.

MiniMax H3 Is Making Noise. My First Check Isn’t the Leaderboard.

Comments
3 min read
A Free Replay Bench for Agent Tool-Call Failures Before You Deploy

A Free Replay Bench for Agent Tool-Call Failures Before You Deploy

Comments
3 min read
Don't Let a Free Model Rewrite Your Code. Let It Generate the Tests You Forgot.

Don't Let a Free Model Rewrite Your Code. Let It Generate the Tests You Forgot.

Comments
5 min read
How do you manage placeholder data without wasting time on manual entry?

How do you manage placeholder data without wasting time on manual entry?

Comments
2 min read
How Much Should We Trust AI-Generated Tests?

How Much Should We Trust AI-Generated Tests?

2
Comments
1 min read
We found a bug that let our test suite write to production. Here's what we did about it.

We found a bug that let our test suite write to production. Here's what we did about it.

Comments 2
2 min read
The monitoring bug that turned out not to be a bug

The monitoring bug that turned out not to be a bug

1
Comments
3 min read
One AI Answer, Eight Brands: Designing a Benchmark Without Multiplying the Evidence

One AI Answer, Eight Brands: Designing a Benchmark Without Multiplying the Evidence

Comments
6 min read
Every Week a New Model Is "Cheaper and Better" — Here's the 30-Minute Harness That Settles It

Every Week a New Model Is "Cheaper and Better" — Here's the 30-Minute Harness That Settles It

Comments
5 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.