DEV Community

Testing

Find those bugs before your users do! 🐛

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Free Coding Models Are Good Enough for Some of Your Tasks — Here's How to Find Which Ones

Free Coding Models Are Good Enough for Some of Your Tasks — Here's How to Find Which Ones

Comments
4 min read
I Thought This Was a Classification Problem. It Wasn't.

I Thought This Was a Classification Problem. It Wasn't.

11
Comments
6 min read
The deliverable isn't the prompt. It's the eval.

The deliverable isn't the prompt. It's the eval.

Comments
7 min read
The AI Exam Author Was Never Wrong. I Still Can't Use Its Exam.

The AI Exam Author Was Never Wrong. I Still Can't Use Its Exam.

3
Comments 2
4 min read
I Rewrote One Exam Question Fifty Ways. First, the Answer Key Was Wrong.

I Rewrote One Exam Question Fifty Ways. First, the Answer Key Was Wrong.

3
Comments
6 min read
Stop Vibe-Checking Models: A Repeatable Comparison Harness You Can Run on Free Compute

Stop Vibe-Checking Models: A Repeatable Comparison Harness You Can Run on Free Compute

Comments 1
6 min read
Why Your Playwright Suite Gets Slower Every Sprint (And the Fix Nobody Talks About)

Why Your Playwright Suite Gets Slower Every Sprint (And the Fix Nobody Talks About)

2
Comments
4 min read
Why OTP Verification Fails (and How to Fix It)

Why OTP Verification Fails (and How to Fix It)

Comments
5 min read
Catch naming drift before an Abaqus run: a deterministic contract-audit pattern in Python

Catch naming drift before an Abaqus run: a deterministic contract-audit pattern in Python

Comments 1
3 min read
13 AI Coding Models Tested: Safety Benchmark Results KDS

13 AI Coding Models Tested: Safety Benchmark Results KDS

Comments
3 min read
The report says verified. Is the change safe to release?

The report says verified. Is the change safe to release?

Comments
5 min read
My Self-Improving Agent Still Couldn't Improve. That Was the Breakthrough.

My Self-Improving Agent Still Couldn't Improve. That Was the Breakthrough.

10
Comments 3
7 min read
I shipped a lock. Then I had to prove it could fail.

Summer Bug Smash: Smash Stories 🐛🛹

I shipped a lock. Then I had to prove it could fail.

1
Comments
5 min read
When should Codex use multiple agents? A benchmark, not a slogan

When should Codex use multiple agents? A benchmark, not a slogan

Comments
5 min read
Run the Cheap Review First

Run the Cheap Review First

1
Comments
4 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.