DEV Community

Testing

Find those bugs before your users do! 🐛

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Three Silent Bugs That Broke Our AI Evaluation Harness

Three Silent Bugs That Broke Our AI Evaluation Harness

Comments
4 min read
I built an answer key for eval suites: six models broken on purpose, exactly.

I built an answer key for eval suites: six models broken on purpose, exactly.

1
Comments
2 min read
How to Build a No-Code AI Test Automation Agent Using RAG + Playwright MCP

How to Build a No-Code AI Test Automation Agent Using RAG + Playwright MCP

Comments
10 min read
Unit Tests Passed. The Feature Never Ran. Three Times in One Session.

Unit Tests Passed. The Feature Never Ran. Three Times in One Session.

1
Comments
8 min read
Testing a Two-Chart Compatibility Engine Without Inventing Missing Data

Testing a Two-Chart Compatibility Engine Without Inventing Missing Data

Comments
6 min read
I Tested a Star Wars API With Playwright, and It Was More Fun Than I Expected

I Tested a Star Wars API With Playwright, and It Was More Fun Than I Expected

7
Comments
6 min read
The First Ten Minutes of Testing a Design on Mobile

The First Ten Minutes of Testing a Design on Mobile

Comments
5 min read
Cyclomatic Complexity Has a Blind Spot — Introducing Coverage Difficulty (CD) and Responsibility Load Factor (RLF)

Cyclomatic Complexity Has a Blind Spot — Introducing Coverage Difficulty (CD) and Responsibility Load Factor (RLF)

Comments
3 min read
How a Test Coverage Ratchet Finally Fixed the Codebase Everyone Was Afraid to Touch

How a Test Coverage Ratchet Finally Fixed the Codebase Everyone Was Afraid to Touch

1
Comments
7 min read
A Safe Outcome Can Hide a Failed Security Control

A Safe Outcome Can Hide a Failed Security Control

Comments
5 min read
Debugging a black box: 36 renders against Claude, and the part where my own data was wrong

Debugging a black box: 36 renders against Claude, and the part where my own data was wrong

1
Comments 1
6 min read
Four High-severity bugs were hiding behind a green test suite in a 7k-star library

Four High-severity bugs were hiding behind a green test suite in a 7k-star library

1
Comments
6 min read
When a Failed Agent Step Looks Finished

When a Failed Agent Step Looks Finished

Comments 1
5 min read
The Two-Hour Gatekeeper for a 'Cheap and Capable' Model Announcement

The Two-Hour Gatekeeper for a 'Cheap and Capable' Model Announcement

1
Comments
4 min read
Testing a data pipeline against the spreadsheets it replaced

Testing a data pipeline against the spreadsheets it replaced

Comments
7 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.