DEV Community

Testing

Find those bugs before your users do! 🐛

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Your test suite is lying to you about which process it's in

Your test suite is lying to you about which process it's in

Comments
5 min read
A failed import that still created three records

A failed import that still created three records

1
Comments 2
5 min read
Feature Flags and Testing: Because QA Deserves Surprises Too

Feature Flags and Testing: Because QA Deserves Surprises Too

Comments
3 min read
Flaky Tests: O Problema Não São os Testes, É a Confiança

Flaky Tests: O Problema Não São os Testes, É a Confiança

Comments
6 min read
What Happens When Your AI Feature Fails?

What Happens When Your AI Feature Fails?

Comments
10 min read
Four small security tools in a month, and the wall each one hit on purpose

Four small security tools in a month, and the wall each one hit on purpose

1
Comments 3
5 min read
25 Findings, One Mobile App: What a Real Pentest Report Taught Me About React Native Security

25 Findings, One Mobile App: What a Real Pentest Report Taught Me About React Native Security

Comments 1
4 min read
Case Study: A Free Model Wrote a C++ Tree Hasher. The Reference Oracle Found Three Bugs.

Case Study: A Free Model Wrote a C++ Tree Hasher. The Reference Oracle Found Three Bugs.

Comments
4 min read
You Benchmarked the Model. Now Benchmark the Server.

You Benchmarked the Model. Now Benchmark the Server.

Comments
5 min read
My Agent's Tests Were Green Because the Model Learned to Cheat

Comments explore fixing broken evaluators

My Agent's Tests Were Green Because the Model Learned to Cheat

24
Picked as gem Comments 22
5 min read
Build a Reproducible AI Agent Evaluation Lab with Docker Compose

Build a Reproducible AI Agent Evaluation Lab with Docker Compose

8
Comments 3
4 min read
Your AI Agent Needs a Cancellation Contract, Not Just a Stop Button

Your AI Agent Needs a Cancellation Contract, Not Just a Stop Button

1
Comments
4 min read
The same detector scores 45.5 or 100 on the OWASP Benchmark. Both are 'true.'

The same detector scores 45.5 or 100 on the OWASP Benchmark. Both are 'true.'

Comments
3 min read
Free Inference Should Not Gate a Merge

Free Inference Should Not Gate a Merge

Comments 2
7 min read
Playwright's fill() threw an error and still inserted the text: the CodeMirror 6 double-paste trap

Playwright's fill() threw an error and still inserted the text: the CodeMirror 6 double-paste trap

Comments
3 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.