DEV Community

Testing

Find those bugs before your users do! 🐛

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Contract Testing a Nutrition API with Millions of Messy Records

Contract Testing a Nutrition API with Millions of Messy Records

5
Comments
4 min read
A benchmark is only as good as the model you use to grade it

Costs twenty-one cents and exposes grading bias

A benchmark is only as good as the model you use to grade it

2
Comments 4
9 min read
My Release Gate Passed. The Model It Shipped Answered 'Neutral' To Everything.

My Release Gate Passed. The Model It Shipped Answered 'Neutral' To Everything.

2
Comments 2
6 min read
Our benchmark was wrong, and a stranger's blog post is why we found out

Our benchmark was wrong, and a stranger's blog post is why we found out

3
Comments
2 min read
Make Copilot Review Configuration Observable From Pull Request to Finding

Make Copilot Review Configuration Observable From Pull Request to Finding

Comments
1 min read
7 Checks Before You Trust an LLM Planner Experiment

7 Checks Before You Trust an LLM Planner Experiment

9
Comments 11
7 min read
A Unified ChatGPT Search Still Needs a Predictable Keyboard Result List

A Unified ChatGPT Search Still Needs a Predictable Keyboard Result List

Comments
1 min read
Test the Copilot Desktop App With a 20-Minute BYOK Exit Drill

Test the Copilot Desktop App With a 20-Minute BYOK Exit Drill

Comments
1 min read
Treat Copilot Code-Review Instructions as Untrusted Policy Input

Treat Copilot Code-Review Instructions as Untrusted Policy Input

Comments
1 min read
Green passed. The fix granted zero seats.

Green passed. The fix granted zero seats.

Comments 1
5 min read
I Let an AI Write My Tests for 30 Days: Coverage Went 38% to 71%

I Let an AI Write My Tests for 30 Days: Coverage Went 38% to 71%

Comments 1
3 min read
The Reliability Math Behind a Green n8n Workflow: Multiply, Don't Average

The Reliability Math Behind a Green n8n Workflow: Multiply, Don't Average

Comments
3 min read
OpenAI-compatible API first-call smoke test before you scale a workflow

OpenAI-compatible API first-call smoke test before you scale a workflow

Comments
2 min read
Stop Counting Devices: Measure Whether an Android Multi-Phone Desk Saves Time

Stop Counting Devices: Measure Whether an Android Multi-Phone Desk Saves Time

Comments
3 min read
I swapped my agent to a 3x smaller model and diffed what actually changed

I swapped my agent to a 3x smaller model and diffed what actually changed

Comments
4 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.