DEV Community

Testing

Find those bugs before your users do! 🐛

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Which parts of this are real? Should be a question with an answer

Which parts of this are real? Should be a question with an answer

1
Comments
7 min read
I Stopped Trusting My Agent's Boundaries Until I Could Break Them in a Throwaway Sandbox

I Stopped Trusting My Agent's Boundaries Until I Could Break Them in a Throwaway Sandbox

Comments
5 min read
Shadow-Gate Your LLM-Generated SQL: A Replay Test Against a Frozen Fixture Database

Shadow-Gate Your LLM-Generated SQL: A Replay Test Against a Frozen Fixture Database

Comments
6 min read
Gate the Toolbelt: A Filesystem-First Smoke Test for Agentic Coding Models

Gate the Toolbelt: A Filesystem-First Smoke Test for Agentic Coding Models

Comments
5 min read
The Question Nobody Asks About Free Coding Models: How Many of Their Patches Break Something Else?

The Question Nobody Asks About Free Coding Models: How Many of Their Patches Break Something Else?

Comments
5 min read
A Reproducible Sandbox Loop for AI-Generated Code: Generate, Isolate, Assert

A Reproducible Sandbox Loop for AI-Generated Code: Generate, Isolate, Assert

Comments
4 min read
Silent Model Drift Will Break Your Prompts: A Weekly Drift Detector You Can Run for Free

Silent Model Drift Will Break Your Prompts: A Weekly Drift Detector You Can Run for Free

Comments
4 min read
Model Swaps Are Boundary Events: Gate Agent Tool Changes With a Deterministic Replay Lane

Model Swaps Are Boundary Events: Gate Agent Tool Changes With a Deterministic Replay Lane

Comments
6 min read
Building a Production AI Agent in Spring Boot: A/B Testing Prompts With an LLM Judge (Part 9)

Building a Production AI Agent in Spring Boot: A/B Testing Prompts With an LLM Judge (Part 9)

Comments
10 min read
Every New Model Gets the Same Six Questions From Me

Every New Model Gets the Same Six Questions From Me

Comments
5 min read
Judge New Models With the Bugs That Already Burned You

Judge New Models With the Bugs That Already Burned You

Comments
7 min read
Python Selenium Architecture

Python Selenium Architecture

Comments
2 min read
I Built a Personal Regression Suite for LLMs — Here's the Design, Not Just the Code

I Built a Personal Regression Suite for LLMs — Here's the Design, Not Just the Code

Comments
7 min read
Selenium and Python

Selenium and Python

Comments
2 min read
I Stopped Reading Model Release Threads and Built a Release-Day Eval Ritual Instead

I Stopped Reading Model Release Threads and Built a Release-Day Eval Ritual Instead

Comments
5 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.