DEV Community

Testing

Find those bugs before your users do! 🐛

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Python Selenium Architecture

Python Selenium Architecture

Comments
2 min read
Judge New Models With the Bugs That Already Burned You

Judge New Models With the Bugs That Already Burned You

Comments
7 min read
I Built a Personal Regression Suite for LLMs — Here's the Design, Not Just the Code

I Built a Personal Regression Suite for LLMs — Here's the Design, Not Just the Code

Comments
7 min read
Selenium and Python

Selenium and Python

Comments
2 min read
I Stopped Reading Model Release Threads and Built a Release-Day Eval Ritual Instead

I Stopped Reading Model Release Threads and Built a Release-Day Eval Ritual Instead

Comments
5 min read
Sanitizer Verdicts Over Demo Gifs: A Spec-Driven C++ Test Rig for Freshly Released Coding Models

Sanitizer Verdicts Over Demo Gifs: A Spec-Driven C++ Test Rig for Freshly Released Coding Models

Comments
7 min read
Treat Every New Open Model Like a Dependency Upgrade: A Pre-Flight Gate

Treat Every New Open Model Like a Dependency Upgrade: A Pre-Flight Gate

Comments
5 min read
Four Sanity Patterns for Distributed-System Testing

Four Sanity Patterns for Distributed-System Testing

Comments
9 min read
A driver that quietly does nothing is worse than one that isn't there

A driver that quietly does nothing is worse than one that isn't there

Comments
8 min read
Your Agent's Sandbox Is a Hypothesis. Here's How I Test Mine

Your Agent's Sandbox Is a Hypothesis. Here's How I Test Mine

Comments
5 min read
The Release-Day Reality Check: A Small Model Evaluation You Can Rerun

The Release-Day Reality Check: A Small Model Evaluation You Can Rerun

5
Comments 1
5 min read
How to test your LLM app for prompt injection: promptfoo vs garak vs Giskard vs PyRIT vs sentinel-scan-cli

How to test your LLM app for prompt injection: promptfoo vs garak vs Giskard vs PyRIT vs sentinel-scan-cli

1
Comments
5 min read
LlamaIndex makes RAG easy to build and hard to debug. Here is how I evaluate it.

LlamaIndex makes RAG easy to build and hard to debug. Here is how I evaluate it.

2
Comments
6 min read
The Gate That Stayed Silent — When a Blocker Count That Drops Reads as Improvement

The Gate That Stayed Silent — When a Blocker Count That Drops Reads as Improvement

12
Comments 5
5 min read
Canary a Free-Tier Model Promotion With Rate Limits, Truncation, and a Shadow Gate

Canary a Free-Tier Model Promotion With Rate Limits, Truncation, and a Shadow Gate

1
Comments
6 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.