DEV Community

Cover image for Your Test Suite Is Green. That Doesn’t Mean You’re Ready to Ship.
David Frei
David Frei

Posted on

Your Test Suite Is Green. That Doesn’t Mean You’re Ready to Ship.

There’s a moment in software development that feels unusually satisfying.

You push the branch.

CI starts.

The little indicators begin turning green.

Unit Tests ............ PASSED
API Tests ............. PASSED
Browser Tests ......... PASSED
Accessibility ......... PASSED
Enter fullscreen mode Exit fullscreen mode

And somebody says:

Looks good. Ship it.

Maybe.

But I think we’ve trained ourselves to put too much emotional weight on the color green.

A green test suite tells you something useful.

It tells you that a collection of checks passed under a specific set of conditions at a specific moment in time.

That’s not the same thing as saying:

This release is safe.

And it definitely doesn’t mean:

If something breaks, we’ll know why.

Those are different problems.

Passing is only one part of testing

A mature testing process really has to answer at least four questions:

Did it pass?
If it failed, why?
Can somebody reproduce the failure?
Do we have enough information to make a release decision?
Enter fullscreen mode Exit fullscreen mode

A surprising number of testing systems are very good at the first one.

And mediocre at the other three.

That’s where things get expensive.

Because the cost of a failed automated test isn’t the 15 seconds it took to execute.

It’s the 45 minutes somebody spends figuring out whether the failure is:

  • a product bug
  • a test bug
  • an environment issue
  • bad test data
  • a deployment difference
  • a third-party problem
  • a timing problem
  • or something that disappeared when they clicked Retry

A useful benchmark for testing platforms should therefore measure more than execution speed.

This benchmark plan for browser testing platforms focuses on failure reproducibility, artifact retention, and release-handoff friction.

I like that framing.

Because a failure that nobody can investigate is barely better than not running the test.

Reproducibility is an underrated feature

Imagine this result:

Checkout test failed.
Step 17.
Element not found.
Enter fullscreen mode Exit fullscreen mode

Cool.

What now?

Compare that with a result containing:

  • the failed step
  • browser and OS
  • screenshot
  • video
  • page source
  • browser console
  • network information
  • timestamp
  • input data
  • previous actions
  • application version
  • environment
  • feature flags

The second failure is obviously more useful.

But we rarely compare testing tools based on this.

Instead we ask whether they have AI.

Or whether test creation takes three minutes instead of seven.

Those things matter.

But if your team runs 30,000 tests a month, failure investigation eventually becomes a much bigger operational cost than test creation.

Your staging environment probably isn't production

One of the easiest traps in QA is believing two environments are equivalent because they look equivalent.

They’re both running version 8.14.2.

They use the same database schema.

They show the same homepage.

So they must be the same.

Right?

Not necessarily.

You could still have differences in:

HTTP headers
Feature flags
CDN behavior
CSP rules
Cookies
Third-party scripts
Environment variables
Caching
Geography
API credentials
Load order
Enter fullscreen mode Exit fullscreen mode

And one tiny difference can completely change behavior.

Imagine a third-party analytics script loading before your authentication code in production but after it in staging.

Or a feature flag enabled for 20% of users.

Or an HTTP security header that only exists behind the production CDN.

Your test isn't necessarily wrong.

Your environments are different.

There’s a useful guide on debugging environment parity drift by comparing headers, feature flags, and third-party script load order.

The article is framed around outsourced QA, but the problem applies to almost every engineering team.

"Works in staging" has probably caused more production incidents than anyone wants to admit.

Test data is part of the environment too

There’s another thing teams casually treat as infrastructure:

Test data.

Suppose your automated test needs:

User: active customer
Plan: Pro
Subscription: valid
Invoices: 2
Feature flag: enabled
Team size: 4
Enter fullscreen mode Exit fullscreen mode

Where does that state come from?

Maybe somebody created the account six months ago.

Maybe a seed script creates it.

Maybe the test itself creates it.

Maybe someone manually edited the database.

Maybe nobody remembers.

Now your staging environment gets recreated.

And 180 tests fail.

This isn't really a test automation problem.

It’s a contract problem.

Your tests depend on data without clearly defining who creates it, what guarantees exist, and how it resets.

This guide on building a reliable test data contract for browser automation when environments are recreated frequently gets into that issue.

I think "test data contract" is a useful way to think about this.

Your test should know what data it can rely on.

And the environment should know what data it's responsible for providing.

Otherwise everybody is depending on mysterious accounts like:

qa-user-final-v2-dont-delete@example.com
Enter fullscreen mode Exit fullscreen mode

Which is usually a sign that civilization has failed.

Traceability sounds boring until you need it

Then you get into test management.

Test management software isn't exactly the most exciting category in software.

Nobody wakes up and says:

I hope I get to evaluate test management platforms today.

But once a product gets large enough, a few very boring questions become important.

For a particular release:

Which requirements were tested?
Which tests cover them?
Which tests actually ran?
What failed?
Who reviewed the failure?
Was it accepted?
What changed?
Enter fullscreen mode Exit fullscreen mode

If answering those questions requires five spreadsheets and a Slack archaeologist, your process probably isn't scaling very well.

This overview of test management platforms for automation traceability, lightweight administration, and clean release reporting covers some of the tradeoffs.

Notice the phrase lightweight admin.

That's important.

A test management system that requires more management than the tests themselves isn't helping much.

You want traceability without turning QA into accounting.

API testing has the same "what are we actually checking?" problem

API tests are another category where people group very different things under one label.

Consider these two checks.

Check A

GET /health

Expected:
HTTP 200
Enter fullscreen mode Exit fullscreen mode

Check B

POST /orders

Validate:
schema
field types
business rules
contract
error responses
authorization
edge cases
Enter fullscreen mode Exit fullscreen mode

They're both API tests.

But they solve very different problems.

One is essentially:

Is the thing alive?

The other is:

Does this interface still obey the contract that other software depends on?

That's why tool comparisons can be misleading when they don't start with the job you're actually trying to accomplish.

The comparison of Assertible vs Karate frames the distinction around CI health checks versus contract-style API validation.

That's a much better question than:

Which API testing tool is best?

Best for what?

Monitoring an endpoint every five minutes and maintaining a complex executable API specification are not the same job.

Responsive testing has become much harder than screenshots

Here's another area where testing tools can create a false sense of confidence.

Responsive visual testing.

At first glance this sounds simple:

Desktop screenshot
Tablet screenshot
Mobile screenshot
Enter fullscreen mode Exit fullscreen mode

Compare images.

Done.

Except modern frontends aren't just changing width anymore.

You have:

  • design tokens
  • variable fonts
  • container queries
  • responsive typography
  • dynamic spacing
  • device pixel ratios
  • browser rendering differences
  • lazy-loaded assets
  • animations
  • hydration
  • sticky components

Two screenshots can represent the same intended layout while differing slightly at the pixel level.

And two layouts can look superficially similar while violating important design rules.

This benchmark plan for measuring responsive screenshot fidelity across browser testing platforms is an interesting way to evaluate the problem.

Especially for applications with mature design systems.

Because the useful question isn't:

Can the tool take screenshots?

Everything can take screenshots.

The question is:

Can it tell me when a visual difference actually matters?

AI testing needs good failure handling even more than traditional automation

Now throw AI into this.

A testing agent creates a test.

It runs.

Something fails.

The AI attempts to explain it.

Maybe it reruns the step.

Maybe it repairs the locator.

Maybe it modifies the test.

This is where I'd start asking slightly uncomfortable questions.

For example:

Can I disable automatic reruns?

Can I see the first failure?

Can I inspect what the AI changed?

Can I compare original and repaired steps?

Can I export the evidence?

Can I hand the result to a developer?

Can I reproduce the exact execution?
Enter fullscreen mode Exit fullscreen mode

AI-generated tests don't make traditional debugging requirements disappear.

If anything, they make them more important.

Because now another system is making decisions between the test definition and the final result.

There’s a useful breakdown of rerun controls, failure explanations, and release handoffs when evaluating AI testing platforms.

I'd put a lot of weight on these capabilities.

An AI that can generate 1,000 tests isn't very useful if every failure turns into:

Something may have changed. Try again.

Managed testing creates a different optimization problem

AI testing is also blurring into managed QA.

There are platforms where your own team primarily owns the tests.

And services where much of the maintenance and review happens outside your team.

Again, neither model is inherently better.

They're optimizing different constraints.

A team with two developers and no QA engineers may value:

less maintenance
less triage
less internal QA work
Enter fullscreen mode Exit fullscreen mode

A larger engineering organization may care more about:

control
custom workflows
fast iteration
internal ownership
deep integrations
Enter fullscreen mode Exit fullscreen mode

That's why something like the Autify vs QA Wolf comparison is more useful when you look at the operating model, not just the feature checklist.

The same distinction appears between different managed testing services.

For example, this comparison of QA Wolf vs Testlio looks at faster release evidence versus broader crowd coverage.

Those aren't necessarily competing goals.

But depending on your product, one could matter much more than the other.

A B2B SaaS team releasing three times a day has very different testing needs from a consumer application trying to validate behavior across hundreds of devices, geographies, and user profiles.

Accessibility testing has exactly the same problem

Tool comparisons often fail because they compare products that solve different depths of the same problem.

Accessibility testing is a perfect example.

One team wants:

Tell me quickly if this page has obvious WCAG problems.

Another wants:

Help developers deeply investigate accessibility issues inside their normal engineering workflow.

Those aren't identical jobs.

This comparison of WAVE vs Accessibility Insights illustrates the distinction between fast triage and deeper developer-oriented review.

And, of course, no automated accessibility checker can completely replace human accessibility testing.

But automation can still catch a lot of obvious regressions very cheaply.

The key is knowing what you're asking the tool to do.

The most useful testing metric might not be pass rate

Here's a metric I think would be interesting for more teams to track:

Mean Time to Understand Failure
Enter fullscreen mode Exit fullscreen mode

Not:

Mean Time to Retry Failure
Enter fullscreen mode Exit fullscreen mode

Actually understand it.

Imagine two testing platforms.

Platform A

10,000 tests run in 30 minutes.

Failures take an average of 22 minutes to investigate.

Platform B

10,000 tests run in 40 minutes.

Failures take an average of 4 minutes to investigate.

Which one is faster?

The dashboard will tell you Platform A.

Your engineering team might tell you Platform B.

Execution time is visible and easy to benchmark.

Investigation time is hidden inside salaries.

That doesn't make it free.

Here's what I'd test before buying a testing platform

Don't give the vendor their favorite demo scenario.

Give them your ugliest workflow.

Then deliberately break it.

For example:

1. Authenticate as a restricted user
2. Enable a feature flag
3. Create unique test data
4. Open the application at mobile width
5. Trigger an API request
6. Cause a visual regression
7. Introduce a backend error
8. Fail an assertion
Enter fullscreen mode Exit fullscreen mode

Then stop looking at the execution.

Look at the aftermath.

Ask:

Can I tell exactly where it failed?

Can I see the relevant application state?

Can somebody else reproduce it?

Can I identify the environment?

Can I identify the data?

Can I export the evidence?

Can I connect it to the release?

Can I hand it to an engineer without explaining everything myself?
Enter fullscreen mode Exit fullscreen mode

That's where testing platforms start looking very different.

Testing isn't primarily about making things green

A green build is satisfying.

And obviously I'd rather have passing tests than failing ones.

But the real purpose of automated testing isn't to generate green rectangles.

It's to reduce uncertainty.

Before a release, you have questions.

Does checkout still work?
Did permissions break?
Did this API contract change?
Is mobile layout okay?
Did the new feature affect existing users?
Enter fullscreen mode Exit fullscreen mode

Good testing gives you evidence.

When something fails, good testing gives you even more evidence.

And eventually that evidence lets someone make a decision:

SHIP
Enter fullscreen mode Exit fullscreen mode

or

DON'T SHIP
Enter fullscreen mode Exit fullscreen mode

That decision is the actual output of the testing process.

Everything else — frameworks, AI, dashboards, screenshots, test management systems, browser clouds — exists to make that decision faster and less risky.

So yes.

Celebrate when CI turns green.

Just don't confuse the color of the dashboard with certainty.

Top comments (0)