DEV Community

Cover image for The Boring Part of Test Automation That Matters More Than AI
Markus Gasser
Markus Gasser

Posted on AI-assisted

The Boring Part of Test Automation That Matters More Than AI

Every few months, test automation gets a new shiny object.

First it was Selenium.

Then Cypress.

Then Playwright.

Now everything has AI attached to it.

AI-generated tests. AI self-healing. AI agents clicking around your application. AI deciding what should be tested. AI deciding whether a test actually failed.

And some of this stuff is genuinely useful.

But after working with test automation for long enough, I've become increasingly convinced that one of the biggest problems isn't creating tests.

It's answering a much less exciting question:

When a test fails, can another human figure out what actually happened?

That sounds boring.

It is boring.

It's also where a surprising amount of engineering time disappears.

A failed test is not the same thing as useful information

Imagine your CI pipeline tells you:

Checkout test failed.

Element not found:
#place-order-button
Enter fullscreen mode Exit fullscreen mode

Okay.

Now what?

Was the button missing?

Was the checkout page still loading?

Was the user redirected?

Did an experiment change the DOM?

Did a cookie banner cover the button?

Did the test end up inside the wrong iframe?

Did the application return a 500?

Did the browser resize and trigger an entirely different layout?

Was there a JavaScript error?

Or did the test simply click too early?

The red FAILED label is usually the least interesting part of the failure.

The interesting part is the evidence around it.

Test automation has an evidence problem

A good failure should ideally give you enough context that you don't need to immediately rerun the test locally.

At minimum, I want some combination of:

  • the exact failed step
  • screenshots
  • browser console output
  • network information
  • page source or DOM state
  • browser and operating system
  • timestamps
  • previous steps
  • test data
  • video when useful
  • enough environment information to reproduce the run

This becomes even more important once the person investigating the failure isn't the person who created the test.

That's common in larger teams.

It's basically guaranteed if you're outsourcing QA.

There's a good breakdown of this problem in The Evidence Standard That Makes Outsourced QA Bugs Reproducible.

The central idea is simple: a bug report shouldn't require a Zoom call to become useful.

The same should be true for an automated test failure.

AI actually makes this problem more important

Here's the funny part.

As AI makes test creation cheaper, we're probably going to create more tests.

Which means:

More executions.

More failures.

More flaky failures.

More unusual edge cases.

And therefore more things somebody has to investigate.

Generating 500 tests sounds impressive until 37 of them fail overnight and your engineering team spends Tuesday morning figuring out which four failures are real.

This is why I'd pay attention to how AI testing platforms handle evidence and review workflows, not just how impressive their test-generation demo looks.

This benchmark plan for AI testing platforms takes an interesting approach by looking at audit-pack completeness, evidence exports, and whether another person can actually reproduce what happened.

That's a much more useful benchmark than:

"We asked the AI to test a todo app and it found three bugs."

Modern web apps are making failures weirder

Web applications aren't getting simpler either.

Consider something as ordinary as an embedded payment form.

It may be:

Your application
    ↓
iframe
    ↓
cross-origin payment widget
    ↓
dynamic authentication flow
Enter fullscreen mode Exit fullscreen mode

When automation fails somewhere inside that stack, "element not found" tells you almost nothing.

Cross-origin frames and embedded widgets deserve their own testing strategy. There's a detailed browser testing benchmark focused specifically on cross-origin iframes, embedded widgets, and frame-switch failure evidence.

These are exactly the kinds of scenarios I'd include when evaluating a testing platform.

Not another login form.

Layout is becoming state

Here's another category of failures that can waste hours.

A test works perfectly.

Then someone introduces:

  • ResizeObserver
  • container queries
  • hydration
  • responsive rendering
  • lazy-loaded components

And suddenly the button that was clickable 100 milliseconds ago is somewhere else.

Or the DOM node gets recreated.

Or a component appears before hydration finishes and then disappears.

If you've encountered tests that mysteriously fail after responsive or hydration changes, this guide on debugging browser tests after ResizeObserver, container queries, or hydration reflow is worth reading.

The broader lesson is that screenshots alone aren't always sufficient.

You need enough evidence to reconstruct how the page got into that state.

The edge cases are where automation platforms separate themselves

Almost every browser automation product can:

  1. open Google
  2. type something
  3. click a button

That's no longer an interesting benchmark.

I want to know what happens when the browser is doing something annoying.

For example, what happens under bandwidth constraints?

Testing prefers-reduced-data and slow connections introduces a completely different set of timing problems. This guide covers testing slow-network and bandwidth-sensitive interfaces without turning the suite into a collection of timing flakes.

Or consider browser storage.

What happens when this happens?

localStorage.setItem("data", veryLargeValue);
Enter fullscreen mode Exit fullscreen mode

And the browser replies:

QuotaExceededError
Enter fullscreen mode Exit fullscreen mode

Does your application recover gracefully?

Does your automated test even know what happened?

There's a useful walkthrough on testing QuotaExceededError, browser storage limits, and graceful fallbacks.

These aren't exotic theoretical scenarios.

They're the exact cases where "works on my machine" starts becoming expensive.

Infinite lists are another great example

Virtualized lists are fantastic for application performance.

They're also wonderfully annoying to automate.

The test may believe it saw 20 rows.

The user thinks there are 2,000.

Then the framework recycles those 20 DOM elements while you scroll.

Now you're testing a UI where:

DOM element ≠ permanent record
Enter fullscreen mode Exit fullscreen mode

That changes how you should think about assertions, scrolling, selectors, and screenshots.

There's a solid guide to testing virtualized infinite lists without missing recycled rows, loading gaps, or offscreen click bugs.

Again, these are the tests I'd run before choosing a tool.

Not whether its landing page says "AI-powered."

And then there's state

State is responsible for an absurd number of test failures.

Cookies.

Local storage.

Accounts.

Feature flags.

Experiments.

Cached resources.

Previous executions.

A test that leaves state behind isn't really an independent test anymore.

Feature flags make this especially interesting.

Suppose you test:

Flag OFF → old checkout
Flag ON  → new checkout
Flag OFF → old checkout again
Enter fullscreen mode Exit fullscreen mode

That third state matters.

You aren't just verifying that a feature flag works.

You're verifying that it rolls back cleanly.

This walkthrough on testing feature-flag rollbacks without carrying state between browser automation runs covers the issue in more depth.

It's a small detail until your release process depends on rolling a bad feature back at 2:00 AM.

Then it stops being a small detail.

Test results should help you make a decision

This is ultimately where I think QA tooling gets evaluated incorrectly.

People compare:

  • recorder quality
  • AI generation
  • number of integrations
  • supported browsers
  • pricing
  • whether there's a nice dashboard

Those matter.

But your testing system ultimately exists to answer:

Can we ship this?

And when the answer is "maybe," it needs to show you why.

A useful failure should be able to move through something like:

Automated test
      ↓
Failure evidence
      ↓
Developer / QA review
      ↓
Jira / GitHub / Slack
      ↓
Decision
Enter fullscreen mode Exit fullscreen mode

If information gets lost between those steps, the automation isn't saving as much time as you think.

There's a useful overview of what to look for in QA platforms used for release sign-off, Jira escalations, and Slack-friendly triage.

I like this framing because it evaluates testing software as part of an engineering process rather than as an isolated test runner.

Autonomy vs. control is another real tradeoff

AI testing products are also starting to split into different philosophies.

Some aim for:

Give us your application and we'll figure everything out.

Others are closer to:

We'll automate a lot, but you still control the scenarios, environments, and browser coverage.

Neither approach is automatically better.

They're solving slightly different problems.

For example, this QA.tech vs Octomind comparison frames the decision around autonomy versus control over browser coverage.

That's the kind of tradeoff teams should actually discuss.

If you're running a small SaaS with three engineers, maximum autonomy may be exactly what you want.

If you're testing a mature application with payments, iframes, multiple tabs, complicated state, permission levels, and several supported browsers, control may matter a lot more.

My boring test for testing tools

So here's the evaluation I'd use.

Take a genuinely ugly flow.

Something with:

  • an iframe
  • multiple browser contexts
  • dynamic content
  • an API call
  • a responsive layout change
  • persisted state
  • a deliberately introduced failure

Then run it.

Don't watch the successful execution.

Break it.

Now hand the failure report to someone who didn't create the test.

And ask:

Can you tell me what happened?

No screen sharing.

No rerunning it five times.

No asking the person who wrote the test.

Just the evidence.

If they can understand the failure and reproduce it, you've got something useful.

If they can't, it doesn't really matter how quickly the AI created the test.

Because test creation is only the first few minutes of a test's life.

You'll be debugging and maintaining it for the next three years.

And that's the boring part nobody puts in the demo.

Top comments (0)