DEV Community

Cover image for What If Your Test Automation Strategy Is Optimizing the Wrong Thing?
Markus Gasser
Markus Gasser

Posted on

What If Your Test Automation Strategy Is Optimizing the Wrong Thing?

Most test automation strategies are built around a few familiar goals:

  • create more tests
  • run them faster
  • increase coverage
  • reduce manual QA
  • catch bugs before production

All reasonable.

But I think a lot of teams are optimizing the wrong part of the system.

Because once your test suite gets large enough, the expensive part isn't necessarily writing the test.

And it isn't always running the test.

It's this:

Something failed. Now how long does it take before a human actually understands why?

That number is usually invisible.

Nobody puts it on the dashboard.

Nobody brags about it in a product demo.

And yet it might be one of the biggest costs in your entire automation strategy.

Test execution is cheap. Confusion is expensive.

Imagine your CI runs 8,000 browser tests overnight.

Forty fail.

The dashboard looks bad, but the actual problem starts the next morning.

Somebody now has to answer:

Is this a real bug?

Is this a test bug?

Is this a browser issue?

Did a third-party script interfere?

Did the environment change?

Did the page load differently?

Did a consent banner appear?

Did a WebSocket reconnect at the wrong moment?

Did the AI "repair" something incorrectly?
Enter fullscreen mode Exit fullscreen mode

If each failure takes 20 minutes to understand, you've burned more than 13 engineering hours before anyone has fixed anything.

The actual test run may have taken 30 minutes.

So which number matters more?

We measure the wrong things because they're easier to measure

Test automation dashboards love clean numbers.

Execution time: 18m 42s
Pass rate: 98.7%
Coverage: 74%
Tests created this month: 312
Enter fullscreen mode Exit fullscreen mode

Those are easy to count.

But what about:

Average time to understand a failed test: 17 minutes
Enter fullscreen mode Exit fullscreen mode

Or:

Percentage of failures another engineer can reproduce
without talking to the test author: 41%
Enter fullscreen mode Exit fullscreen mode

Those metrics are much harder to measure.

They're also much closer to the real cost of maintaining automation.

There's an interesting benchmark plan for measuring how quickly teams can reconstruct a failed browser test from evidence packs, logs, and replay data.

I think that's a much more interesting benchmark than shaving another three seconds off a test run.

The browser is no longer just your application

This is where things get messy.

People still talk about browser tests as if the browser is loading your page.

It's not.

It's loading something more like:

Your application
+ analytics
+ tag manager
+ cookie consent
+ feature flags
+ chat widget
+ monitoring
+ payment scripts
+ A/B testing
+ whatever marketing added last Thursday
Enter fullscreen mode Exit fullscreen mode

Now imagine one of those scripts loads slowly.

Or not at all.

Or in a different order.

Or injects a banner.

Or moves the layout.

Or blocks a click.

Your test fails with:

Element not clickable
Enter fullscreen mode Exit fullscreen mode

And everybody starts looking at the selector.

The selector may be perfectly fine.

The button may simply have a consent banner sitting on top of it.

This guide on investigating browser test failures caused by third-party scripts, tag managers, and consent banners covers this category of failure in much more detail.

The important lesson is bigger than cookie banners.

A modern browser test is often testing an ecosystem, not just your frontend.

"Page loaded" is becoming a meaningless concept

Here's another assumption that quietly breaks automation.

The page finished loading.

So the page is ready.

Not necessarily.

A modern application might do this:

Navigation complete
    ↓
React hydration
    ↓
Feature flags loaded
    ↓
User state fetched
    ↓
WebSocket connected
    ↓
Live data arrives
    ↓
UI updates
Enter fullscreen mode Exit fullscreen mode

At which point is the application actually ready?

That depends on what you're testing.

If your test starts asserting against the UI between steps four and six, you can get failures that look random but are actually completely deterministic.

You just happened to observe the application in a valid intermediate state.

Real-time applications make this much worse

WebSocket-heavy applications are a good example.

Suppose your dashboard says:

Open incidents: 12
Enter fullscreen mode Exit fullscreen mode

Then a WebSocket event arrives.

Now it says:

Open incidents: 13
Enter fullscreen mode Exit fullscreen mode

Fine.

Now imagine this sequence:

Socket disconnects
Reconnect begins
Two events arrive
UI renders one
Another event arrives
React rerenders
Assertion runs
Enter fullscreen mode Exit fullscreen mode

You now have a race that may happen only occasionally.

The test fails.

Someone reruns it.

It passes.

"Flaky."

Done.

Except nothing was random.

Your test just didn't model the application lifecycle correctly.

There's a useful guide on testing WebSocket-driven UI updates, reconnect states, and live data races in Playwright and Selenium.

These are the kinds of scenarios I'd use to evaluate an automation stack.

Not whether it can fill in a login form.

More AI-generated tests may actually make this problem worse

This is the part I think a lot of teams are going to discover over the next couple of years.

AI makes test creation cheaper.

That's useful.

But if you make something cheaper, people usually do more of it.

So instead of:

500 tests
Enter fullscreen mode Exit fullscreen mode

you get:

5,000 tests
Enter fullscreen mode Exit fullscreen mode

Great.

Now what happens when 83 fail after a deploy?

The bottleneck moves.

You didn't remove the maintenance problem.

You moved it from:

How do we write enough tests?

to:

How do we understand what all these tests are telling us?

That's why platform capabilities around failure evidence, CI stability, ownership, and handoff become much more important as test volume grows.

This rubric on choosing a testing platform when you need stable CI gates, shareable failure evidence, and low-maintenance ownership looks at the problem from that angle.

That sounds much less exciting than "AI writes your tests."

But it's probably closer to the problem mature teams eventually have.

A passing test can be more dangerous than a failing one

This is another uncomfortable part of AI-assisted testing.

Let's say you have this step:

Click "Delete Account"
Enter fullscreen mode Exit fullscreen mode

The UI changes.

The locator breaks.

An AI system repairs the step by pointing it at:

Deactivate Account
Enter fullscreen mode Exit fullscreen mode

The test runs.

Everything passes.

Your dashboard is green.

Except the test is no longer testing the same thing.

A failed test tells you:

Something changed.

A wrongly repaired test tells you:

Nothing changed.

That's worse.

So if your platform supports automatic repair, I want answers to questions like:

What changed?
Why was this replacement selected?
What was the original locator?
What is the new one?
Can I review the repair?
Can I reject it?
Can I see the first failure before the rerun?
Enter fullscreen mode Exit fullscreen mode

If the answer is no, then "self-healing" might be reducing visible failures while increasing invisible risk.

That's not a trade I'd make blindly.

This is why "best testing tool" is usually the wrong question

People love asking:

What's the best test automation platform?

There usually isn't one.

Different teams are optimizing for different things.

One team may care about:

  • fast authoring
  • easy onboarding
  • low maintenance
  • fewer QA specialists

Another may care about:

  • governance
  • permissions
  • reporting
  • reusable components
  • centralized administration
  • broader QA workflows

That's why a comparison like mabl vs Katalon is more useful when it frames the question as faster authoring versus broader governance.

That's an actual tradeoff.

A score like this:

Tool A: 9.1
Tool B: 8.8
Enter fullscreen mode Exit fullscreen mode

is almost meaningless without knowing what the team cares about.

Software teams don't buy scores.

They buy operating models.

Feature matrices are becoming less useful too

Most modern testing platforms can claim some version of this:

AI ✓
Cross-browser ✓
CI/CD ✓
Assertions ✓
Reports ✓
Self-healing ✓
Enter fullscreen mode Exit fullscreen mode

Okay.

Now show me what happens when:

  • a third-party script blocks the page
  • the WebSocket disconnects
  • the UI hydrates slowly
  • an iframe changes origin
  • the browser opens another tab
  • test data changes
  • the locator repair picks the wrong element
  • the test fails after 37 successful steps

That tells me much more.

There are some broader comparisons trying to test these tools against more realistic scenarios. For example, this independent review of 10 AI-powered test automation tools in 2026 looks at practical flows including iframes, file uploads, multiple tabs, email validation, variables, API calls, and language changes.

That's the kind of evaluation I find more useful than another grid of checkmarks.

Here's the question I'd ask your QA team tomorrow

Not:

How many tests do we have?

Not:

What's our pass rate?

Not even:

How fast does CI run?

Ask this:

When a browser test fails, how long does it usually take us to know what actually happened?

And then ask:

Could someone who didn't write the test figure it out from the artifacts alone?

If the answer to the second question is no, your automation strategy has a scaling problem.

Because every failure now depends on tribal knowledge.

That's fine with 50 tests.

It's painful with 500.

It's expensive with 5,000.

And AI can get you from 500 to 5,000 surprisingly quickly.

A better automation benchmark

If I were evaluating a testing platform today, I'd build one deliberately nasty scenario.

Something like:

1. Start with an authenticated user
2. Trigger a consent banner
3. Load a third-party script
4. Connect to a WebSocket
5. Receive a live data update
6. Change network conditions
7. Disconnect the socket
8. Reconnect
9. Trigger a layout change
10. Intentionally cause a failure
Enter fullscreen mode Exit fullscreen mode

Then I'd stop looking at the test runner.

I'd give the output to someone else.

Someone who didn't create the test.

No walkthrough.

No hints.

No rerunning it five times.

And I'd ask:

What broke?

Then I'd start the clock.

That number might tell you more about the value of the platform than most of the metrics you're tracking today.

Maybe the goal isn't more automation

This is the part I think is worth questioning.

A lot of teams treat "more automation" as automatically better.

But automation isn't the goal.

The goal is reducing uncertainty.

A test that fails and creates 30 minutes of confusion isn't necessarily saving you much.

A test that fails and immediately explains:

The checkout button became inaccessible
after the consent layer loaded.
Here is the screenshot.
Here is the DOM.
Here are the console logs.
Here is the network activity.
Here is the exact sequence that reproduced it.
Enter fullscreen mode Exit fullscreen mode

is incredibly valuable.

So maybe your next test automation project shouldn't be:

How do we generate 2x more tests?

Maybe it should be:

How do we make every failure 5x easier to understand?

Because once test creation gets cheap enough, that's probably where the real bottleneck moves.

And if your current strategy doesn't account for that, you may be optimizing a number that stopped mattering a while ago.

Top comments (0)