DEV Community

Cover image for Most Flaky Tests Aren’t Flaky. Your Test Environment Is Lying to You.
Antoine Dubois
Antoine Dubois

Posted on

Most Flaky Tests Aren’t Flaky. Your Test Environment Is Lying to You.

We use the word flaky way too casually in test automation.

A test passes nine times.

Then it fails.

Someone reruns it.

It passes.

And the verdict is:

"Eh. Flaky test."

Ticket closed.

Or worse, somebody adds another two-second wait and commits it.

await page.waitForTimeout(2000);
Enter fullscreen mode Exit fullscreen mode

Excellent.

Software engineering.

The uncomfortable reality is that a lot of supposedly flaky tests aren't random at all.

They're deterministic failures caused by conditions we forgot to control.

Browser state.

Cache state.

Session state.

Network timing.

Webhooks.

Timezone.

Locale.

Scroll position.

Service workers.

And increasingly, AI-generated changes to tests that nobody has really inspected.

The browser knows exactly why the test failed.

We just didn't collect enough information to figure it out.

State is probably guilty until proven innocent

Here's a test automation debugging rule I've developed over the years:

If something passes locally, fails in CI, passes when retried, and nobody can explain why...

Look at state.

Not selectors.

State.

Browser automation is basically one giant experiment where we keep trying to pretend every execution starts from the same universe.

It usually doesn't.

Consider something as ordinary as authentication.

A browser context may contain:

  • cookies
  • local storage
  • session storage
  • IndexedDB
  • cached assets
  • service worker data
  • authentication tokens

Change one of those and you can get a completely different application.

This gets especially interesting with Playwright because browser contexts make isolation convenient enough that it's easy to assume isolation is perfect.

It isn't always.

This guide on debugging Playwright tests that fail after browser context storage changes covers exactly this class of problem.

And it's worth understanding because the failure usually doesn't look like a storage failure.

It looks like:

Expected element to be visible.
Element was not found.
Enter fullscreen mode Exit fullscreen mode

The selector wasn't necessarily wrong.

The user may simply have landed in a different application state.

Service workers make this even more fun

If you really want to spend an afternoon questioning your career choices, debug a cache-related failure involving a service worker.

You can have:

Browser
   ↓
Service Worker
   ↓
Cache
   ↓
Network
Enter fullscreen mode Exit fullscreen mode

Which means what your automation sees isn't necessarily what your server just deployed.

Maybe the new JavaScript bundle exists.

Maybe the browser is still using the old one.

Maybe the service worker has updated but isn't controlling the current page.

Maybe the cache invalidation worked on Chrome 1 but not Chrome 2.

Maybe your "offline" test isn't actually offline because something was already cached.

These are all legitimate application behaviors.

They're also easy to mislabel as automation flakiness.

There's a useful walkthrough on testing service worker updates, cache invalidation, and offline fallbacks without carrying flaky browser state between runs.

This is exactly the kind of scenario I'd include when evaluating a browser automation platform.

Not:

"Can it click the Login button?"

Everything can click the Login button.

Until a sticky banner covers it.

Speaking of things covering buttons...

Modern browser UIs are surprisingly hostile to automation.

Sticky headers.

Floating chat widgets.

Cookie banners.

Animated navigation.

Lazy-loaded content.

Infinite scroll.

Responsive layouts.

A perfectly valid element can exist in the DOM and still be impossible to click.

You might have:

Target element:
x = 420
y = 87
Enter fullscreen mode Exit fullscreen mode

Unfortunately, your sticky navigation bar occupies:

y = 0 → 96
Enter fullscreen mode Exit fullscreen mode

Your automation framework sees the button.

The user visually sees the button.

The click lands somewhere else.

Then somebody "fixes" the problem with:

await page.locator('#save').click({ force: true });
Enter fullscreen mode Exit fullscreen mode

Which is occasionally correct.

And occasionally just hides a real user-facing problem.

There's a practical guide to testing sticky headers, scroll jank, and element occlusion without creating flaky click assertions.

The important distinction is:

"The automation couldn't click it"

and

"A user couldn't reliably click it"

are not always the same bug.

Good testing infrastructure should help you distinguish between them.

Sleep is not synchronization

Another common source of "flakiness":

sleep(5000);
Enter fullscreen mode Exit fullscreen mode

This means:

"I have absolutely no idea when the thing will happen, but five seconds feels emotionally safe."

Sometimes five seconds is enough.

Sometimes the event takes six.

Congratulations.

You now have a flaky test.

Webhooks are a perfect example.

Imagine:

Browser action
     ↓
Backend request
     ↓
Payment provider
     ↓
Webhook
     ↓
Worker
     ↓
Database update
     ↓
UI refresh
Enter fullscreen mode Exit fullscreen mode

How long does that take?

The wrong answer is:

Probably three seconds.

The better answer is:

Wait for evidence that the expected event happened.

That could mean polling an API, waiting for a database condition through a test helper, checking an event endpoint, or verifying a known downstream effect.

There's a good breakdown of testing webhooks in QA without relying on fragile sleep-based assertions.

This sounds obvious when written out.

And yet I still regularly see mature automation suites containing:

Wait 5 seconds
Wait 10 seconds
Wait 3 seconds
Enter fullscreen mode Exit fullscreen mode

At some point you're not testing the application anymore.

You're testing whether your timing guesses were lucky.

Time itself is state

Here's another fun one.

A test passes in Bucharest.

It fails on a cloud browser in Virginia.

The page says:

October 3
Enter fullscreen mode Exit fullscreen mode

The assertion expects:

October 2
Enter fullscreen mode Exit fullscreen mode

Both can be correct.

Welcome to timezones.

International applications introduce several hidden inputs:

  • browser timezone
  • server timezone
  • operating system locale
  • browser language
  • daylight saving rules
  • currency formatting
  • decimal separators
  • 12-hour vs 24-hour time
  • date ordering

Even a simple price can become:

$1,234.56
€1.234,56
1 234,56 €
Enter fullscreen mode Exit fullscreen mode

depending on context.

Hardcoding rendered strings is a wonderful way to create tests that pass everywhere except where your users live.

This guide on testing timezone, locale, and currency rendering without flaky browser assertions gets into these problems in more detail.

Locale testing is one of those things teams postpone until they acquire international customers.

Then suddenly it's a release blocker.

Role-based applications multiply the problem

Now take everything above and add permissions.

Imagine an admin platform with:

Super Admin
Admin
Manager
Editor
Viewer
External Collaborator
Enter fullscreen mode Exit fullscreen mode

Each role may see a different navigation.

Different buttons.

Different API responses.

Different data.

Different redirects.

Different session rules.

And now you've multiplied your state space significantly.

Testing these systems isn't just about whether your automation tool supports Chrome.

It's about whether you can reliably create isolated sessions and understand failures when one role behaves differently from another.

This breakdown of browser testing platforms for role-based admin panels, session state, and debuggable failures highlights some useful criteria.

For complex business applications, I'd care a lot about this.

A demo recording a checkout flow tells you very little about how the platform behaves with six permission levels and stale sessions.

And after deployment, you're still testing

There's another mistake I see fairly often.

Teams think:

CI passed → deployment successful
Enter fullscreen mode Exit fullscreen mode

Those are different statements.

CI verifies one environment.

Your users interact with another.

After deployment you can still encounter:

  • bad environment variables
  • broken CDN assets
  • DNS issues
  • expired credentials
  • missing migrations
  • third-party outages
  • routing failures
  • region-specific problems

Which is why synthetic monitoring is basically test automation that refuses to go home after deployment.

A useful synthetic check should do more than scream:

CHECKOUT FAILED
Enter fullscreen mode Exit fullscreen mode

You want to know where it failed and whether rollback might actually help.

There's a useful overview of selecting synthetic monitoring for post-deploy smoke checks, rollback verification, and API failure triage.

For important flows, the lifecycle really looks more like:

Build
  ↓
Automated tests
  ↓
Deploy
  ↓
Synthetic checks
  ↓
Real users
Enter fullscreen mode Exit fullscreen mode

Your release doesn't become correct just because GitHub Actions turned green.

AI test repair adds an interesting economic problem

AI is now starting to "solve" flaky automation too.

A locator changes.

The AI repairs it.

The test keeps running.

Sounds great.

Sometimes it is.

But there's a question I think gets overlooked:

How much autonomous repair should happen before a human should review what's changing?

Imagine an agent repairs 100 tests.

Maybe 97 repairs are perfect.

Maybe three subtly changed what the test means.

For example:

Original:
Click "Delete User"

Repair:
Click "Disable User"
Enter fullscreen mode Exit fullscreen mode

Both controls might exist in roughly the same place.

The test could continue successfully.

And you've now created something much worse than a failed test:

A passing test that's testing the wrong thing.

That's why I like the framing in this benchmark plan comparing autonomous test repair cost against human review time.

The metric shouldn't simply be:

How many tests did the AI repair?

It should also include:

repair cost
+
review cost
+
incorrect repair risk
+
time saved
Enter fullscreen mode Exit fullscreen mode

That's a much more useful business equation.

The test itself is becoming an asset

There's another question teams don't ask early enough:

What happens if we leave this testing platform?

This mattered less when nearly everything was code.

A Selenium test was Python, Java, C#, JavaScript, whatever.

You owned it.

Now we're seeing more proprietary:

  • visual workflows
  • AI-generated models
  • natural-language tests
  • hosted execution formats
  • platform-specific steps
  • proprietary repair metadata

Some of these abstractions are valuable.

But abstraction has a price.

If you create 4,000 tests over three years, switching platforms becomes an entirely different conversation.

That's why measuring test artifact portability across browser clouds, open-source frameworks, and AI testing platforms is a worthwhile benchmark.

I'd want to know:

  • Can I export my tests?
  • In what format?
  • Can another system execute them?
  • Can humans understand them?
  • What happens to assertions?
  • What happens to screenshots and historical results?
  • What happens to AI-generated metadata?

Nobody worries about vendor lock-in when they have 20 tests.

They start worrying when they have 20,000.

Managed QA versus self-service AI is becoming another dividing line

We're also seeing two very different approaches emerge in AI testing.

One says:

Give us responsibility for the outcome.

The other says:

Here's powerful automation. You operate it.

Those are fundamentally different products even if both put "AI testing" on their homepage.

A managed service can remove maintenance work.

A self-service platform can give your team more direct control and faster iteration.

Neither model universally wins.

This comparison of QA Wolf vs QA.tech frames the difference around managed maintenance versus self-service browser coverage.

And if you're trying to understand the broader market, this list of the 10 best AI-powered test automation tools also gives a useful overview of how different AI testing products are approaching the problem.

Personally, I think category labels like "AI testing platform" are becoming less useful.

The more important questions are:

Who creates the tests?
Who maintains them?
Who investigates failures?
Who owns the artifacts?
Who decides when an AI repair is correct?
Enter fullscreen mode Exit fullscreen mode

Those answers tell you much more about a platform than whether it has a chatbot.

My favorite flaky-test debugging trick

When a test fails intermittently, don't immediately modify it.

First try to classify the hidden variable.

Ask:

What changed between the passing run
and the failing run?
Enter fullscreen mode Exit fullscreen mode

Check:

  1. Browser state
  2. Application state
  3. Network state
  4. Test data
  5. Browser version
  6. Viewport
  7. Locale and timezone
  8. Timing
  9. Third-party dependencies
  10. Deployment version

Then compare evidence from the runs.

The difference is usually somewhere in that list.

And if you don't have enough evidence to compare two executions?

That's the actual problem you should fix first.

Because adding:

await page.waitForTimeout(5000);
Enter fullscreen mode Exit fullscreen mode

might make your dashboard greener.

But it doesn't mean your test suite became more reliable.

It just means you taught the failure to hide better.

Flakiness is information

This is probably the biggest mental shift.

A flaky test isn't merely annoying noise.

It's information.

Sometimes it tells you your test is poorly synchronized.

Sometimes it tells you your application behaves unpredictably.

Sometimes it tells you environments aren't isolated.

Sometimes it tells you your infrastructure behaves differently under load.

Sometimes it tells you that you've accidentally discovered a real race condition.

So before deleting the test, increasing the timeout, or clicking Retry for the sixth time, ask:

What condition makes this failure possible?

Because computers aren't usually rolling dice behind your back.

Something changed.

The job of good test automation isn't just to detect the failure.

It's to leave enough evidence behind that you can figure out what changed.

Top comments (0)