DEV Community

Cover image for Why Your Playwright Tests Are Flaky (And Why Retries Won’t Save You)
David Frei
David Frei

Posted on

Why Your Playwright Tests Are Flaky (And Why Retries Won’t Save You)

Playwright has auto-waiting.

It has isolated browser contexts.

It has auto-retrying assertions.

It has traces.

It has screenshots.

It has videos.

It has one of the best debugging experiences of any browser automation framework I've used.

So why do mature Playwright suites still end up with this?

412 passed
7 flaky
3 failed
Enter fullscreen mode Exit fullscreen mode

Then:

Retrying...
Enter fullscreen mode Exit fullscreen mode

And five minutes later:

422 passed
Enter fullscreen mode Exit fullscreen mode

Green build.

Everyone moves on.

Until tomorrow.

The uncomfortable answer is that Playwright solved a lot of browser automation flakiness.

It did not solve application flakiness, data flakiness, environment flakiness, dependency flakiness, or bad assumptions in your test architecture.

And once your suite gets large enough, those become the real problem.

Auto-waiting does not mean "the application is ready"

Playwright's actionability checks are genuinely good.

Before a normal click, Playwright checks things like whether the element is visible, stable, enabled, and able to receive events.

That's a huge improvement over the old pattern of guessing how long to sleep.

The Playwright documentation on auto-waiting explains those checks in detail.

But there's a limitation hiding in plain sight.

Playwright understands the browser.

It does not understand your business logic.

Consider this sequence:

Dashboard renders
    ↓
"Save" button becomes enabled
    ↓
Background request finishes
    ↓
Application recalculates permissions
    ↓
React rerenders the form
    ↓
Final state is ready
Enter fullscreen mode Exit fullscreen mode

Playwright might correctly decide the button is clickable at step two.

Your workflow may not actually be safe until step six.

The element is ready.

The application isn't.

This is one of the biggest sources of flaky tests in modern frontend applications.

The wrong fix

await page.waitForTimeout(3000);
Enter fullscreen mode Exit fullscreen mode

This works right up until CI needs 3.2 seconds.

Then someone changes it to 5 seconds.

Six months later your test suite contains 140 seconds of accumulated superstition.

The better fix

Wait for the state that actually matters.

For example:

await expect(page.getByTestId('invoice-status')).toHaveText('Paid');
Enter fullscreen mode Exit fullscreen mode

or wait for the relevant response:

const response = page.waitForResponse(
  r => r.url().includes('/api/invoices/') && r.status() === 200
);

await page.getByRole('button', { name: 'Refresh' }).click();
await response;
Enter fullscreen mode Exit fullscreen mode

The key idea is simple:

Wait for meaning, not time.

Hydration creates a particularly nasty version of this problem

Modern frameworks can render something that looks ready before it really is.

The browser may show:

Submit
Enter fullscreen mode Exit fullscreen mode

The button exists.

It's visible.

Playwright can find it.

Then hydration finishes and the framework replaces the node.

Or event handlers attach slightly later.

Or data arrives and the component rerenders.

Now a test that looked perfectly reasonable fails once every 40 runs.

This gets worse with:

  • server-side rendering
  • React hydration
  • lazy components
  • Suspense
  • client-side routing
  • animation libraries
  • container queries
  • virtualized content

These failures often get mislabeled as "Playwright being flaky."

Usually Playwright is doing exactly what you told it to do.

The test just observed the application during a transitional state.

Shared state is still one of the fastest ways to destroy a test suite

Playwright creates isolated browser contexts for tests by default.

That's great.

Cookies, local storage, and session storage are isolated between those contexts. The Playwright isolation documentation is explicit about why this matters.

But your backend data isn't automatically isolated.

Imagine two parallel tests:

Test A:
Log in as qa@example.com
Delete all notifications
Create notification

Test B:
Log in as qa@example.com
Expect 3 notifications
Enter fullscreen mode Exit fullscreen mode

Each browser is isolated.

The account isn't.

Now increase the worker count.

Congratulations, you built a slot machine.

Common shared-state mistakes

I've seen all of these:

  • every test uses the same user
  • every test edits the same project
  • tests depend on "latest" records
  • tests assume a clean database
  • tests depend on execution order
  • teardown from one test deletes another test's data
  • parallel tests use the same uploaded filename
  • tests reuse the same shopping cart
  • email tests read from one shared inbox

The usual symptom is:

It passes when I run it alone.

That's one of the most useful clues you can get.

A stable test should own its data

Where practical, give the test unique state.

Instead of:

qa@example.com
Enter fullscreen mode Exit fullscreen mode

create:

qa+run-84217@example.com
Enter fullscreen mode Exit fullscreen mode

Instead of:

Project: Automation Test
Enter fullscreen mode Exit fullscreen mode

use:

Project: Automation Test 84217
Enter fullscreen mode Exit fullscreen mode

Instead of:

page.getByText('Latest Order')
Enter fullscreen mode Exit fullscreen mode

assert against the order ID created by the test.

You're trying to make this true:

Test outcome
=
application behavior
Enter fullscreen mode Exit fullscreen mode

rather than:

Test outcome
=
application behavior
+
whatever seven other workers happened to do
Enter fullscreen mode Exit fullscreen mode

Third-party dependencies quietly inherit your test reliability

Here's another trap.

Your signup flow works like this:

Create account
    ↓
Send verification email
    ↓
Email provider accepts message
    ↓
Message arrives
    ↓
Test opens link
    ↓
Account becomes verified
Enter fullscreen mode Exit fullscreen mode

Your application can be healthy.

Playwright can be behaving perfectly.

The email can still arrive 45 seconds late.

Same problem with:

  • Stripe test environments
  • OAuth providers
  • SMS services
  • analytics scripts
  • CDNs
  • fraud detection
  • CAPTCHA services
  • embedded support widgets
  • rate-limited APIs

If the dependency is live, its reliability becomes part of your test reliability.

You need to decide which external systems you're actually testing.

Three sensible strategies

1. Mock it

Best when the third-party behavior itself is not what you're trying to validate.

2. Test the integration separately

Keep a smaller set of slower integration checks for the real provider.

3. Accept the dependency

Sometimes a true end-to-end check is worth the instability. Just don't pretend it has the same reliability characteristics as an isolated regression test.

What usually doesn't work is accidentally combining all three approaches in one giant suite and then wondering why the pass rate is weird.

Parallel execution finds race conditions you didn't know you had

Parallelization is one of Playwright's strengths.

It's also extremely good at exposing poor test architecture.

A suite might look stable with:

workers: 1
Enter fullscreen mode Exit fullscreen mode

and collapse with:

workers: 8
Enter fullscreen mode Exit fullscreen mode

That doesn't necessarily mean parallel execution is the problem.

It usually means parallel execution revealed a hidden dependency.

Look for:

  • shared accounts
  • shared filesystem paths
  • shared downloads
  • mutable global fixtures
  • rate limits
  • database cleanup
  • environment-wide feature flags
  • reusable test records

Turning the workers back down to one may make the suite green.

It can also hide the actual bug in the suite design.

"It only fails in CI" is usually useful information

Local:

✓ 632 passed
Enter fullscreen mode Exit fullscreen mode

CI:

✗ 6 failed
Enter fullscreen mode Exit fullscreen mode

Rerun CI:

✓ 632 passed
Enter fullscreen mode Exit fullscreen mode

It's tempting to shrug.

Don't.

CI differences often expose assumptions your local machine accidentally satisfies.

Look at:

  • CPU pressure
  • memory pressure
  • viewport
  • operating system
  • browser version
  • locale
  • timezone
  • fonts
  • animation speed
  • network latency
  • environment variables
  • feature flags
  • proxy behavior

A test that fails only under slower execution may contain a race that was always there.

Your MacBook simply outran it.

Don't ignore timezone and locale

Here's a wonderfully boring source of flaky tests:

await expect(date).toHaveText('10/09/2026');
Enter fullscreen mode Exit fullscreen mode

Looks fine.

Until the runner uses another locale.

Or the execution crosses midnight UTC.

Or daylight saving changes.

Or your cloud browser is running in a different timezone.

Playwright lets you configure locale and timezone explicitly.

Use that when the scenario depends on them.

Environment assumptions should be configuration, not luck.

Retries can turn a broken suite into a green dashboard

Playwright supports retries for failed tests. The official retry documentation even classifies tests that fail initially and then pass on retry as "flaky."

That's an important distinction.

This:

Attempt 1: FAIL
Attempt 2: PASS
Enter fullscreen mode Exit fullscreen mode

does not mean:

PASS
Enter fullscreen mode Exit fullscreen mode

It means:

FLAKY
Enter fullscreen mode Exit fullscreen mode

Those are different states.

Unfortunately, dashboards have trained us to mostly care about red versus green.

So a retry becomes a convenient way to make red disappear.

Retries are useful

I absolutely use retries diagnostically.

For example, capturing a trace on the first retry can be useful.

Playwright explicitly supports that pattern.

Retries are dangerous

The danger is letting this become permanent:

"If it passes eventually, it's fine."
Enter fullscreen mode Exit fullscreen mode

Because once your team stops investigating flaky tests, the suite starts accumulating uncertainty.

Soon you don't know whether:

flaky
Enter fullscreen mode Exit fullscreen mode

means:

bad test
Enter fullscreen mode Exit fullscreen mode

or:

real intermittent product bug
Enter fullscreen mode Exit fullscreen mode

And that distinction matters.

Sometimes the flaky test found a real bug

This gets overlooked.

A test that fails once every 100 runs might be exposing:

  • a race condition
  • duplicate requests
  • stale cache state
  • websocket ordering
  • token refresh behavior
  • eventual consistency
  • backend concurrency
  • rendering under load

Deleting the test because it's "annoying" can remove the only automated signal telling you the product occasionally breaks.

A good rule is:

Prove the test is wrong before calling the failure meaningless.

Trace Viewer should be your first stop, not your last resort

Playwright's Trace Viewer is excellent.

The official Trace Viewer documentation shows why.

You can inspect:

  • actions
  • before/after DOM snapshots
  • screenshots
  • network requests
  • console messages
  • source code
  • timing
  • errors

For CI failures, I'd rather have a trace than 500 lines of logging somebody added after the fact.

A useful config is:

use: {
  trace: 'retain-on-failure',
  screenshot: 'only-on-failure'
}
Enter fullscreen mode Exit fullscreen mode

or a retry-based trace strategy if the storage/performance tradeoff makes more sense for your suite.

The goal is to preserve enough evidence to answer:

What was different in the failing run?

That's the actual debugging question.

Locator problems still exist

Playwright's locator system is better than brittle XPath everywhere.

But people can still write terrible locators.

This:

page.locator('div:nth-child(4) > div:nth-child(2) > button')
Enter fullscreen mode Exit fullscreen mode

is basically a resignation letter disguised as test code.

Prefer semantic locators when possible:

page.getByRole('button', { name: 'Create account' })
Enter fullscreen mode Exit fullscreen mode

or stable test IDs for application-specific elements.

Avoid basing tests on:

  • DOM depth
  • generated classes
  • styling hooks
  • arbitrary nth-child selectors
  • text that changes frequently

And remember that "self-healing" a bad locator doesn't necessarily solve the problem.

It can make the test click the wrong thing more confidently.

Network idle is not a universal definition of ready

Another anti-pattern is treating network inactivity as application readiness.

Modern applications may have:

  • analytics
  • websockets
  • polling
  • background refresh
  • prefetching
  • long-lived requests

There may never be a clean "nothing is happening" moment.

Again, wait for the actual application outcome you care about.

The more specific your readiness condition is, the less room the test has to guess.

Animations create tiny race windows

A menu can be:

visible
Enter fullscreen mode Exit fullscreen mode

while still moving.

A button can exist while an overlay is fading out.

A drawer can technically be open before the final layout settles.

Playwright's actionability logic handles a lot of this, but application-level animation state can still introduce edge cases.

For test environments, many teams reduce or disable non-essential animation.

That's not cheating.

You're testing behavior, not whether CSS can spend 240 milliseconds easing between coordinates.

Unless animation behavior itself is what you're testing.

Virtualized lists create another category of false assumptions

Suppose the UI shows 10,000 records.

The DOM might contain 30 rows.

Scroll down.

Those same DOM nodes get recycled with different content.

Now imagine a test that:

  1. finds a row
  2. stores the element
  3. scrolls
  4. tries to interact with the stored element

The user's mental model is:

10,000 rows
Enter fullscreen mode Exit fullscreen mode

The browser's actual model may be:

30 reusable DOM nodes
Enter fullscreen mode Exit fullscreen mode

Tests have to understand the application model they're interacting with.

This is another example of why "the element exists" doesn't necessarily mean "the workflow is stable."

Feature flags can make the same test mean two different things

You run a test today:

PASS
Enter fullscreen mode Exit fullscreen mode

Tomorrow:

FAIL
Enter fullscreen mode Exit fullscreen mode

The code didn't change.

The test didn't change.

The feature configuration did.

Modern applications can vary by:

  • account
  • region
  • experiment cohort
  • role
  • plan
  • release ring

If the test depends on a flag, control it.

Don't let your regression suite accidentally participate in an A/B test.

Authentication is another hidden source of state

A lot of teams optimize Playwright tests by reusing authentication state.

Makes sense.

Logging in 1,000 times is wasteful.

But now your suite depends on:

  • token expiration
  • session lifetime
  • role state
  • cookie configuration
  • refresh behavior

The more state you reuse, the more carefully you need to define who owns it and when it gets refreshed.

Optimization can quietly reintroduce coupling.

A practical flakiness triage workflow

Here's the process I'd use before adding another retry.

1. Reproduce it repeatedly

Run the individual test many times.

npx playwright test checkout.spec.ts --repeat-each=50
Enter fullscreen mode Exit fullscreen mode

If it fails predictably under repetition, that's good.

Now you have something you can investigate.

2. Compare isolated vs. full-suite execution

If it only fails in the full suite, suspect shared state.

3. Compare workers

Try one worker versus several.

If parallelism changes the result, look for collisions.

4. Keep the first failing trace

Don't only trace the retry that passes.

The first failure is the evidence you care about.

5. Look for application transitions

Ask whether the test is waiting for:

element ready
Enter fullscreen mode Exit fullscreen mode

instead of:

workflow ready
Enter fullscreen mode Exit fullscreen mode

6. Identify external dependencies

List every external system involved.

You may discover the "Playwright flake" is actually an email provider flake.

7. Make the environment explicit

Timezone.

Locale.

Viewport.

Feature flags.

Browser version.

Data.

Stop depending on defaults you didn't choose.

Track flakes like defects

If your suite is important, measure flakiness.

At minimum:

test
first-failure rate
retry-pass rate
browser
environment
failure category
first seen
last seen
Enter fullscreen mode Exit fullscreen mode

Then you can distinguish:

"We occasionally have flakes."
Enter fullscreen mode Exit fullscreen mode

from:

"8.7% of runs contain at least one non-deterministic failure."
Enter fullscreen mode Exit fullscreen mode

Those are very different conversations.

The bigger problem: Playwright flakiness becomes an ownership problem

At some point you've done everything correctly.

You have:

  • good locators
  • isolated data
  • deterministic fixtures
  • sensible mocks
  • controlled environments
  • trace collection
  • stable CI
  • careful assertions

And you're still spending real engineering time maintaining all of it.

That's when I think teams should zoom out.

The question stops being:

How do we fix another flaky Playwright test?

and becomes:

How much of our testing infrastructure do we actually want to own?

Because a mature Playwright setup isn't just test files.

It's often:

Playwright
+
fixtures
+
data factories
+
authentication helpers
+
mock servers
+
CI configuration
+
reporting
+
artifact storage
+
flake triage
+
custom utilities
+
ongoing maintenance
Enter fullscreen mode Exit fullscreen mode

That can absolutely be the right architecture.

But it isn't free just because Playwright is open source.

There is interesting research comparing Endtest with Playwright + Claude

If the maintenance burden is the part you want to reduce, there are now some fairly detailed public comparisons worth reading.

One is We Spent 6 Months Comparing Playwright + Claude vs. Endtest. Here’s What We Found.

According to the authors, eight software engineers and two test automation engineers compared the approaches for roughly six months using real workflows including:

  • authentication
  • CRUD operations
  • uploads and downloads
  • email flows
  • APIs
  • accessibility
  • visual regression
  • PDF validation

They reported:

Average completed-test creation

Endtest:              ~3 minutes
Playwright + Claude: ~27 minutes
Enter fullscreen mode Exit fullscreen mode

They define "completed" as created, successfully executed, verified, and approved, not simply code generated by AI.

They also reported:

9x faster completed-test creation
4x lower maintenance effort
100x growth in automated coverage
~500 hours/month of manual testing eliminated
Enter fullscreen mode Exit fullscreen mode

Those are the authors' reported results, not universal benchmarks.

But they're interesting precisely because Playwright was allowed to use Claude.

The comparison wasn't:

AI
vs.
no AI
Enter fullscreen mode Exit fullscreen mode

It was closer to:

code-based automation + AI assistance
vs.
a managed AI testing platform
Enter fullscreen mode Exit fullscreen mode

The authors say Claude could produce initial Playwright code quickly, but engineers still spent time executing it, correcting assumptions, adjusting locators and assertions, troubleshooting, and verifying the final test.

That's the part of the lifecycle that matters when you're thinking about maintenance.

Another team says they eventually stopped using Playwright entirely

A second detailed writeup, We Replaced Playwright with Endtest. Four Months Later, We Don’t Use Playwright at All, describes a four-month migration away from Playwright.

The interesting part is that the author says they initially expected to keep Playwright for the complicated scenarios.

Instead, they report eventually moving those as well.

Their published numbers include:

Automated scenarios:
85 → 410

Average completed-test creation:
30.8 minutes → 4.6 minutes

Monthly maintenance:
61 hours → 17 hours

Regression executions:
~20/month → ~60/month
Enter fullscreen mode Exit fullscreen mode

Again, those numbers belong to that team.

Your application may produce completely different results.

But both reports point to the same underlying question:

Is the best way to reduce flaky-test maintenance to keep improving the Playwright infrastructure, or to own less of the infrastructure in the first place?

That's a question worth asking.

Endtest is not interesting because Playwright can't automate the scenarios

Playwright can automate an enormous range of workflows.

That's not the issue.

The Endtest argument is about the amount of machinery your team has to own around those workflows.

The six-month comparison describes using several ways to create and maintain tests in Endtest:

  • browser recording
  • AI-assisted creation
  • autonomous exploration
  • manual editable steps
  • importing existing tests

That is a very different operating model from maintaining a code-first Playwright suite.

If your team enjoys owning test infrastructure, Playwright may still be the better fit.

If your actual problem is:

too many flakes
too much maintenance
too much debugging
not enough coverage
Enter fullscreen mode Exit fullscreen mode

then an alternative where more of that machinery is handled by the platform deserves a serious benchmark.

How I would compare them

Don't migrate because somebody wrote a Medium article.

Don't reject the idea because Playwright is popular either.

Take your worst 20 tests.

Not login.

Not search.

The horrible ones.

Include:

  • an iframe
  • an upload
  • a download
  • email verification
  • multiple tabs
  • dynamic data
  • API setup
  • a flaky third-party dependency
  • a role change
  • something that regularly fails in CI

Build the same scenarios both ways.

Then measure:

Time to first working test
Time to verified test
First-run pass rate
Retry rate
Flake rate
Maintenance hours
Human investigation time
90-day survival rate
Enter fullscreen mode Exit fullscreen mode

The last metric matters a lot.

A demo can tell you how fast a test can be created.

Ninety days tells you what you actually bought.

The goal is not zero flaky tests

That sounds strange, but hear me out.

You can probably reach zero flakes by deleting every difficult test.

That's not useful.

The goal is:

Failures should mean something.

When a test goes red, the team should assume:

something needs investigation
Enter fullscreen mode Exit fullscreen mode

not:

eh, probably Playwright
Enter fullscreen mode Exit fullscreen mode

Once engineers stop trusting the suite, the automation has lost most of its value.

People stop investigating.

Retries increase.

Failures become background noise.

Eventually CI is green and nobody believes it.

That's much worse than having fewer tests.

Playwright solved a lot. It didn't solve ownership.

Playwright deserves its popularity.

It fixed a huge amount of browser automation pain.

But a serious test suite still needs someone to understand:

  • application readiness
  • state
  • data
  • dependencies
  • parallelism
  • CI
  • environment configuration
  • failures
  • maintenance

That's why flaky Playwright tests are still a thing.

Not because Playwright forgot how to wait for buttons.

Because the difficult part of testing moved up a layer.

And once your team spends enough time fighting that layer, it becomes reasonable to ask whether another retry is really the answer.

Maybe the better question is whether you should keep owning all of it.

Top comments (2)

Collapse
 
arhancanli profile image
Arhan Canli •

The "fails once every 40 runs" test from the hydration section is a good size for a number that retries hide. If a test fails 1 in 40 independently, one retry makes the chance of a red build 1/1600 (0.06%), so it stays green for years while the 2.5% underneath it is still there. And to see it on purpose you need volume: 100 repeats surface at least one failure only 92% of the time, and about 91 repeats are needed for 90%. One --repeat-each run of 10 would show it roughly 22% of the time.

So the useful split is probably two jobs: the normal run with retries for the merge signal, and a scheduled job that runs only the tests marked flaky in the last N reports, repeated 100-200 times, and records the failure rate per test. A rate you can trend (2.5% to 0.3% after a fix) tells you whether the fix worked; a single green rerun after a change cannot, because it would also have been green most of the time before.

Collapse
 
sgaggjhkjh profile image
Arjun Sharma •

The distinction between browser automation flakiness and application flakiness is the sharpest point in this article. The section about accumulating waitForTimeout resonated, I have seen suites with literal minutes of accumulated sleeps where everyone was afraid to touch them. The one addition I would make is that retries are fine as a temporary diagnostic signal, but the retry pass rate you mention tracking is what separates a team that manages flakiness from a team that just paints the dashboard green. Also agreed on Trace Viewer, retain on failure changed how we debug CI failures more than any other single setting.