Playwright has auto-waiting.
It has isolated browser contexts.
It has auto-retrying assertions.
It has traces.
It has screenshots.
It has videos.
It has one of the best debugging experiences of any browser automation framework I've used.
So why do mature Playwright suites still end up with this?
412 passed
7 flaky
3 failed
Then:
Retrying...
And five minutes later:
422 passed
Green build.
Everyone moves on.
Until tomorrow.
The uncomfortable answer is that Playwright solved a lot of browser automation flakiness.
It did not solve application flakiness, data flakiness, environment flakiness, dependency flakiness, or bad assumptions in your test architecture.
And once your suite gets large enough, those become the real problem.
Auto-waiting does not mean "the application is ready"
Playwright's actionability checks are genuinely good.
Before a normal click, Playwright checks things like whether the element is visible, stable, enabled, and able to receive events.
That's a huge improvement over the old pattern of guessing how long to sleep.
The Playwright documentation on auto-waiting explains those checks in detail.
But there's a limitation hiding in plain sight.
Playwright understands the browser.
It does not understand your business logic.
Consider this sequence:
Dashboard renders
↓
"Save" button becomes enabled
↓
Background request finishes
↓
Application recalculates permissions
↓
React rerenders the form
↓
Final state is ready
Playwright might correctly decide the button is clickable at step two.
Your workflow may not actually be safe until step six.
The element is ready.
The application isn't.
This is one of the biggest sources of flaky tests in modern frontend applications.
The wrong fix
await page.waitForTimeout(3000);
This works right up until CI needs 3.2 seconds.
Then someone changes it to 5 seconds.
Six months later your test suite contains 140 seconds of accumulated superstition.
The better fix
Wait for the state that actually matters.
For example:
await expect(page.getByTestId('invoice-status')).toHaveText('Paid');
or wait for the relevant response:
const response = page.waitForResponse(
r => r.url().includes('/api/invoices/') && r.status() === 200
);
await page.getByRole('button', { name: 'Refresh' }).click();
await response;
The key idea is simple:
Wait for meaning, not time.
Hydration creates a particularly nasty version of this problem
Modern frameworks can render something that looks ready before it really is.
The browser may show:
Submit
The button exists.
It's visible.
Playwright can find it.
Then hydration finishes and the framework replaces the node.
Or event handlers attach slightly later.
Or data arrives and the component rerenders.
Now a test that looked perfectly reasonable fails once every 40 runs.
This gets worse with:
- server-side rendering
- React hydration
- lazy components
- Suspense
- client-side routing
- animation libraries
- container queries
- virtualized content
These failures often get mislabeled as "Playwright being flaky."
Usually Playwright is doing exactly what you told it to do.
The test just observed the application during a transitional state.
Shared state is still one of the fastest ways to destroy a test suite
Playwright creates isolated browser contexts for tests by default.
That's great.
Cookies, local storage, and session storage are isolated between those contexts. The Playwright isolation documentation is explicit about why this matters.
But your backend data isn't automatically isolated.
Imagine two parallel tests:
Test A:
Log in as qa@example.com
Delete all notifications
Create notification
Test B:
Log in as qa@example.com
Expect 3 notifications
Each browser is isolated.
The account isn't.
Now increase the worker count.
Congratulations, you built a slot machine.
Common shared-state mistakes
I've seen all of these:
- every test uses the same user
- every test edits the same project
- tests depend on "latest" records
- tests assume a clean database
- tests depend on execution order
- teardown from one test deletes another test's data
- parallel tests use the same uploaded filename
- tests reuse the same shopping cart
- email tests read from one shared inbox
The usual symptom is:
It passes when I run it alone.
That's one of the most useful clues you can get.
A stable test should own its data
Where practical, give the test unique state.
Instead of:
qa@example.com
create:
qa+run-84217@example.com
Instead of:
Project: Automation Test
use:
Project: Automation Test 84217
Instead of:
page.getByText('Latest Order')
assert against the order ID created by the test.
You're trying to make this true:
Test outcome
=
application behavior
rather than:
Test outcome
=
application behavior
+
whatever seven other workers happened to do
Third-party dependencies quietly inherit your test reliability
Here's another trap.
Your signup flow works like this:
Create account
↓
Send verification email
↓
Email provider accepts message
↓
Message arrives
↓
Test opens link
↓
Account becomes verified
Your application can be healthy.
Playwright can be behaving perfectly.
The email can still arrive 45 seconds late.
Same problem with:
- Stripe test environments
- OAuth providers
- SMS services
- analytics scripts
- CDNs
- fraud detection
- CAPTCHA services
- embedded support widgets
- rate-limited APIs
If the dependency is live, its reliability becomes part of your test reliability.
You need to decide which external systems you're actually testing.
Three sensible strategies
1. Mock it
Best when the third-party behavior itself is not what you're trying to validate.
2. Test the integration separately
Keep a smaller set of slower integration checks for the real provider.
3. Accept the dependency
Sometimes a true end-to-end check is worth the instability. Just don't pretend it has the same reliability characteristics as an isolated regression test.
What usually doesn't work is accidentally combining all three approaches in one giant suite and then wondering why the pass rate is weird.
Parallel execution finds race conditions you didn't know you had
Parallelization is one of Playwright's strengths.
It's also extremely good at exposing poor test architecture.
A suite might look stable with:
workers: 1
and collapse with:
workers: 8
That doesn't necessarily mean parallel execution is the problem.
It usually means parallel execution revealed a hidden dependency.
Look for:
- shared accounts
- shared filesystem paths
- shared downloads
- mutable global fixtures
- rate limits
- database cleanup
- environment-wide feature flags
- reusable test records
Turning the workers back down to one may make the suite green.
It can also hide the actual bug in the suite design.
"It only fails in CI" is usually useful information
Local:
✓ 632 passed
CI:
✗ 6 failed
Rerun CI:
✓ 632 passed
It's tempting to shrug.
Don't.
CI differences often expose assumptions your local machine accidentally satisfies.
Look at:
- CPU pressure
- memory pressure
- viewport
- operating system
- browser version
- locale
- timezone
- fonts
- animation speed
- network latency
- environment variables
- feature flags
- proxy behavior
A test that fails only under slower execution may contain a race that was always there.
Your MacBook simply outran it.
Don't ignore timezone and locale
Here's a wonderfully boring source of flaky tests:
await expect(date).toHaveText('10/09/2026');
Looks fine.
Until the runner uses another locale.
Or the execution crosses midnight UTC.
Or daylight saving changes.
Or your cloud browser is running in a different timezone.
Playwright lets you configure locale and timezone explicitly.
Use that when the scenario depends on them.
Environment assumptions should be configuration, not luck.
Retries can turn a broken suite into a green dashboard
Playwright supports retries for failed tests. The official retry documentation even classifies tests that fail initially and then pass on retry as "flaky."
That's an important distinction.
This:
Attempt 1: FAIL
Attempt 2: PASS
does not mean:
PASS
It means:
FLAKY
Those are different states.
Unfortunately, dashboards have trained us to mostly care about red versus green.
So a retry becomes a convenient way to make red disappear.
Retries are useful
I absolutely use retries diagnostically.
For example, capturing a trace on the first retry can be useful.
Playwright explicitly supports that pattern.
Retries are dangerous
The danger is letting this become permanent:
"If it passes eventually, it's fine."
Because once your team stops investigating flaky tests, the suite starts accumulating uncertainty.
Soon you don't know whether:
flaky
means:
bad test
or:
real intermittent product bug
And that distinction matters.
Sometimes the flaky test found a real bug
This gets overlooked.
A test that fails once every 100 runs might be exposing:
- a race condition
- duplicate requests
- stale cache state
- websocket ordering
- token refresh behavior
- eventual consistency
- backend concurrency
- rendering under load
Deleting the test because it's "annoying" can remove the only automated signal telling you the product occasionally breaks.
A good rule is:
Prove the test is wrong before calling the failure meaningless.
Trace Viewer should be your first stop, not your last resort
Playwright's Trace Viewer is excellent.
The official Trace Viewer documentation shows why.
You can inspect:
- actions
- before/after DOM snapshots
- screenshots
- network requests
- console messages
- source code
- timing
- errors
For CI failures, I'd rather have a trace than 500 lines of logging somebody added after the fact.
A useful config is:
use: {
trace: 'retain-on-failure',
screenshot: 'only-on-failure'
}
or a retry-based trace strategy if the storage/performance tradeoff makes more sense for your suite.
The goal is to preserve enough evidence to answer:
What was different in the failing run?
That's the actual debugging question.
Locator problems still exist
Playwright's locator system is better than brittle XPath everywhere.
But people can still write terrible locators.
This:
page.locator('div:nth-child(4) > div:nth-child(2) > button')
is basically a resignation letter disguised as test code.
Prefer semantic locators when possible:
page.getByRole('button', { name: 'Create account' })
or stable test IDs for application-specific elements.
Avoid basing tests on:
- DOM depth
- generated classes
- styling hooks
- arbitrary nth-child selectors
- text that changes frequently
And remember that "self-healing" a bad locator doesn't necessarily solve the problem.
It can make the test click the wrong thing more confidently.
Network idle is not a universal definition of ready
Another anti-pattern is treating network inactivity as application readiness.
Modern applications may have:
- analytics
- websockets
- polling
- background refresh
- prefetching
- long-lived requests
There may never be a clean "nothing is happening" moment.
Again, wait for the actual application outcome you care about.
The more specific your readiness condition is, the less room the test has to guess.
Animations create tiny race windows
A menu can be:
visible
while still moving.
A button can exist while an overlay is fading out.
A drawer can technically be open before the final layout settles.
Playwright's actionability logic handles a lot of this, but application-level animation state can still introduce edge cases.
For test environments, many teams reduce or disable non-essential animation.
That's not cheating.
You're testing behavior, not whether CSS can spend 240 milliseconds easing between coordinates.
Unless animation behavior itself is what you're testing.
Virtualized lists create another category of false assumptions
Suppose the UI shows 10,000 records.
The DOM might contain 30 rows.
Scroll down.
Those same DOM nodes get recycled with different content.
Now imagine a test that:
- finds a row
- stores the element
- scrolls
- tries to interact with the stored element
The user's mental model is:
10,000 rows
The browser's actual model may be:
30 reusable DOM nodes
Tests have to understand the application model they're interacting with.
This is another example of why "the element exists" doesn't necessarily mean "the workflow is stable."
Feature flags can make the same test mean two different things
You run a test today:
PASS
Tomorrow:
FAIL
The code didn't change.
The test didn't change.
The feature configuration did.
Modern applications can vary by:
- account
- region
- experiment cohort
- role
- plan
- release ring
If the test depends on a flag, control it.
Don't let your regression suite accidentally participate in an A/B test.
Authentication is another hidden source of state
A lot of teams optimize Playwright tests by reusing authentication state.
Makes sense.
Logging in 1,000 times is wasteful.
But now your suite depends on:
- token expiration
- session lifetime
- role state
- cookie configuration
- refresh behavior
The more state you reuse, the more carefully you need to define who owns it and when it gets refreshed.
Optimization can quietly reintroduce coupling.
A practical flakiness triage workflow
Here's the process I'd use before adding another retry.
1. Reproduce it repeatedly
Run the individual test many times.
npx playwright test checkout.spec.ts --repeat-each=50
If it fails predictably under repetition, that's good.
Now you have something you can investigate.
2. Compare isolated vs. full-suite execution
If it only fails in the full suite, suspect shared state.
3. Compare workers
Try one worker versus several.
If parallelism changes the result, look for collisions.
4. Keep the first failing trace
Don't only trace the retry that passes.
The first failure is the evidence you care about.
5. Look for application transitions
Ask whether the test is waiting for:
element ready
instead of:
workflow ready
6. Identify external dependencies
List every external system involved.
You may discover the "Playwright flake" is actually an email provider flake.
7. Make the environment explicit
Timezone.
Locale.
Viewport.
Feature flags.
Browser version.
Data.
Stop depending on defaults you didn't choose.
Track flakes like defects
If your suite is important, measure flakiness.
At minimum:
test
first-failure rate
retry-pass rate
browser
environment
failure category
first seen
last seen
Then you can distinguish:
"We occasionally have flakes."
from:
"8.7% of runs contain at least one non-deterministic failure."
Those are very different conversations.
The bigger problem: Playwright flakiness becomes an ownership problem
At some point you've done everything correctly.
You have:
- good locators
- isolated data
- deterministic fixtures
- sensible mocks
- controlled environments
- trace collection
- stable CI
- careful assertions
And you're still spending real engineering time maintaining all of it.
That's when I think teams should zoom out.
The question stops being:
How do we fix another flaky Playwright test?
and becomes:
How much of our testing infrastructure do we actually want to own?
Because a mature Playwright setup isn't just test files.
It's often:
Playwright
+
fixtures
+
data factories
+
authentication helpers
+
mock servers
+
CI configuration
+
reporting
+
artifact storage
+
flake triage
+
custom utilities
+
ongoing maintenance
That can absolutely be the right architecture.
But it isn't free just because Playwright is open source.
There is interesting research comparing Endtest with Playwright + Claude
If the maintenance burden is the part you want to reduce, there are now some fairly detailed public comparisons worth reading.
One is We Spent 6 Months Comparing Playwright + Claude vs. Endtest. Here’s What We Found.
According to the authors, eight software engineers and two test automation engineers compared the approaches for roughly six months using real workflows including:
- authentication
- CRUD operations
- uploads and downloads
- email flows
- APIs
- accessibility
- visual regression
- PDF validation
They reported:
Average completed-test creation
Endtest: ~3 minutes
Playwright + Claude: ~27 minutes
They define "completed" as created, successfully executed, verified, and approved, not simply code generated by AI.
They also reported:
9x faster completed-test creation
4x lower maintenance effort
100x growth in automated coverage
~500 hours/month of manual testing eliminated
Those are the authors' reported results, not universal benchmarks.
But they're interesting precisely because Playwright was allowed to use Claude.
The comparison wasn't:
AI
vs.
no AI
It was closer to:
code-based automation + AI assistance
vs.
a managed AI testing platform
The authors say Claude could produce initial Playwright code quickly, but engineers still spent time executing it, correcting assumptions, adjusting locators and assertions, troubleshooting, and verifying the final test.
That's the part of the lifecycle that matters when you're thinking about maintenance.
Another team says they eventually stopped using Playwright entirely
A second detailed writeup, We Replaced Playwright with Endtest. Four Months Later, We Don’t Use Playwright at All, describes a four-month migration away from Playwright.
The interesting part is that the author says they initially expected to keep Playwright for the complicated scenarios.
Instead, they report eventually moving those as well.
Their published numbers include:
Automated scenarios:
85 → 410
Average completed-test creation:
30.8 minutes → 4.6 minutes
Monthly maintenance:
61 hours → 17 hours
Regression executions:
~20/month → ~60/month
Again, those numbers belong to that team.
Your application may produce completely different results.
But both reports point to the same underlying question:
Is the best way to reduce flaky-test maintenance to keep improving the Playwright infrastructure, or to own less of the infrastructure in the first place?
That's a question worth asking.
Endtest is not interesting because Playwright can't automate the scenarios
Playwright can automate an enormous range of workflows.
That's not the issue.
The Endtest argument is about the amount of machinery your team has to own around those workflows.
The six-month comparison describes using several ways to create and maintain tests in Endtest:
- browser recording
- AI-assisted creation
- autonomous exploration
- manual editable steps
- importing existing tests
That is a very different operating model from maintaining a code-first Playwright suite.
If your team enjoys owning test infrastructure, Playwright may still be the better fit.
If your actual problem is:
too many flakes
too much maintenance
too much debugging
not enough coverage
then an alternative where more of that machinery is handled by the platform deserves a serious benchmark.
How I would compare them
Don't migrate because somebody wrote a Medium article.
Don't reject the idea because Playwright is popular either.
Take your worst 20 tests.
Not login.
Not search.
The horrible ones.
Include:
- an iframe
- an upload
- a download
- email verification
- multiple tabs
- dynamic data
- API setup
- a flaky third-party dependency
- a role change
- something that regularly fails in CI
Build the same scenarios both ways.
Then measure:
Time to first working test
Time to verified test
First-run pass rate
Retry rate
Flake rate
Maintenance hours
Human investigation time
90-day survival rate
The last metric matters a lot.
A demo can tell you how fast a test can be created.
Ninety days tells you what you actually bought.
The goal is not zero flaky tests
That sounds strange, but hear me out.
You can probably reach zero flakes by deleting every difficult test.
That's not useful.
The goal is:
Failures should mean something.
When a test goes red, the team should assume:
something needs investigation
not:
eh, probably Playwright
Once engineers stop trusting the suite, the automation has lost most of its value.
People stop investigating.
Retries increase.
Failures become background noise.
Eventually CI is green and nobody believes it.
That's much worse than having fewer tests.
Playwright solved a lot. It didn't solve ownership.
Playwright deserves its popularity.
It fixed a huge amount of browser automation pain.
But a serious test suite still needs someone to understand:
- application readiness
- state
- data
- dependencies
- parallelism
- CI
- environment configuration
- failures
- maintenance
That's why flaky Playwright tests are still a thing.
Not because Playwright forgot how to wait for buttons.
Because the difficult part of testing moved up a layer.
And once your team spends enough time fighting that layer, it becomes reasonable to ask whether another retry is really the answer.
Maybe the better question is whether you should keep owning all of it.
Top comments (2)
The "fails once every 40 runs" test from the hydration section is a good size for a number that retries hide. If a test fails 1 in 40 independently, one retry makes the chance of a red build 1/1600 (0.06%), so it stays green for years while the 2.5% underneath it is still there. And to see it on purpose you need volume: 100 repeats surface at least one failure only 92% of the time, and about 91 repeats are needed for 90%. One
--repeat-eachrun of 10 would show it roughly 22% of the time.So the useful split is probably two jobs: the normal run with retries for the merge signal, and a scheduled job that runs only the tests marked flaky in the last N reports, repeated 100-200 times, and records the failure rate per test. A rate you can trend (2.5% to 0.3% after a fix) tells you whether the fix worked; a single green rerun after a change cannot, because it would also have been green most of the time before.
The distinction between browser automation flakiness and application flakiness is the sharpest point in this article. The section about accumulating waitForTimeout resonated, I have seen suites with literal minutes of accumulated sleeps where everyone was afraid to touch them. The one addition I would make is that retries are fine as a temporary diagnostic signal, but the retry pass rate you mention tracking is what separates a team that manages flakiness from a team that just paints the dashboard green. Also agreed on Trace Viewer, retain on failure changed how we debug CI failures more than any other single setting.