DEV Community

Cover image for Most Test Suites Only Work in a World That Doesn't Exist
Antoine Dubois
Antoine Dubois

Posted on

Most Test Suites Only Work in a World That Doesn't Exist

A lot of automated tests begin in an unusually pleasant universe.

The user is logged in.

The session is valid.

The database contains exactly the right data.

The feature flags are known.

The browser has no weird state.

The network works.

Nothing unexpected is covering the screen.

And the deployment that happened five minutes ago is assumed to be fine because CI passed.

Then the test runs.

Green.

Fantastic.

Unfortunately, real users don't live in that universe.

They leave tabs open overnight.

Their access tokens expire halfway through a workflow.

They get assigned to experiments.

They lose network connectivity.

A modal opens while another component still has focus.

A deployment changes one service but not another.

And sometimes the application quietly logs them back in while they're doing something important.

If your test automation strategy assumes those things don't happen, you might have excellent coverage of a product that doesn't actually exist.

The clean-session assumption is particularly dangerous

Here's a very normal test:

Login
Open Dashboard
Create Project
Invite User
Verify Invitation
Enter fullscreen mode Exit fullscreen mode

Looks good.

Now try this:

Login
Open Dashboard
Wait until access token expires
Start creating project
API returns 401
Application silently refreshes token
Request retries
Continue workflow
Enter fullscreen mode Exit fullscreen mode

Did the project get created once?

Twice?

Not at all?

Did the UI recover?

Did the application dump the user back onto the login page?

Did it preserve what they had typed?

Did the test framework even notice that authentication changed underneath it?

This is one of those situations where a test can pass hundreds of times simply because you're always starting with a fresh session.

There's a useful benchmark plan for session expiry recovery, silent re-login, and token-refresh evidence that treats this as something browser testing platforms should be evaluated against.

I agree.

Authentication isn't just a setup step.

It's application state.

Users don't politely expire between tests

Most automation frameworks encourage isolation.

That's generally a good thing.

Fresh context.

Fresh cookies.

Fresh test data.

Predictable environment.

But isolation can accidentally remove the exact states that make real applications difficult.

A real user might:

09:00  Log in
09:15  Open a form
09:22  Switch tabs
10:04  Come back
10:07  Submit
Enter fullscreen mode Exit fullscreen mode

Your automated test might do the entire thing in 14 seconds.

Those are technically the same workflow.

They're not the same conditions.

If your product has session expiration, refresh tokens, idle timeouts, or role changes, time itself becomes part of the test.

Feature flags create a similar illusion

Feature flags are great.

They let you roll things out gradually.

Test with internal users.

Run experiments.

Disable broken features quickly.

But they also create a new dimension of application state.

The same URL can now represent two different products.

checkout_v2 = false

vs.

checkout_v2 = true
Enter fullscreen mode Exit fullscreen mode

And if your test suite handles this by copying flag setup into 80 scenarios, you eventually get something like:

Set checkout_v2 = true
Set new_nav = false
Set pricing_test = variant_b
Set recommendations = true
Enter fullscreen mode Exit fullscreen mode

Repeated everywhere.

Then somebody changes the feature configuration and half the suite becomes archaeological evidence.

This guide on testing feature flag overrides without hardcoding the same flag values into every scenario tackles that exact problem.

The larger lesson is useful even if you're not using feature flags heavily.

Test setup is part of your architecture.

If every test independently recreates the world it needs, maintenance eventually becomes expensive.

UI state is more than "element visible"

Here's another situation I think browser automation often oversimplifies.

Modals.

Testing a modal usually looks something like:

Click "Delete"
Verify confirmation modal is visible
Click "Confirm"
Enter fullscreen mode Exit fullscreen mode

Okay.

But accessibility changes what "working" actually means.

If a modal is open, what happens to the rest of the page?

Can keyboard focus escape?

Can you tab into controls behind the modal?

Can screen-reader users interact with background content that visually appears locked?

The relatively new inert attribute gives browsers a cleaner way to make parts of a page non-interactive.

But testing it properly requires more than:

modal.isVisible() === true
Enter fullscreen mode Exit fullscreen mode

This walkthrough on testing the inert attribute, modal background locking, and tabbable leakage is a good example of how much behavior can hide behind something that looks visually correct.

A screenshot can show you a perfect modal.

The keyboard experience can still be completely broken.

This is why I don't love "end-to-end coverage" as a metric

Teams will say:

We have 80% of our critical user journeys automated.

That sounds great.

But what does it mean?

If you test:

Login → Dashboard → Checkout → Confirmation
Enter fullscreen mode Exit fullscreen mode

do you also test:

Login
    ↓
Session expires
    ↓
Checkout begins
    ↓
Token refreshes
    ↓
Feature flag changes rendering
    ↓
Modal opens
    ↓
Keyboard navigation continues
    ↓
Purchase completes
Enter fullscreen mode Exit fullscreen mode

Probably not.

And I'm not saying you should automate every imaginable combination.

That would be ridiculous.

The point is that "journey coverage" can give you false confidence if your journeys are modeled as clean sequences instead of changing states.

AI-generated tests can amplify this problem

AI makes it easier to generate a lot of automation.

Type:

Test checkout.

And an agent can potentially create the flow for you.

That's useful.

But the first generated version will often represent the obvious path.

Open checkout
Enter card
Click Pay
Verify success
Enter fullscreen mode Exit fullscreen mode

The interesting part is what happens later.

Can you edit the steps?

Can you add recovery behavior?

Can another person understand why something failed?

Can you attach notes explaining why a strange step exists?

Can the test survive being handed to a different team six months from now?

Those questions matter much more once AI starts generating a significant percentage of the suite.

This rubric on editable steps, recovery notes, and team handoffs in AI testing platforms focuses on exactly that.

I think "editable" is an underrated feature in AI systems.

Autonomy sounds cooler.

Editability is what saves you when reality gets complicated.

Fully autonomous QA sounds great until you ask what happens at 3 AM

There's a lot of interest now in autonomous QA agents.

And I understand why.

The idea is compelling:

Give agent application
    ↓
Agent explores
    ↓
Agent creates tests
    ↓
Agent runs continuously
    ↓
Agent finds failures
Enter fullscreen mode Exit fullscreen mode

That's a very different model from manually maintaining hundreds or thousands of scripts.

There's a demonstration of that approach in Endtest Bot: Your Autonomous QA Agent.

I think agentic testing will become a meaningful part of QA.

But the important question isn't whether the agent can click around your product.

It's:

What happens when the application behaves unexpectedly?

An autonomous system becomes much more valuable when it can recognize:

Session expired
Token refreshed
Expected state restored
Continue scenario
Enter fullscreen mode Exit fullscreen mode

rather than simply reporting:

Step failed
Enter fullscreen mode Exit fullscreen mode

That difference is huge.

Recovery may be more important than self-healing

"Self-healing" became one of the favorite phrases in test automation.

Usually it means:

Locator broke
    ↓
Tool found another locator
    ↓
Test continued
Enter fullscreen mode Exit fullscreen mode

Useful.

But that's a pretty narrow definition of healing.

Real applications need recovery from things like:

  • expired sessions
  • intermittent APIs
  • reauthentication
  • WebSocket reconnects
  • temporary loading states
  • feature changes
  • interrupted transactions
  • background refreshes

Changing #submit-button to [aria-label="Submit"] is the easy version.

Recovering application state without changing the meaning of the test is much harder.

That's where I think the next generation of testing tools will either become genuinely useful or just generate impressive demos.

And then we deploy

There's another weird boundary in a lot of testing strategies.

Everything before deployment gets enormous attention.

Everything after deployment gets:

We'll check Sentry.

But a successful deployment pipeline doesn't mean the product actually works.

You can still have:

Bad environment variable
Broken CDN asset
Missing migration
Expired production credential
Misconfigured feature flag
Third-party API failure
DNS issue
Regional outage
Enter fullscreen mode Exit fullscreen mode

CI may have been completely green.

Which raises a practical question:

What exactly should we run after deployment?

There isn't one answer.

Sometimes a fast API check is enough.

Sometimes you need a real browser.

Sometimes an important customer journey should be monitored continuously.

This decision guide for API smoke checks, full browser flows, and synthetic monitoring gives a useful framework for deciding where each belongs.

I'd think about them as different layers.

API Smoke Check
"Is the service basically alive?"

Browser Flow
"Can a user actually complete this?"

Synthetic Monitoring
"Does it keep working after we leave?"
Enter fullscreen mode Exit fullscreen mode

Those are related questions.

They're not interchangeable.

Your production test strategy probably deserves more attention than your CI strategy

Suppose your checkout service deploys successfully.

CI is green.

Your API smoke check passes.

But the production cookie-consent configuration changed.

Now the checkout button is covered for European users.

The backend is healthy.

The API is healthy.

The deployment is healthy.

Checkout is broken.

This is why I increasingly think the traditional line between:

Testing
Enter fullscreen mode Exit fullscreen mode

and:

Monitoring
Enter fullscreen mode Exit fullscreen mode

is becoming artificial.

If an automated browser checks checkout before deployment, we call it a test.

If the same browser checks checkout every five minutes after deployment, we call it synthetic monitoring.

The user doesn't care what we call it.

They care whether checkout works.

Try auditing your suite for assumptions instead of tests

Here's an exercise that may be more useful than counting test cases.

Take ten important automated tests.

For each one, write down its assumptions.

Something like:

User has valid session
Feature X enabled
Feature Y disabled
Network stable
No third-party overlay
API responds within 5 seconds
User role unchanged
No background token refresh
Desktop viewport
English locale
Enter fullscreen mode Exit fullscreen mode

Now ask:

Which of these assumptions are actually guaranteed?

You might find that your tests aren't really validating a workflow.

They're validating a workflow under one carefully curated version of reality.

That's still useful.

But it's important to know what you're getting.

I wouldn't try to test every combination

There's an obvious objection.

If we start considering:

sessions
flags
roles
locales
networks
devices
browser state
accessibility state
deployment state
Enter fullscreen mode Exit fullscreen mode

the number of combinations explodes.

Absolutely.

You shouldn't test all of them.

The solution isn't combinatorial madness.

It's risk-based selection.

Ask:

Which state changes have broken before?

Which ones affect money?

Which ones affect authentication?

Which ones affect permissions?

Which ones are difficult to recover from?

Which ones differ between staging and production?

Which ones would be embarrassing if a customer found them first?
Enter fullscreen mode Exit fullscreen mode

Test those.

The uncomfortable question

So here's the question I'd ask your team:

If our application behaved exactly like a normal production application instead of a freshly reset test environment, how much of our automation would still work?

Expired sessions.

Changing flags.

Half-finished interactions.

Browser history.

Modal focus.

Background refreshes.

Post-deployment configuration.

Real user state.

If the answer is "not much," that's worth knowing.

Because maybe your biggest automation gap isn't another missing test case.

Maybe your suite is simply making assumptions that production never agreed to.

Top comments (2)

Collapse
 
pm25coder profile image
pm25coder •

The cases you list are all application state. There's a class next to them that's harder, because the suite doesn't set it — it inherits it and never writes it down. The ambient world: the clock, the encoding, the spelling of a path.

Three receipts, same codebase.

The clock. A regression test pinned a date. The code had a clock injection point, but it fed only the timestamp while the retention window was computed from the wall clock. The same commit was green on the CI leg (UTC) and red on the developer's host (+08) — the test's universe had a fixed "now" the code didn't share. It has since gone red everywhere, which is the time bomb going off. You write that time itself becomes part of the test; this is the version where it's part of the test and not part of the fixture.

The directory that existed. A suite precondition skipped unless node_modules was present. It was present, and empty — 0 entries, no .bin, no runner binary. The check passed, the runner was missing, and the host reported a failure where the docstring promised a skip. "Present" and "usable" are two different worlds, and a presence check can only speak about the first.

The decoder. A tool's test harness captured stdout as UTF-8; production decoded the same bytes in the console code page. 56 of 6,236 lines were unrepresentable in the real decoder. The suite was green in a world — a UTF-8 stdout — that doesn't exist on the machine that runs the tool, and the defect it hid was in the failure-reporting path, the one path you can least afford to have tested against a different world.

Threaded through all three, plus the usual one: a test that compared two spellings of the same directory (ADMINI~1 from the temp path builder vs Administrator from the tool's own output) passed on the unfixed code, because "the path is not in the listing" was vacuously true.

The pattern: your clean-session examples are state the test chooses; these are state the test assumes, and assumptions don't show up in a diff when they change. Same fix shape as yours — make the fixture name the ambient state it depends on, the way it already names the login state. A fixed clock, the expected code page, the expected path form are all fixture fields. If they aren't written down, the suite is green about a machine nobody has.

Collapse
 
cubl9snp71hm profile image
cubl9snp71hm •

Bài viết chạm đúng vào nỗi đau của nhiều team: test suite xanh trên CI nhưng production vẫn vỡ vì data race, network timeout, hoặc state dính từ session trước.

Điểm mấu chốt thường bị bỏ qua là test isolation thực sự. Không chỉ cleanup DB giữa các case, mà là giả lập lại toàn bộ context: clock skew, partial failure của downstream service, rate limit, thậm chí cả user bấm back/refresh lúc request đang pending.

Một pattern hiệu quả mình thấy: viết "chaos fixture" — một tập hợp scenario xấu (token hết hạn giữa chừng, DB connection drop, third-party trả 500) và inject random vào suite chạy nightly. Không cần phức tạp, chỉ cần một wrapper bọc test case gốc và throw fault ngẫu nhiên. Số bug regression bắt được từ đây thường gấp 3-4 lần so với happy-path-only.

Còn về flaky test: thay vì quarantine rồi quên, hãy treat nó như production incident — root cause, fix tại source (thường là timing assumption hoặc shared state), rồi thêm assertion chắc chắn hơn. Flaky test không phải "test kém may mắn", nó là test đang nói thật về hệ thống PS: the tool I meant is on labagent .tech