Most teams can tell you how many automated tests they have.
Some can tell you their pass rate.
A few can tell you their code coverage.
But ask this:
How many meaningful state transitions in your product are actually tested?
And the room usually gets quiet.
That's a much harder question.
It's also probably a better one.
Because users don't experience your product as a collection of pages.
They experience it as a sequence of state changes.
Logged out
↓
Logged in
↓
Trial
↓
Paid
↓
Payment failed
↓
Grace period
↓
Canceled
Or:
Draft
↓
Submitted
↓
Approved
↓
Rejected
↓
Edited
↓
Resubmitted
The bugs that hurt usually aren't:
The Settings page doesn't exist.
They're more like:
A user who downgraded, then upgraded, then changed teams still has the wrong permissions.
That's not a page problem.
That's a state-transition problem.
A lot of "coverage" is really screenshot coverage
Here's how browser automation often grows.
Someone automates login.
Then signup.
Then checkout.
Then settings.
Then the admin panel.
Eventually you have a few hundred tests and a dashboard that says:
742 automated tests
97.8% passing
Looks healthy.
But those numbers don't tell you whether you've tested the dangerous combinations.
For example:
Trial user → paid user
Paid user → canceled user
Canceled user → reactivated user
Admin → downgraded member
Invited user → expired invitation
Feature OFF → feature ON → feature OFF
Those transitions are where systems tend to get weird.
The UI may look correct in each isolated state.
The bug happens because the transition left something behind.
A stale permission.
An old token.
A cached plan.
A lingering feature flag.
A half-updated record.
This guide on mapping user-journey test coverage to state transitions makes a good case for treating transition coverage as a first-class QA problem.
I think that idea applies whether QA is outsourced or entirely internal.
Happy paths hide structural gaps
Take a subscription flow.
You probably test:
Signup
↓
Choose plan
↓
Pay
↓
Dashboard
Great.
Now test:
Signup
↓
Choose plan
↓
Payment fails
↓
Retry
↓
Payment succeeds
↓
Cancel
↓
Reactivate
↓
Change card
Different system.
Same product.
The first scenario proves the happy path works.
The second starts probing whether your application handles state.
And this is where I've become skeptical of raw test counts.
Ten carefully chosen transition tests can sometimes tell you more than 100 page-oriented checks.
API testing has the same problem
The same mistake happens at the API layer.
A team may have lots of tests like:
POST /users → 201
GET /users/123 → 200
DELETE /users/123 → 204
Useful.
But that doesn't necessarily tell you whether the API still behaves correctly as a contract between systems.
There's a meaningful difference between:
Does this request return what I expect?
and:
Are the producer and consumer still honoring the same contract?
That's the distinction behind comparisons like Pact vs REST Assured.
They're not interchangeable ways of doing the same thing.
One pushes you toward contract-first compatibility between services.
The other is very good at request-level API checks.
Which one matters more depends on how your system is built.
But if your frontend, backend, mobile app, and third-party integrations are all evolving independently, testing isolated requests may not be enough.
Frontend teams are quietly becoming QA infrastructure teams
There's another shift happening.
Frontend teams increasingly own:
- component tests
- browser tests
- visual checks
- accessibility checks
- CI workflows
- test data setup
- mock APIs
- release validation
Which means "frontend testing" isn't really just frontend testing anymore.
The frontend team may now own a substantial part of the release confidence system.
That's a different job.
If a team owns production UI code and the automation around it, the tooling needs to fit their development workflow rather than behave like a separate QA island.
This guide on browser testing platforms when frontend teams own the automation too explores that operating model.
I think this matters more than a generic list of features.
A tool can be excellent in isolation and still be a bad fit if developers hate touching it.
Ownership matters more than people admit
Here's a question I rarely see in tool comparisons:
Who is expected to maintain the tests six months from now?
That answer changes everything.
Imagine two models.
Model A
Your internal team:
Creates tests
Maintains tests
Reviews failures
Updates selectors
Handles flaky runs
Owns CI
Model B
A service or platform handles much of:
Maintenance
Triage
Review
Repair
Those are not small differences.
They're completely different operating models.
This BlinqIO vs QA.tech comparison frames the decision around managed maintenance versus faster self-serve browser coverage.
That is a much better way to think about it than asking which product has more AI.
The important question is:
What work is my team still responsible for after we buy this?
AI testing tools are splitting into different philosophies
The same thing is happening across AI testing platforms.
Some products want to give you:
Fast self-service automation.
Others lean more toward:
Guided workflows, review, and assistance.
Those approaches can look very similar in a demo.
Type an instruction.
Watch the browser move.
Test created.
Done.
But they create different day-to-day workflows.
A comparison like Reflect vs Octomind is more useful when you look at it through that lens: self-serve browser coverage versus more guided review workflows.
Neither is automatically better.
A five-person startup may want to move as fast as possible with minimal process.
A large QA organization may care much more about reviewability, governance, and consistency.
Same category.
Different problem.
I think we're entering the second phase of AI test automation
The first phase was:
Can AI create a test?
That's basically settled.
Yes.
It can.
Sometimes surprisingly well.
The more interesting questions now are:
Can it maintain the test?
Can humans understand what it changed?
Can it model complicated state?
Can it handle multiple roles?
Can it deal with non-happy paths?
Can it produce evidence a developer trusts?
Can teams own the result long term?
Those are much harder problems.
There are broader comparisons of the current market, including this review of 10 AI-powered test automation tools in 2026, that look at practical scenarios rather than only feature checklists.
That's the direction I think tool evaluation needs to go.
"Has AI" isn't a meaningful differentiator anymore.
Almost everybody has AI.
The question is what happens after the first impressive demo.
Your test model might be wrong
This is the bigger issue.
A lot of automation suites are modeled like this:
Page
↓
Actions
↓
Assertions
But many real applications behave more like this:
State A
↓
Transition
↓
State B
↓
Transition
↓
State C
The second model forces you to ask different questions.
For example:
What states can this user be in?
What transitions are allowed?
Which transitions should be impossible?
What data changes during the transition?
What permissions change?
What happens if the transition is interrupted?
What happens if the user repeats it?
Can the transition be reversed?
That's where some of the nastiest bugs live.
Here's a simple exercise
Take one important workflow in your application.
Not the whole product.
Just one.
Maybe:
Subscription
Write down every meaningful state.
For example:
Trial
Active
Past due
Canceled
Expired
Reactivated
Now draw the transitions.
Trial → Active
Trial → Expired
Active → Past due
Past due → Active
Past due → Canceled
Canceled → Reactivated
Then compare that graph with your automated tests.
Which edges are actually covered?
You may discover something uncomfortable.
You have excellent coverage of each page.
But weak coverage of the transitions that connect them.
This changes how I think about "critical paths"
Teams often say:
We automate our critical paths.
Usually they mean:
Signup
Login
Checkout
Create project
Invite user
Those are important.
But they're not really paths.
They're destinations.
A true critical path includes how users get into and out of important states.
For example, inviting a user isn't just:
Admin sends invitation
User accepts
It's also:
Admin sends invitation
Invitation expires
Admin resends
User accepts old link
User accepts new link
Admin changes role
User logs in with stale session
That's where "works in QA" turns into "why does this one customer have admin access?"
Better coverage can mean fewer tests
This is the counterintuitive part.
If you model transitions properly, you may not need more tests.
You may need better-selected tests.
Instead of 30 slight variations of:
Open page
Click button
Check confirmation
you might want 10 scenarios deliberately chosen to cross the risky boundaries in your system.
This is why I don't think test-count growth is necessarily a sign of a better automation program.
Sometimes it means the strategy is becoming more thorough.
Sometimes it means you're accumulating redundant scripts.
The number itself tells you almost nothing.
The question I'd ask instead of "How much is automated?"
I'd stop asking:
What percentage of QA is automated?
That's too vague.
I'd ask:
Which important state transitions can still reach production without being exercised automatically?
That's a much harder question.
It's also actionable.
Because once you know the answer, you can decide whether each gap deserves:
- a browser test
- an API test
- a contract test
- a component test
- a manual exploratory test
- or no test at all
Now you're designing a testing strategy.
Not just collecting automation.
The goal isn't maximum automation
It's useful to remember what all of this is for.
The goal isn't:
10,000 tests
The goal is:
Fewer important surprises in production
Those are not the same thing.
You can have a massive automation suite that repeatedly confirms the obvious.
Or a smaller one that aggressively probes the transitions where your system is most likely to break.
If I had to choose, I'd take the second one.
So before you add another 200 tests this quarter, try something simpler.
Draw the important states in your product.
Draw the transitions between them.
Then put your current automated coverage on top.
You might discover that the gaps aren't where you thought they were.
And you might also discover that your impressive test count has been measuring the wrong thing.
Top comments (0)