When I joined a previous company, the test suite lived in its own repo, completely separate from the app and more or less invisible to the developers it was supposed to help. Running it was entirely a manual process. Someone would kick off the pipeline whenever a change needed testing, which meant deploying that change to a shared environment first. Since everyone was sharing the same environments, half-finished branches were constantly stepping on each other and muddying the results.
And the results themselves… only one QA engineer could see them. They'd pass failures along to the developers in what was basically a very slow game of telephone, so by the time anyone heard about a problem, they'd usually moved on to something else entirely. The suite was slow, it depended on a long list of things going right, and honestly, it was only reliable for that one person.
I don't say any of that to knock the people who built it. Every piece made sense to somebody when it was added. The problem was that, over time, nobody had ever stopped to ask a pretty basic question: who is this actually for?
In my experience, that's not unusual at all. Most teams I've worked with hold their customer-facing software to real standards. We care whether it's reliable, whether it's easy to use, whether it's fast, and whether we can maintain it. And of course we do. Happy customers pay the bills. Then we turn around and build test automation like a side project, something we keep adding to but rarely stop to actually design. We add tests out of habit, because automation is "good" and it's part of the software development lifecycle.
But a test suite has users too. It's the developer opening a PR at 4pm, the QA engineer triaging a nightly failure, the person deciding whether it's safe to ship, and the PM asking if we're good to release. And more and more, it's the AI agent writing code alongside your team, relying on your tests to tell it whether it broke something.
Every one of those decisions eventually reaches a real customer.
With that in mind, let's raise the bar together: your test suite is a product, and it's time we started treating it like one.
The Double Standard
Would you ship a checkout page that failed at random for one out of every seven customers? Of course not. But that's roughly the level of reliability plenty of teams accept from their tests, and I'm not just talking about small teams without the infrastructure to do better. Here's what some of the best engineering orgs in the world have published about their own suites:
- Google: about 1.5% of all test runs report a flaky result, and almost 16% of tests show some level of flakiness. When a test flips from pass to fail post-submit, a flaky test is involved about 84% of the time. In other words, most of their "failures" weren't real. (Google Testing Blog, summary)
- Atlassian: flakiness caused up to 21% of master build failures in the Jira Frontend repo, and the reruns cost more than 150,000 developer hours a year. (Atlassian Engineering)
- Slack: in July 2020, 56.76% of test jobs were failing. (Slack Engineering)
I imagine these companies have entire teams dedicated to automation and CI, and their suites still hit numbers none of them would ever tolerate from a feature. To be fair though, there's a lot of missing context behind those numbers, and these companies invested heavily to bring them down. But that's kind of the point. If it happens at their scale, with their resources, it's happening on most teams. I don't think the issue is effort or talent. I just think we just quietly hold our tests to a lower standard.
What Makes a Product "Good"?
Think about a product you genuinely love using. Maybe it's your banking app, your favorite music app, or the tool you'd fight your team to keep. Whatever it is, I'd bet it gets most of these things right:
- It's reliable. You trust what it tells you. When your banking app shows your balance, you don't refresh three times to make sure. The best products even protect you when things go wrong. Stripe, for example, lets you safely retry a payment request so a network hiccup never charges a customer twice.
- It's easy to use. You figured it out without a manual, and when something does go wrong, it tells you what happened and what to do next. The best products also seem to know what you need before you even think to ask.
- It's fast. Fast enough that you never think about speed at all. You only notice performance when it's missing.
- It's maintainable. Updates make it better without breaking the things you rely on.
- It's measured. The team behind it knows which features people use, where they get stuck, and what nobody touches.
- It's planned. They're deliberate about what to build next and just as deliberate about what to retire.
So Why Bother?
Now, fair question: why should a test suite get the same treatment?
Plenty of teams ship successful software with a messy suite. But that's a bit like driving for years with the check-engine light on. The car still gets you where you need to go. You've just stopped knowing when something's actually wrong, and you're quietly paying for it every day until the engine finally gives out.
We write automated tests for two reasons: to raise product quality and confidence, and to help the team develop and ship faster. Every time a suite falls short of either one, someone pays. Developers wait on slow runs, re-run flaky ones, or dig through vague failures. Real bugs slip past because nobody trusted the test results. And the decision to ship gets made on gut feel instead of real signal. Atlassian's 150,000 developer hours a year is what that bill looks like when someone finally adds it up. Maybe we don't have to live with that?
Applying Product Standards to Your Suite
So let's take those six qualities and hold your suite to every one of them.
Reliability: Can People Trust It?
If you had to prioritize one of these first, it would be this one, because every other quality depends on it. People will put up with a slow suite or clunky failure messages for a surprisingly long time. They won't put up with a suite they can't trust. They'll either stop running it, or worse, keep it around and quietly learn to ignore it. And a test suite nobody believes is just noise and busy work with an engineering and CI bill attached.
A reliable product does the same thing every time. For a test suite, that means the same code produces the same result on every run, on every machine, in any order. It sounds obvious, but most suites break that promise in lots of small ways.
The good news is that flakiness usually isn't mysterious. When researchers dug into 235 flaky UI tests from real projects, nearly half came down to timing, and most of the rest traced back to test logic, environment differences, or how tests interact with the page. (Romano et al., 2021) So the habits that make a suite reliable line up with those causes pretty directly:
- Tests wait on real signals. A network request finishing, an element becoming visible, a toast appearing. Never a hard-coded sleep or hope. Timing is the single biggest source of flakiness, so invest in this one and you've hardened a big chunk of your suite.
- Every test is data-independent. Whether it creates its own data, uses a seeded fixture, or resets state before it runs, it never depends on what another test left behind, and it passes no matter what order the specs run in.
- The environment is boring and consistent. Tests run the same way everywhere, locally and in CI. And every run makes it obvious exactly what was tested: which commit, which environment, which configuration. When a test fails, nobody should have to guess what code it actually ran against.
- Tests find elements the same way every time. Stable, dedicated selectors that follow one explicit pattern make tests predictable to read and troubleshoot, and give AI a clear pattern to follow instead of guessing.
- A flaky test is a bug, not a "quirk". It gets an owner and a ticket, and it gets fixed or quarantined that week or day. It doesn't get retried until it passes, and it doesn't get ignored.
That last habit matters most. The moment a team learns that the answer to a failing suite is "just re-run it," you've undermined the reliability of the entire suite, and people start taking the results for granted. Don't train your team to expect noise. Flaky tests are worse than having no tests at all for this very reason.
Also notice that none of these habits depend on your framework. Cypress, Playwright, Selenium, whatever comes next: the tool doesn't make a suite reliable. The practices do. And that's the part most teams skip. Writing an automated test is easy, and with AI it's easier than ever. Writing a good one is hard. If your team can't explain why one test is reliable and another isn't, you can keep adding tests, but you won't be able to scale them with any confidence. Invest in the practices first. The tests get easier after that.
Usability: Can People Actually Use It?
This is where I think most suites struggle to scale, because it's easy to forget that a suite has three completely different user experiences. Someone has to write tests, someone has to run them, and someone has to figure out what went wrong when one fails. At small scale almost anything works. It's only as the suite grows that the friction starts to compound. (I covered this as Phase 3 of the Automation Maturity Pyramid, if you want the longer version.)
Writing a test. Could a developer who has never touched the suite add a test for their feature in an afternoon? That mostly comes down to structure. Every spec follows the same pattern, and honestly, which pattern you pick matters less than picking one. Any type of code organization beats spaghetti-code chaos, and you can refine it over time. The boring parts, like logging in, seeding data, or navigating to a page, live in reusable helpers. There's one selector convention, not five, and a handful of good example tests people can copy from. Add some light docs on the common gotchas too. We all know writing docs sucks, but ask yourself what past you would have wanted on day one. Those same consistent patterns are also what let AI tooling generate tests that actually fit your suite. If the only way to write a new test is to ask the QA engineer, that's a long-term usability bug.
Running a test. Running the whole suite locally should be one command, and running a single spec should be just as easy. No secret environment variables, no "oh, you need to start these three services first," no Confluence page last updated in 2021. Tests should also run as close to the change as possible, ideally pre-merge, so the person who caused a failure is the one who sees it while the context is still fresh. If developers can't easily run tests themselves, they won't, and they'll find out about failures from CI an hour after they've moved on.
Debugging a failure. This is the big one, because failures will happen, and how quickly someone can respond makes or breaks a suite's value. A good failure tells you what the test was trying to do, what it expected, what actually happened, and where to look next. Screenshots, video, network logs, and the app's own errors should be one click away from the red build. Readability matters just as much. Smaller, focused tests are far easier to diagnose than one giant test that checks twelve things, so choose clarity over cleverness and always ask, "will this make sense to me in six months?" (I dug into this more in Optimize Debugging Automation Tests.)
And the tooling only works if the culture backs it up. Build a habit of immediate triage, where every failure gets looked at with the same urgency, so nothing sits red long enough to become normal.
Performance: How Fast Is the Feedback?
Feedback speed quietly decides how people use your suite. When a run takes five minutes, developers wait for it. When it takes forty, they open Slack, pick up another ticket, and come back two hours later having half forgotten what they changed. By the time the results show up, nobody cares anymore.
The tricky part is that suites rarely get slow all at once. It's death by a thousand small inefficiencies. One extra second per test doesn't sound like much, until you multiply it across 300 tests and realize you just added five minutes to every run. So making a suite fast is rarely one big fix. It's a lot of small product decisions:
- Test at the right level. Not everything needs to be an end-to-end test. If a component or API test can prove the same thing in a fraction of the time, that's usually the better call.
- Set up state through the back door. Only one test needs to prove the signup form works. The other two hundred can start already logged in, with their data seeded through an API instead of clicked together through the UI.
- Parallelize, and balance it. Splitting across machines only helps if the shards are even, because your pipeline is only as fast as its slowest spec.
- Get the cheap, relevant signal first. Run linting, type checks, and quick static checks before you ever spin up a browser, and let the tests covering what just changed report before anything else. If something is knowable in one second, don't pay fifteen minutes to learn it.
- Trim the pipeline, not just the tests. Cache dependencies between runs, and only keep videos and screenshots for tests that fail. A surprising amount of "test time" is really setup and teardown around the tests.
- Measure, and watch the trend. Know which tests are your slowest and start with the outliers. Then keep an eye on total runtime, because it creeps up a few seconds at a time until one day it's thirty minutes and nobody can say when that happened.
If you're a Cypress user, I went much deeper on the specifics in Your Cypress Tests Are Slower Than You Think.
Maintainability: What Does It Cost to Keep?
Every test you write is something the team has to support for as long as it exists. Plenty of tests are absolutely worth that. Some aren't.
Google's testing book has the best definition of this I've found. It's written about unit tests, but it applies even more to end-to-end tests, since they touch so much more of the system: the ideal test only changes when the requirements change. Refactoring the code, adding a new feature, or fixing a bug shouldn't require touching existing tests. Only a genuine change in behavior should. (Software Engineering at Google) That's a high bar, but it's the right one to aim for, and it mostly comes down to three things.
How tests are written.
- Test user behavior, not implementation. Assert what a user would actually see or do, not the internal structure that happens to produce it today.
- Keep one source of truth. Selectors, common flows, and test data each live in one place, whether that's helpers, page objects, or data factories, so a UI change is a one-file fix instead of a twelve-file hunt.
- Stay readable over clever. Share logic, but don't abstract so aggressively that nobody can read a test top to bottom. A little duplication is fine if it makes a test clearer.
Where tests live.
- Tests live with the code. Same repo, updated in the same PR as the feature, ideally by the person making the change. A suite that lives somewhere else, maintained by someone else, will always be one step behind the app.
- Test code gets reviewed like app code. Not rubber-stamped because "it's just a test," but reviewed with the same care and standards: Is it readable? Does it follow the patterns? Will it hold up when the UI changes? Is it performant?
How the suite stays healthy.
- Automate the guardrails. Linting and static checks for test code catch problems at PR time instead of in a failed run. My selector check is a good example: it catches renamed or deleted selectors in about a second.
- Prune regularly. Deleting tests is maintenance too. Duplicate coverage, tests for features that no longer exist, and assertions that can't fail anymore all cost time without protecting anything. A regular audit, even an AI-assisted one, keeps the suite honest.
- Keep dependencies current. Upgrading your test framework should be a routine chore.
Analytics: Where Do We Focus?
Product teams know which features people use, which ones nobody touches, and where users get stuck.
That matters because analytics is what ties everything else together. Reliability, usability, performance, and maintainability are all great goals, but you can't improve what you can't see. Without data, you end up fixing whatever annoyed you most recently instead of whatever is actually costing the team the most. With it, you can point your time at the real pain.
You don't need a fancy platform to start. A few simple questions, each mapped to the areas above, go a long way:
- Reliability: which tests fail the most, and why? Tag every failure as a real bug, a flaky test, or a broken test. After a month, the pattern usually jumps right out, and your flake rate tells you how much people can actually trust a red build.
- Usability: how long does it take to write a new test, or to figure out why one failed? If adding coverage for a simple feature takes a day, or failures routinely sit for hours before anyone knows what went wrong, something is making the suite hard to work with. Ask the people using it where they get stuck. Your engineers are your user research, and they'll usually tell you exactly what's painful if you ask.
- Performance: which tests are the slowest, and is total runtime creeping up? Knowing your slowest tests tells you exactly where to start, and tracking the trend means you'll catch a slowdown the week it happens, not six months later.
- Maintainability: how much of your time goes into keeping existing tests alive? If more effort goes into updating, fixing, and refactoring old tests than into adding new coverage, the suite is costing more than it should.
- Value: what is the suite actually catching, and what can't it see? Which bugs did automation catch before release, and which ones slipped through?
Coverage can be part of this too, but treat it as directional, not definitive. A high number tells you code was executed, not that anything meaningful was checked. And whatever you track, use the numbers to guide investment, not to assign blame. The moment metrics become a way to punish people, the data stops being honest. If you can put the trends on a dashboard, even better. Bar graphs are fun, line graphs always look convincing, and don't even threaten me with a good time by bringing up pie charts.
Roadmap: What Should We Work On Next?
A healthy suite has a real roadmap, with a dedicated backlog of automation improvements. Not a graveyard, but an honest list of where the suite falls short today and a plan to close those gaps one at a time. Everything from the analytics above feeds into it: the flakiest tests, the slowest specs, the failures that take too long to debug, the brittle areas that keep breaking, and the blind spots you wrote down.
Then treat those tickets exactly like product tickets. Give them clear acceptance criteria, size them, prioritize them against everything else, and spike the ones you don't fully understand yet. Keep adding new coverage, because that still matters, but carve out some time every sprint or cycle for improvement work too. It might be tooling that makes writing and debugging easier, speed work, reliability fixes, monitoring for the things a pre-merge suite can't see, or deleting tests that cost more than they're worth. A little every cycle beats a big cleanup project that never quite gets scheduled.
AI makes this much easier than it used to be. I have an automated job that files tickets every Monday for our flakiest and most-regressed tests, and once a month I have AI review the entire suite and write up improvement tickets based on what it finds in performance, usability, and maintainability. I review every one, approve the ones that make sense, and they go straight into the backlog like any other work. The suite basically tells me what it needs at this point.
What This Looks Like in Practice
Fair question at this point: does any of this actually work, or does it just sound nice on paper?
Here's where things stand for me today. I'm the only QA engineer at my company, and our suite runs about 650 end-to-end tests on every PR in roughly 12 minutes, with a flake rate under 1%. A visual suite runs weekly on top of that.
I'm proud of that, but I'm not sharing it to brag. The point is that one person can only realistically support a suite that size if it's built like a product rather than maintained like a chore. Every habit in this post is something I lean on daily. Every failure gets treated the same, because a test that's only sometimes reliable isn't useful. Readability is non-negotiable, especially now that I use AI to help write, debug, and plan tests, because consistent patterns are what keep its output consistent too. And the suite has its own backlog, fed by real data and real feedback from the developers who use it.
As I wrote in How to Be a 10x Engineer, that stability isn't a trophy, it's a budget. Every hour I'm not spending firefighting the suite is an hour I get to put into full-stack work, DevOps, code reviews, and customer support. And coffee. Gotta love coffee time.
Final Thoughts
So what happened to that hidden suite at my previous company?
Funny enough, the first thing I did wasn't write a single test. I moved the whole suite into the app repo and wired it into CI/CD so the existing tests ran on every push. For the first time, every developer could see how the tests were doing against their own work, without waiting on a message from someone else. The first runs were a disaster, an absolute massacre of test results, but it was exciting nonetheless.
Then we held a team meeting about it. I half expected some pushback, but people were genuinely excited. The suite had been gated off for so long that just being invited into the conversation and getting to have an opinion felt like an opportunity. From there, we started triaging the performance and reliability problems one by one until the tests were something people could trust long term. It took a lot of hand-holding at the start, and that's okay.
The most important fix wasn't more tests. It was giving the suite actual users. Once people could see it, question it, and care about it, everything else finally had a reason to get better.
If you're not sure where to start with your own suite, start the same way: get people in a room. Grab a couple of developers, maybe someone from product, and ask someone who doesn't normally work in the suite to write a small test, run it locally, and debug a failure. Watch where they get stuck. Then ask the uncomfortable questions. Is the first instinct on a red build to investigate or to re-run? Does everyone know what a green build actually covers? What can't the suite see at all? Whatever you can't answer, or don't love the answer to, goes straight into the backlog. One honest conversation like that is a far better plan than "write more tests."
That's really the whole point of this post. A mature automation strategy optimizes for usefulness over test count: signals people trust, failures people can act on, and feedback that shows up while it still matters.
Your suite already has users, a job to do, and a cost to keep around. It's a product whether we treat it like one or not, so we might as well start.
With that in mind, as always, happy testing.
Top comments (0)