I Built a Tool to Find the Tests Developers Have Stopped Trusting
A test fails in CI.
You look at it.
You run it again.
It passes.
You run it again.
It passes.
You rerun the pipeline.
Everything is green.
So you merge the pull request.
A few hours later, the same test fails again.
At some point, the team stops treating that failure as useful information.
That's the real problem with flaky tests.
They don't just make CI unreliable.
They make developers stop trusting CI.
The problem isn't writing tests
Modern JavaScript projects already have excellent testing tools.
Playwright is great for browser testing.
Vitest and Jest are great for unit and integration testing.
CI platforms can execute thousands of tests.
The problem starts after the tests exist.
A growing project eventually accumulates:
- Tests that fail intermittently
- Tests that take too long
- Tests that haven't failed in months
- Tests that fail only in CI
- Tests that fail because of timing
- Tests that produce almost identical errors
- Tests that nobody wants to touch
The test suite becomes a second system that needs maintenance.
And most teams don't have a good way to measure its health.
So I built TestOps Kit
The idea is simple:
Don't replace the test runner. Analyze what it produces.
TestOps Kit consumes machine-readable test results and builds a reliability layer around them.
The workflow looks like this:
Playwright / Vitest / Jest
↓
Test results
↓
TestOps Kit
↓
┌──────────┼───────────┐
↓ ↓ ↓
History Reliability Performance
↓ ↓ ↓
Flaky Quarantine Slow tests
tests
↓
Dashboard
The existing testing workflow stays intact.
TestOps simply adds another layer of information.
The first feature I wanted was flaky-test detection
A single failed test doesn't necessarily mean a test is flaky.
You need history.
For example:
Run 1 → PASS
Run 2 → PASS
Run 3 → FAIL
Run 4 → PASS
Run 5 → FAIL
That pattern is much more interesting than a single failure.
TestOps tracks test behavior across runs and calculates reliability metrics from the available history.
The goal isn't to say:
"This test is definitely broken."
The goal is to say:
"This test has demonstrated inconsistent behavior and deserves investigation."
That's a much more useful signal.
Then came quarantine
Once you identify unreliable tests, the next problem is operational.
What do you do with them?
TestOps can generate a quarantine manifest based on a configurable flakiness threshold.
For example:
Flaky tests
checkout/payment.spec.ts 42%
account/login.spec.ts 31%
orders/create.spec.ts 27%
Instead of relying on someone's memory, the team now has a concrete list.
The important part is that quarantine is treated as a temporary reliability workflow, not a way to permanently hide failures.
Slow tests are another form of test debt
A test doesn't have to fail to become expensive.
Imagine a suite containing:
1,200 tests
and a handful of tests account for a significant portion of the runtime.
Those tests affect every developer.
Every pull request.
Every CI run.
Every deployment.
So TestOps also tracks duration and highlights slow tests.
The goal is straightforward:
Find the tests that are costing the team time.
Visual regression belongs in the same conversation
Testing isn't only about pass and fail.
UI changes can also introduce unexpected regressions.
TestOps includes deterministic snapshot checking so a CI pipeline can detect changed snapshot content.
For example:
npx testops snapshot snapshots --strict
If a snapshot changes unexpectedly, the command returns a failure.
That makes it possible to use the same reliability workflow in CI.
I also wanted a boring CLI
Developer tools don't need complicated installation procedures.
The basic workflow should be something like:
npx testops analyze --input test-results.json
Then:
npx testops report
And you get a dashboard.
The dashboard focuses on the questions developers actually care about:
How many tests do we have?
How many are failing?
Which ones are flaky?
Which ones are slow?
What happened in previous runs?
Testing the tester
One of the most important parts of building a developer tool is testing it against something real.
A synthetic demo can prove that the software works in theory.
It doesn't prove that developers can use it in their projects.
So the validation workflow is:
Existing project
↓
Existing tests
↓
Generate real test report
↓
Feed report into TestOps
↓
Compare metrics
↓
Create intentional flaky test
↓
Run repeatedly
↓
Verify detection
↓
Create slow test
↓
Verify performance detection
↓
Modify snapshot
↓
Verify regression detection
↓
Run inside CI
This is the test that matters.
Not whether the landing page looks good.
Not whether the dashboard has impressive charts.
Whether it can survive a real repository.
What I learned
The interesting part of test infrastructure isn't necessarily generating more tests.
It's understanding the tests you already have.
A team can have 5,000 tests and still have a reliability problem.
More tests don't automatically mean more confidence.
Sometimes the real question is:
Which tests can we trust?
That's the problem TestOps Kit is designed around.
The goal
The long-term idea is bigger than a CLI.
A mature version could understand:
- Test history
- CI environments
- Failure patterns
- Code changes
- Test ownership
- Dependency relationships
- Visual changes
- Performance regressions
- Failure clusters
Eventually, the system could answer questions like:
Which tests became unreliable after this deployment?
Or:
Which changed files are responsible for most of the tests we need to run?
Or:
Which tests have consumed the most CI time this month?
But the first version starts with something much simpler:
Measure test reliability instead of guessing about it.
Because once developers stop trusting their test suite, the test suite has already become technical debt.
And that's a problem worth measuring.
Top comments (0)