A flaky test passes sometimes and fails sometimes, with no change to the code. One flaky test is an annoyance. Twenty of them destroy trust in your pipeline: developers start clicking "re-run" by reflex, and real failures slip through because everyone assumes it's "just flakiness again."
After many years working on QA and release pipelines, I've found that most flaky tests come from a small number of root causes. Here are six practical fixes, roughly in the order you should try them.
1. Measure before you fix
You can't fix what you can't see. Before changing any tests, find out which tests are flaky and how often.
A simple approach: re-run your suite several times on the same commit and record which tests produce different results.
# Run the suite 10 times and keep each JUnit report
for i in $(seq 1 10); do
pytest tests/ --junitxml=reports/run-$i.xml || true
done
Any test that both passed and failed across those runs is flaky. Many CI platforms and test-reporting tools can also track this automatically over time. Rank the flaky tests by failure rate and start with the worst offenders.
2. Replace fixed sleeps with explicit waits
This is the single most common cause I see, especially in UI and API tests:
# Flaky: assumes the page loads within 2 seconds
time.sleep(2)
button = driver.find_element(By.ID, "submit")
On a busy CI runner, two seconds is sometimes not enough. Wait for a condition instead of a duration:
# Stable: waits up to 10 seconds, but continues as soon as the button is clickable
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
button = WebDriverWait(driver, 10).until(
EC.element_to_be_clickable((By.ID, "submit"))
)
The same principle applies to services: poll a health endpoint until it responds rather than sleeping for a fixed time after startup.
3. Make tests independent of each other
If test B only passes when test A runs first, you have hidden shared state. It usually shows up when tests run in parallel or in a different order.
Common culprits:
- Tests that share a database row or user account
- Global variables or singletons modified by one test
- Files written to a shared temporary folder
Fix: each test creates its own data and cleans it up. In pytest, fixtures are the natural tool:
import uuid
import pytest
@pytest.fixture
def new_customer(api_client):
customer = api_client.create_customer(name=f"test-{uuid.uuid4()}")
yield customer
api_client.delete_customer(customer.id)
A quick way to expose order dependencies is to run your suite in random order (for example with the pytest-randomly plugin) and see what breaks.
4. Control time, randomness and external services
Tests that depend on things outside your control will eventually fail:
- Time: a test that checks "created today" can fail when run near midnight. Freeze or inject the clock.
- Randomness: seed random generators so failures are reproducible.
- External APIs: a third-party sandbox that is slow or down will fail your build. Mock it in unit tests, and keep a small number of real integration tests in a separate stage.
5. Quarantine, don't ignore
When a flaky test can't be fixed immediately, move it into a quarantine group that still runs but does not block the pipeline:
@pytest.mark.quarantine
def test_payment_callback_timeout():
...
# Blocking stage: everything except quarantined tests
pytest -m "not quarantine"
# Non-blocking stage: quarantined tests, results reported but not enforced
pytest -m quarantine || true
Two rules make quarantine work:
- Every quarantined test has an owner and a ticket.
- The quarantine list is reviewed regularly. Quarantine is a waiting room, not a graveyard.
6. Treat automatic retries as a last resort
Retry plugins can make the pipeline green again quickly, but they hide the problem. If you do use retries, limit them (one retry at most), log every retry, and track retried tests as flaky so they still get fixed.
The payoff
Fixing flakiness is unglamorous work, but the return is large. A pipeline that developers trust gets faster feedback, fewer "re-run and hope" cycles, and real failures that get investigated instead of ignored.
Start small: measure this week, fix your top five flaky tests next week, and add a quarantine process so new flakiness never quietly piles up again.
Top comments (0)