Retries made my Playwright jobs worse, not better: failure rate went from 3.1% to 6.7% in one week. The culprit wasn't flaky code — it was retrying things that should never be retried.
I found this the hard way after wiring a blanket retry wrapper around every step of my account-automation pipeline. Here's what broke and what actually fixed it.
1. Retries turned "already done" into "done twice"
The worst failure class had nothing to do with the network. A step would time out on the response, my wrapper would retry, and the step would run a second time — because the first attempt had actually succeeded server-side.
In two days I double-posted to the same account twice and sent a duplicate signup to one platform. No exception, no error log. Just silent duplicates.
# WRONG: retry on any exception, no idempotency key
def step(action):
for _ in range(3):
try:
return action()
except Exception:
time.sleep(2 ** _)
raise
2. The rule I actually needed: retry only idempotent actions
I split every step into two buckets:
- Safe to retry — reads, GETs, checks. Retry freely.
- Unsafe to retry — writes, posts, sends. Retry only with an idempotency key, or not at all.
The write steps got an explicit key (account_id + action + date) so a replay would be rejected instead of duplicated.
3. Backoff without jitter synchronized my retries
Every failed job retried on the same exponential schedule, so a brief API hiccup would make 20 workers hammer the endpoint at the exact same second. Adding jitter — a random 0–500 ms — spread them out and cut the second-order failures by more than half.
time.sleep(2 ** attempt + random.uniform(0, 0.5))
4. Numbers after the fix
| Metric | Before | After |
|---|---|---|
| Failure rate | 6.7% | 2.9% |
| Duplicate writes | 5 in 2 days | 0 in 7 days |
| Retry storms | ~2/day | 0 |
The change that mattered most wasn't the retry count or the backoff curve. It was deciding what deserves a retry at all.
I kept the idempotency-key helper in the shared utils module — it's now the default for every write step in the pipeline. Next I'm applying the same split to the signup flow, which still has one hot spot where a captcha timeout re-submits the whole form.
What's the one place in your automation you retry blindly?
Top comments (1)
The thing about AI is, it’ll never tell you your product is good to go no matter how perfect it is. It’ll always find issues that lead to a bunch of other issues when fixed and ends up going in an endless loop of fixes. That’s why I do the testing myself nowadays