DEV Community

Tech Auto Lab
Tech Auto Lab

Posted on

I Added Retries to My Automation and It Doubled the Failures

Retries made my Playwright jobs worse, not better: failure rate went from 3.1% to 6.7% in one week. The culprit wasn't flaky code — it was retrying things that should never be retried.

I found this the hard way after wiring a blanket retry wrapper around every step of my account-automation pipeline. Here's what broke and what actually fixed it.

1. Retries turned "already done" into "done twice"

The worst failure class had nothing to do with the network. A step would time out on the response, my wrapper would retry, and the step would run a second time — because the first attempt had actually succeeded server-side.

In two days I double-posted to the same account twice and sent a duplicate signup to one platform. No exception, no error log. Just silent duplicates.

# WRONG: retry on any exception, no idempotency key
def step(action):
    for _ in range(3):
        try:
            return action()
        except Exception:
            time.sleep(2 ** _)
    raise
Enter fullscreen mode Exit fullscreen mode

2. The rule I actually needed: retry only idempotent actions

I split every step into two buckets:

  1. Safe to retry — reads, GETs, checks. Retry freely.
  2. Unsafe to retry — writes, posts, sends. Retry only with an idempotency key, or not at all.

The write steps got an explicit key (account_id + action + date) so a replay would be rejected instead of duplicated.

3. Backoff without jitter synchronized my retries

Every failed job retried on the same exponential schedule, so a brief API hiccup would make 20 workers hammer the endpoint at the exact same second. Adding jitter — a random 0–500 ms — spread them out and cut the second-order failures by more than half.

time.sleep(2 ** attempt + random.uniform(0, 0.5))
Enter fullscreen mode Exit fullscreen mode

4. Numbers after the fix

Metric Before After
Failure rate 6.7% 2.9%
Duplicate writes 5 in 2 days 0 in 7 days
Retry storms ~2/day 0

The change that mattered most wasn't the retry count or the backoff curve. It was deciding what deserves a retry at all.

I kept the idempotency-key helper in the shared utils module — it's now the default for every write step in the pipeline. Next I'm applying the same split to the signup flow, which still has one hot spot where a captcha timeout re-submits the whole form.

What's the one place in your automation you retry blindly?

Top comments (1)

Collapse
 
moekhair profile image
Moe Khair •

The thing about AI is, it’ll never tell you your product is good to go no matter how perfect it is. It’ll always find issues that lead to a bunch of other issues when fixed and ends up going in an endless loop of fixes. That’s why I do the testing myself nowadays