DEV Community

Stephano kambeta
Stephano kambeta

Posted on

Retry with Exponential Backoff in Python, Done Properly

A script that creates invoices runs at 2 a.m. and calls an API. For about thirty seconds, the API returns errors.

Without retries, tonight's invoices never go out. With careless retries, a customer might get the same invoice twice.

Both are real problems, and they pull in opposite directions. Retrying sounds simple, but doing it properly means deciding what to retry, how long to wait, when to stop, and what happens if a retry causes a duplicate.

This guide builds a small helper that handles all of that. It's under 90 lines, uses only what Python already includes, and works with any API you can reach with a normal web request.

What a good retry does

Some no-code tools retry failed steps for you. A script doesn't, so it has to do that work itself. (I compared the two approaches in Zapier vs GitHub Actions for small business workflows.)

A good retry follows five rules:

  1. It retries only temporary problems. Some errors go away on their own. Others never will.
  2. It waits longer after each failure. This gives a struggling server room to recover.
  3. It adds a little randomness. That stops many scripts from retrying at the exact same moment.
  4. It listens to the server. When a service says "come back in 30 seconds," it waits 30 seconds.
  5. It knows when to stop. Retries are for blips, not for outages.

There's a sixth rule that matters just as much: don't retry something that could happen twice, unless you've made it safe. We'll get to that.

Which errors are worth retrying?

Not every failure deserves another try. Sending the same broken request five times just gives you five errors.

Problem Retry? Why
Timeout, connection reset Yes Network blips are usually brief
429 Too Many Requests Yes, after waiting You're going too fast, and the server may say how long to wait
500, 502, 503, 504 Yes The server is having trouble, often only briefly
408 Request Timeout Yes The server gave up waiting for you
400, 422 No The request itself is wrong, so it will fail the same way again
401, 403 No A wrong password or missing permission won't fix itself
404 No The thing isn't there
Certificate error No It won't fix itself either

A 500 error is sometimes caused by something in your request, which no amount of retrying will fix. That's one more reason to limit the number of attempts.

Where retries help most

A few places where this pays off quickly:

  • Calls to an API. Reading data, saving results or posting updates. Most outages last seconds.
  • Website checks. One failed check is often a blip. Retrying once or twice before raising an alarm cuts false alerts. I built a check like this in my Python uptime monitor guide.
  • Posting to a platform. If you've followed my Bluesky automation bot guide, wrapping the posting call in this helper is a good first use. Keep duplicates in mind (more on that below), since a double post is visible to everyone.

Step 1: Add the helper

Create a file called retry.py:

"""Retry temporary failures with exponential backoff, and nothing else."""
import email.utils
import logging
import random
import ssl
import time
import urllib.error
import urllib.request
from datetime import datetime, timezone

log = logging.getLogger("retry")

RETRY_CODES = {408, 425, 429, 500, 502, 503, 504}
IDEMPOTENT_METHODS = {"GET", "HEAD", "PUT", "DELETE"}  # doing these twice is harmless


def is_temporary(error):
    """Is this the kind of failure that might not happen again?"""
    if isinstance(error, urllib.error.HTTPError):
        return error.code in RETRY_CODES
    if isinstance(error, urllib.error.URLError):
        return not isinstance(error.reason, ssl.SSLCertVerificationError)
    return isinstance(error, (TimeoutError, ConnectionError))


def server_wait(error):
    """How many seconds the server asked us to wait (Retry-After), if it said."""
    value = None
    if isinstance(error, urllib.error.HTTPError) and error.headers:
        value = error.headers.get("Retry-After")
    if not value:
        return 0.0
    if value.strip().isdigit():
        return float(value)
    try:
        date = email.utils.parsedate_to_datetime(value)
        return max(0.0, (date - datetime.now(timezone.utc)).total_seconds())
    except (TypeError, ValueError):
        return 0.0


def backoff(attempt, base_delay, max_delay):
    """Wait about base_delay at first, double it each time, and never wait more than max_delay."""
    cap = min(max_delay, base_delay * 2 ** (attempt - 1))
    return cap / 2 + random.uniform(0, cap / 2)  # a little randomness, so clients don't all retry together


def retry(action, attempts=5, base_delay=1.0, max_delay=30.0, total_timeout=120.0):
    """Call action() until it works. Gives up on permanent errors, too many attempts, or too much time."""
    started = time.monotonic()
    for attempt in range(1, attempts + 1):
        try:
            result = action()
        except Exception as error:
            if not is_temporary(error):
                raise  # a 404 or a wrong password won't fix itself
            if attempt == attempts:
                if attempts > 1:
                    log.error("Giving up after %d attempts: %s", attempts, error)
                raise
            wait = max(backoff(attempt, base_delay, max_delay), server_wait(error))
            if time.monotonic() - started + wait > total_timeout:
                log.error("Giving up, because waiting again would pass %s seconds: %s", total_timeout, error)
                raise
            log.warning("Attempt %d of %d failed (%s). Retrying in %.1f seconds", attempt, attempts, error, wait)
            time.sleep(wait)
        else:
            if attempt > 1:
                log.info("Succeeded on attempt %d", attempt)
            return result


def request(url, data=None, method=None, headers=None, idempotency_key=None, timeout=15, **retry_options):
    """Make an HTTP request and retry temporary failures, when it's safe to do so."""
    method = method or ("POST" if data is not None else "GET")
    headers = {"User-Agent": "retry-demo/1.0", **(headers or {})}
    if idempotency_key:
        headers["Idempotency-Key"] = idempotency_key
    if method not in IDEMPOTENT_METHODS and not idempotency_key:
        retry_options["attempts"] = 1  # it may have worked even though we saw an error

    def action():
        req = urllib.request.Request(url, data=data, method=method, headers=headers)
        with urllib.request.urlopen(req, timeout=timeout) as response:
            return response.read()

    return retry(action, **retry_options)
Enter fullscreen mode Exit fullscreen mode

Here's what each part does:

  • is_temporary() applies the table above. It says yes to timeouts, dropped connections and the temporary status codes, and no to everything else.
  • server_wait() reads the Retry-After header, which a server can send to say how long to wait. It can be a number of seconds or a date, and this handles both.
  • backoff() works out the wait. The first retry waits about a second, the next about two, then four, then eight, up to a limit. The wait is picked randomly from the upper half of that range, so scripts don't move in lockstep.
  • retry() is the loop. It stops when the error is permanent, the attempts run out, or waiting again would pass the total time limit. When it gives up, it raises the original error, so nothing is hidden.
  • request() is the web request itself. It sets a timeout (more on that below) and decides whether retrying is safe.

The defaults are 5 attempts, waits that top out at 30 seconds, and a total limit of 2 minutes. You can change any of them: request(url, attempts=3, max_delay=10).

Step 2: Use it

Here's how it looks in a real script:

import json
import logging

from retry import request

logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")

# Reading data is safe to repeat
report = request("https://api.example.com/reports/latest")

# Creating something needs more care (see the next section)
order = request(
    "https://api.example.com/orders",
    data=json.dumps({"customer": "A. Smith", "total": 120}).encode(),
    headers={"Content-Type": "application/json"},
    idempotency_key="invoice-2026-10-06-1042",
)
Enter fullscreen mode Exit fullscreen mode

The first call is a read. If the API hiccups, the helper quietly tries again.

The second call creates something, so it carries an idempotency_key. That word needs an explanation.

The duplicate problem

Here's how a retry can go wrong:

  1. Your script sends "create invoice 1042".
  2. The server creates the invoice.
  3. The reply gets lost on the way back, so your script sees a timeout.
  4. Your script retries, and the server creates the invoice again.

Now there are two invoices, and the script thought it was being careful.

An action is idempotent if doing it twice has the same effect as doing it once. Reading a page is. Setting a customer's email address to a certain value is. Creating a new invoice isn't.

The fix is an idempotency key: a unique label that you attach to the request. If the server sees the same key twice, it knows it's a repeat and returns the original result instead of creating another invoice. Many payment and billing APIs support this, Stripe being a well-known example. The header is often called Idempotency-Key, but check your API's documentation.

Two rules for using it:

  • Build the key from something stable. An invoice number works well. A random value generated on every attempt defeats the purpose. A stable key also protects you if the whole script is run again later.
  • Only use it if the API supports it. An API that doesn't know about the header will ignore it, and you'll get duplicates anyway.

The helper follows these rules. For a request like POST that isn't safe to repeat, it makes only one attempt unless you give it a key. If the API has no idempotency support, the options are to look up the item before creating it, or to try once and handle a failure by hand.

The same problem exists on the receiving side. Services that send you webhooks often retry when they don't get a quick reply, so the same event can arrive twice. Handling that is the same idea in reverse: recognize what you've already processed.

How long should it keep trying?

Retries are for blips, not for outages.

With the defaults, the helper waits roughly 1, 2, 4 and 8 seconds across five attempts, so about 8 to 15 seconds in total. That's plenty for a short hiccup. If the service is still failing after that, it's probably down for a while.

At that point, let the script fail. If it runs on a schedule, the next run is a fresh attempt.

Long waits cost something, too. A script that's sleeping is still running, and on GitHub Actions the minutes keep counting. If you use a private repository, keep your waiting budget small.

One more thing: set a timeout on every request. Python's urlopen has no timeout by default, so a connection that hangs can wait forever. The helper sets 15 seconds, and a timeout counts as a temporary error that gets retried.

Step 3: Try it on purpose

The best way to trust a retry helper is to watch it work. Create a file called flaky_server.py, a tiny web server that fails twice and then succeeds, over and over:

from http.server import BaseHTTPRequestHandler, HTTPServer

count = 0


class Flaky(BaseHTTPRequestHandler):
    def do_GET(self):
        global count
        count += 1
        code = 200 if count % 3 == 0 else 503  # two failures, then a success, over and over
        self.send_response(code)
        self.end_headers()
        self.wfile.write(b"hello" if code == 200 else b"try again later")


HTTPServer(("127.0.0.1", 8000), Flaky).serve_forever()
Enter fullscreen mode Exit fullscreen mode

Then create demo.py:

import logging
from retry import request

logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
print(request("http://127.0.0.1:8000/"))
Enter fullscreen mode Exit fullscreen mode

Start the server in one terminal with python3 flaky_server.py, then run python3 demo.py in another. You should see something like this:

2026-10-06 22:09:34,942 WARNING Attempt 1 of 5 failed (HTTP Error 503: Service Unavailable). Retrying in 0.9 seconds
2026-10-06 22:09:35,852 WARNING Attempt 2 of 5 failed (HTTP Error 503: Service Unavailable). Retrying in 1.4 seconds
2026-10-06 22:09:37,213 INFO Succeeded on attempt 3
b'hello'
Enter fullscreen mode Exit fullscreen mode

The waits grow from one attempt to the next and aren't identical each time you run it. Try a few experiments: change the server to always return 404 and watch the helper give up at once, or set attempts=2 and see it fail faster.

Don't hide the failure

A retry that works leaves no mark, and that can be a problem. A script that needs three attempts every night looks perfectly healthy while the API gets steadily worse.

The helper logs every retry, so read your logs now and then. A lot of "Succeeded on attempt 3" lines is an early warning.

When retries run out, the error should reach something that tells you. The helper raises the original error, so it fails the run and your alerts can fire. If you haven't set those up yet, my guide on monitoring your automations is a good place to start.

What can go wrong

Retrying things that aren't safe. This is how duplicates happen. Use idempotency keys, or don't retry.

Retrying forever. Cap both the attempts and the total time.

Everyone retrying at once. If many scripts fail together and retry after exactly the same delay, they hit the server together again. The randomness in backoff() spreads them out.

Ignoring Retry-After. When a server tells you how long to wait, believe it. Some services may block clients that keep hammering them.

Swallowing the final error. A helper that returns None after giving up will make a failure look like success. This one raises instead.

Retrying a bug. If your request is wrong, retrying won't help. That's why 4xx errors are never retried.

Quick answers

Why exponential instead of a fixed delay?
A fixed delay is either too short for a real outage or too slow for a blip. Doubling gives you a fast first retry and a patient later one.

Why add randomness?
Without it, scripts that fail together retry together. A little randomness spreads the load.

How many attempts should I use?
Three to five is typical. More than that rarely helps and just makes slow failures slower.

Should I use a library instead?
It's a fair choice. If you already use requests, its underlying urllib3 library has a built-in retry option, and tenacity is a popular general-purpose package. This helper is useful when you want no dependencies, or when you want to understand what's happening.

Does this work for things other than HTTP?
The retry() function works with any action, such as a database call or a file download. Only is_temporary() is specific to web requests, so adapt that part.

The short version

Retrying is easy to get wrong in both directions. Too little, and a thirty-second blip costs you a night's work. Too much, and a small problem turns into duplicates or a blocked account.

Retry only temporary errors, wait a bit longer each time, add some randomness, respect what the server says, and know when to stop.

You can find more practical automation guides on Procwire.

Top comments (0)