DEV Community

mattleeee
mattleeee

Posted on Originally published at hkcode.dpdns.org

Fail-Fast Retry Design for Scheduled Scripts: Three Attempts with Backoff

Unattended scripts fail for the least interesting reasons. A DNS lookup times out at 02:00. An API returns 429 because a cron job on another box happened to fire at the same second. A TCP connection resets halfway through a response. None of these are bugs in your logic, and all of them will page you if the job is important enough.

The instinct is to wrap everything in a try/except and loop until it works. That instinct produces a worse failure mode: jobs that hang forever, hide real errors, and leave the scheduler unable to tell whether the previous run finished. This post describes a retry wrapper that stays small, fails fast, and gives the Windows Task Scheduler layer enough information to act.

The failure modes worth designing for

Before writing the decorator, it helps to be explicit about what counts as retryable. In practice, the set is narrow:

  • Network-level errors: connection reset, timeout, DNS failure.
  • HTTP 5xx responses.
  • HTTP 429 with a Retry-After header or a known rate limit.
  • Transient file locks on a shared drive (common on Windows when another process holds a handle).

Everything else should fail immediately. A 401 means your token expired — retrying three times just wastes thirty seconds before the same error. A 400 means your payload is malformed — retrying will never fix it. A KeyError in your parsing code means the upstream schema changed, and no amount of backoff will help.

The practical rule: retry on exceptions that represent transport problems, not logic problems. In Python, that usually means requests.RequestException and its subclasses, plus a small allowlist of OS-level errors.

A decorator that does one thing

Here is the core implementation. It retries up to three times with linear backoff — 10, 20, then 30 seconds — and logs each attempt with its failure reason.

import functools
import logging
import time

import requests

logger = logging.getLogger(__name__)

RETRYABLE = (requests.RequestException,)


def retry(max_attempts: int = 3, base_delay: int = 10):
    """Retry a function on transport errors with linear backoff.

    Delays are base_delay * attempt: 10s, 20s, 30s for the default.
    KeyboardInterrupt and SystemExit propagate untouched.
    """
    def decorator(func):
        @functools.wraps(func)
        def wrapper(*args, **kwargs):
            last_exc = None
            for attempt in range(1, max_attempts + 1):
                try:
                    return func(*args, **kwargs)
                except RETRYABLE as exc:
                    last_exc = exc
                    logger.warning(
                        "%s attempt %d/%d failed: %s: %s",
                        func.__name__, attempt, max_attempts,
                        type(exc).__name__, exc,
                    )
                    if attempt < max_attempts:
                        delay = base_delay * attempt
                        logger.info("Sleeping %ds before retry", delay)
                        time.sleep(delay)
            logger.error(
                "%s exhausted %d attempts, last error: %s",
                func.__name__, max_attempts, last_exc,
            )
            raise last_exc
        return wrapper
    return decorator
Enter fullscreen mode Exit fullscreen mode

Three details matter here.

except RETRYABLE, not except Exception. Catching Exception would swallow KeyboardInterrupt? No — KeyboardInterrupt inherits from BaseException, not Exception, so a bare except Exception would actually let it through. But except Exception would catch SystemExit? Also no. The real danger is catching AttributeError, KeyError, ValueError — logic errors that retrying cannot fix. Those should crash the job immediately so you see them in the log at 02:01, not at 02:06 after three pointless sleeps.

Explicit re-raise after the loop. If you let the loop fall through, the function returns None, and the caller has no idea anything went wrong. Re-raising last_exc preserves the original traceback context and lets the exit-code layer above decide what to do.

Logging on every attempt, not just the final failure. When you're debugging a flaky 3 AM job, the pattern matters. Three timeouts in a row points at the network. A 429 followed by two successes points at rate limiting. A 500 followed by a 200 followed by a 500 points at a load balancer routing to a bad backend. You only see these patterns if you log each attempt.

Backoff shape: linear vs exponential

Linear backoff (10, 20, 30) is the right default for scheduled jobs that run on a fixed cadence. If your job runs every five minutes, the total retry window of 60 seconds fits comfortably inside the interval. Exponential backoff (10, 20, 40, 80…) is better for interactive or high-frequency callers where you want to shed load quickly, but for a job that fires every N minutes, linear keeps the worst-case runtime predictable.

If you need to honor a server-provided delay, extract it before sleeping:

def _retry_after(exc: requests.RequestException) -> int | None:
    resp = getattr(exc, "response", None)
    if resp is None or resp.status_code != 429:
        return None
    header = resp.headers.get("Retry-After")
    if header and header.isdigit():
        return int(header)
    return None
Enter fullscreen mode Exit fullscreen mode

Then in the loop, prefer _retry_after(exc) over the computed delay when it's present. This is the one case where the server knows better than you do.

Exit codes: the contract with the scheduler

Windows Task Scheduler records the exit code of every run, and you can filter on it in the history view. More importantly, a wrapper script can branch on it. The useful convention is small and unambiguous:

Code Meaning
0 Success — work was done
1 Failed after retries — investigate
2 Nothing to do — no work this cycle
3 Configuration or setup error — fix before next run

The distinction between 1 and 2 is the one people skip, and it's the one that saves the most time. A job that polls an API for new records will legitimately find nothing most cycles. If "nothing to do" returns 1, your monitoring is useless — every quiet night looks like an outage. If it returns 0, you can't tell from the logs whether the job actually ran.

Here's how the wrapper looks:

import sys

EXIT_OK = 0
EXIT_FAILED = 1
EXIT_NOTHING = 2
EXIT_CONFIG = 3


def main() -> int:
    try:
        records = fetch_pending()
    except requests.RequestException:
        return EXIT_FAILED
    except ConfigError:
        return EXIT_CONFIG

    if not records:
        logger.info("No pending work")
        return EXIT_NOTHING

    process(records)
    return EXIT_OK


if __name__ == "__main__":
    sys.exit(main())
Enter fullscreen mode Exit fullscreen mode

The fetch_pending function is decorated with @retry(). When it exhausts its attempts, it re-raises, main catches the specific exception type, and returns EXIT_FAILED. The scheduler sees 1, your alert rule fires, and the log already contains three timestamped failure reasons.

Why retry state must live on disk

This is the part that bites people who test their retry logic interactively and then deploy it. Consider a job that processes a queue of items:

@retry()
def process_batch():
    items = load_queue()  # reads from disk
    for item in items:
        call_api(item)
        mark_done(item)
Enter fullscreen mode Exit fullscreen mode

The retry state here is the queue on disk. If attempt 1 processes items 1–5 and then fails on item 6, attempt 2 reloads the queue and sees items 6 onward. That's correct. The progress is durable because it lives in a file, not in a variable.

Now consider the in-memory version:

@retry()
def process_batch():
    items = load_queue()
    for item in items:
        call_api(item)
        mark_done(item)
    return items  # cached in memory
Enter fullscreen mode Exit fullscreen mode

If you cache items at module level to avoid re-reading, a retry after a partial failure will re-process items 1–5. Depending on the API, that's either wasted work or a duplicate side effect. Worse, if the process is killed between attempts — a reboot, a taskkill, a power event — any in-memory state is gone entirely, and the next scheduled run has no idea what happened.

The rule: the retry loop should re-read its input from the durable source on every attempt. The decorator doesn't know or care what the function does; it just calls it again. It's the function's responsibility to start from a consistent, on-disk state each time.

This has a second benefit. If the job crashes hard — no retry, just a blue screen — the next scheduled run picks up from the same disk state. The retry decorator handles the fast failures; the disk state handles the slow ones.

Wiring it into Task Scheduler

The scheduler itself has retry settings, and they interact with yours. Under the task's Settings tab, "If the task fails, restart every" is a separate mechanism from your in-process retry. Keep them distinct:

  • In-process retry handles transient errors within a single run. Fast, cheap, no scheduler involvement.
  • Scheduler restart handles process-level failures: a crash, a hung process killed by a timeout, a machine that rebooted mid-run.

A reasonable pairing is three in-process attempts with linear backoff, and one scheduler restart after a short delay. That gives you four total chances spread over a few minutes without turning a broken job into an infinite loop.

One Windows-specific gotcha: if your task uses "Stop the task if it runs longer than", set that duration longer than your worst-case retry window. Three attempts with 10/20/30 second backoff plus request timeouts can easily exceed two minutes. If the scheduler kills the process at 60 seconds, you never see the third attempt, and the exit code is a generic timeout rather than your EXIT_FAILED.

What this design deliberately doesn't do

It doesn't retry on HTTP 4xx (except 429). It doesn't use jitter, because scheduled jobs don't have the thundering-herd problem that interactive clients do. It doesn't persist retry counts across runs — if the job fails three times and the scheduler restarts it, that's a fresh three attempts, and the log will show eight failures total, which is a signal in itself.

It also doesn't try to be clever about partial success. If a batch of 100 items processes 60 and then fails, the retry starts from item 61 because that's what the disk says. If your API doesn't support idempotent writes, you need a different design — one that records the exact operation before calling out, not after.

The whole wrapper is about forty lines. Most of the value is in the parts that don't retry: the narrow exception tuple, the explicit re-raise, the exit code that separates "failed" from "empty." Fail fast, log everything, and let the disk be the source of truth.

More notes like this ship every week on this site.


Daily Picks

The following pairs are selected from the multi-timeframe trend scanner (Gate.io futures) and are for technical-analysis study only — not investment advice.
Data updated: 2026-10-06 12:36:33

Long

Pair Signal Price Take Profit Stop Loss R/R
SKYAI $0.0416 $0.0433 $0.0406 1:1.6
RE $0.4991 $0.5181 $0.4866 1:1.5

2 picks selected. Scanner runs every 15 minutes.

Top comments (0)