You should freeze a retry budget before any generated client is allowed to sleep, because unbounded backoff can pin a worker for hours. A small in-process budget can express attempt caps, elapsed caps, and delay clamps as data you test without a network. This case study walks that budget from a frozen decision table to commands that prove each stop condition. The files are runnable teaching code, not a report of a live outage or a measured production improvement.
You are looking at one module and one unittest file, which is enough to expose the failure without a platform narrative. The worker calls a flaky read and then updates a local row, so it must stop when the frozen budget says stop. A generated draft that retries every exception will hide a permanent validation error behind a long and useless sleep. You prevent that outcome by writing the contract first and by treating every later draft as an untrusted patch.
Background
The pressure usually arrives as a short task that says to wrap the call, add retries, and move on to the next ticket. A sample loop from a coding assistant often uses an open-ended loop, doubles the delay without a ceiling, and catches a broad exception. That shape can survive a happy-path demo, and then it can hold a thread through a long upstream outage. It also retries mistakes you cannot fix by waiting, such as a rejected identifier or a failed local precondition.
You do not need a new framework, a queue, or a hosted workflow engine to close that particular gap. You need a policy object, a wrapper that accepts injected time, and tests that fail if sleep is called after a cap. This case study uses those three pieces and leaves webhook signatures, money rounding, and export idempotency for other write-ups. Keeping the scope that small is what makes the contract reviewable in one sitting by someone who did not write the loop.
Goal
The goal is a wrapper that returns the operation result or raises a budget error when a frozen limit is hit. You should be able to explain that behavior to a teammate without asking them to read the loop line by line. A reviewer should be able to reject a patch by pointing at one row in the decision table below. The tests should stay on the standard library so a clean Python process is enough to reproduce the result.
Frozen rules
You freeze the rules before implementation, and you keep them stable while any draft is proposed or edited. Changing a rule means changing the table and the tests in the same edit, not silently fixing a red assertion. The list below is the whole policy for this case study, and anything absent from it is out of scope on purpose.
- The attempt cap counts every try, must be at least one, and includes the first call rather than only the retries.
- After failure number n, the delay is the base delay times two to the n minus one, clamped at the configured maximum.
- If time already spent plus the next delay would pass the elapsed ceiling, you raise the budget error and skip sleep.
- Only exception types you listed as retryable may sleep, and every other error propagates on the attempt that raised it.
- Sleep and the clock are injected arguments, so the tests never call the real sleep function or read the wall clock.
Invalid construction is part of the same contract, not a later cleanup you leave for a second patch. A zero attempt cap, a negative duration, or a maximum delay below the base delay raises ValueError immediately. That early failure matters because a generated client should not run at all with a policy you already know is nonsense.
Decision table
You review generated code against the rows below, not against how confident or fluent the draft happens to sound. Each row names the inputs that matter and the only result you should accept for that situation. If a draft needs a new row, you add the row before you accept the code, so the table stays the source of truth.
| Situation | Attempts used | Next delay | Elapsed plus delay | Required result |
|---|---|---|---|---|
| Success on first try | 1 | none | not used | return the value |
| Retryable error with room left | below cap | clamped exponential | within ceiling | sleep, then retry |
| Retryable error at attempt cap | equals cap | ignored | not used | raise budget error |
| Retryable error at time cap | below cap | computed | above ceiling | raise budget error, no sleep |
| Non-retryable error | any | none | not used | re-raise the original error |
| Invalid policy values | none | none | not used | raise ValueError at construction |
Implementation
The module below is the reference implementation for this worked example, and you can save it as retry_budget.py. You should change it only when the decision table changes, because the tests encode those rows rather than a style preference. The listing has no narrative comments, so a reviewer has to match behavior to the table instead of trusting prose in the file. Treat the listing as example code you can run locally, not as a versioned library with compatibility promises.
class RetryBudgetExceeded(Exception):
"""Raised when a frozen attempt cap or elapsed cap blocks another try."""
class RetryPolicy:
def __init__(self, max_attempts, max_elapsed_ms, base_delay_ms, max_delay_ms):
if max_attempts < 1:
raise ValueError("max_attempts must be at least 1")
if max_elapsed_ms < 0 or base_delay_ms < 0:
raise ValueError("time values must be non-negative")
if max_delay_ms < base_delay_ms:
raise ValueError("max_delay_ms is below base_delay_ms")
self.max_attempts = max_attempts
self.max_elapsed_ms = max_elapsed_ms
self.base_delay_ms = base_delay_ms
self.max_delay_ms = max_delay_ms
def delay_for_attempt(policy, attempt):
scaled = policy.base_delay_ms * (2 ** (attempt - 1))
if scaled > policy.max_delay_ms:
return policy.max_delay_ms
return scaled
def run_with_budget(policy, operation, sleep, clock, retryable=(TimeoutError,)):
started = clock()
attempt = 0
while True:
attempt += 1
try:
return operation()
except retryable as exc:
if attempt >= policy.max_attempts:
raise RetryBudgetExceeded("attempt cap reached") from exc
wait = delay_for_attempt(policy, attempt)
if clock() - started + wait > policy.max_elapsed_ms:
raise RetryBudgetExceeded("elapsed cap reached") from exc
sleep(wait)
Three details are where generated drafts usually drift, even when the function names look identical to this listing. The delay exponent uses the attempt that just failed, so the first wait equals the base delay rather than double that value. The elapsed check reads the clock after the failed call, which means a slow attempt can exhaust the budget by itself. Sleep runs only after both checks pass, which is why the time-cap test can assert that the sleep log stayed empty.
Tests you can run
Save the next listing as test_retry_budget.py in the same directory as the module so the import stays boring. The tests record sleeps in a list and feed the clock from a short script, which keeps the file deterministic on a shared runner. Nothing in this file needs a network, a clock permission, or a third-party package, which is intentional for a clean environment. You can add cases later, but you should not delete an assertion to help a draft turn green.
import unittest
from retry_budget import (
RetryBudgetExceeded,
RetryPolicy,
delay_for_attempt,
run_with_budget,
)
def raise_error(exc):
raise exc
class ScriptedClock:
def __init__(self, readings):
self.readings = list(readings)
def __call__(self):
if not self.readings:
raise AssertionError("clock exhausted")
return self.readings.pop(0)
class RetryBudgetTests(unittest.TestCase):
def test_success_does_not_sleep(self):
slept = []
policy = RetryPolicy(3, 1000, 100, 400)
value = run_with_budget(policy, lambda: "ok", slept.append, lambda: 0)
self.assertEqual(value, "ok")
self.assertEqual(slept, [])
def test_attempt_cap_sleeps_only_between_tries(self):
calls = {"n": 0}
def fail():
calls["n"] += 1
raise TimeoutError("slow upstream")
slept = []
policy = RetryPolicy(2, 10000, 100, 400)
with self.assertRaises(RetryBudgetExceeded):
run_with_budget(policy, fail, slept.append, lambda: 0)
self.assertEqual(calls["n"], 2)
self.assertEqual(slept, [100])
def test_elapsed_cap_does_not_sleep(self):
slept = []
policy = RetryPolicy(5, 1000, 100, 400)
clock = ScriptedClock([0, 950])
with self.assertRaises(RetryBudgetExceeded):
run_with_budget(
policy,
lambda: raise_error(TimeoutError("late")),
slept.append,
clock,
)
self.assertEqual(slept, [])
def test_validation_error_is_not_retried(self):
slept = []
policy = RetryPolicy(4, 5000, 50, 200)
with self.assertRaises(ValueError):
run_with_budget(
policy,
lambda: raise_error(ValueError("bad sku")),
slept.append,
lambda: 0,
)
self.assertEqual(slept, [])
def test_delay_clamps_and_policy_rejects_bad_values(self):
policy = RetryPolicy(3, 5000, 100, 150)
self.assertEqual(delay_for_attempt(policy, 3), 150)
with self.assertRaises(ValueError):
RetryPolicy(0, 1000, 10, 10)
if __name__ == "__main__":
unittest.main()
From the directory that holds both files, you run the standard-library commands in the next block. You should see five tests pass, with no installer step and no outbound call required by the example itself. If a draft removes the elapsed check, the elapsed-cap test fails because the sleep list is no longer empty. If a draft retries ValueError, the validation test fails because the wrapper no longer preserves the original error type.
python -m py_compile retry_budget.py test_retry_budget.py
python -m unittest test_retry_budget.py -v
python test_retry_budget.py -v
If the runner exposes only python3, you swap the command name and you leave the module arguments unchanged. A compile failure is a syntax problem, while a unittest failure is a contract problem, and you should not mix those two signals. You keep the verbose flag so the printed names match the review notes in the next section.
What a red test means
A red elapsed-cap test means the wrapper slept, or raised the wrong type, after the scripted clock crossed the ceiling. A red validation test means a permanent error was retried or wrapped so tightly that the original type disappeared. A red clamp test means the delay formula forgot the maximum and let exponential growth pick the wait by itself. You fix the wrapper or the table, and you do not loosen the assertion to manufacture a pass on a remote runner.
How you review a generated patch
You read the diff before you celebrate a green test run, because a pass can hide a weakened test. A draft can also pass for the wrong reason if it retries less by accident while breaking a row you care about later. Your review looks for a short list of drifts that show up often when a model writes retry loops from memory. You comment on the table row that broke, and you ask for a smaller patch rather than a broader rewrite.
- Reject a bare exception handler that retries validation errors, cancellations, or an interrupt you did not mark as retryable.
- Reject a call to real sleep, a random source, or the wall clock inside the wrapper, even if the tests still pass.
- Reject an elapsed check that counts only the planned delay and ignores time already spent inside the operation.
- Reject a first-wait formula that doubles too early, so the first failure sleeps twice the configured base delay.
- Reject a test edit that deletes the empty-sleep assertion or widens the retryable tuple to the base exception type.
- Reject a new dependency when the standard library already compiles and runs both files on a clean interpreter.
You paste the failed assertion back to the assistant instead of asking it to try harder without evidence. That habit keeps the case study short enough to finish in one sitting, even if the first draft is wrong. It also keeps a small remote runner useful, because the feedback is a unittest name rather than a screenshot or a log dump. You stop when the five tests pass and the diff still matches the table, not when the draft adds extra features.
A practical drafting workflow
You can implement the module by hand, and that remains a valid ending if you never call an assistant at all. When you do want a draft, you keep the assistant on a short leash and you follow the order below. The order is the method, and skipping ahead to a finished client is how the open-ended sleep loop returns.
What you hand the assistant
- You paste the decision table and the unittest file, and you withhold the reference module so the draft cannot copy it.
- You ask for one dependency-free implementation that satisfies the tests and does not call the real sleep function.
- You diff that draft against the frozen rules before you run it, and you discard edits that invent new policy.
- You run the unittest command, then you reject any change that weakens an assertion merely to obtain a pass.
- You keep the table in the review notes so the next generated tweak has the same fence as the first.
Where a free drafting setup fits
MonkeyCode's free model access fits the drafting step, because you can request the patch without a paid seat for this exercise. Disclosure: This article was prepared as part of MonkeyCode's product outreach. The free server option fits the unittest step, because that same command can judge the patch on a clean Python process. You should confirm current access terms on the project page before you depend on them for anything beyond this exercise.
This article states no model names, token quotas, hardware shapes, or permanent limits, since those details change and were not verified here. You still review the diff yourself, and you still discard a green run that came from a weakened test. A free runner does not make the table optional, and a fluent draft does not outrank a failed assertion. If the current free access is unavailable, you run the same command on any local Python 3 interpreter and the case still holds.
Results you should expect
If you run the files as written, the expected result is a local pass of the five tests in this case study. The attempt-cap test shows one sleep between two tries, and then a budget error instead of a third call. The elapsed-cap test shows a budget error with an empty sleep list when the scripted clock is already at 950 milliseconds. The clamp check shows a third-failure delay of 150 milliseconds rather than the unclamped value of 400 milliseconds.
These checks describe the contract, not throughput, hosting cost, or the behavior of any live inventory service. They also do not prove that a particular model will draft a correct patch, because that outcome depends on the prompt and the review. You should record a failure as a failed assertion name, not as a story about the assistant being careless or clever. When the command prints five successful tests, you have reproduced the case study, and you do not need further metrics.
Limitations
This wrapper is single-process and single-threaded, and it does not cancel a call that has already started. It does not share a budget across workers, and it does not add jitter to spread out a thundering herd. It also does not make the following write idempotent, so a successful retry can still double-apply a non-idempotent update. The injected clock must move forward in milliseconds, because a backward jump can mask an elapsed-cap failure.
You should not use this module for card charges, device control, or any loop where a wrong retry is irreversible. You should not ask a model to invent the caps from a vague prompt that only says to make the client resilient. If you cannot fill the decision table with real numbers, you are not ready to generate the loop at all. Teams that need cross-host coordination should write a shared limiter with its own contract, not a longer copy of this file.
Lessons
You get a worker that fails visibly when the stop rule is data you can assert in a test name. You get a calmer review when the table is small enough to read without a meeting or a design document. You get portable checks when sleep and time are arguments, which is why a modest remote process can host them. The product note is optional to the lesson, because the module, the table, and the command still stand if you delete it.
If you lack a local interpreter this week, you can rehearse the same two-file case on a free coding setup. MonkeyCode's free model access and free server are enough for that rehearsal, provided you still keep the table as the fence. Merge only the diff that preserves the five tests, and leave every other generated convenience out of the branch. That is the whole invitation: try the workflow on the exercise above, not a claim that a tool replaces the contract.
Top comments (1)
Rule 3 is the one I would test hardest: "time spent plus the next delay would pass the ceiling" refuses a sleep that fits but leaves no time for the attempt after it. With a 10 s elapsed cap, 4 s already spent and a 5 s delay, the check passes, you sleep to 9 s, and the next call has 1 s of budget but may run for 30. Either subtract an expected call duration or put a deadline on the call itself, otherwise the cap bounds the sleeping, not the total.
Two table rows I did not see: jitter, since a deterministic base x 2^(n-1) schedule means every client that failed together retries together, and a server-sent Retry-After, which should replace the computed delay when it is larger and trigger the budget error when it exceeds the cap. A test with a fake clock that jumps by more than the delay (a call that blocks for 20 s) would also show whether elapsed is measured around the call or only around the sleeps.
Did the frozen table include the case where the first attempt itself already exceeds the elapsed ceiling?