Characterize retry classes before any delay helper moves. Integer wait rows are the gate for that split. A green table beats a cleaner function that drifts.
Class and delay must not move in one edit. Tests pin which errors retry and which waits occur. Only then does delay math leave the caller.
A cleanup that also changes policy is not a refactor. It is a behavior change with a style excuse. Keep those two kinds of commits fully apart.
Why this module breaks quietly
Retry helpers hide three choices in one loop. They choose the error class, the wait, and the stop rule. A rename can flip any one of them.
Silent drift shows up as extra posts or missing posts. A timeout may stop retrying after a type check. A validation error may start retrying after a broad except.
Fixture rules
The sample below is a synthetic teaching fixture. It is not a production incident report or metric. These numbers exist only to make assertions exact.
The code blocks below are labeled unexecuted teaching examples. Run them in your own checkout before you trust them. Do not treat the lists as measured latency.
Treat the four rows as the locked contract. Do not add any rows during this extract. Add rows only in a later change with new tests.
The mixed function
import time
def deliver(post, payload, attempts=4, sleep=time.sleep, jitter_ms=lambda: 50):
delay_ms = 200
for _ in range(attempts):
try:
return post(payload)
except TimeoutError:
sleep(delay_ms + jitter_ms())
delay_ms *= 2
except ConnectionError:
sleep(delay_ms)
delay_ms *= 2
except ValueError:
raise
raise RuntimeError("exhausted")
This function hides three decisions inside one body. Which exception retries is the first decision here. How long to wait is the second decision.
Whether sleep actually runs is a separate decision. Random jitter makes repeated local runs fully non-deterministic. A characterization test cannot lock delays with a live clock.
The fixture passes milliseconds straight into the fake sleep. A real time.sleep call expects seconds, not milliseconds. Convert units only in an adapter outside this extract.
Step 1. Write the rows before the seam
Write the expected table before editing any helpers. Each row names one exception and one outcome. Delay values always use a fixed jitter substitute.
| Row | Error script | Posts | Sleeps (ms) | Result |
|---|---|---|---|---|
| 1 | TimeoutError, TimeoutError, ok | 3 | 250, 450 | body |
| 2 | ConnectionError, ok | 2 | 200 | body |
| 3 | ValueError | 1 | none | ValueError |
| 4 | OSError | 1 | none | OSError |
Row 1 fails twice, then returns a body. Row 2 fails once, then returns a body. Row 3 aborts on the first raised error.
Row 4 never enters an except branch at all. Base OSError is not a retry class here. A later policy can add it only with a new row.
These numbers are fixture choices, not measured service data. They define the only gate for this extract. Change them only with a new test commit.
Check the integer path by hand before coding. The first timeout waits 200 plus 50 milliseconds. The second timeout waits 400 plus 50 milliseconds.
Those two waits yield 250 and then 450. A connection wait uses the base 200 only. No float equality is required for these rows.
Step 2. Add seams without changing defaults
Replace both time.sleep and the jitter source with explicit callables. Pass them as arguments with the old defaults. Behavior stays the same for all current callers.
The listing above already shows that injected seam. This edit is a seam, not the extract. Defaults still preserve the old external call path.
Tests pass a list append instead of real sleep. Tests pass a lambda that always returns 50. Callers that omit both arguments keep prior behavior.
Step 3. Lock every row in tests
Use a fake poster that raises a scripted sequence. Record every sleep argument inside one local list. Assert the list and the final result together.
def scripted(errors, body):
pending = list(errors)
def post(_payload):
if pending:
raise pending.pop(0)
return body
return post
def test_timeout_row():
slept = []
result = deliver(
scripted([TimeoutError("slow"), TimeoutError("slow")], {"ok": True}),
{"id": 7},
sleep=slept.append,
jitter_ms=lambda: 50,
)
assert result == {"ok": True}
assert slept == [250, 450]
def test_connection_row():
slept = []
result = deliver(
scripted([ConnectionError("reset")], {"ok": True}),
{"id": 7},
sleep=slept.append,
jitter_ms=lambda: 50,
)
assert result == {"ok": True}
assert slept == [200]
def test_value_error_row():
slept = []
try:
deliver(
scripted([ValueError("bad")], {"ok": True}),
{"id": 7},
sleep=slept.append,
jitter_ms=lambda: 50,
)
except ValueError:
assert slept == []
else:
raise AssertionError("expected ValueError")
def test_oserror_is_not_retried():
slept = []
try:
deliver(
scripted([OSError("other")], {"ok": True}),
{"id": 7},
sleep=slept.append,
jitter_ms=lambda: 50,
)
except OSError:
assert slept == []
else:
raise AssertionError("expected OSError")
The test catch is wider than the product clause. The empty sleep list is what pins non-retry. A successful return would fail the else branch.
Run the full table before any further move. The command below stays small and fully local. A single red row blocks the whole extract.
python -m pytest tests/test_deliver_rows.py -q --tb=line
Fix the fixture or the code in a separate commit. Do not fix a red row by weakening the assert. Do not skip a row to force a passing run.
Step 4. Move only the wait function
After the table is passing, extract delay math. Leave all classification inside the original deliver function. The new function returns the next base and the sleep value.
def next_wait(kind, delay_ms, jitter_ms):
if kind == "timeout":
slept_ms = delay_ms + jitter_ms()
elif kind == "connection":
slept_ms = delay_ms
else:
raise ValueError(kind)
return delay_ms * 2, slept_ms
Wire the helper without changing any branch order. A timeout still doubles after the added jitter. A connection error still doubles without any jitter.
except TimeoutError:
delay_ms, slept_ms = next_wait("timeout", delay_ms, jitter_ms)
sleep(slept_ms)
except ConnectionError:
delay_ms, slept_ms = next_wait("connection", delay_ms, jitter_ms)
sleep(slept_ms)
except ValueError:
raise
A ValueError still never calls this new helper. Re-run the same four rows after the move. The sleep list must match the frozen table.
If a recorded list drifts, stop and revert the helper. Exact list equality is the pass rule here. Do not switch asserts to ranges in this change.
Step 5. Prove the test file did not move
Record the full characterization output before the helper extract. Record the same output again after the extract. The diff should be empty for all four rows.
python -m pytest tests/test_deliver_rows.py -q --tb=line
git diff -- tests/test_deliver_rows.py
An empty test diff is part of the gate. A changed assert means the extract is not safe. Restore the test file and revise the helper.
Do not add property tests in this commit. Do not rename public arguments in this commit. Do not change attempt counts in this commit.
What this commit refuses
| Change | In this extract | Needs a new row |
|---|---|---|
| Inject sleep and jitter | yes | no |
| Assert four class rows | yes | extend later |
| Extract next_wait | yes | no |
| Retry base OSError | no | yes |
| Retry ValueError | no | yes, with sign-off |
| Read a Retry-After header | no | yes |
| Cap the max wait | no | yes |
The smallest safe change is the helper move. Policy changes need their own separate new rows. Mixing them hides the cause of a red test.
BrokenPipeError remains a ConnectionError subclass in Python 3. The current except clause already catches that subclass. Do not narrow that clause while moving math.
That subclass path is not in the four rows. Add it later if callers depend on it. Until that row exists, leave the except line untouched.
Read the Python 3 exception class list before narrowing except. That page is the primary source for the subclass claim.
Use a free model only as a drafter
Disclosure: This article was prepared as part of MonkeyCode's product outreach. A free model can draft row tests from the table above. A free server can run that same pytest file in a clean checkout.
You still compare the recorded sleep lists yourself. Use the model to propose seams, not to approve the extract. Paste the red traceback and ask for a smaller diff.
Reject any patch that edits an assertion to go green. MonkeyCode's free model access and free server option fit that loop. They do not replace the characterization test gate.
If a plan rewrites deliver in one shot, discard it. Allowed model edits stay narrow in this pass. The table below is the accept and reject rule.
| Model output | Accept | Reject |
|---|---|---|
| A test that encodes one table row | yes | no |
| A seam with defaulted callables | yes | no |
| A helper that preserves both wait formulas | yes | no |
| A broader except OSError | no | yes |
| A deleted or weakened assert | no | yes |
| A new retry policy | no | yes |
The operator supplied only two availability facts here. Free model access is the first supplied fact here. A free server option exists for a clean run.
No quota, hardware, or duration is stated here. No model name is required for the method. No speed claim is part of the gate.
Limits you should keep visible
Four rows do not cover partial writes or duplicate posts. The fixture poster is synchronous and fully in-process. Real clients can raise subclasses that miss these branches.
Integer asserts stay exact only because jitter is fixed. A live clock source will fail this test style. Keep the injected jitter in tests for stable lists.
This pattern does not prove any thread safety. It does not prove deadline behavior under load. It proves one function kept the pinned rows.
Retries can duplicate a post that already succeeded. These rows do not prove the payload is idempotent. Pin idempotency keys in a separate test file.
The exhausted path is also outside this gate. Four timeouts would sleep 250, 450, 850, then 1650. Add that row only when you intend to lock the stop rule.
Who should not use this approach
Skip it when the module has no stable outcomes to pin. A prototype that changes rules hourly will fight the table. Write the product rule first, then the rows.
Skip it when failure handling is a one-line library call. A wrapper test around that library adds little signal. Pin your own branches, not the library internals.
Skip it when you cannot inject sleep or randomness. A hidden clock makes the characterization table flaky. Add the seam before you trust any green run.
Skip it when the next edit must change product policy. Characterization tests block that policy change on purpose. Write the new rule, then replace the rows.
Review checklist
- Four rows existed before the helper extract began.
- Public caller defaults still match the old path.
- Sleep lists match the table after the move.
- No assertion was edited to hide a red row.
- The test file diff is empty after the extract.
- Policy changes stay absent from this single commit.
If any item fails, revert the helper and keep the tests. The table is the required artifact for this change. The helper stays optional until that table holds.
Next cut
Pick one messy function that both classifies and waits. Write two rows that must never retry at all. Run them on a clean machine before any extract.
Draft those two rows through MonkeyCode's free model access. Use the free server as a clean checkout runner. Run pytest on your machine before you merge.
Top comments (0)