Your nightly embed job pages you at 02:14 local time.
The free slot still reports spare capacity on the board.
The batch deadline has already slipped by eleven minutes.
Treat this page as a composite drill, not a named outage.
Which action follows from that evidence, not from the green badge?
You do not add workers just because the badge stays green.
You price the free slot before you admit the next batch.
Spare capacity is the bait, not the decision itself.
What spare capacity hid
Spare capacity is not the same thing as a cheap path.
A free token price can still waste your nightly window.
Retries duplicate work you already spent on a dead attempt.
This note is an ops cost drill, not a pricing brochure.
You will run a local ledger against a declared workload.
You will reject free-path work when retry tax breaks budget.
Topology before any code
Put four boxes on the diagram before you write a script.
- The batch client submits each job with a deadline and budget.
- The admission ledger scores that job before any enqueue call.
- The free slot is one route, not the only route available.
- The reserved worker is the rollback path when tax runs high.
Declare the workload in config, not in a chat thread later.
window_seconds: 900
jobs: 40
input_tokens_per_job: 1200
max_attempts: 3
inject_fail_after_attempt: 1
fail_fraction: 0.25
budget_units: 60000
queue_wait_unit_per_second: 200
seed: 7
service_seconds: 1.5
Those numbers are declared test inputs, not live measurements.
Change them to match your window before you trust a reject.
Keep the seed fixed so a teammate can replay your rows.
Fields that earn the page
Record these fields on every attempt, including the killed one.
-
job_idties the attempt back to one batch item. -
attemptstarts at one and climbs only on retry. -
queue_wait_msis delay before that attempt actually starts. -
tokens_inandtokens_outcome from your own meter. -
outcomeisok,killed, orrejectedfor that row. -
deadline_slack_msis remaining window time, not host utilization. -
retry_taxis the admission number you will compare to budget.
Separate what you observed from what you infer next.
Observed means the ledger row the script just wrote down.
Inferred means a vendor story you did not measure here.
How retry tax is defined
Retry tax is the waste a zero price tag can hide.
Count tokens already spent on attempts that did not finish.
Add a time charge for queue wait inside this window.
Use this definition only inside the drill, and label it local.
retry_tax = spent_tokens_on_failed_attempts
+ queue_wait_seconds * queue_wait_unit_per_second
admit if projected_tax + new_job_tokens <= budget_units
Do not treat that formula as a cloud invoice.
It is an admission signal for this declared workload only.
A zero token price does not zero the time charge.
That weight is declared so wait can rival a token chunk.
It is not a dollar rate and not a vendor price.
Why not utilization
CPU idle time can sit next to a blown token budget.
A free worker can look empty while your window is gone.
Utilization answers a different question than retry tax does.
The local ledger
This script is a proposal you execute on your own laptop.
It does not call a remote model and does not spend quota.
Failure injection is a coin flip you control in process.
#!/usr/bin/env python3
"""Local retry-tax ledger. Proposal until you run it."""
import json
import random
from pathlib import Path
CFG = {
'window_seconds': 900,
'jobs': 40,
'input_tokens_per_job': 1200,
'output_tokens_on_ok': 180,
'max_attempts': 3,
'inject_fail_after_attempt': 1,
'fail_fraction': 0.25,
'budget_units': 60000,
'queue_wait_unit_per_second': 200,
'seed': 7,
'service_seconds': 1.5,
}
def retry_tax(spent_failed, queue_wait_s):
wait_units = int(queue_wait_s * CFG['queue_wait_unit_per_second'])
return spent_failed + wait_units
def row(job_id, attempt, queue_wait_s, tokens_in, tokens_out, outcome, clock, tax):
return {
'job_id': job_id,
'attempt': attempt,
'queue_wait_ms': int(queue_wait_s * 1000),
'tokens_in': tokens_in,
'tokens_out': tokens_out,
'outcome': outcome,
'deadline_slack_ms': int((CFG['window_seconds'] - clock) * 1000),
'retry_tax': tax,
}
def main():
rng = random.Random(CFG['seed'])
rows = []
window_spent = 0
clock = 0.0
rejected = 0
finished = 0
killed_rows = 0
abandoned = 0
for job_id in range(CFG['jobs']):
spent_failed = 0
done = False
for attempt in range(1, CFG['max_attempts'] + 1):
queue_wait_s = rng.uniform(0.4, 6.0)
tax = retry_tax(spent_failed, queue_wait_s)
projected = tax + CFG['input_tokens_per_job']
slack_ms = int((CFG['window_seconds'] - clock) * 1000)
remaining = CFG['budget_units'] - window_spent
if projected > remaining or slack_ms <= 0:
rows.append(row(
job_id, attempt, queue_wait_s, 0, 0, 'rejected', clock, tax,
))
rejected += 1
done = True
break
clock += queue_wait_s + CFG['service_seconds']
fail_now = (
attempt == CFG['inject_fail_after_attempt']
and rng.random() < CFG['fail_fraction']
)
if fail_now:
spent_failed += CFG['input_tokens_per_job']
window_spent += CFG['input_tokens_per_job']
killed_rows += 1
tax_after = retry_tax(spent_failed, queue_wait_s)
rows.append(row(
job_id, attempt, queue_wait_s,
CFG['input_tokens_per_job'], 0, 'killed', clock, tax_after,
))
continue
window_spent += (
CFG['input_tokens_per_job'] + CFG['output_tokens_on_ok']
)
finished += 1
rows.append(row(
job_id, attempt, queue_wait_s,
CFG['input_tokens_per_job'], CFG['output_tokens_on_ok'],
'ok', clock, tax,
))
done = True
break
if not done:
abandoned += 1
payload = {
'config': CFG,
'finished': finished,
'killed_rows': killed_rows,
'rejected': rejected,
'abandoned': abandoned,
'window_spent': window_spent,
'budget_left': CFG['budget_units'] - window_spent,
'rows': rows,
}
out = Path('/tmp/retry-tax-ledger.json')
out.write_text(json.dumps(payload, indent=2))
print(
'finished=%s killed_rows=%s rejected=%s abandoned=%s budget_left=%s'
% (finished, killed_rows, rejected, abandoned, payload['budget_left'])
)
print('wrote %s' % out)
if __name__ == '__main__':
main()
Run the file only after you read the config block above.
python3 retry_tax_ledger.py
Expected output is labeled, because this draft did not execute it.
You should see one summary line plus a JSON ledger file.
The file path is /tmp/retry-tax-ledger.json on that laptop.
# expected shape, not a measured production result
finished=<int> killed_rows=<int> rejected=<int> abandoned=<int> budget_left=<int>
wrote /tmp/retry-tax-ledger.json
If rejected stays at zero, your budget is too loose.
Tighten budget_units and run the same seed again.
If every job is rejected, your time charge is too harsh.
What the injection is for
The drill kills a fraction of first attempts on purpose.
A killed row still records tokens_in for that attempt.
The next attempt must pay that spend inside retry_tax.
That behavior is the point of the injection, not a bug.
A free slot that dies mid-job is not a free retry.
You already burned input tokens a later attempt will repeat.
On a killed row, retry_tax includes tokens from that failed attempt.
The admit check used the tax from before that spend.
Watch deadline_slack_ms on the killed row before you retry.
Negative slack means the window is already lost for that job.
Reject it instead of spending a second attempt on a corpse.
Build that sample by hand so the fields stay obvious.
The next block is an illustration, not seed output.
Use the numbers below to check your own addition only.
job_id: 3
attempt: 1
queue_wait_ms: 2100
tokens_in: 1200
tokens_out: 0
outcome: killed
deadline_slack_ms: 840000
retry_tax: 1620
Wait units are 2.1 seconds times the declared weight of 200.
Add the killed input tokens and you get 1620.
If your hand total disagrees, stop and fix the formula first.
When the free path is the wrong bet
Use this table before you route the nightly window.
| Signal you can read | Free path bet | Action you take |
|---|---|---|
| Retry tax under budget and slack still positive | Reasonable for this batch | Admit to the free slot |
| First attempt killed and tax near the budget | Wrong bet tonight | Reject or switch path |
| Queue wait alone consumes the remaining window | Wrong bet tonight | Shed new jobs now |
| An interactive caller is waiting on this queue | Wrong bet always | Do not use this drill |
Read the table as admission policy, not as a benchmark.
No latency percentile is claimed from any vendor here.
You measured only the local rows your script wrote down.
Where a free path still fits
Disclosure: This article was prepared as part of MonkeyCode's product outreach.
MonkeyCode exposes free model access and a free server option.
Those two routes can sit in the diagram as the free slot.
You still reject the job when projected tax breaks budget.
Do not hard-code a token grant, a duration, or a hardware size.
Those limits were not measured in this drill and can change.
Meter the path you actually call, then apply the same threshold.
A free label does not replace the admission check above.
If the meter cannot see killed-attempt tokens, do not route there.
Missing tokens_in is how a free path looks cheaper than it is.
If you already meter each attempt, try that check on the free server option.
The threshold you defend on the call
Ask for one threshold before the nightly window opens.
This drill uses budget_units, not a utilization target.
Queue age alone misses tokens you are about to pay twice.
Deadline slack alone misses budget you already spent once.
You would rather reject a job than duplicate a killed attempt.
The budget is the cap on that duplication for this window.
Pick the number from your window, not from the example config.
Start from input tokens times jobs, then cut for fail fraction.
Write that rationale beside the config so the next on-call sees it.
If you cannot explain the cut, you are not ready to reject.
Actions after the page
Follow this order when the badge is green and the window is not.
- Read the latest ledger rows with
outcome=killedbefore you scale. - Sum
retry_taxfor this window, not for the whole day. - Compare that sum with
budget_unitsyou declared at the start. - Reject new free-path jobs when the next projection exceeds budget.
- Drain in-flight jobs only while their slack is still positive.
- Switch the route to the reserved worker if the window must finish.
Do not restart the free slot as your first move tonight.
A restart hides the killed rows you needed for the tax math.
Capture the JSON file, then decide on admit or reject.
Cleanup and rollback
Cleanup is part of the drill, not a follow-up ticket later.
Confirm the process name before you run the stop command.
Then remove the ledger file and stop only that local script.
rm -f /tmp/retry-tax-ledger.json
pkill -f retry_tax_ledger.py || true
Rollback the route with a config flag, not a redeploy guess.
admission_mode: reserved_path
free_path_enabled: false
Apply that flag, then stop admitting new work to the free slot.
Let in-flight reserved jobs finish inside the same declared window.
Re-enable the free path only after a clean ledger hour.
A clean hour means zero killed rows and tax under budget.
If either check fails, leave free_path_enabled set to false.
Write the reason in the alert comment so the next shift sees it.
Limits to say out loud
This drill is a single process with a fixed random seed.
It does not model multi-tenant noise on a shared free slot.
It does not read a real invoice or a live quota counter.
Free access terms can change, so do not hard-code a promise.
The script invents queue wait with random.Random, not a trace.
Service time is fixed at one and a half seconds in this script.
Replace that constant with a trace before you trust a reject.
A fixed service time will lie once your real calls vary.
Who should skip this approach entirely:
- Do not put interactive user traffic on this admission path.
- Do not use it where a missed job is a safety incident.
- Do not use it if you cannot meter tokens on each attempt.
- Do not treat the example budget as your real production cap.
- Do not use the drill to justify a cost number on an invoice.
If you lack per-attempt token counts, stop and fix metering first.
A ledger without tokens_in on killed rows will under-count tax.
That under-count is how spare capacity fools the nightly page.
What you keep with the alert
Keep the script, the config, and the rollback flag in one folder.
Run the seed once before the window, not while you are paging.
Store the JSON beside the alert so the next person sees rows.
If the projection exceeds budget, reject the job before enqueue.
Leave the rollback flag in the alert until a clean hour passes.
That is the whole loop from measured tax to cleanup.
Top comments (0)