A Thursday page with an idle host
Your pager fires at 02:14 on a Thursday. A batch summarizer has waited nine minutes in queue. The free worker still looks idle on the host graph.
Token counters climb because each retry resends the full prompt. Your deadline slack has fallen to forty seconds. You need an action that follows from that evidence.
Public posts this week cheer fast AI demos. Your batch queue does not share that mood. You still have to price each retry in time.
This opening is a drill scenario, not a production incident report. The numbers above are inputs you can change. Do not treat them as measured fleet telemetry.
The contradiction in the graphs
The host graph says the worker has spare CPU. The queue graph says age is already past slack. Those two pictures do not describe the same risk.
CPU idle does not mean the job will finish on time. A free invoice does not mean the retry is cheap. Tokens leave the budget on every resend, even at zero dollars.
You are buying time with a hidden token meter. Queue age is the clock you should trust first. Utilization is only a hint about the host.
Topology you can sketch on one page
Put four boxes on the diagram before you tune anything. Number them so the runbook matches the code. No second region appears in this drill.
- The producer stamps each job with a deadline.
- The queue stores age, retries, and the idempotency key.
- The gate chooses stay, shed, or fail closed.
- The worker calls the model only on stay.
No cache tier hides a slow first token. You are testing admission, not a full platform.
Declared workload for the local run
Fix the workload before you compare any result. Each job carries a prompt of eighteen hundred tokens. The deadline slack starts at ninety seconds flat.
You allow two retries after the first attempt. A retry resends the same prompt with the same key. The local token budget for this class is eight thousand.
You run four scripted cases, not a live customer mix. You do not claim a throughput number from this script. You only claim the branch the gate selects.
When free capacity is the wrong bet
Free capacity is the wrong bet in four cases. Read them as policy, not as vendor facts. Time, tokens, and retries decide the branch together.
- Queue age has reached deadline slack, so you shed the pull.
- Retry count has hit the cap, so you fail closed.
- Estimated tokens have crossed the remaining class budget.
- The caller cannot shed, so this gate does not fit.
Use this table when you review a sample by hand.
| Check | Stay when | Leave when |
|---|---|---|
| Queue age | Age is under slack | Age meets slack, then shed |
| Retry count | Count is under the cap | Cap is hit, then fail closed |
| Token estimate | Estimate is under budget | Budget is hit, then shed |
| Shed allowed | Caller can accept a skip | Caller cannot skip, do not use |
Stay only when age is under slack and budget remains. A zero dollar line does not override those checks. That is the cost decision this note is about.
Fields you should log
Log the fields that justify the branch you took. Guesses about the model do not belong in the alert. Separate the sample you saw from the rule you applied.
- Record queue age in seconds at the pull.
- Record deadline slack in seconds at the pull.
- Record the retry count beside the idempotency key.
- Record prompt tokens and the estimated resend total.
- Record the branch name and the threshold that fired.
The script prints a branch from your inputs. It does not observe a remote vendor queue. Label that print as computed, not collected.
Config you pin before the run
Pin the class name, the budget, and the flag default. The snippet below is a drill config, not a cluster manifest. Keep units in the key names so a reviewer cannot mix them.
class: batch-summarizer
prompt_tokens: 1800
deadline_slack_s: 90
max_retries: 2
token_budget: 8000
primary_signal: queue_age_s
secondary_signal: token_estimate
flag: cost_gate_enabled
flag_default: false
A gate you can run locally
The following program is a labeled local drill. It is not a client for any hosted control plane. Run it on your laptop before you touch production.
Use Python 3.9 or newer for this file. You need only the standard library for this file. Do not add network calls while you learn the branches.
#!/usr/bin/env python3
'''Local cost gate. Drill file, not a vendor meter.'''
import os
from dataclasses import dataclass
from typing import List, Optional
@dataclass
class Sample:
queue_age_s: float
deadline_slack_s: float
retry_count: int
prompt_tokens: int
max_retries: int
token_budget: int
def token_estimate(sample: Sample) -> int:
# Model only: each attempt resends the full prompt.
return sample.prompt_tokens * (sample.retry_count + 1)
def decide(sample: Sample) -> str:
estimate = token_estimate(sample)
if sample.queue_age_s >= sample.deadline_slack_s:
return 'shed'
if sample.retry_count >= sample.max_retries:
return 'fail_closed'
if estimate >= sample.token_budget:
return 'shed'
return 'stay'
def emit(sample: Sample) -> None:
estimate = token_estimate(sample)
branch = decide(sample)
print(f'{branch} estimate={estimate} age={sample.queue_age_s:g}')
def from_env() -> Optional[Sample]:
raw_age = os.getenv('COST_AGE')
if raw_age is None:
return None
return Sample(
queue_age_s=float(raw_age),
deadline_slack_s=float(os.getenv('COST_SLACK', '90')),
retry_count=int(os.getenv('COST_RETRIES', '0')),
prompt_tokens=int(os.getenv('COST_PROMPT_TOKENS', '1800')),
max_retries=int(os.getenv('COST_MAX_RETRIES', '2')),
token_budget=int(os.getenv('COST_BUDGET', '8000')),
)
def built_in_cases() -> List[Sample]:
return [
Sample(120, 40, 3, 1800, 2, 8000),
Sample(10, 90, 0, 1800, 2, 8000),
Sample(15, 90, 4, 1800, 2, 8000),
Sample(20, 90, 2, 1800, 5, 5000),
]
def main() -> None:
single = from_env()
cases = [single] if single else built_in_cases()
for case in cases:
emit(case)
if __name__ == '__main__':
main()
Save the file as cost_gate.py beside your notes. If COST_AGE is set, the file prints one sample only. If COST_AGE is absent, the file prints four built-in cases.
Expected output, labeled as expected
This block is expected output from the inputs above. It is not a benchmark from a shared cluster. If your print differs, trust your run and diff the inputs.
shed estimate=7200 age=120
stay estimate=1800 age=10
fail_closed estimate=9000 age=15
shed estimate=5400 age=20
Case one sheds because age already beat slack. Case two stays because age, retries, and tokens fit. Case three fails closed because retries passed the cap.
Case four sheds because the resend estimate crossed budget. That fourth case is the free-capacity trap in miniature. The invoice can read zero while the budget is already gone.
Age wins when two gates would fire together. The function checks slack before retries and tokens. Read that order as policy, not as physics.
Inject the failure on purpose
You inject delay by editing queue age in the sample. You do not need a chaos tool for this lesson. Change only one field, then rerun the same file.
Commands
The shell lines below pass those injections as environment variables. They call the same file you already saved. Run them one at a time so the branch stays obvious.
python3 cost_gate.py | tee /tmp/cost-gate-cases.txt
COST_AGE=90 python3 cost_gate.py
COST_AGE=10 COST_RETRIES=2 python3 cost_gate.py
COST_AGE=10 COST_BUDGET=1000 python3 cost_gate.py
What those commands should print
Set queue age to the slack value and expect shed. Set retry count to the cap and expect fail closed. Set the budget under the estimate and expect shed.
The first injection sets age equal to slack and sheds. The second injection hits the retry cap and fails closed. The third injection sets a budget under the estimate and sheds.
shed estimate=1800 age=90
fail_closed estimate=5400 age=10
shed estimate=1800 age=10
Write the three results into your runbook as examples. Mark those results expected, not observed from production traffic. Keep the command output next to the config snippet.
A future reader should reproduce the branch without guessing. Leave the scratch file until you finish that diff. Remove that scratch file in the cleanup step below.
Thresholds and the reason you pick them
Prefer queue age against deadline slack for this class. Utilization can stay low while age still misses the turn. A busy CPU percent will not page the right owner.
Why slack beats utilization
Use token estimate against a class budget as the second gate. Retries multiply the prompt, so the next attempt is not free. A cap on retry count stops a loop that never converges.
Pick slack as the primary threshold for batch summarizers. The job is late when age meets slack, even if the host is idle. That rationale beats a utilization alert for this failure.
Do not copy these numbers into another class without a review. Interactive work needs a different clock than batch age. Probe traffic needs a different budget than summarizers.
Failure handling when the gate fires
Shed means you cancel this attempt and emit a terminal skip. Fail closed means you stop retries and surface the idempotency key. Stay means you pull once and send the prompt a single time.
On shed, write a skip event with age and slack. On fail closed, write a stop event with the retry count. On stay, write a start event before the model call.
If the caller cannot accept a skip, do not use this gate. Forcing a stay on a late job hides the miss. Page a human when the skip rate crosses the class budget.
Cleanup and rollback
Cleanup is part of the drill, not an afterthought. Delete the local sample output when the review ends. Clear any scratch queue you created for the injection.
rm -f /tmp/cost-gate-cases.txt
Roll back the gate if it sheds healthy jobs. Widen slack only after you show the false shed samples. Ship the old pull path behind the same config flag.
Keep the flag default off until the runbook is reviewed. A reviewed flag is your rollback, not a redeploy scramble. Name an owner for the flag before the first enable.
Where a free endpoint fits
You can point the stay branch at a free model endpoint. MonkeyCode is the open source project named in this outreach. Disclosure: This article was prepared as part of MonkeyCode's product outreach.
The outreach brief cites a free token allowance of ten million tokens. It also cites a free server option for hands-on runs. Treat both lines as claims you must re-check today.
Open the current project docs before you budget a class. Grants change, and this note does not freeze a quota. Do not paste a token grant into an alert threshold.
Use that project here only as a place to exercise stay. Remove the name and the gate still teaches the same decision. Do not invent install flags from this article.
Follow the current project docs for setup and cleanup. If the project docs disagree with the brief, believe the docs. A free worker is optional substrate, not the decision itself.
Who should skip this approach
Skip this gate if the job cannot be shed. Skip it if you cannot log tokens per attempt. Skip it if a human must read every summary.
Skip it for interactive chat with a hard latency target. A free worker can miss that clock without a loud CPU graph. This note does not claim a latency number for any model.
Skip it if your team will treat a free grant as a contract. Free capacity is a lab bet, not a capacity promise. When that bet is wrong, paid capacity is the honest path.
Limits of the drill
The script models resend cost on your laptop. It does not measure provider latency or cold start. It does not prove that a free server stays up.
The story numbers are fixtures, not fleet observations. Architectural guesses stay outside the printed branch text. This draft does not include a hosted queue run.
This draft does not time a free server or count a grant. You should run the file and compare the expected block. If the project docs disagree with the brief, believe the docs.
Update your budget from those docs, not from this article. Then rerun the injection cases before you enable the flag. That order keeps a free endpoint from becoming a silent queue.
What you do next
Replay last week's queue samples through the same decide function. Compare shed counts with the skips your users can accept. Enable the flag only for one batch class.
If you want a free endpoint for the stay branch, stop here. Read the current MonkeyCode docs and confirm the live grant. Only then should you point stay at that endpoint.
Top comments (0)