DEV Community

Quinn Li
Quinn Li

Posted on

Stamp a Deadline Before You Pick a Lane

A zero token price does not make a job free. If the work has a deadline, the queue wait and the retry fan-out are the invoice, and a free lane is the wrong place to send it.

You already know the cheap thrill: a prompt returns, the dashboard shows nothing owed, and the next attempt feels harmless. Then the job sits, a worker restarts, and the same prompt goes out again because nobody stamped a stop. The token counter stays polite. The clock does not.

Treat that gap as an admission problem, not a model problem. Before you pick a lane, write down the deadline, the worst wait you will tolerate, and how many attempts you will allow. If those three numbers cannot coexist, do not send the job. A later, smarter model will not repair a queue you never bounded.

Picture a ferry that boards for nothing and a taxi that prints a fare. The ferry is fine when you are carrying a sample and you can miss the evening, and it is a bad vehicle when someone is waiting on the other shore and the boat has no posted departure. A zero-price lane works like that ferry, while a lane you can actually hold works like the taxi. You match the vehicle to the promise you already made, not to whichever ride looks cheaper.

What you are actually buying

Time, tokens, and retries are one budget wearing three names. Queue time is time you do not control. Generation time is time you half control. Tokens are the residue of both, because a retry sends the prompt again and often sends a longer context than the first try.

A free price zeroes only the third name. It does not zero the first two. If the machine you did not hold is busy, preempted, or simply slow to schedule, you can miss a deadline while spending no money at all. That is still a failed job, because your caller does not grade you on the invoice.

Retries make the math worse. Two attempts are not twice as innocent as one when each attempt waits in line. Four attempts with a cold start can cost more wall clock than one attempt on a lane you held, even when the held lane has a nonzero token rate. Compute that comparison before you feel loyal to the zero.

This is also how a casual rerun becomes an incident. The second run is not a copy of the first. It is a new wait, a new context, and sometimes a new failure mode. If you cannot point at a hold, or at an explicit decision that this job may miss its window, you are guessing with someone else's afternoon.

A smaller model can still be the right model. That choice is about output quality per token, and it belongs in a different column from lane admission. A cheap model that starts immediately can beat a free model that sits. Do not collapse those columns, or you will celebrate a price you never paid while the queue eats the window.

A ledger you can run locally

The snippet below is a proposal. It is not a benchmark, and it was not executed against a live provider for this article. It only encodes the rule: deadline work does not ride an unreserved lane, and the retry fan-out must fit inside the window.

from dataclasses import dataclass

@dataclass(frozen=True)
class Attempt:
    lane: str
    queue_wait_s: float
    gen_s: float
    prompt_tokens: int
    completion_tokens: int
    succeeded: bool

def summarize(attempts, prompt_rate, completion_rate):
    tokens = 0
    wall = 0.0
    money = 0.0
    for a in attempts:
        tokens += a.prompt_tokens + a.completion_tokens
        wall += a.queue_wait_s + a.gen_s
        if a.lane != "free":
            money += a.prompt_tokens * prompt_rate
            money += a.completion_tokens * completion_rate
    return {"tokens": tokens, "wall_s": wall, "money": money}

def admit(deadline_s, predicted_wait_s, max_attempts, predicted_gen_s, lane):
    if deadline_s is not None and lane == "free":
        return False, "deadline jobs do not ride the free queue"
    worst = max_attempts * (predicted_wait_s + predicted_gen_s)
    if deadline_s is not None and worst > deadline_s:
        return False, "retry fan-out breaks the deadline"
    if max_attempts < 1:
        return False, "nothing to run"
    return True, "ok"
Enter fullscreen mode Exit fullscreen mode

Save it as lane_admit.py. Then check the branches with the standard library. No network, no key, and no account are required.

python - <<'PY'
from lane_admit import Attempt, summarize, admit

probe = [
    Attempt(lane="free", queue_wait_s=40, gen_s=6,
            prompt_tokens=900, completion_tokens=180, succeeded=False),
    Attempt(lane="free", queue_wait_s=55, gen_s=7,
            prompt_tokens=900, completion_tokens=40, succeeded=False),
]
held = [
    Attempt(lane="reserved", queue_wait_s=1.5, gen_s=5,
            prompt_tokens=900, completion_tokens=180, succeeded=True),
]
assert summarize(probe, 0.000002, 0.000008)["money"] == 0.0
assert summarize(probe, 0.000002, 0.000008)["wall_s"] > summarize(held, 0.000002, 0.000008)["wall_s"]
assert admit(30, 40, 2, 6, "free")[0] is False
assert admit(12, 8, 2, 5, "reserved")[0] is False
assert admit(None, 40, 1, 6, "free")[0] is True
print("lane checks passed")
PY
Enter fullscreen mode Exit fullscreen mode

Read the three decisions in order. The free lane is refused because a deadline was set, and the held lane is refused too, in this made-up example, because two attempts still overflow twelve seconds. Only the last call is admitted, because you passed deadline_s=None, and that call is a probe. It may wait or be dropped, but it must not sit on the path that pages a human.

The rates in the command are placeholders so the arithmetic is visible. Do not treat them as a vendor price. Swap in the figures from your current billing page when you care about money. When the lane is free, the function already forces the money term to zero, so you can still lose on wall_s.

Keep the inputs honest. predicted_wait_s should be a high percentile from your own recent attempts, not a number you liked in a demo. If you have no samples, pass a wait you are willing to be wrong about, and treat an admit as provisional. A function cannot discover a queue it has never seen.

Failed rows still count

You will want to delete the failures before you summarize. Don't, because a timeout still occupied a slot and still attached every prompt token you sent. If the server returned a partial completion, those tokens belong in the row. Drop them and the two charts disagree, so you fix the wrong one.

When you debug a slow evening, sort by queue_wait_s before you sort by answer length. A long answer is a generation problem, and a long wait is a lane problem. Mixing them is how a team swaps models for a week and never touches scheduling. The log has to separate those, because a dashboard that reports no token spend will not.

python - <<'PY'
import json
rows = [json.loads(line) for line in open("attempts.jsonl")]
rows.sort(key=lambda r: r["queue_wait_s"], reverse=True)
for r in rows[:5]:
    print(r["job_id"], r["lane"], r["queue_wait_s"], r["succeeded"])
PY
Enter fullscreen mode Exit fullscreen mode

Give each attempt an idempotency key before it leaves your process. A retry must be the same job, not a new job with a fresh context and a second copy of the prompt. Without that key, this ledger under-counts, and your notebook will disagree with the platform about what you owe in time. Store the key, the lane, the wait, the token estimates, and whether you intended a deadline.

Run the zero-price path on purpose, as a bounded sample, not as the default route. Cap it at a count you can read, such as one hundred attempts that carry no deadline, and record wait, generation time, and whether the call returned. After that sample, put a high percentile wait into predicted_wait_s. If that percentile already exceeds the tightest deadline you owe a caller, that lane is disqualified for that caller no matter what the token price says.

Where a zero price earns a place

Disclosure: This article was prepared as part of MonkeyCode's product outreach. MonkeyCode, as the operator described it for this note, includes a free model path and a free server option. Those are availability claims, not a measured quota, a hardware class, or a promise that the queue is short. Use them for the probe half of the workflow: prompt diffs, parser checks, and replaying recorded waits through admit.

Keep caller-visible work, deploy gates, and anything you will retry inside a window on a lane you can hold, or shed the job in your own code. Do not paste a token ceiling into the runbook from memory. Limits, model names, and windows move, and a stale figure is worse than a blank. Read the project docs on the day you enroll, write down what you observed, and attach that note to the job id.

If the docs are silent, record the free lane as zero price and unknown wait. Unknown wait is not a hold. If your attempt logs already exist, replay a week of them through admit before you move a user-facing job onto that free path. That replay is the only next step this note asks for.

What this will not do

The function does not know why a queue grew. It will not see a provider incident, a bad prompt, or a client that resubmits on every non-200. You still own the stop condition. A refusal inside admit is only useful if the caller actually stops.

It will not turn yesterday's samples into today's capacity. A quiet Tuesday does not hold Wednesday. If predicted_wait_s is a wish, the function will approve a lane you merely labeled reserved. Feed it measurements, or feed it a ceiling you can afford to miss.

Skip this approach if you need a contractual latency target and you have no lane you can hold. A refusal cannot manufacture capacity. Skip it if your team will treat a false result as advice and send the job anyway. Skip it, too, if the plan is to drop secrets, customer text, or production identifiers into a free probe because the token price is zero.

A free process is still a process you must account for, so check the current docs before you decide what it is allowed to see. Label every number you have not measured. The waits in the sample command are illustrations, picked so the branch is obvious, and your queue will not match them. If you publish a chart from this ledger, keep estimated tokens apart from invoiced tokens, or the fiction will win the meeting.

Stamp the deadline first. If the wait and the retries cannot fit inside that stamp, do not pick a lane. Learn the wait on the free path, and hold a lane or shed the job when a person is actually waiting on you.

Top comments (0)