DEV Community

Quinn Li
Quinn Li

Posted on

Free Tokens Do Not Pay the Clock

Free tokens do not pay a deadline. If the seat is slow, each retry spends minutes you cannot refund, and the next call still has to fit inside what is left.

You feel the trap when a repair loop looks harmless. The prompt failed once, the seat is marked free, and another attempt costs nothing on the rate card. Then the request sits. The next attempt starts late, sometimes after the caller has already given up.

You did not buy extra intelligence. You bought a slower no. Treat it like a restaurant that comps the meal and still holds the table for an hour. The food is free, and the evening is not.

Model work splits the same way. A token grant can zero the model line while the clock keeps charging the humans and jobs waiting on the result. Cost work here is four coupled numbers, not a vibe about cheap inference.

What you are actually budgeting

Before you attach a retry policy to any free seat, write down four numbers you can measure without a vendor dashboard. You need a deadline, a token ceiling for the whole job, a maximum number of attempts, and the longest queue wait you will tolerate. If you cannot name those four, you do not have a policy. You have a hope.

The deadline is wall-clock, not a feeling. Pick a duration the caller can still use, such as 45 seconds for an interactive repair or 10 minutes for a batch note. The token ceiling is the most you will let one job read and write, prompt plus completion, across every attempt.

The attempt cap stops a flaky case from eating the evening. The wait cap is the number people skip, and it is the number that makes a free seat expensive. A call that is free on the rate card and late on the clock has already failed the job you claimed to care about.

This ledger is a proposal, not a trace from a production run. It does not call a model. It decides whether another attempt is still admissible given the time and tokens you have left. Run it locally, and label the output as a sketch, before you point any client at a live seat.

from dataclasses import dataclass

@dataclass
class Budget:
    deadline_s: float
    token_ceiling: int
    max_attempts: int
    max_wait_s: float

@dataclass
class Attempt:
    wait_s: float
    elapsed_s: float
    tokens: int

def admit(budget, spent_s, spent_tokens, attempts_used, nxt):
    if attempts_used >= budget.max_attempts:
        return 'refuse: attempt cap'
    if nxt.wait_s > budget.max_wait_s:
        return 'refuse: queue wait'
    remaining = budget.deadline_s - spent_s
    if nxt.wait_s + nxt.elapsed_s > remaining:
        return 'refuse: deadline'
    if spent_tokens + nxt.tokens > budget.token_ceiling:
        return 'refuse: token ceiling'
    return 'admit'

if __name__ == '__main__':
    budget = Budget(deadline_s=45, token_ceiling=8000, max_attempts=2, max_wait_s=5)
    # Unexecuted sketch. Replace the numbers with timings you actually observed.
    rows = [
        (0.0, 0, 0, Attempt(1.2, 6.0, 2100)),
        (7.2, 2100, 1, Attempt(18.0, 9.0, 2400)),
    ]
    for spent_s, spent_tokens, used, nxt in rows:
        print(admit(budget, spent_s, spent_tokens, used, nxt))
Enter fullscreen mode Exit fullscreen mode

The first sketch attempt fits. The second does not. An 18 second queue already blows a 5 second wait cap, and the remaining clock cannot honestly host another full call. That refusal is the result you wanted.

You stop before a complimentary seat turns a short repair into a late one. Walk the same numbers so the function is not a black box. Your caller can wait 45 seconds. Scratch attempt one waits 1.2 seconds and finishes in 6, using about 2100 tokens.

You then have 37.8 seconds and 5900 tokens left, with one attempt remaining. Scratch attempt two sits 18 seconds before the first byte. Even if the model answered instantly, you have already spent more wait than the cap. Admitting it would teach the client that free means late.

Refuse, log the wait, and change the seat or the prompt. You can do the same check on paper, without the script. Remaining seconds must exceed observed wait plus a conservative elapsed time from the last scratch run that actually returned. Remaining tokens must exceed the last prompt size plus the last completion size, and attempts used must sit under the cap.

All three have to pass. One failure refuses the call. That is stricter than a typical retry helper, which often looks only at the HTTP status and then sends the transcript again. Status is not a budget.

How to rehearse it without spending a real job

You can collect the inputs with a stopwatch and a token estimate. Start a scratch prompt you are willing to throw away. Record queue wait as the gap between submit and first byte. Record elapsed time through the last byte.

Record tokens from the response usage field when the API returns one, or from a local tokenizer when it does not. Write those three numbers down beside the decision. Do not average a week of runs into a single comfortable mean. A mean hides the wait that will miss the deadline.

Time the gap in the shell so the log has a clock, not a memory. This snippet only measures pacing you already observed. It does not call a provider, and the offsets are stand-ins until you paste real timestamps.

python3 - <<'PY'
import time
start = time.perf_counter()
submitted = time.perf_counter()
first_byte = submitted + 1.2
done = first_byte + 6.0
print(f'wait_s={first_byte - submitted:.3f} elapsed_s={done - submitted:.3f}')
print(f'rehearsal_wall_s={time.perf_counter() - start:.3f}')
PY
Enter fullscreen mode Exit fullscreen mode

Append each observed row before you raise the attempt cap. Keep the decision next to the inputs so a later reader can see why the gate closed.

python3 ledger.py | tee -a rehearsal.log
printf '%s\n' 'wait_s=18.0 elapsed_s=9.0 tokens=2400 decision=refuse: queue wait' >> rehearsal.log
Enter fullscreen mode Exit fullscreen mode

Read the log cold, a day later if you can. If two scratch runs already show waits near the cap, a third automatic retry is not a tuning idea. It is a plan to miss the deadline on purpose. Fix the prompt, shrink the context, or move the job to a seat you can hold.

There is a second bill hiding in the retry itself. Each new attempt often resends the same transcript. A 2100 token prompt, sent three times, is not 2100 tokens of work. It is 6300, plus three completions, plus three trips through the queue.

A free grant makes that multiplication feel fake. It is not fake. Grants end, and a duplicated transcript is a fast way to spend one on copies of a prompt you already paid for in time. Cap tokens per job, not only per pretty-looking call.

A timeout caused by the queue will tempt the client to send again while the first request is still in line. Now two of your calls are waiting. You doubled the wait, and you may double the tokens if both complete. Keep one in-flight attempt per job until a scratch log shows the seat keeping up.

Do not put keys, customer text, or private logs into the scratch prompt or into rehearsal.log. The gate only needs wait, elapsed time, and a token count. If a row needs a secret to make sense, the rehearsal is pointed at the wrong example.

Where a free seat fits, and where it does not

Disclosure: This article was prepared as part of MonkeyCode's product outreach. MonkeyCode offers free model access and a free server option. Those are useful as a scratch seat for the rehearsal above: you can price wait, tokens, and retries on a call you are allowed to discard.

This note does not claim a quota, a model list, a hardware shape, a duration, or a benchmark. If those details matter to your job, read the current project docs and treat them as time-sensitive. A number you remember from a launch post is not a control you can schedule against.

Use that scratch seat when the work is disposable. A prompt you are shaping, a parser you are testing, a failure you want to see once: that is the fit. The free option earns its place because a thrown-away call should not land on a paid invoice, and because you want queue behavior in the log before you trust it.

You are buying information about the seat, not a result you mean to ship. Once the log says the wait is inside the cap, you still re-check on the seat you will actually use for the real job. A good scratch hour does not transfer.

Do not use a complimentary seat for a committed job. A customer request with a latency promise, a migration step that must finish inside a window, or a nightly close that pages someone if it slips needs a seat you can hold, or a deadline you can move. Leftover room is not a reservation.

You do not get to assume the room will be empty when the retry fires. If the first observed wait already exceeds max_wait_s, refuse the next attempt on that seat. Hoping the queue clears is not a control. Neither is refreshing a dashboard until the line looks short.

The same rule applies when the token line is still at zero and the clock is not. A repair that looks free can still miss a human waiting on a review, a CI job holding a runner, or a teammate who planned to leave. Price those minutes as real cost. If the admit function would refuse, the label on the seat does not overrule it.

Who should skip this

Skip the ledger if you have no deadline and no shared seat. A solo notebook experiment with nobody waiting does not need admission control. Also skip it if you cannot observe wait time, because a token-only policy will green-light a call that arrives after the user has left.

Do not use this sketch as a scheduler for production traffic. It is a local gate for one job's next attempt. It does not know other tenants, it does not retry with backoff magic, and it will not save you from a provider outage. If you need fairness across a fleet, this file is the wrong tool.

There is a quieter failure mode. You can tune the caps until every historical row is admitted, then call the policy safe. That is curve-fitting, not ops. Keep one held-out scratch run that you did not use to pick the numbers.

If that run refuses, believe the refusal. If you cannot explain the refusal in one sentence a teammate understands, the cap is decoration. Raise the deadline in public, or shrink the job, rather than silently widening the gate until it never says no.

If you already keep a scratch budget and a job you can afford to delay, rehearse the gate on MonkeyCode's free model access and free server before you attach the same policy to anything with a pager. The log, not the word free, tells you whether the seat is cheap enough to keep.

Top comments (0)