You do not control spend when the invoice shows up. You control it in the quiet second before the next model round leaves your process. A meter that has not moved is not permission to widen the loop. If that lane is free, the round still costs time, context, and the tool you are about to fire.
Picture a kitchen that has already called the next plate. The check in your hand describes the last plate, not the one on the grill. An agent run behaves the same way. The person typed one message. Behind it sits a chain of rounds: prompt, completion, tool result, prompt again.
That chain is the unit you should fence. A user turn is a poor proxy, because one turn can hide several model calls, and the later prompt is carrying tool text it did not start with. The round tends to get heavier, not lighter. You feel this when a short lookup starts pasting file bodies back into the prompt, and the next answer is mostly a recap of what the tool just returned.
The mistake is to treat a quiet meter as slack. Free model access feels like an empty pan, so you add another tool call, then another, because the dollar field stayed at zero. Free capacity is the wrong bet when the thing you are buying with it is a wider loop. Your caller still waits. Your tools still run.
Fence the unsent round
When the run starts, write down a ceiling in tokens and in milliseconds. After each completed round, keep the observation: prompt tokens, completion tokens, elapsed milliseconds, and the size of any tool payload you are about to attach. Before you send the next round, project it from the last observation plus that payload, and add the projection to what you have already spent. If the sum crosses the ceiling, return the partial result. Do not send the call and hope the other side is gentle.
The module below is a proposal you can run on your own machine. It is not a provider client, and it is not a bill. This article does not include a live model run. Token fields are observations you pass in. The byte divisor is a stand-in for a moment when you do not have that model's tokenizer, not a rule any vendor promised you.
"""round_fence.py
Proposal: refuse the next model round before you send it.
Counts are caller-supplied observations, not a provider invoice.
"""
from __future__ import annotations
from dataclasses import dataclass
@dataclass(frozen=True)
class Ceiling:
max_tokens: int
max_millis: int
@dataclass(frozen=True)
class Observation:
prompt_tokens: int
completion_tokens: int
elapsed_millis: int
tool_result_bytes: int = 0
@dataclass(frozen=True)
class Decision:
allow: bool
reason: str
projected_tokens: int
projected_millis: int
def project_next(last, pending_tool_bytes, bytes_per_token=4):
if bytes_per_token <= 0:
raise ValueError("bytes_per_token must be positive")
if pending_tool_bytes < 0:
raise ValueError("pending_tool_bytes must be non-negative")
extra = (pending_tool_bytes + bytes_per_token - 1) // bytes_per_token
if last is None:
return extra, 0
prompt = last.prompt_tokens + extra
return prompt + last.completion_tokens, last.elapsed_millis
def admit(
ceiling,
spent_tokens,
spent_millis,
last,
pending_tool_bytes,
bytes_per_token=4,
):
projected_tokens, projected_millis = project_next(
last, pending_tool_bytes, bytes_per_token
)
next_tokens = spent_tokens + projected_tokens
next_millis = spent_millis + projected_millis
if last is None:
return Decision(
False,
"no observation yet; calibrate before you loop",
next_tokens,
next_millis,
)
if next_tokens > ceiling.max_tokens:
return Decision(
False,
"next round would cross the token ceiling",
next_tokens,
next_millis,
)
if next_millis > ceiling.max_millis:
return Decision(
False,
"next round would cross the time ceiling",
next_tokens,
next_millis,
)
return Decision(True, "within ceiling", next_tokens, next_millis)
def drive(observations, ceiling, bytes_per_token=4):
"""First item is a calibration send. Later items must pass admit().
For a later item, tool_result_bytes is the payload you would attach
before that send, not a bill for a call you already made.
"""
spent_tokens = 0
spent_millis = 0
last = None
sent = []
for index, obs in enumerate(observations):
if last is not None:
decision = admit(
ceiling,
spent_tokens,
spent_millis,
last,
obs.tool_result_bytes,
bytes_per_token,
)
if not decision.allow:
return {
"sent": sent,
"refused_index": index,
"reason": decision.reason,
"projected_tokens": decision.projected_tokens,
"projected_millis": decision.projected_millis,
}
spent_tokens += obs.prompt_tokens + obs.completion_tokens
spent_millis += obs.elapsed_millis
last = obs
sent.append(index)
return {
"sent": sent,
"refused_index": None,
"reason": "loop exhausted",
"projected_tokens": spent_tokens,
"projected_millis": spent_millis,
}
Read admit and drive as two different doors. admit will not bless a cold start, because an elapsed time of zero is not a measurement. drive lets the first observation through so you have something to project from, then asks admit before every later round. That first send is a hole. Completion length is unknown until it comes back, so the first round can overshoot a ceiling you only enforce from round two. If you cannot afford one unbounded completion, do not start. Cap the prompt you already know, and keep the first tool read-only.
What the local tests lock
The tests use fake observations. They do not call a network. They lock the refusal shape: a heavier tool payload trips the token ceiling, a slow observation trips the time ceiling even when tokens would fit, a short loop inside both limits is allowed to finish, and a cold admit refuses.
"""test_round_fence.py — local proposal, no network."""
from round_fence import Ceiling, Observation, admit, drive
def test_refuses_when_next_round_crosses_token_ceiling():
ceiling = Ceiling(max_tokens=100, max_millis=10_000)
rounds = [
Observation(40, 10, 200, tool_result_bytes=0),
Observation(40, 10, 200, tool_result_bytes=80),
]
result = drive(rounds, ceiling)
assert result["refused_index"] == 1
assert result["sent"] == [0]
assert "token ceiling" in result["reason"]
# spent 50; next prompt 40 + ceil(80/4) plus last completion 10 => 70
# 50 + 70 = 120, which is the projected total on the refusal
assert result["projected_tokens"] == 120
def test_refuses_on_time_even_if_tokens_fit():
ceiling = Ceiling(max_tokens=10_000, max_millis=500)
rounds = [
Observation(10, 5, 400, 0),
Observation(10, 5, 400, 0),
]
result = drive(rounds, ceiling)
assert result["refused_index"] == 1
assert "time ceiling" in result["reason"]
assert result["projected_millis"] == 800
def test_allows_a_short_loop_inside_both_ceilings():
ceiling = Ceiling(max_tokens=500, max_millis=5_000)
rounds = [
Observation(20, 10, 100, 0),
Observation(20, 10, 100, 16),
]
result = drive(rounds, ceiling)
assert result["refused_index"] is None
assert result["sent"] == [0, 1]
def test_cold_start_is_not_a_permit():
decision = admit(Ceiling(100, 1000), 0, 0, None, 40)
assert decision.allow is False
assert decision.projected_tokens == 10
assert decision.projected_millis == 0
assert "calibrate" in decision.reason
Put both files in the same directory and run the tests, then a three-round sketch, then the cold-start door on its own. You want the second round refused and the third never considered.
python -m pytest test_round_fence.py -q
python - <<'PY'
from round_fence import Ceiling, Observation, drive
ceiling = Ceiling(max_tokens=100, max_millis=10_000)
rounds = [
Observation(40, 10, 200, 0),
Observation(40, 10, 200, 80),
Observation(40, 10, 200, 80),
]
print(drive(rounds, ceiling))
PY
python - <<'PY'
from round_fence import Ceiling, admit
print(admit(Ceiling(100, 1000), 0, 0, None, 40))
PY
If the files match the listings, pytest should pass, the sketch should print refused_index 1 with a token-ceiling reason and projected_tokens 120, and the cold-start call should refuse with projected_tokens 10 under the default divisor. That printout is the save. The unsent round never left the process, so it could not stall, and it could not call a tool. Compare the projection with a real usage line only after a separate calibration send. If they diverge, fix the divisor or stop projecting. Do not paper over the gap by moving the loop onto a lane whose price reads zero.
Keep the ceiling beside the run, not in a dashboard you will open tomorrow. A small file is enough for a single-operator script. The numbers below are placeholders for your own limit, not a recommendation and not a benchmark.
{
"max_tokens": 8000,
"max_millis": 20000,
"bytes_per_token": 4
}
Load them into Ceiling and log every decision, including the ones you allow. A week later the interesting field is refused_index. If you never refuse, the ceiling is decoration. If you refuse on the second round every time, the task is the wrong shape for a loop. Split the ticket. Hunting a cheaper model will not fix a fan-out you refused to count.
Where a free lane fits
You still need a place to learn what one round costs. That is a measurement job, not a production fan-out. Disclosure: This article was prepared as part of MonkeyCode's product outreach. MonkeyCode's free model access and free server option can host that measurement, if you already have them: one script, read-only tools, a ceiling written down, and a stop.
This note does not name models, quotas, hardware, or how long that access lasts. Those facts move, and a post that invents them is stale on arrival. Read the current terms before you depend on the lane. If it stalls or vanishes mid-run, the fence in your process is what stops the next call. The lane will not stop it for you.
Use that lane for a single-threaded calibration whose tools only read. Do not use it for a caller who cannot slip a deadline. Do not use it to widen the loop because the dollar field is empty. Do not use it as a second home for a round you already decided was too expensive. Relocating a round is not the same as refusing it. Do not aim a tool that writes, charges, or deletes at a lane you do not control, then act surprised when a slow completion applies the write late. The fence cannot unwind a tool that already committed.
Limits, and who should skip it
The projection is a heuristic. It assumes the next prompt is at least as large as the last one, plus a crude byte split. Real tokenizers disagree with a divide-by-four rule, sometimes by a wide margin. A completion can run longer than the previous one if the model starts a list and will not stop. Elapsed time on a free server is not a sample of any other lane, and one sample is not a distribution.
This fence does not reserve capacity. It does not replace a provider spend cap. It does not see tokens burned inside a tool that calls a model on its own. If a run can spawn a child loop, put the same check in the child, and pass the remaining budget down. Otherwise the parent ceiling is theater: the child spends, and the parent only counts the wrapper.
Skip this pattern when you need invoice-grade reconciliation. Pull the provider usage export and match that. Do not pretend the dataclass is a ledger. Skip it when a late refusal is worse than an over-ceiling round, for example a migration that must finish once it has locked a row. In that case use a hard provider cap and a human checkpoint. Skip it when you have no observation and you are tempted to invent a default, then ship the default as if it were a measurement. The cold-start refusal is the point. Calibrate, then loop.
If you already use MonkeyCode, run the local tests before you touch the network. Then send one calibration script through the free model access on the free server option, write down what a round actually returned, and leave production fan-out behind the fence in your own code. A quiet meter is still not a wider loop.
Top comments (0)