A free model is a draft engine, not a release gate. Ship a local budget check before any network call. Kill the loop if that check cannot refuse an over-budget prompt.
The only hypothesis
This spike tests one hypothesis and nothing else. A local gate must refuse an over-budget prompt before any send. Ninety minutes is the entire window for evidence.
A pass means three local checks go green. A fail means you stop and write the kill note. A fluent model answer does not count as a pass.
Why this angle
DEV's recent front page leans toward profile toys and portfolio pieces. Those posts show craft, not a spend kill switch. This account already covered bad oracles, empty collects, and weakened tests.
This note stays on a different failure mode. The failure is sending a prompt before you can refuse. The artifact is a gate, a fixture pair, and a decision table.
Agent demos often start inside the prompt box. This spike starts on the refuse path instead. If you cannot show a killed call, you have a demo, not evidence.
Claims you must re-check
MonkeyCode is an optional runner for the later minute block. Disclosure: This article was prepared as part of MonkeyCode's product outreach.
The outreach brief cites free model access and a free server option. The same brief cites a free allowance of ten million tokens. This draft did not open a primary pricing page.
Do not plan capacity from that sentence alone. Open the current project docs and copy the live number into your log. If the page disagrees with the brief, the page wins.
No model name is assumed in this protocol. No hardware size is assumed in this protocol. No uptime promise is assumed in this protocol.
Clock
Use one timer and do not extend it. A slipped clock turns the spike into a demo. Write the start time before you read anything.
- Minutes 0 to 15: read the live free-tier page and record every visible limit.
- Minutes 15 to 30: write one short fixture and one over-ceiling fixture.
- Minutes 30 to 50: run the gate on both fixtures and save logs.
- Minutes 50 to 70: send at most one draft, then save the raw body.
- Minutes 70 to 85: fill the decision table from logs, not from memory.
- Minutes 85 to 90: mark ship or kill, then close the editor.
If minute 70 arrives with no saved response, skip the model call. The local gate can still ship without that call. Do not borrow time from the table block.
Arithmetic you can audit
Take a local budget of 180 tokens and a margin of 0.2. The ceiling becomes 144 tokens under that setting. The heuristic maps one token to about four characters.
A 100-character prompt estimates to 25 tokens. That case should be allow, with exit code 0. An 800-character prompt estimates to 200 tokens.
Two hundred tokens sits over a 144 ceiling. That case should be kill, with exit code 2. These figures are arithmetic, not a recorded lab run.
Use the same budget in every command for this spike. Changing the budget mid-run voids the decision table. Record the formula beside the log so a reviewer can recompute it.
The gate
The script below is a proposed tool, not a measured benchmark. The token count is a character heuristic, not a vendor tokenizer. Run it yourself before you trust any exit code.
#!/usr/bin/env python3
"""Budget gate for a 90-minute drafting spike.
Heuristic only: about four characters per token.
Not a vendor tokenizer. Not a quota client.
"""
import argparse
import json
import sys
from pathlib import Path
REQUIRED = {"summary", "risk", "next_check"}
def estimate_tokens(text: str) -> int:
return max(1, (len(text) + 3) // 4)
def main() -> int:
parser = argparse.ArgumentParser()
parser.add_argument("--prompt", required=True)
parser.add_argument("--budget", type=int, required=True)
parser.add_argument("--margin", type=float, default=0.2)
parser.add_argument("--response", default="")
parser.add_argument("--log", default="spike-log.json")
args = parser.parse_args()
prompt = Path(args.prompt).read_text(encoding="utf-8")
estimate = estimate_tokens(prompt)
ceiling = int(args.budget * (1.0 - args.margin))
decision = "kill" if estimate > ceiling else "allow"
record = {
"estimate_tokens": estimate,
"budget_tokens": args.budget,
"margin": args.margin,
"ceiling_tokens": ceiling,
"decision": decision,
"response_checked": False,
"missing_keys": [],
}
if decision == "allow" and args.response:
payload = json.loads(Path(args.response).read_text(encoding="utf-8"))
missing = sorted(REQUIRED - set(payload))
record["missing_keys"] = missing
record["response_checked"] = not missing
if missing:
record["decision"] = "kill"
Path(args.log).write_text(
json.dumps(record, indent=2) + "\n",
encoding="utf-8",
)
print(json.dumps(record))
return 0 if record["decision"] == "allow" else 2
if __name__ == "__main__":
raise SystemExit(main())
Commands
Run all three commands from a clean directory. Point the budget at 180 so the over fixture must die. Keep the shell history together with the logs.
python3 budget_gate.py --prompt under.txt --budget 180 --log under.json
python3 budget_gate.py --prompt over.txt --budget 180 --log over.json
python3 budget_gate.py --prompt under.txt --budget 180 --response draft.json --log response.json
Read the outcomes from files, not from the terminal scrollback. Match the exit code to the decision field. A green terminal color is not evidence here.
- The under log shows allow when the short fixture sits under the ceiling.
- The over log shows kill, and the process exits 2.
- The response log stays allow only when all three required keys exist.
Fixture rules
Write under.txt as one task in under 100 characters. State the file to read, the output shape, and the stop rule. Do not include secrets, tokens, or customer rows.
Write over.txt by repeating a harmless sentence until the estimate clears 144. Count characters before you run the gate. If the over file still allows, the fixture is wrong, not the hypothesis.
Store both fixture files in the same folder as the script. A reviewer should rebuild the estimate without opening a chat window. That rebuild is part of the spike evidence.
Draft contract
Use the free model only inside minutes 50 to 70. Ask for one JSON object and no markdown fence. Limit the keys to summary, risk, and next_check.
Save the raw body as draft.json before you edit anything. If the body still has a fence, the check fails. Do not repair it by hand and then claim a pass.
{
"summary": "one sentence on the change",
"risk": "one sentence on how the gate can miscount",
"next_check": "recompute ceiling from budget and margin"
}
That object is a shape example, not a model output from a live call. Your file must come from the run you actually made. If you skip the call, omit the file and mark the row skipped.
Decision table
Fill every row from a file on disk. An empty evidence cell means kill for that row. Do not backfill any cell from memory after the timer ends.
| Check | Evidence file | Ship bar | Kill bar |
|---|---|---|---|
| Live allowance copied | notes.md |
Number, source URL, and read time | Figure taken only from this article |
| Under fixture | under.json |
decision is allow
|
Missing log or a nonzero exit |
| Over fixture | over.json |
decision is kill and exit 2 |
Gate allows the long prompt |
| Optional draft | draft.json |
Three keys, then a local recheck | Prose, a fence, or a missing key |
| Scope | timer note | Stopped by minute 90 | Extra retries after the bell |
What the free server is for
Use a free server only to run this script or one draft. The server does not decide the ship mark. Your local JSON log decides the ship mark.
Free model access matters only in the optional block. One model call is the hard cap here. A second call is a scope break, even if the first looked weak.
If the live docs show no free server, skip that block. The three local checks still answer the hypothesis. Write not offered in the notes instead of guessing a host name.
Ship or kill
Ship only if the three local bars pass. The optional draft may be absent from the folder. A missing draft is a recorded limit, not a silent pass.
Kill the spike if any line below is true. Treat each bullet below as a hard stop. Do not average a weak row into a pass.
- The over fixture is allowed through the gate.
- The live allowance was never copied from a current page.
- The response file is prose, or it lacks a required key.
- The hypothesis changed after minute 30 on the clock.
- The result needs a second model to look valid.
A kill note is a valid spike result. Name the failed check in one short line. Then stop, and do not open a follow-up prompt.
Limits
Counting error
This heuristic miscounts code, CJK text, and tool schemas. A 20 percent margin only reduces that risk. It does not remove the counting risk itself.
The default margin of 0.2 is a spike choice, not a product setting. Change it only before minute 30, and record the new ceiling. A silent margin edit after a kill is a failed spike.
Network gap
The script is written never to call a network API. That network gap is intentional in this design. A gate that must phone home cannot prove a refuse-before-send claim.
Shape versus truth
Key presence is not the same as truth. A filled risk field can still be wrong. Pair this gate with your own tests before any merge.
Moving terms
Free access can change without a notice in this article. A log from today is not a contract for next week. Re-read the docs on the day you actually run.
The character heuristic is not the same as billed usage. Vendors count tokens with their own tokenizer rules. If you need a hard invoice match, use that tokenizer instead.
Who should skip
Skip this spike if you need audited billing or a signed quota letter. Skip it if the prompt may hold secrets, production writes, or customer data. Skip it if you cannot stop when the timer hits 90.
Also skip it when the question is a public leaderboard rank. This protocol does not rank models against each other. It checks refusal, payload shape, and elapsed time only.
Skip it if your team already has a vendor tokenizer in preflight. Use that vendor tokenizer for the budget decision. Do not add a second, weaker estimator beside a trusted one.
After the bell
Archive the logs, both fixtures, and the filled table. Keep the live-doc note with a URL and a timestamp. That small file bundle is the whole spike.
If you want a free runner later, read the current MonkeyCode docs first. Repeat the same three local checks on that later day. Do not raise the budget until the over fixture still dies.
Top comments (0)