You ship a summarize button on a Saturday afternoon. It works on the first try. By Wednesday the usage bill is larger than the project will ever earn.
That gap is the real risk in AI side projects. The model call is a five-line function. The spend is unbounded unless you bound it.
So this build has one rule: the weekend demo must be unable to surprise you. The artifact is small. It is a token budget guard with a fallback ladder.
Scope: three things, one evening
The build ships exactly three pieces:
- One function that summarizes a fixed block of text.
- A guard that reserves tokens before each call.
- A fallback ladder that still returns something at the ceiling.
Everything else was cut on purpose. The cuts are listed at the end.
Step 1: Decide where the free tier fits
Before writing the guard, decide what it protects. A free tier changes the shape of that decision.
MonkeyCode is an open-source AI coding project. Its operator says the free plan includes free model access, a 10M token allowance, and a free server option. I have not audited those numbers. Treat them as a starting estimate and check the current terms yourself.
Disclosure: This article was prepared as part of MonkeyCode's product outreach.
The free server matters more than it sounds. It removes the "where does this run" question from a weekend build. You lose one moving part and gain one hard constraint: the allowance is finite.
Step 2: Turn the allowance into a ceiling
Do the arithmetic before you write any code. The math below is mine, not a vendor promise.
Assume 1,500 tokens per request, prompt plus completion. That number is a placeholder. Measure yours on day one.
| Requests per day | Tokens per day | Days on a 10M token allowance |
|---|---|---|
| 50 | 75,000 | ~133 |
| 200 | 300,000 | ~33 |
| 1,000 | 1,500,000 | ~6 |
Pick a daily ceiling well under the monthly figure. For this build I used 40,000 tokens per day. That is roughly 26 requests. Small on purpose.
Step 3: Reserve tokens before the call
The guard checks the budget before it calls the model. Never after.
# guard.py — template, adapt to your provider's response shape
import json, os, time
from dataclasses import dataclass
from pathlib import Path
STATE = Path("var/budget.json") # one file, one number, easy to read
class BudgetExceeded(RuntimeError):
pass
@dataclass
class Budget:
ceiling: int
spent: int = 0
@classmethod
def load(cls, ceiling):
if STATE.exists():
return cls(ceiling, json.loads(STATE.read_text())["spent"])
return cls(ceiling)
def remaining(self):
return max(self.ceiling - self.spent, 0)
def charge(self, used):
self.spent += used
tmp = STATE.with_suffix(".tmp")
tmp.write_text(json.dumps({"spent": self.spent, "ts": time.time()}))
os.replace(tmp, STATE) # atomic enough for one process
def guarded_call(client, prompt, budget, estimate=1500):
if budget.remaining() < estimate:
raise BudgetExceeded(f"{budget.remaining()} < {estimate}")
resp = client(prompt)
budget.charge(resp["usage"]["total_tokens"]) # check your provider's field
return resp["text"]
Two choices here are deliberate. The ceiling check uses an estimate, so the guard trips before the overspend instead of after. And the state is a single JSON file, not a log stream. One process, one number.
Step 4: Define the fallback ladder
A guard without fallbacks just breaks the feature. Decide each level in advance.
| Level | Trigger | Behavior | The user sees |
|---|---|---|---|
| 0 | remaining > 40% | Full call | Full summary |
| 1 | 10–40% | Shorter prompt, capped output | Shorter summary |
| 2 | remaining < estimate | Return the last cached result | Slightly stale summary |
| 3 | Two provider errors | Freeze the feature for the hour | Plain retry message |
Write this table before you write the code. It becomes the demo script, not documentation.
Step 5: Prove the guard with three commands
The test needs no network call. A stub client is enough.
# tests/test_guard.py
import pytest
from guard import Budget, BudgetExceeded, guarded_call
def test_guard_trips_before_overspend():
b = Budget(ceiling=300)
stub = lambda p: {"text": "ok", "usage": {"total_tokens": 120}}
guarded_call(stub, "hello", b, estimate=120)
guarded_call(stub, "hello", b, estimate=120)
with pytest.raises(BudgetExceeded):
guarded_call(stub, "hello", b, estimate=120)
Run the three checks below. Each one answers a single question.
# 1. Does the guard trip at the right call?
pytest tests/test_guard.py -q
# 2. Does a full day of traffic stay under the ceiling?
python scripts/dry_run.py --prompts fixtures/prompts.txt --ceiling 40000
# 3. What does the state file say after the run?
python -m json.tool var/budget.json
If check two fails, your estimate is wrong. Lower the ceiling and rerun. No model call required.
What I skipped
- Auth. One local user, one key, no accounts.
- Streaming. Buffered responses are easier to count.
- Retry with backoff. The guard makes blind retries expensive.
- RAG and embeddings. A fixed text block proves the guard.
- Evals. Replaced by one assertion about the trip point.
- A database. One JSON file is enough.
Limitations and who should skip this
Free allowances and free servers change. Verify the terms before you depend on them.
This guard is per-process. Two workers mean two counters and a real overspend. Move the state to Redis or Postgres first.
Estimates drift. A long prompt can burn three times your placeholder. Record the estimate and the actual usage together, then tune.
Skip this approach if you need multi-tenant quotas, streaming chat, or audited billing. Build the boring version first.
Ship it, then watch the number
The weekend goal was simple. Build a demo that cannot surprise you on Wednesday.
The guard is about twenty lines. The ladder table is ten. Together they turn a free allowance into a bounded experiment.
If you want a free tier to test against, MonkeyCode's operator advertises free model access and a free server option. Keep the ceiling low for the first day. Raise it only when the numbers justify it.
Top comments (0)