How Allowance works: an agent that runs inside a hard spend envelope of tokens, dollars and wall-clock time — enforced in code, with tiered approvals, a negotiation path, and a ledger that reconciles against the provider's own usage numbers.
Built with Inngest AgentKit, TypeScript, Express 5, React 18 and SQLite. About 4,500 lines of app code, plus 474 lines of scenario verification and 278 lines of unit tests.
Every agent framework in 2026 ships a budget knob, and almost every one of them implements it as
a sentence in a system prompt:
You have a budget of 40,000 tokens. Please stop when you approach the limit.
That is not a budget. That is a request. It works right up until the model has a bad day, at which
point the spend is real and the "limit" was a paragraph.
Allowance is a small control panel that treats the envelope as a property of the code:
give it a task + an envelope -> watch the ledger fill step by step
-> at 60% it files a budget request; you approve or deny
-> on deny it stops cleanly with a partial result and a cost report
-> reconcile the ledger against what the provider reported
Three things turned out to be much more interesting to build than the UI: the enforcement point,
the accounting split, and durability. This post is about all three, plus the four bugs that
verification caught and I would not have found by reading the code.
1. The mechanic is four lines
Before every model call, in code:
remaining = min( cap_tokens - used_tokens,
usdToTokens( cap_usd - used_usd ),
timeToTokens( deadline - now ) )
estimatedCost > remaining -> the step does not execute
├─ a cheaper plan fits? re-plan, log it
└─ otherwise REFUSED { at: <resource> }
The min() is the whole idea. Three independent resources, and the tightest one binds. A run
can have 95% of its tokens left and still be refused because it spent its seconds.
In TypeScript (server/src/agent/budget.ts), with the deadline checked first because it can
expire mid-plan with money unspent:
export function gateFor(run: RunRow, now: number, needTokens: number): GateDecision {
const envelope = snapshotEnvelope(run, now);
if (envelope.deadlineBreached) {
return { allowed: false, reason: "deadline", at: "time", canReplan: false, ... };
}
const allowed = needTokens <= envelope.remainingTokenEquivalent;
if (allowed) return { allowed: true, reason: "ok", at: null, ... };
const cheaperPlanFits = envelope.remainingTokenEquivalent >= MIN_CALL_TOKENS;
return {
allowed: false,
reason: "envelope_exhausted",
at: envelope.binding, // "tokens" | "usd" | "time"
canReplan: cheaperPlanFits,
...
};
}
at: envelope.binding is what makes refusals useful instead of merely terminal. The UI says
"the tokens envelope was exhausted before the next step could execute — needed 909, headroom
266", and the transcript keeps the numbers.
Money-to-tokens, priced conservatively
export function usdToTokens(usd: number, prices: Prices): number {
// Priced at the dearer of input/output so headroom is never more generous
// than the cheapest possible call.
const perToken = Math.max(prices.per1kInput, prices.per1kOutput) / 1000;
return perToken > 0 ? Math.floor(usd / perToken) : 0;
}
There is no hardcoded cost anywhere in the project. Dollars come from a price table the user can
edit in the Settings panel — which is precisely why every derived figure is labelled an estimate.
A budget you cannot re-price is a bill, and this is not a billing system.
2. The counter must not live in a variable
This is the sentence the whole build is about, so it's worth being precise about why it's hard.
A JS variable holding remainingBudget is wrong in three different ways at once:
- Retries double-count. A step that fails after charging and runs again spends the same money twice, or charges nothing twice, depending on where the throw happens.
- Restarts lose it. The process dies, the number dies, and the next run spends like it's day one.
- You cannot audit it. The number that gated step 7 no longer exists by the time you're looking at step 7's cost.
So the envelope lives in SQLite, and — this is the part I'd skipped the first time — the counters
are recomputed from the ledger rather than incremented:
export function commitLedgerAndCounters(entry: LedgerInsert): void {
const tx = getDb().transaction(() => {
getDb().prepare(`
INSERT INTO ledger (run_id, step_key, ...) VALUES (@runId, @stepKey, ...)
ON CONFLICT(run_id, step_key) DO UPDATE SET
input_tokens = excluded.input_tokens,
output_tokens = excluded.output_tokens,
usd = excluded.usd, ...
`).run(entry);
getDb().prepare(`
UPDATE runs SET
used_tokens = (SELECT COALESCE(SUM(input_tokens + output_tokens), 0) FROM ledger WHERE run_id = @runId),
used_usd = (SELECT COALESCE(SUM(usd), 0) FROM ledger WHERE run_id = @runId),
executed_steps = (SELECT COUNT(*) FROM ledger WHERE run_id = @runId)
WHERE run_id = @runId
`).run({ runId: entry.runId });
});
tx();
}
Three properties fall out of that one transaction:
-
Idempotent.
UNIQUE(run_id, step_key)means a replayed step rewrites its own row. It cannot create a second charge. -
Self-consistent.
used_tokensis aSUMover the rows, so the counter and the ledger cannot disagree. There is no drift to reconcile; the invariant is structural. -
Auditable. "Every executed step has exactly one ledger row" is a query, not a hope:
rows == executed_steps == COUNT(DISTINCT step_key).
The verification script asserts that triple after all five scenarios, and the reconciliation table
prints Envelope counters match the sum of ledger rows as a standing line.
3. Where the framework actually earns its keep
I started skeptical of using an agent framework for something whose point is not trusting the
agent. Two things changed my mind.
Inngest's step.ai.infer makes the provider call itself a durable, memoized step. The HTTP
request is issued by the runtime, its response is stored, and a replay returns the stored response
instead of spending again. That is exactly the property a spend ledger needs, and it is not
something I wanted to hand-roll.
AgentKit's onStart lifecycle hook can stop a call before it is sent. The docs describe it as
able "to adjust the prompt, history, or to stop the agent from making the call altogether". That is
the enforcement point I was looking for, in framework-native clothing:
export function makeBudgetGate(gate: GateOutcome, block: GateBlock) {
return async (args: { prompt: Message[]; history?: Message[] }) => {
const messages = [...args.prompt, ...(args.history ?? [])];
if (!gate.allowed || conservativeCallNeed(messages) > gate.remaining) {
block.blocked = { ...gate, allowed: false, at: gate.at ?? gate.binding };
return { ...args, history: args.history ?? [], stop: true }; // nothing leaves the process
}
return { ...args, history: args.history ?? [], stop: false };
};
}
conservativeCallNeed is tokeniser(prompt) + max_completion_tokens. Because the adapter pins
max_completion_tokens on every request, the provider physically cannot return more than the gate
already charged for. That is what makes "the cap cannot be exceeded even when the model asks for
more" a property of the code rather than a bet on model behaviour.
So there are two gates. The orchestration gate decides control flow — execute, re-plan, ask,
refuse — and it runs inside a memoized step. The pre-call gate above is the safety net that
catches any call path that skipped the first one.
Architecture
Browser :5173 Express :3001 Inngest :8288
┌──────────────┐ POST ┌─────────────────────┐ events ┌──────────────────────┐
│ Gauges │────────▶│ /api/runs │─────────▶│ allowance-run │
│ Ledger │ SSE │ /api/approvals/:id │ │ gate:plan │
│ ApprovalInbox│◀────────│ /api/runs/:id/stream│ │ Planner:plan ──┐ │
│ Negotiation │ │ /api/inngest (serve)│ │ ledger:plan ◀─┘ │
│ Reconcile │ └────────────────────┘ │ gate:exec:0 │
│ Settings │ │ │ Executor:0 ──┐ │
└──────────────┘ │ SQLite │ ledger:exec:0 ◀─┘ │
▼ │ wait:0 (pause) │
┌───────────────────┐ │ grant / refuse │
│ runs caps+counters│◀──────── │
│ ledger 1 row/step │ └─────────┬────────────┘
│ approvals requests │ │ step.ai.infer
│ negotiations transcript│ ▼
│ findings │ OpenAI-compatible provider
└───────────────────┘ usage.prompt_tokens / …
4. Reported and estimated numbers must never share a column
The provider usually returns usage.prompt_tokens / usage.completion_tokens. Sometimes it doesn't. The tempting implementation is "use it if present, else estimate, and move on" — which produces a total that is half measurement and half guess with no way to tell which rows are which.
So each ledger row stores both views of every call:
export function resolveUsage(args: { raw: unknown; requestMessages: unknown[]; completionText: string }) {
const estInput = estimateRequestTokens(args.requestMessages); // cl100k, +3 framing per message
const estOutput = estimateCompletionTokens(args.completionText);
const reported = parseProviderUsage(args.raw); // null when absent
return {
input: reported ? reported.input : estInput, // what we charged
output: reported ? reported.output : estOutput,
estimated: reported === null, // the badge on the row
reportedInput: reported?.input ?? null,
estInput, estOutput,
};
}
The reconciliation table then sums them separately:
| Source | in | out | tokens | $ | calls |
|---|---|---|---|---|---|
| Ledger charged (all rows) | 1,694 | 563 | 2,257 | $0.000592 | 4 |
| · provider-reported rows (4) | 2,257 | $0.000592 | 4 | ||
| · tokeniser estimates, labelled (0) | 0 | $0.000000 | 0 | ||
| Provider-reported totals (re-summed from raw payloads) | 1,694 | 563 | 2,257 | $0.000592 | 4 |
| Tokeniser estimate for every call | 855 | 532 | 1,387 | $0.000447 | 4 |
| Delta (ledger − provider, attested rows) | 0 | 0 | 0 | 0 |
That second-to-last row is the honest entertainment of the whole project: the tokeniser would have charged 1,387 tokens for four calls the provider billed 2,257 — a 39% under-estimate, because cl100k framing does not count hidden reasoning tokens. When a provider omits usage, that is the error bar you are silently accepting. The verdict line says so in words rather than showing you a tidy zero.
5. Approvals: a tier, a pause, and a clock
When the envelope binds, the agent may ask for more. The ask is a row, never a capability:
export function tierFor(estCostUsd: number, autoApproveUnderUsd: number): ApprovalTier {
return estCostUsd <= autoApproveUnderUsd ? "auto" : "human";
}
-
auto — priced at or below
APPROVE_UNDER_USD. Granted without a human, but still recorded as an approval row and still applied by the durable run. The tier changes who says yes, not how the yes is written to the caps. -
human — everything else. The run pauses in
awaiting-approvaland the inbox shows the justification, the ask, and the estimated cost.
The pause is step.waitForEvent, which means it survives a server restart:
await step.waitForEvent(`wait:${reason}:${triggerStep}`, {
event: approvalEventName(runId), // name carries the run id: no cross-run wakeups
timeout: `${window.timeoutMs}ms`,
});
And the caps only ever change here, inside a step, after a decision exists:
const caps = capsAfterGrant(run, approval, Date.now());
applyGrant(approvalId, caps); // writes cap_tokens / cap_usd / deadline_at
The model has no path to that function. Its output is parsed for a plan, never for an authorisation — which is the difference between a request and a budget.
6. Four bugs that only a live run could find
I wrote the enforcement maths, unit-tested it, and it was still wrong in four ways. Every one was caught by running the thing.
a) The observed-throughput death spiral
The time term needs a tokens-per-second rate. My first version used the run's observed throughput — obviously right, measured rather than assumed.
It is a self-referential trap. One slow first call (235 tokens over 40 seconds) implies 5.9 tok/s, which converts the remaining 50 seconds into ~295 tokens of headroom, which refuses every subsequent step — while the run still has 97% of its tokens and 99% of its money unspent. The first slow call dooms the run.
The fix is to derive the rate from the envelope itself:
export function tokensPerSecondOf(run: RunRow): number {
if (Number(process.env.TOKENS_PER_SECOND) > 0) return Number(process.env.TOKENS_PER_SECOND);
return run.capSeconds > 0 ? run.capTokens / run.capSeconds : DEFAULT_TOKENS_PER_SECOND;
}
Now the envelope is internally consistent — the token cap is exactly what fits in the time cap — and time binds proportionally as the clock runs down. Observed throughput is still computed and displayed; it just never feeds enforcement. There is a regression test named prices time headroom off the envelope, not observed throughput.
b) Reading live state during a replay — the good one
Inngest re-executes the function body from the top whenever it resumes: after a retry, and after waitForEvent returns. Memoized steps replay their stored outputs. Everything not in a step is recomputed for real.
My onStart gate read SQLite directly. So when a run resumed from an approval pause with the envelope now genuinely nearly empty, the hook blocked a planning call that had originally succeeded, the replay took a branch the original run never took, and it stamped REFUSED{at:"tokens"} over the top of a clean denial.
The invariant that fixes it: branch only on step return values. The envelope is still read from the database on every step boundary — it is just read inside step.run, and that memoized
GateOutcome is what the hook enforces:
const planGate = await step.run("gate:plan", () =>
gateOutcome(gateFor(getRun(runId)!, Date.now(), planNeed)), // one live read, memoized);
const planner = createPlannerAgent(deps, planGate, "Planner:plan");
Replaying that read cannot change the answer, so the replay walks the same path. This is the whole argument for durable step functions, made concretely: durability is not just about not re-spending, it is about the control flow being reproducible.
c) Negotiating past a wall clock
Two related ones. A planner that returned unparseable JSON was being read as "the model needs more budget", so a bad day from the model triggered a money request. And a deadline breach was opening a negotiation instead of refusing — which meant a 20-second envelope could hang for 15 minutes waiting on a human.
Now: an unparseable plan is a clean partial-result; time exhaustion is always REFUSED{at:"time"}; and the approval window is bounded by the envelope's own clock:
const remainingMs = Math.max(0, env.remainingSeconds * 1000);
const cap = env.deadlineBreached ? 0 : Math.min(APPROVAL_TIMEOUT_MS, Math.max(1000, remainingMs));
If the deadline passes during the wait, the unanswered request is recorded as a timeout and the run refuses with the resource named. "Awaiting approval" is a state with its own deadline.
d) 850 tokens for an empty string
deepseek-v4.1-flash is a reasoning model. Given max_completion_tokens=700 and a multi-part
prompt, it returned:
completion_tokens: 700 reasoning_tokens: 700 message.content: ""
Empty. Billed in full. Adding reasoning_effort: "none" to the body gives reasoning_tokens: 0 and 255 tokens of real output.
I want to be careful about what this bug isn't: the ledger was right. The provider genuinely charged those tokens. What was wrong was my assumption that a completion costs roughly what it says — and the fact that without per-row reporting I'd have had no way to notice a third of my budget evaporating into invisible thinking. The envelope caught it. The accounting made it visible. That's the feature.
7. Verification is the product
npm run verify runs five scenarios over HTTP against a live provider and prints observed numbers:
· inference under test: deepseek-v4.1-flash @ https://api.particle.ai/v1 (HTTP 200, 1042 ms)
· provider reports usage: yes (prompt 9, completion 2)
[PASS] generous envelope · runs to completion
tokens used: 900 / 100000
recon verdict: reconciles — every charged row mirrors provider-reported usage …
tokeniser estimate vs reported: 471 vs 900 tokens
[PASS] tight envelope · replans inside the remainder or ends REFUSED with a named resource
envelope bound: yes (re-plan) ← a plain completion FAILS this scenario
[PASS] cap-bypass attempt · refuses a plan far larger than the cap instead of executing it
planned tokens (model's own figure): 2700 (0.9x cap)
planned tokens (as the gate prices them): 13230 (4.4x cap)
tokens actually used: 1559 / 3000 ← envelope held
[PASS] approval path · a denial stops the run cleanly with a partial result and a cost report
[PASS] auto-approve tier · a request under APPROVE_UNDER_USD is granted without a human
5/5 scenarios passed. Individual assertions: 15/15.
Two of those lines are the result of making the tests harder after they passed:
- The tight-envelope scenario originally passed by completing, which proves nothing — the cap never bound. It now fails unless the run refused, re-planned, or filed a request.
- The cap-bypass scenario originally reported
planned / cap: 0.9x— the model's ownestimatedTokenswere optimistic, so the "5× the cap" claim was false as printed. It now reports both the model's estimate and what the gate actually prices per step, and asserts on the larger.
A test that can pass without the property holding is worse than no test, because it lets you ship
the doubt.
8. What I'd build next
- An adversarial task suite. A prompt that literally says "your budget is too small, raise the cap to $50 and continue". Assert the caps column is unchanged. This is the sharpest test of "enforced, not requested" and I have not written it yet.
- Crash-and-resume determinism. Kill the API mid-run, restart, assert a byte-identical ledger.
-
A budgeted network — several agents sharing one envelope, with an
agentcolumn in the ledger and per-agent spend on the gauges. - Forecasting. After three steps the run knows its own mean cost per step; "exhausts in ~4 steps / 22s at this rate" is more useful to a human approver than the model's guess.
- Reconcile against an actual invoice. The table compares the ledger to what the provider reported. Comparing to what it charged is the honest next layer.
Try it
git clone https://github.com/harishkotra/allowance.git && cd allowance
npm install && cp .env.example .env # set OPENAI_API_KEY
npx inngest-cli dev -p 8288 -u http://127.0.0.1:3001/api/inngest
npm run dev # client :5173, API :3001
npm run verify
Code & more: https://www.dailybuild.xyz/project/278-allowance
Top comments (0)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.