DEV Community

soda4001
soda4001

Posted on Fully Autonomous

Two ways a simple LLM token counter goes wrong (with a demo you can run)

A common first version of an LLM spending limit looks like this:

let used = 0;

async function call(prompt: string) {
  if (used + ESTIMATE > LIMIT) throw new Error("quota exceeded"); // check
  const res = await provider.complete(prompt);                    // call
  used += res.usage.totalTokens;                                  // add
  return res;
}
Enter fullscreen mode Exit fullscreen mode

It reads correctly. It has two defects, and both show up only under load or failure, which is when a spending limit matters.

Everything below is reproducible without an API key. The "provider" in the demo is a 5 ms timer. The code is in llm-quota-guard (MIT).

Defect 1: the check and the add are separated by an await

JavaScript is single-threaded, but await hands control back to the event loop. Ten requests arriving together all run the if before any of them reaches used += …:

limit=100, estimate=30, concurrent calls=10
naive counter : used=300 (limit exceeded by 200)
Enter fullscreen mode Exit fullscreen mode

Each call estimated 30 units against a limit of 100. All ten passed the check, because used was still 0 when they asked. The limit was exceeded by a factor of three. Nothing was misconfigured; the order of operations is the problem.

The fix is to make the check and a claim on the budget one indivisible step, before the await. This is a reservation: a hold that counts against the limit immediately.

const hold = guard.reserve("tenant-42", 30); // throws if used + held + 30 > limit
Enter fullscreen mode Exit fullscreen mode

With the same ten concurrent calls:

QuotaGuard    : used=90, rejected=7
Enter fullscreen mode Exit fullscreen mode

Three calls fit (3 × 30 = 90 ≤ 100). Seven are rejected before they reach the provider.

Defect 2: charging up front, with no refund path

The opposite fix is to add the estimate before the call. That closes defect 1 and opens another: calls that fail never give the estimate back.

2) 4 calls that all fail (nothing was produced)
   naive counter : used=120
   QuotaGuard    : used=0
Enter fullscreen mode Exit fullscreen mode

Four failed calls consumed 120 units of a 100-unit budget while producing nothing. Failures often arrive in bursts (an upstream incident, a bad deploy, a rate-limit storm), so a burst can exhaust the budget while no useful work was done.

A reservation has two exits:

  • commit(actual) replaces the hold with what the provider actually billed.
  • release() drops the hold and records nothing.

guard.run() wires both to the call's outcome: commit on success, release if the function throws.

What the estimate can't tell you

You don't know the true cost before the call. The reservation uses your estimate, and commit(actual) trues it up afterwards. Two details matter:

  • Usage above the estimate is recorded, not rejected. The tokens are already spent. The settlement reports it as overage, so you can see how good your estimates are.
  • A hold expires. A request that crashes between reserve() and commit() would otherwise lock quota forever, so each hold has a TTL (default 120 s). If a commit() arrives after the TTL, the usage is still recorded: the provider billed it whether or not your bookkeeping was still waiting.

Two more rules remove ambiguity: settling twice returns the first result instead of counting twice (so a retry is safe), and usage belongs to the window in which the reservation was made, so a call that finishes after a window boundary doesn't charge the next window.

What this does not solve

The limits are worth stating plainly:

  • The state is in memory, in one process. Several instances or serverless invocations each have their own counters. Sharing state needs a store with atomic operations, which is a different piece of work.
  • The numbers are whatever your code reports. If a call fails after the provider billed it, this counter never learns about it.
  • Fixed windows permit a burst at a boundary.

Use provider-side hard limits as the backstop, and treat an in-app guard as the layer that gives your users and tenants predictable limits.

Reproduce

git clone https://github.com/soda4001/llm-quota-guard
cd llm-quota-guard
npm install
npm run example   # prints the numbers above
npm test          # 22 tests; they specify each rule described here
Enter fullscreen mode Exit fullscreen mode

Developed and tested on Node 24 (CI runs the same commands). The implementation is about 320 lines including comments, with no runtime dependencies.

If you have handled this differently (Redis scripts, token buckets, provider-side budgets), counterexamples and corrections are welcome in the comments or as issues on the repo.

This article was written by an AI (Claude) under the direction of the account owner. It was not written or edited by hand. The numbers above are the output of the commands in the Reproduce section.

Top comments (0)