DEV Community

Cover image for A request limit cannot cap your LLM bill. Reserve tokens in Redis
Rakesh Singh
Rakesh Singh

Posted on

A request limit cannot cap your LLM bill. Reserve tokens in Redis

One request can be a one-line question. Another starts an agent run that makes a dozen model calls, each resending a longer context than the last. A request limit counts both as one.

The obvious fix is to count tokens after each response. It fails, because concurrent requests all see "under the limit" and all pass.

The rule: limit tokens, reserve an estimate before the call, and correct it after. This post covers what to count, when to count it, where the counter lives, and what an agent loop changes.

Why does a request limit not control spending?

The system answers questions over telecom specification documents. Every answer ends in a paid model call, and the hard questions go to an agent that makes several.

Caching took care of the questions that repeat:

For every question that does not repeat, the only thing between a user and my bill is the rate limit.

The provider bills by tokens, so tokens are what you limit. A requests-per-minute limit still has a job: it stops floods and protects your own servers. It cannot control spending.

Tokens are not all the same price. Output tokens usually cost more than input tokens, models are priced differently, and cached input is cheaper. So there are two counters with two purposes:

  • Tokens per minute is for capacity and fairness between users.
  • Cost per day or month is for money.

Why can't you count tokens after the call?

The wrong model is "count what was used."

A request limit can be checked before the work starts, because the cost of a request is one. A token limit cannot: the output does not exist yet.

If you only count after the response, concurrent requests all see the same "under the limit" and all pass. The overshoot is bounded only by how many requests are in flight.

LiteLLM documents this trade-off directly. Its request limit is a hard limit, exact at the threshold. Its token limit is best-effort: it may slightly exceed, and it blocks once already over.

What do you count, and when?

The stricter design counts twice.

Estimate. Count the input tokens and add the output cap you set on the call. This is why every model call needs an output cap: it bounds the estimate and the worst case.

Reserve. In one atomic step, check that the estimate fits under the limit and add it to the counter. If it does not fit, reject it before the model is called.

Reconcile. Read the actual usage from the provider's response and correct the counter by the difference. If the window has rolled over, apply the correction to the current window.

Flow diagram of six steps: request, estimate, reserve in Redis, model call, actual usage, and reconcile in Redis. A branch from the reserve leads to a 429 rejection where the model is never called.
The limit is enforced before the call with an estimate and corrected after it with the actual usage.

The reserve step is one. Redis script:

-- KEYS[1] = token counter for this user and window
-- ARGV[1] = estimate (input tokens + output cap)
-- ARGV[2] = limit, ARGV[3] = window in seconds
local used = tonumber(redis.call('GET', KEYS[1]) or '0')
local estimate = tonumber(ARGV[1])
if used + estimate > tonumber(ARGV[2]) then
  return 0  -- reject: the model is never called
end
redis.call('INCRBY', KEYS[1], estimate)
if redis.call('TTL', KEYS[1]) < 0 then
  redis.call('EXPIRE', KEYS[1], tonumber(ARGV[3]))
end
return 1
Enter fullscreen mode Exit fullscreen mode

The line that matters is return 0. The request is rejected before the model is called, so a rejected request costs nothing. Reconcile is one INCRBY after the call, with the difference between actual and estimated tokens.

Which counter you reserve against depends on the job.

A token bucket holds a capacity and refills at a steady rate. A user can spend a burst up to the capacity and then is held to the refill rate. It suits interactive traffic, where a burst is normal behaviour. In Redis it is two fields in a hash, the tokens left and the time of the last refill, updated by one script that takes its clock from Redis.

A window quota allows a fixed amount per period and no more. It suits budgets, where the point is a hard cap. In Redis it is a counter per window with a TTL, like the script above.

Traefik's AI gateway ships both for exactly this split: a token bucket for traffic spikes and a sliding-window quota for budget control and hard spending caps.

Where does the counter live?

A limit only exists if every instance reads and writes the same counter.

Two diagrams. On the left, three instances each allow 100,000 tokens, giving an effective limit of 300,000. On the right, three instances share one Redis counter, giving an effective limit of 100,000.
A limit counted in each instance's memory is multiplied by the number of instances. One counter in Redis keeps it at the configured value.

This is easy to get half right. A bug reported against LiteLLM in 2026, and since fixed, was exactly that. The write path updated Redis. The check before the call read only in-process memory. With several replicas, the effective token limit was the limit multiplied by the replica count.

The same thing happens one level up. LiteLLM's multi-region guide says it plainly: with one Redis per region, a limit of 100 requests per minute means 100 per region, not globally. A single shared Redis puts a cross-region round trip into every check. A limit is only as global as its counter.

Three rules for the counters:

  1. Check and add in one atomic script. Read, decide, then write as three steps lets two concurrent requests both pass.
  2. Check every limit that applies in the same script: the run, the user, the workspace, and the provider. Otherwise a request is charged to one counter and then rejected by another.
  3. Take the time from Redis, not from the app server. Otherwise, clock skew between instances becomes part of the limit.

The limits themselves run in two directions.

Inbound limits protect you from your users. They are keyed by the authenticated user, the workspace, or the API key. They answer, "How much may this caller spend?"

Outbound limits protect you from your provider. The provider gives your account a quota per model, and every instance of your service draws from it. That counter is keyed by provider and model. It answers, "How much may all of us send?"

Your outbound counter is an estimate of the provider's counter, and theirs is the truth. When the provider returns a 429 or its own remaining-quota headers, believe those over your number. When the outbound limit is reached, the options are to queue, to back off, or to route to another model.

What does an agent loop change?

One user message becomes many model calls. Each step resends the conversation so far plus the latest tool results. The input grows with every step, and the cost of a run grows faster than its step count. A per-minute limit sees this too late.

So a run gets its own budget: a maximum number of steps, of tokens, of cost, and of wall-clock time. It is enforced by the orchestrator in code, before every step. A limit written into the prompt is a request, not a limit.

In my system every agent loop has a step cap and a cost cap. Budget exhaustion is one of the cases in my adversarial test suite: a logged-in client trying to spend as much as

Top comments (0)