If your team runs Claude Code on a company Anthropic API key, two questions come up sooner or later. Who spent what last week? And what stops an agent stuck in a loop at 2am from spending the rest of the month?
With one shared key, the honest answers are "we don't know" and "nothing". The fix is boring and effective: give each engineer their own key, give each engineer a hard daily cap, and check that cap before the request reaches Anthropic. This guide covers why that works, the options you have, and the exact setup with TokenRouter, the gateway we build.
Why a shared key fails
A shared key has three problems.
No attribution. Anthropic's usage console shows spend per key. If everyone uses the same key, you see one number. Anthropic's own Claude apps gateway docs say it plainly: with one shared upstream credential, "your provider's bill attributes everything to that credential, not to individual developers."
No per-person limit. Any limit on a shared key is a limit on the whole team. One runaway session can use up everyone's budget.
Offboarding means rotating the key for everyone.
What a useful cap looks like
A cap is only useful if it does four things:
- It blocks. It doesn't just report. An alert that arrives after the money is spent is a receipt.
- It's per person, not per company.
- It fails in a way tools can see. An agent should get a clear error, not something that looks like a normal reply. There's an open Claude Code issue (anthropics/claude-code#98441) where a plan cutoff arrives as an ordinary assistant turn, so subagents can't tell it apart from a real answer.
-
It's checked before the call, not after. Claude Code's
--max-budget-usdflag is checked after each call returns, and an open issue (#100111) measured a $1 cap stopping at $1.38. That's fine for a single headless run. It's a weak bound for a whole day of agent traffic.
Your options
-
--max-budget-usd: a per-run bound for headless runs, not a per-person budget. - Claude Team or Enterprise seats: per-member limits for seat users, but not for raw API-key usage in scripts and CI.
- Anthropic's Claude apps gateway: self-hosted, with per-user, per-group and org spend limits, for Claude traffic.
- A self-hosted proxy with virtual keys and budgets, such as the open-source LiteLLM.
- A hosted gateway with per-person budgets, which is where TokenRouter fits.
Setup with TokenRouter
1. Connect your Anthropic key. In the console, open Provider Keys → Add Provider Key, pick Anthropic, and paste the key. Provider keys are AES-encrypted at rest, decrypted only in memory at request time, and never written to logs. Anthropic keeps billing you directly at your rates. We charge a flat subscription and take nothing per token.
2. Create a team and invite engineers. The structure is organization → teams → members → keys. Create a team under Teams → Create team (for example platform or mobile). Then go to Members → Invite member, enter the engineer's email, leave the role as Member, and pick the team. Seats depend on the plan: Free has 1 seat and 1 team, Starter has 5 seats, and Team has 25.
3. Give each engineer a daily hard cap. Open Budgets → Create budget and set:
- Scope: Member, then pick the engineer
- Period: Daily
- Limit (USD): your number (see "Choosing the number" below)
- Alert thresholds: 50%, 80% and 100% are preselected, and you can change them
- Hard cap: on. The dialog spells out what that means: "Requests beyond the limit are rejected with a 429 until the period resets."
A member budget covers every key that engineer owns. You can also put a team budget on top as an overall ceiling, and a request has to fit under both.
When does "daily" reset? Budget periods use UTC day boundaries, so a daily budget resets at 00:00 UTC. That's 8 PM US Eastern during daylight time and 7 PM in winter. Your organization's time zone setting changes how the console displays time, not when budgets reset.
4. Each engineer creates their own key. Each engineer signs in, opens API Keys → Create key, picks the team, and names it something like claude-code-alice. Keys a member creates are tied to that member automatically, so their member budget applies and analytics attribute spend to them. Keys an admin creates in the console belong to the team but no particular member, so only team budgets apply to those. For per-person caps, let each person create their own.
5. Add a rate limit as a safety valve. The same dialog has RPM limit and TPM limit fields, plus an optional expiry in days. Going over a rate limit returns 429 rate_limit_exceeded with a Retry-After header. A budget catches the slow burn over a day. A TPM limit catches a retry loop within minutes.
6. Optionally, restrict models. Each team can have a model allowlist (under Model Access, or the team's own page). A request for a model outside it fails with 403 model_not_allowed.
7. Point Claude Code at the gateway. On each engineer's machine:
export ANTHROPIC_BASE_URL=https://api.tokenrouter.io
export ANTHROPIC_AUTH_TOKEN=tr_your_key_here
claude
There's no /v1 on the base URL. Claude Code appends /v1/messages itself, so adding /v1 produces /v1/v1/messages and a 404. ANTHROPIC_API_KEY=tr_... also works. Put the exports in your shell profile.
8. Verify.
claude -p "Reply with exactly: routed via TokenRouter"
If the reply comes back, you're routed through the gateway. The request shows up under Logs within seconds, attributed to that engineer's key. Claude requests pass through verbatim, so extended thinking, tools and vision behave as they do with Anthropic directly.
What happens at the cap
The cap is checked before the request is forwarded. The gateway reserves the request's worst-case cost: an estimate of the prompt plus the request's maximum output tokens, priced at that model's rates. If spend so far, plus other requests still in flight, plus that reservation would go over the limit, the request is refused with 429 budget_exceeded and never reaches Anthropic. Once the response finishes, the reservation is replaced with the actual cost, which is computed from exact token counts and the provider's published pricing.
Two practical consequences:
- Close to the cap, large requests are refused first, while small ones may still fit.
- The prompt part of the reservation is an estimate. Treat the cap as a tight bound, not to-the-cent accounting.
Gotchas to test before rolling out
These apply to any gateway or proxy, including ours:
- The desktop app's Code tab currently ignores
ANTHROPIC_BASE_URLfrom~/.claude/settings.json, while the terminal CLI honors it (anthropics/claude-code#97574). - An open issue reports that any non-default base URL, even a byte-for-byte pass-through, drops message threads and sharply reduces parallel tool calls in headless runs (#98464). If parallelism matters to your workflow, measure before and after.
- Plugins that make their own model calls may not send custom headers (#99857).
Choosing the number
There's no universal right daily cap, and we won't pretend there is. A reasonable way to pick one: run a week with the hard-cap switch off (alerts only), look at each engineer's heaviest normal day, then turn the hard cap on with a limit comfortably above that. The cap's job is to stop runaway loops, not to ration normal work. Keep the TPM limit tight enough that a loop can't burn the whole day in minutes.
Offboarding
Revoke that person's keys on the API Keys page and remove them under Members. Everyone else keeps working, and nobody rotates anything.
We build TokenRouter. The free plan needs no card and includes the same budgets and hard caps as paid plans. The Claude Code setup is in the docs: tokenrouter.io/docs/claude-code.
Top comments (0)