TL;DR
- A Claude Code gateway sits between Claude Code and model providers so developers hold gateway credentials instead of provider keys, and usage, budgets, and audit trails live in one place.
- The deciding criteria are forwarding fidelity (streaming, beta headers, error bodies, prompt caching), governance depth, provider coverage, and where the gateway runs.
- Bifrost connects to Claude Code with two settings.json values,
ANTHROPIC_BASE_URLandANTHROPIC_AUTH_TOKENset to a virtual key, and needs no Anthropic account login. - Bifrost routes Claude Code to Claude on Amazon Bedrock, Vertex AI, or Azure through stable deployment names, and keeps each session on one provider key so prompt cache hits continue.
- Enterprise rollout follows five steps: deploy the gateway, issue one virtual key per developer, distribute managed settings, verify, then lock the endpoint.
A Claude Code gateway is a service that sits between Claude Code and model providers, authenticating each developer, enforcing budgets, and forwarding every request to Anthropic, Amazon Bedrock, Google Vertex AI, or another provider from one endpoint. Bifrost, the open-source AI gateway built in Go by Maxim AI, fills that role for engineering organizations running Claude Code at scale, with 11 microseconds of overhead per request at 5,000 RPS. This guide explains what to evaluate in a Claude Code gateway, why teams searching for a "Claude Code proxy" usually need a gateway, and how to configure, govern, and roll one out.
What Is a Claude Code Gateway?
A Claude Code gateway is an Anthropic Messages-compatible endpoint that Claude Code calls instead of api.anthropic.com. Developers authenticate with a gateway-issued credential, and the gateway applies access rules, budgets, and routing before forwarding the request with the organization's provider credential. Provider keys never reach developer laptops.
Claude Code selects the gateway through ANTHROPIC_BASE_URL. Anthropic's guide to other LLM gateways lists what a gateway centralizes (credentials, usage tracking, cost controls, audit logging, provider switching) and states that Anthropic does not endorse, maintain, or audit third-party gateway products, and does not support routing Claude Code to non-Claude models through any gateway, so model substitution is a choice your team validates and owns.
Figure 1: Developers hold gateway credentials; provider keys, budgets, and routing decisions stay in one place the platform team operates.
The same gateway can serve Codex CLI; see Codex CLI routing across any model. The AI gateway architecture explainer covers the general pattern.
Claude Code Proxy vs Gateway: Which One Teams Need
Teams search for a "Claude Code proxy" for three reasons: running Claude Code against non-Anthropic models, sending traffic through a corporate HTTPS proxy, or adding cost and access control. Only the first two are proxy problems; cost control and access policy need identity, budgets, and logs, which is gateway work.
Open-source Claude Code proxy projects typically translate Anthropic-format requests into another provider's format for one developer. A corporate HTTPS_PROXY is a network control that sits between Claude Code and every server, including the gateway. Neither knows which developer sent a request or what it cost.
| Need | Translation proxy | Corporate HTTPS proxy | Claude Code gateway (Bifrost) |
|---|---|---|---|
| Use other providers' models | Yes, usually one developer | No | Yes, 25+ providers |
| Per-developer credentials | No | No | Virtual key per developer |
| Budgets and rate limits | No | No | Virtual key, team, customer |
| Request logs with tokens and cost | Varies | No | Built in |
| Central provider switching | No | No | Yes, no laptop changes |
Bifrost is an AI gateway rather than a proxy: it holds governance state, resolves models across providers, and records every request. The 5 best AI gateways in 2026 roundup compares self-hosted and managed options.
Why Claude Code Enterprise Rollouts Need a Gateway
Claude Code enterprise deployments need per-developer attribution, enforceable spend limits, provider redundancy, and one place to revoke access. Without a gateway, each developer authenticates to Anthropic directly and spend is visible only in aggregate on the invoice.
Cost scales with agent activity rather than seat count. Four problems appear at team scale:
- Attribution: finance cannot tell which team or developer drove a spike.
- Spend control: no per-developer cap exists until someone reads the invoice.
- Offboarding: revoking access means rotating shared provider keys.
- Provider dependency: a rate limit or outage on one account blocks everyone.
Bifrost addresses each through virtual keys: one per developer, revocable on its own, with budgets checked before every request. The Bifrost governance overview describes how keys, teams, and customers fit together, and the Bifrost Enterprise tier adds RBAC, SSO, and in-VPC deployment for regulated teams.
Key Criteria for Evaluating a Claude Code Gateway
Evaluate a Claude Code gateway on forwarding fidelity first, then governance, provider coverage, observability, and deployment model. Claude Code adds request fields and beta headers with each release, and Anthropic's gateway compatibility guide states that a gateway which strips or rewrites them breaks the matching feature.
| Criterion | What to test | Failure symptom in Claude Code |
|---|---|---|
| Streaming | Events relayed as they arrive, SSE pings kept | Stalls, idle-timeout aborts |
| Beta header pass-through |
anthropic-beta and anthropic-version forwarded verbatim |
400 errors naming unknown fields |
| Error bodies | Upstream errors forwarded unmodified | Automatic retry and recovery stop working |
| Prompt caching |
cache_control preserved; sessions stay on one key |
High input_tokens, low cache reads |
| Tool-call streaming | Tool arguments streamed intact by the upstream | Empty arguments, failed file edits |
| Governance | Budgets and rate limits per developer, team, customer | Spend found after the fact |
| Deployment | Self-hosted, in-VPC, or air-gapped | Source code leaves your network |
| Overhead | Published figure with instance type and load | Latency compounds over long sessions |
Claude Code's retry logic matches on the upstream's error wording, so a gateway that wraps errors in its own envelope breaks recovery even when it keeps the status code. Prompt caching depends on routing too: a session that moves between provider keys misses the cache it warmed a turn earlier.
On overhead, Bifrost publishes sustained benchmarks of 11 µs per request at 5,000 RPS with a 100% success rate. The production-ready LLM gateway comparison scores general-purpose gateways on these criteria, and what an AI gateway does and why it matters covers the concepts.
How Bifrost Works as a Claude Code Gateway
The Bifrost AI gateway exposes an Anthropic-compatible endpoint at /anthropic, authenticates Claude Code with a virtual key, checks budgets and rate limits, resolves the model through routing rules, and pins the session to one provider key. Claude Code needs no code changes, only a base URL and a credential.
Figure 2: Policy checks run before any provider is called, and session affinity keeps a session on the provider whose prompt cache it already warmed.
The capabilities that matter most for Claude Code:
-
Session affinity: Claude Code sends
x-claude-code-session-idon every request, and session affinity adopts it with no configuration, so a session and its subagents stay on one provider and key and keep hitting the same prompt cache. -
Model aliasing: routing rules can rewrite an alias such as
sonnet-modelto any configured provider and model, matched on theclaude-cliuser agent. - Provider coverage: 25+ providers and 10,000+ models through one API, listed on the supported providers page.
-
Failover: automatic fallbacks rotate keys on
429and retry5xxerrors with backoff. - Request logs: built-in observability records inputs, outputs, tokens, cost, and latency for each request, asynchronously.
Since Claude Code 2.1.212, Anthropic enforces anthropic-version and related headers. In Bifrost, set Settings > Client Settings > Allowed Headers to *, or add the list from the Bifrost Claude Code integration guide.
Claude Code Environment Variables for a Gateway
Claude Code reads gateway configuration from the env block of settings.json. With Bifrost, two variables are required: ANTHROPIC_BASE_URL pointing at the /anthropic endpoint and ANTHROPIC_AUTH_TOKEN holding a virtual key. The token is sent as a bearer header, so no Anthropic account login is needed.
Merge the block into the most specific settings file that applies: ~/.claude/settings.json (user), .claude/settings.json (project), or .claude/settings.local.json (personal override). Remove any top-level model field, because it overrides the environment-based model selection.
"env": {
"ANTHROPIC_BASE_URL": "http://localhost:8080/anthropic",
"ANTHROPIC_AUTH_TOKEN": "your-bifrost-virtual-key",
"ANTHROPIC_DEFAULT_SONNET_MODEL": "claude-sonnet-4-6",
"ANTHROPIC_DEFAULT_HAIKU_MODEL": "claude-haiku-4-5"
}
| Variable | Value with Bifrost | What it does |
|---|---|---|
ANTHROPIC_BASE_URL |
http://localhost:8080/anthropic |
Sends Claude Code traffic to the gateway |
ANTHROPIC_AUTH_TOKEN |
Bifrost virtual key | Bearer credential; replaces subscription login |
ANTHROPIC_CUSTOM_HEADERS |
x-bf-vk: <key> |
Alternative; still requires an Anthropic account login |
ANTHROPIC_DEFAULT_SONNET_MODEL, _HAIKU_, _OPUS_
|
provider/model or an alias |
Maps each Claude Code tier to a model |
CLAUDE_CODE_ENABLE_GATEWAY_MODEL_DISCOVERY |
1 |
Fills /model from the gateway's model list (2.1.129+) |
CLAUDE_CODE_ENABLE_FINE_GRAINED_TOOL_STREAMING |
1 |
Restores a default that is off behind custom base URLs |
Gateway model discovery only lists model IDs beginning with claude or anthropic. The Bifrost CLI skips manual setup entirely: npx -y @maximhq/bifrost-cli configures the base URL, key, and model, stores the virtual key in the OS keyring, and launches Claude Code. Pinning tiers to GPT, Gemini, or local models is covered in the comparison of AI gateways for running Claude Code with non-Anthropic models.
Running Claude Code on Amazon Bedrock Through Bifrost
To run Claude Code on Bedrock through Bifrost, map Bifrost deployment names to Bedrock model IDs on the AWS Bedrock provider key, allow the bedrock provider on the developer's virtual key, and set Claude Code's tier variables to names such as bedrock/claude-sonnet-5. AWS credentials stay on the gateway.
Figure 3: Developers only ever see stable deployment names; the Bedrock model ID and AWS credentials live on the Bifrost provider key.
{
"env": {
"ANTHROPIC_BASE_URL": "https://gateway.example.com/anthropic",
"ANTHROPIC_AUTH_TOKEN": "your-bifrost-virtual-key",
"ANTHROPIC_DEFAULT_OPUS_MODEL": "bedrock/claude-opus-4-8",
"ANTHROPIC_DEFAULT_SONNET_MODEL": "bedrock/claude-sonnet-5",
"ANTHROPIC_DEFAULT_HAIKU_MODEL": "bedrock/claude-haiku-4-5-20251001"
}
}
If the provider key restricts models, list every deployment name there; mappings do not expand the allowlist. Test with /model, then confirm in Bifrost Logs that the request resolved to the expected Bedrock model ID. The full steps and a troubleshooting table are in the Claude Code with Bedrock runbook, and a walkthrough lives in routing Claude Code through AWS Bedrock with Bifrost.
Compared with Claude Code's native Bedrock mode (CLAUDE_CODE_USE_BEDROCK=1), this path keeps per-developer keys, budgets, and logs, and lets one settings file switch a tier to Vertex AI or Azure. Through ANTHROPIC_BASE_URL, Claude Code sends its full capability set; if a Bedrock upstream rejects pre-release fields with 400 errors, set CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS=1.
Connecting Claude Code MCP Servers Through Bifrost
Bifrost also acts as an MCP gateway for Claude Code, aggregating every connected MCP server behind one /mcp endpoint. One claude mcp add command replaces a config entry per server, and the virtual key in the header scopes which tools each developer can call.
claude mcp add --transport http bifrost http://localhost:8080/mcp \
--header "Authorization: Bearer your-virtual-key" \
--scope user
Run /mcp inside Claude Code to confirm bifrost is connected. When inference also routes through Bifrost, turn off auto tool injection to keep the two paths separate. Per-key MCP tool filtering controls which tools appear.
Tool definitions matter for cost behind any gateway. Claude Code turns MCP tool search off when ANTHROPIC_BASE_URL points to a non-first-party host, per Anthropic's MCP documentation, so every tool definition loads upfront unless tool search is re-enabled. Bifrost Code Mode cut input tokens by 92.8% in a 508-tool benchmark by exposing tools through four meta-tools.
The Bifrost MCP gateway page and the guide to reducing MCP token costs for Claude Code cover the details.
Budgets and Rate Limits for Claude Code Teams
Bifrost enforces Claude Code budgets at four levels (customer, team, virtual key, and provider config) and rate limits at the virtual key and provider config levels. Every applicable budget is checked before a request is forwarded, and any single failure blocks it, so a developer cannot exceed a team cap through their own key.
| Level | Budget | Rate limits | Typical Claude Code use |
|---|---|---|---|
| Customer | Yes | No | Business unit or cost center |
| Team | Yes | No | Platform, mobile, data teams |
| Virtual key | Yes | Requests and tokens | One per developer |
| Provider config | Yes | Requests and tokens | Cap Opus-tier spend on one provider |
Budgets reset on 1d, 1w, 1M, 1Q, or 1Y windows and can be calendar-aligned to UTC boundaries. Virtual keys are deny-by-default for providers.
In Bifrost Enterprise, access profiles auto-issue a governed virtual key to each user from a role, and alerting notifies Slack, Microsoft Teams, PagerDuty, or a webhook when a budget crosses a threshold.
For dashboards and per-developer reporting, see monitoring Claude Code token usage with an AI gateway. Teams running both agents can apply the same keys to Codex CLI token spend.
Rolling Out a Claude Code Gateway Across the Organization
A Claude Code gateway rollout has five steps: deploy the gateway with provider credentials, issue one virtual key per developer, distribute the base URL and credential through managed settings, verify each machine, then restrict Claude Code to the gateway endpoint. The first four match Anthropic's rollout checklist.
Figure 4: Lock the endpoint last, after every machine has proven it reaches the gateway, so a distribution gap never blocks a developer.
-
Deploy: run Bifrost with
npx -y @maximhq/bifrostor Docker, then move to clustering or an in-VPC deployment for production. - Issue keys: one virtual key per developer, or access profiles issued from SSO roles.
-
Distribute: push the
envblock through a managed settings file via MDM; a managedANTHROPIC_BASE_URLcannot be overridden by a developer's shell. -
Verify: have each developer run
/status, then confirm their requests appear in Bifrost logs. -
Lock: set
allowedProvidersto["customEndpoint"]in managed settings (Claude Code 2.1.285+) so sessions pointed anywhere else are refused.
For machines where configuring each tool is impractical, AI Gateway + Bifrost Edge extends the same governance to the endpoint: Bifrost remains the policy engine, and Bifrost Edge, currently in alpha, routes Claude Code traffic from each laptop through it with no settings.json changes. The guide to governing coding agents in enterprise rollouts covers policy design across agents.
Frequently Asked Questions
What is the difference between a Claude Code proxy and a Claude Code gateway?
A Claude Code proxy typically forwards or translates requests for one developer, often to reach a non-Anthropic model. A Claude Code gateway adds identity, budgets, rate limits, request logs, and central provider switching for a whole organization. Bifrost is an AI gateway: each developer authenticates with their own virtual key, and every request is checked against budgets and recorded with tokens and cost.
What LLM does Claude Code use?
Claude Code uses Anthropic's Claude models by default: Sonnet as the main tier, Opus for complex tasks, and Haiku for lightweight work. Behind a gateway, each tier maps to whatever the ANTHROPIC_DEFAULT_*_MODEL variables specify, including Claude on Bedrock or Vertex AI. Anthropic does not support non-Claude models through gateways; running Claude Code with non-Anthropic models covers what teams test.
Does a gateway slow Claude Code down?
Bifrost adds 11 microseconds per request at 5,000 RPS in sustained benchmarks, which is negligible next to model response times. The larger performance factor is prompt caching: Bifrost session affinity keeps each Claude Code session on one provider key, so cache reads continue across turns instead of resetting when routing changes.
Do developers still need a Claude subscription behind a gateway?
No, not when the gateway credential is set. With ANTHROPIC_AUTH_TOKEN holding a Bifrost virtual key, requests carry that credential instead of a claude.ai login, and usage is billed to the provider account behind the gateway. Setting only ANTHROPIC_BASE_URL without a credential keeps the subscription login active, so its usage limits and billing still apply.
Can Claude Code use Amazon Bedrock through a gateway?
Yes. In Bifrost, map deployment names such as claude-sonnet-5 to Bedrock model IDs on the AWS Bedrock provider key, allow bedrock on the virtual key, and set Claude Code's tier variables to bedrock/claude-sonnet-5. AWS credentials stay on the gateway, and each request is logged with the resolved Bedrock model ID for verification.
How do you set budgets for a Claude Code team?
Issue each developer a virtual key with its own budget and rate limits, attach the keys to a team with a team budget, and optionally group teams under a customer budget. Bifrost checks every applicable budget before forwarding a request. Alerting can notify the team before a cap is reached.
Start Running Claude Code Through Bifrost
A Claude Code gateway gives platform teams per-developer credentials, enforceable budgets, routing to Bedrock and Vertex AI, and request-level logs without changing how developers use Claude Code. Bifrost does this with two settings.json values and 11 µs of overhead. The Bifrost Claude Code resource page has setup details, and to plan a Claude Code enterprise rollout with the team, book a Bifrost demo.




Top comments (0)