Excerpt: An AI gateway is the infrastructure layer that routes, governs, secures, and observes all traffic between AI applications, models, and tools. This guide explains AI gateway architecture, the request path, MCP and coding-agent support, governance, and how to evaluate one.
TL;DR
- An AI gateway is the control layer between AI callers (applications, agents, coding tools) and the model providers and tool servers they use, applying routing, access, cost, and safety policy in one place.
- AI gateway architecture splits into a control plane, where keys, budgets, and routing rules are defined, and a data plane, where every request is authenticated, checked, cached, routed, and logged.
- An AI gateway differs from a traditional API gateway because it meters tokens rather than requests, fails over across equivalent models, caches by meaning, and inspects prompts for LLM-specific risks.
- Modern AI gateways also govern agent tool calls over the Model Context Protocol (MCP) and the traffic from coding agents such as Claude Code and Codex CLI.
- Bifrost, the open-source AI gateway built by Maxim AI, connects 25+ providers and 10,000+ models through one OpenAI-compatible API and adds 11 microseconds of overhead per request at 5,000 RPS.
An AI gateway is the infrastructure layer that sits between AI applications and the model providers, cloud AI platforms, and tool servers they call, and it enforces routing, access, cost, and safety policy for all of that traffic in one place. Teams adopt one once a second provider, a second team, or a first agent reaches production, because those are the points where failover, cost attribution, and audit trails can no longer live in application code. Bifrost, the open-source AI gateway written in Go and built by Maxim AI, implements this layer for enterprise teams, with 25+ providers behind one OpenAI-compatible API and 11 microseconds of added overhead at 5,000 requests per second. This guide covers what an AI gateway is, its architecture, how a request moves through it, how it differs from an API gateway, and how to evaluate one.
What Is an AI Gateway?
An AI gateway is a control layer that gives every AI caller one endpoint for every model and tool, while centralizing authentication, routing, failover, budgets, guardrails, and logging. Applications call the gateway instead of provider SDKs, and the gateway picks the provider, model, and key for each request under the caller's policies.
The category carries several names. Vendors and analysts also call it an AI API gateway, a GenAI gateway, an AI model gateway, or, when the scope is limited to model calls, an LLM gateway. What separates a full AI gateway from a narrower model router is scope: it governs chat, embeddings, images, audio, and agent tool calls under one identity and one budget.
Figure 1: One layer in the middle is the only place a single policy can reach every caller and every provider.
As Figure 1 shows, four kinds of callers now share the same layer:
- Applications and services calling models for chat, search, extraction, or classification features.
- Coding agents such as Claude Code, Codex CLI, Cursor, and Gemini CLI, which issue many model calls per task.
- AI agents and copilots that call tools through MCP as well as models.
- Unmanaged AI usage, where individuals connect desktop apps or browser AI directly to providers and bypass every policy.
Bifrost exposes one OpenAI-compatible endpoint for all of these callers and translates each request to the provider's native API, covering the providers listed in the supported providers matrix, from hosted APIs to cloud platforms and self-hosted models.
Why an AI Gateway Matters for Production AI
An AI gateway matters because direct provider integrations offer no shared mechanism for failover, spend control, access policy, or audit, and each team rebuilds those controls differently. Centralizing them in one layer turns provider choice into configuration and makes AI traffic operable, attributable, and reviewable like any other production dependency.
Four pressures push teams toward the gateway pattern:
- Provider fragmentation. Every LLM vendor ships a different API contract, auth scheme, error model, and rate-limit structure. Supporting several providers directly creates integration sprawl, and supporting one creates lock-in.
- Token-priced, unpredictable cost. Spend scales with tokens, not requests, so one long-context prompt can cost far more than expected.
- Provider incidents and rate limits. A 429 or a 5xx burst from one provider becomes an application outage unless something reroutes traffic, which is the job described in how AI gateways tackle rate limiting.
- Compliance and security review. Regulated teams must show who called which model, with what data, under which policy, and that evidence has to come from one place.
Figure 2: The gateway stops being optional at the second stage, when a provider outage or a second team first touches production.
Different stakeholders bring different problems to the same layer, which is why AI gateway governance usually ends up owned by the platform team:
| Role | Problem they bring | What the AI gateway gives them |
|---|---|---|
| Platform engineering | Every team integrates providers separately | One endpoint, shared failover, caching, and routing |
| Engineering leadership | Agents and coding tools multiply model calls | Per-team budgets, model allow-lists, usage visibility |
| Security | Prompts and outputs carry sensitive data | Guardrails, secrets detection, access control per caller |
| Finance and FinOps | AI spend arrives as one provider invoice | Token-level cost attribution per team, customer, and project |
| Compliance | Auditors ask for evidence of control | Request logs per call and signed audit events for configuration changes |
AI Gateway Architecture and Core Components
AI gateway architecture has two halves. The control plane stores configuration: virtual keys, budgets, rate limits, routing rules, provider credentials, and model pricing. The data plane sits in the request path and applies that configuration to each call through a unified API, policy hooks, a router, a cache, and a logging pipeline.
Figure 3: Policy is written once in the control plane and enforced on every request in the data plane.
This gateway architecture separates who sets policy from where it runs: administrators change budgets or routing weights in the control plane, and the data plane applies them without application redeploys. Most production AI gateways include these components:
| Component | What it does | How Bifrost implements it |
|---|---|---|
| Unified API | One request format for all providers | OpenAI-compatible endpoint; a drop-in replacement changes only the SDK base URL |
| Router | Picks provider, model, and key per request | Weighted routing on virtual keys plus CEL-based routing rules |
| Retries and failover | Recovers from 429s, 5xx errors, and outages | Retries with backoff and key rotation, then fallback chains across providers |
| Load balancing | Spreads traffic across keys and accounts | Weighted API key distribution with model-specific filtering |
| Cache | Replays answers without a provider call | Direct hash matching plus embedding-based semantic matching |
| Policy hooks | Auth, budgets, rate limits, guardrails | Virtual keys, hierarchical budgets, Enterprise guardrails |
| Observability | Logs, metrics, traces | Async request logs with tokens, cost, and latency; Prometheus and OpenTelemetry |
| Model catalog | Prices and capabilities per model | Pricing sheet synced every 24 hours for cost calculation |
Three components behave differently from their counterparts in traditional infrastructure.
Routing works on model equivalence, not endpoint health. A gateway can send reasoning-heavy requests to one model and high-volume classification to a cheaper one, then move both to another provider during an incident. Bifrost evaluates routing rules in scope order (virtual key, team, customer, global), and the patterns are covered in LLM routing strategies every AI gateway needs.
Caching works on meaning, not bytes. Bifrost semantic caching runs a direct hash lookup first and an embedding similarity search on a miss, so a reworded question can reuse an earlier answer. Requests opt in with a cache key (or a configured default key), and the cache covers chat, text completions, the Responses API, embeddings, transcription, speech, and image generation, including streaming.
Observability is token-aware. Bifrost built-in request logging records inputs, outputs, tokens, cost, and latency asynchronously, so logging adds no latency to the request. Metrics export through Prometheus and traces through OpenTelemetry, which is how gateway data reaches existing LLM observability stacks.
How an AI Gateway Works: The Request Path
An AI gateway processes each request as an ordered pipeline: identify the caller, apply limits and guardrails, check the cache, route to a provider with failover, then log the response with token counts and cost. Cheap checks run first, so a rejected prompt or a cache hit ends the request before any provider is billed.
Figure 4: Cheap checks run first, so a blocked prompt or a cache hit never costs a provider call.
Following Figure 4, a single request in Bifrost moves through these stages:
- Ingress. The caller sends an OpenAI-, Anthropic-, or Gemini-format request to the gateway endpoint instead of the provider.
-
Identity and limits. The request carries a virtual key (for example in the
x-bf-vkorAuthorizationheader), which resolves to allowed providers and models, a budget, and request and token rate limits. - Input guardrails. Enterprise guardrail rules inspect the prompt for secrets, PII, or policy violations and can block or redact it before it leaves the network.
- Cache lookup. A direct or semantic hit returns a stored response without a provider call.
- Routing and failover. The router selects the provider, model, and key. If the call fails, retries and fallbacks rotate keys or move to the next provider in the chain, each with its own retry budget.
- Response handling. Output guardrails run, the response is normalized to the caller's format, and the log entry records tokens, cost, latency, and the virtual key.
The application sends one request and receives one response; the gateway absorbs provider differences, outages, and policy checks.
AI Gateway vs API Gateway
An AI gateway and an API gateway both sit in front of services and handle authentication, routing, and rate limiting. The difference is the traffic. An API gateway manages request-priced REST calls with predictable payloads; an AI gateway manages token-priced model and tool calls with streaming responses, long latencies, large contexts, and prompt-level security risks.
LLM traffic breaks several API gateway assumptions: one request can stream for minutes, cost depends on tokens rather than request count, and failure modes include prompt injection and data leakage. The OWASP Top 10 for LLM Applications lists prompt injection, sensitive information disclosure, and unbounded consumption among the top risks, and each one is easiest to control at a central gateway.
| Dimension | Traditional API gateway | AI gateway |
|---|---|---|
| Traffic | Synchronous REST, predictable payloads | Model and tool calls, streaming, large context windows |
| Unit of cost | Request | Input and output tokens, per model and per caller |
| Rate limiting | Requests per second | Requests and tokens per key, team, or customer |
| Caching | Byte-exact responses | Exact hash plus semantic similarity |
| Failover | Healthy instance of the same service | Equivalent model on another provider |
| Security focus | Authentication, quotas | Prompt injection, PII and secrets, content safety |
| Agent support | None | MCP tool discovery, execution, and filtering |
Many teams run both, with the AI layer behind the API layer. The trade-offs between those deployment choices are compared in AI gateway vs API gateway options for LLM traffic.
The AI Gateway for Agents and MCP
Agents call tools as well as models, so the gateway now governs tool traffic too. Acting as an MCP gateway, it connects to MCP servers once, controls which tools each caller can see and execute, manages upstream credentials, and logs tool calls in the same trail as model calls.
Anthropic introduced the Model Context Protocol in November 2024 as an open standard for connecting AI applications to tools and data. Without a gateway, every agent embeds its own MCP clients and credentials, and security teams lose track of which tools each agent can reach; some vendors call the combined layer an AI agent gateway.
Bifrost acts as both an MCP client and an MCP server. It connects to upstream MCP servers over STDIO, HTTP, or SSE, and it exposes the tools it aggregates at a single /mcp endpoint that Claude Desktop, Cursor, and other MCP clients can use. The Bifrost MCP gateway adds these controls:
- Tool filtering at the client, request, and virtual key levels, so each consumer sees only an allow-listed set of tools.
- Six MCP authentication types: None, Headers, Per-User Headers, OAuth 2.0, Per-User OAuth, and Token Exchange (Enterprise).
- Agent Mode for autonomous tool execution with configurable auto-approval.
- Code Mode, where the model writes Python (Starlark) against tool stubs that Bifrost executes in a sandbox, instead of loading every tool definition into context.
Code Mode targets a problem Anthropic's engineering team has described: tool definitions loaded upfront consume context and raise cost as tool counts grow. In the Bifrost Code Mode benchmark, input tokens fell by 58.2% with 96 tools across 6 servers and by 92.8% with 508 tools across 16 servers, with pass rates unchanged at 100%. The full methodology is in the MCP gateway access control and cost governance write-up, and the concepts are covered in depth in what an MCP gateway is.
Governing Coding Agents Through an AI Gateway
Coding agents issue many model calls per task and run on every developer machine, and an AI gateway is how teams give them budgets, model choice, and visibility. Pointing an agent's base URL at the gateway lets the platform team attribute token spend per developer, enforce limits, and route the agent to any provider.
Bifrost has documented integrations for CLI agents and editors, including Claude Code, Codex CLI, Gemini CLI, Cursor, Opencode, Qwen Code, Zed, and Roo Code. The Bifrost CLI (npx -y @maximhq/bifrost-cli) launches Claude Code, Codex CLI, Gemini CLI, or Opencode against the gateway with configuration applied automatically and virtual keys stored in the OS keyring.
For Claude Code, the integration sets ANTHROPIC_BASE_URL to the gateway's /anthropic endpoint and ANTHROPIC_AUTH_TOKEN to a Bifrost virtual key. The full decision guide is choosing an AI gateway for Claude Code, and two narrower guides cover how to monitor Claude Code token usage per developer and how to run Claude Code with non-Anthropic models such as GPT, Gemini, or local models.
For Codex CLI, a named model_providers entry in config.toml points Codex at the gateway's OpenAI-compatible endpoint with a virtual key as the API key. From there, teams can route Codex CLI to any model, including Anthropic, Gemini, and Ollama models, and manage Codex CLI token spend with per-developer budgets and rate limits.
The pattern is the same across agents: one virtual key per developer or team, a model allow-list, and request logs showing which agent spent what.
AI Governance at the Gateway Layer
AI governance at the gateway means every model and tool call is tied to an identity, checked against a budget and rate limit, filtered by allow-lists and guardrails, and recorded. The gateway is the natural enforcement point because it is the only component that sees all AI traffic before it reaches a provider.
Bifrost organizes governance around virtual keys and a budget and rate-limit hierarchy. A customer can hold teams, a team can hold virtual keys, and a virtual key can hold per-provider configurations, with an independent budget at each level and request- and token-based limits on keys and provider configs. A request must pass every level's check before it is served.
Bifrost Enterprise extends this for large organizations:
- Identity and access. OIDC login with Okta, Microsoft Entra, Keycloak, Zitadel, or Google Workspace, plus role-based access control with custom roles.
- Policy at scale. Access Profiles define reusable provider, model, budget, rate-limit, and MCP policies that auto-allocate virtual keys to users.
- Guardrails. Bifrost-managed Prompt Guardrails, Custom Regex, and Secrets Detection, plus external providers such as AWS Bedrock Guardrails, Azure Content Safety, Google Model Armor, and Patronus AI.
- Audit and alerting. HMAC-signed audit logs of administrative activity, archivable to S3 or GCS, and alert rules that notify Slack, Microsoft Teams, PagerDuty, or webhooks when budgets or rate limits cross thresholds.
Request logs show which caller sent which prompt to which model at what cost; audit logs show who changed a key, budget, or policy. Compliance reviews usually need both, and the Bifrost Enterprise deployment adds in-VPC hosting and clustering for teams in regulated industries. How these controls work together in practice is covered in governing LLM usage in the enterprise.
Shadow AI and the endpoint
A gateway governs only the traffic configured to flow through it. Shadow AI is the remainder: desktop chat apps, browser AI, coding agents, and MCP servers that employees connect directly to providers.
The combined model is AI Gateway + Bifrost Edge. Bifrost, the AI gateway, stays the control plane and policy engine, and Bifrost Edge runs on macOS, Windows, and Linux machines to route that endpoint traffic through the gateway, so the same virtual keys, budgets, guardrails, and audit logs apply. Edge is currently in alpha. The full pattern is described in closing the last mile of AI governance.
How to Choose an AI Gateway
Choosing an AI gateway comes down to six questions: how much latency it adds, where it can be deployed, which providers it supports, how deep its governance goes, whether it handles MCP and agents, and whether it is open source. The weight of each depends on whether a developer team or a platform team owns the decision.
| Criterion | What to check | Why it matters |
|---|---|---|
| Overhead | Added latency per request at your peak RPS | The gateway sits in the hot path of every call |
| Deployment | Self-hosted, in-VPC, air-gapped, or managed only | Regulated data often cannot leave your network |
| Provider coverage | Native providers, model count, custom endpoints | Determines how much lock-in you actually remove |
| Governance depth | Virtual keys, hierarchical budgets, RBAC, SSO, audit logs | Decides whether finance and security can sign off |
| Agent readiness | MCP client and server, tool filtering, MCP auth | Agents call tools as well as models |
| Observability | Request logs, Prometheus, OpenTelemetry, log export | Gateway data must reach existing monitoring |
| License | Open source or proprietary | Affects inspection, extensibility, and exit cost |
Two guides apply these criteria to specific products. The roundup of the best AI gateways in 2026 is organized by use case, from self-hosted to managed, and the production-ready LLM gateway comparison scores tools on overhead, failover, governance, MCP, and compliance.
Observability often decides enterprise evaluations and has its own comparison of enterprise AI gateways for LLM observability. For a structured scoring template, the LLM Gateway Buyer's Guide lists the capabilities to verify during a proof of concept.
How Bifrost Implements the AI Gateway
The Bifrost AI gateway is open source, written in Go, and licensed under Apache 2.0. It unifies 25+ providers and 10,000+ models behind one OpenAI-compatible API, and in sustained benchmarks at 5,000 requests per second it adds 11 microseconds of overhead per request with a 100% success rate.
The published benchmarks measured that 11 µs figure on a t3.xlarge instance (4 vCPU, 16 GB) and 59 µs on a t3.medium (2 vCPU, 4 GB), both at 5,000 RPS with no failed requests.
Bifrost runs in two forms. The HTTP gateway ships with a web UI for providers, virtual keys, and routing rules, and starts with zero configuration:
npx -y @maximhq/bifrost
# or
docker run -p 8080:8080 -v $(pwd)/data:/app/data maximhq/bifrost
The Go SDK embeds the same engine directly in a Go service. Existing applications migrate by changing the base URL in the OpenAI, Anthropic, Google GenAI, or LiteLLM SDK, as shown in the gateway setup guide.
A single open-source instance handles roughly 3,000 to 5,000 RPS; larger deployments move to Bifrost Enterprise, which adds clustering with gossip-based state sync and zero-downtime rolling updates, adaptive load balancing based on live provider health, and in-VPC deployments for teams that need traffic to stay inside private infrastructure.
Frequently Asked Questions
What does AI gateway mean?
An AI gateway is a middleware layer that sits between AI applications and the models and tools they call. It gives every caller one API, then handles routing, failover, budgets, access control, guardrails, caching, and logging centrally. The term covers model traffic and, increasingly, agent tool calls over MCP, so it is broader than an LLM gateway that only forwards model requests.
Do I need an AI gateway?
A team needs an AI gateway once it runs more than one model provider, more than one team or product calling models, or any agent with tool access in production. At that point failover, per-team budgets, cost attribution, and audit trails cannot be maintained reliably in each application. A single prototype calling one provider can usually wait.
How much does an AI gateway cost?
AI gateway cost has two parts: the gateway itself and the model tokens it routes. Open-source gateways such as Bifrost are free to self-host under Apache 2.0, so the cost is the compute to run them. Enterprise editions add clustering, SSO, guardrails, and support under a commercial license. Token spend usually dominates, and caching and routing reduce it.
What are the most popular AI gateways?
Popular AI gateways fall into three groups: open-source self-hosted gateways such as Bifrost, managed edge gateways from CDN and hosting platforms, and API-management vendors that added AI plugins. Bifrost is the choice for enterprises that need low overhead, MCP support, and self-hosting. Managed gateways trade control for lower operational effort.
What is AI gateway vs API gateway?
An API gateway manages request-priced REST traffic with predictable payloads, handling authentication, routing, and quotas. An AI gateway manages token-priced model and tool traffic: it meters input and output tokens, streams long responses, caches by semantic similarity, fails over to equivalent models on other providers, and inspects prompts for injection, PII, and secrets. Many enterprises run both, with the AI gateway behind the API gateway.
What are the disadvantages of using an AI gateway?
An AI gateway adds a network hop, another service to operate, and a component that must stay highly available because every AI call depends on it. Teams mitigate these by choosing a gateway with low measured overhead, running it in a cluster with zero-downtime updates, and self-hosting it close to the applications. Bifrost adds 11 microseconds per request at 5,000 RPS.
Getting Started with an AI Gateway
An AI gateway gives every application, agent, and coding tool one governed path to every model and tool. Bifrost delivers that layer as open source, with 11 microseconds of overhead at 5,000 RPS and enterprise options for clustering, in-VPC deployment, and identity-based governance. Browse the Bifrost resources hub for architecture and governance guides, or book a demo with the Bifrost team to see how a production AI gateway fits your stack.




Top comments (0)