Every AI team hits the same wall. The prototype talks to one provider with one key. Then production arrives: five teams, three providers, a finance lead asking where the money went, and a 2am outage that takes your chatbot down with it. Or worse, a compromised dependency in the one layer every request flows through, quietly shipping your provider keys to someone else.
That is the job of an AI gateway. And picking one in 2026 is less about feature lists than about three questions: where your traffic is allowed to live, which features you will actually have to pay for, and what happens when a provider falls over.
TL;DR
- Bifrost: best for enterprises that need self-hosted performance, governance and MCP in one gateway, open source at the core with an Enterprise tier for air-gapped, VPC and on-prem deployments.
- Kong AI Gateway: best if you already run Kong and want AI traffic under the same policies as your APIs.
- LiteLLM: best for Python teams that want the widest provider coverage and are happy to tune it.
- Cloudflare AI Gateway: best for zero-ops observability, caching and one bill, at the edge.
- Vercel AI Gateway: best for teams on Vercel and the AI SDK that want broad model access with no token markup.
What an enterprise AI gateway does (and how it differs from an API gateway)
An AI gateway is a reverse proxy for LLM traffic. Your apps call one API, and the gateway handles the rest: which provider, which key, what happens on failure, who is allowed to spend what.
The difference from a classic API gateway is what it understands.
An API gateway sees requests: paths, auth headers, request counts. An AI gateway sees tokens, models, prompts and dollars. It can cap a team's monthly spend, retry a failed call on a different provider, serve a repeated question from cache, and, in 2026, govern the tool calls your agents make through MCP.
Most enterprises end up running both. The API gateway guards the front door. The AI gateway guards the bill and the models.
API gateway vs AI gateway. The API gateway counts requests and checks paths; the AI gateway counts tokens and dollars, routes between models, caches, applies guardrails and governs agent tool calls.
How I evaluated them
I scored each gateway on six things that decide whether it survives an enterprise rollout:
- Deployment. What I checked: self-hosted, VPC, on-prem, or managed only. Why it matters: decides whether regulated traffic can use it at all.
- Licensing. What I checked: what is open source and what sits behind a paid tier. Why it matters: shows where the real cost starts.
- Governance. What I checked: virtual keys, budgets, rate limits, SSO, audit logs. Why it matters: who can spend what, and a record of who did what.
- Reliability. What I checked: failover, load balancing, retries. Why it matters: what users feel when a provider goes down.
- Cost control. What I checked: caching, spend visibility, token markup. Why it matters: whether the bill stays predictable as usage grows.
- Agents. What I checked: MCP gateway support. Why it matters: governing tool calls, not just model calls.
Two of the five, Bifrost and LiteLLM, are open-source gateways you run yourself, so I ran both on my own machine against the same mock provider, with the same load, including a forced provider outage. Kong, Cloudflare and Vercel are covered from their current docs (checked September 2026) and, where they exist, their own published benchmarks.
1. Bifrost
Bifrost is the open-source AI gateway from Maxim AI. It is written in Go, licensed Apache 2.0, and puts 1000+ models behind one OpenAI-compatible API. It is a drop-in replacement for the OpenAI, Anthropic, AWS Bedrock and Google GenAI SDKs, so moving an app over is a base-URL change. Version 2.2.1 shipped on September 18, 2026.
What stands out
- Speed. Maxim's published benchmark shows 11 µs of added overhead at 5,000 requests per second on a t3.xlarge. On my laptop it added about 2.6 ms at the median while serving ~10,000 requests per second, with request logging switched on.
- Governance in the free tier. Virtual keys with budgets and rate limits, assignable to teams or customers, are open source.
- Semantic caching with Redis or Valkey, Qdrant, Weaviate or Pinecone as the store, so near-identical prompts skip the model call. See the semantic caching docs.
- A real MCP gateway. Tool filtering per virtual key, an agent mode that auto-executes approved tools, and a code mode that cuts tool-definition tokens.
- Failover that stays fast. With the primary provider dead in my test, every request succeeded through the fallback at an ~8 ms median.
On the enterprise side, Bifrost Edge (currently in alpha) extends the same gateway policies out to employee devices, so the gateway governs server traffic and Edge governs the AI apps on laptops.
Best for: Bifrost is built for enterprises running mission-critical AI workloads that require best-in-class performance, scalability, and reliability. It serves as a centralized AI gateway to route, govern, and secure all AI traffic across models and environments with ultra low latency. Bifrost unifies LLM gateway, MCP gateway, and Agents gateway capabilities into a single platform. Designed for regulated industries and strict enterprise requirements, it supports air-gapped deployments, VPC isolation, and on-prem infrastructure. It provides full control over data, access, and execution, along with robust security, policy enforcement, and governance capabilities.
Watch out for
- The paid line is real. Clustering, adaptive load balancing, guardrails, SSO with RBAC and SCIM, signed audit logs, vault integrations and in-VPC deployments are Enterprise features with custom pricing.
- A younger ecosystem. About 8,000 GitHub stars against LiteLLM's ~59,000, and the repo only dates back to March 2025.
- Vendor benchmarks. The headline numbers on its site are Maxim's own. My test points the same direction, but it is a laptop, not a lab.
Pricing: open source is free. Enterprise is custom, with a 14-day trial. Maxim states SOC 2 Type 2, ISO 27001, HIPAA and GDPR compliance.
2. Kong AI Gateway
Kong started with AI plugins on top of Kong Gateway. On September 1, 2026, Kong AI Gateway 2.0 went GA as its own product, with its own runtime, control plane and release cycle, available self-managed or through Kong Konnect. It covers three kinds of traffic: calls to LLMs, access to MCP servers, and agent-to-agent (A2A).
What stands out
- Policy depth. Seven load-balancing algorithms including semantic and lowest-latency routing, token and cost-based rate limits, semantic caching, and guardrail integrations with AWS, Azure, Google and Lakera.
- Serious MCP support. An MCP proxy that can turn REST APIs into MCP tools, OAuth2 for MCP, and MCP Server Bundling, where each caller only sees the tools it is authorized to use.
- Raw proxy speed. Kong's own benchmark (July 2025, mocked LLM, no policies enabled) measured about 23,400 requests per second against about 2,700 for LiteLLM on the same hardware.
- One control plane for APIs and AI, which is the whole point if Kong already runs your API estate.
Best for: organizations that already run Kong for their APIs and want LLM, MCP and agent traffic under the same policies.
Watch out for
- Open source covers the basics only. The Apache 2.0 repo includes the basic AI proxy and prompt plugins. AI Proxy Advanced, semantic caching, advanced AI rate limiting, PII sanitizing and the MCP proxy are Enterprise.
- A migration is coming. AI plugins become opt-in in Kong Gateway 3.18, so 3.x users should plan a move to the 2.0 runtime.
- Per-model pricing. The Konnect Plus plan lists AI Gateway at $100 per month per LLM model, capped at five, on top of gateway fees of $200 per month per hybrid control plane.
Pricing: a 30-day free trial, Plus as above, Enterprise custom and billed annually.
3. LiteLLM
LiteLLM is a Python SDK plus a proxy server from BerriAI. The core is MIT licensed, it has the biggest community in this list (about 59,000 GitHub stars), and it exposes 100+ LLM APIs in the OpenAI format.
What stands out
- Coverage and community. If a provider exists, LiteLLM probably supports it, and someone has probably written about it.
- A generous open-source tier. Virtual keys, teams, budgets, rate limits, fallbacks, exact and semantic caching, a long list of guardrail integrations, an MCP gateway, an A2A agent gateway and an admin UI.
- Pricing that is not per token. Enterprise quotes are sized by request capacity.
Best for: Python teams that want the widest provider coverage and a feature-rich open-source tier, and are ready to tune and patch it.
Watch out for
- Throughput. In my test, a single LiteLLM worker served about 380 requests per second at a 114 ms median. With eight workers it reached about 1,200 requests per second at a ~28 ms median. LiteLLM's own docs report around 2 ms of median overhead on a scaled-out setup, and the team is building a Rust gateway (in beta) to close the gap.
-
Retry defaults. With the primary provider down, the default config spent about 4.7 seconds per request retrying before it fell back. Setting
num_retries: 0brought that down to a 46 to 87 ms median. Check this before you ship. - A rough 2026 on supply chain. On March 24, malicious versions 1.82.7 and 1.82.8 were live on PyPI for about 40 minutes, carrying a credential stealer. The official Docker proxy image was not affected, and a clean 1.83.0 shipped on March 30 from a rebuilt CI/CD pipeline. Pin versions and prefer the official image.
- Enterprise-only essentials. SSO beyond five users, audit logs, JWT auth and per-key guardrails need the paid tier.
Pricing: open source is free. Enterprise is quote-based.
4. Cloudflare AI Gateway
Cloudflare AI Gateway is a managed proxy running on Cloudflare's network. The docs list 24 providers, it offers OpenAI- and Anthropic-compatible endpoints, and in August 2026 Cloudflare merged it with Workers AI into a single control plane.
What stands out
- The core is free. Analytics, caching and rate limiting only need a Cloudflare account.
- Unified Billing. Pay OpenAI, Anthropic, Google, xAI and Groq through Cloudflare credits, with provider prices passed through and a 5% fee on credits.
- Spend limits by model, provider or user metadata, plus configurable retries and dynamic routing rules.
- Guardrails and DLP at the edge, using Llama Guard and Cloudflare One detection profiles.
Best for: teams already on Cloudflare that want free observability, caching and rate limiting with almost no setup, and one bill across providers.
Watch out for
- Managed only. No self-hosting, Regional Services is not supported, and log storage cannot be restricted by jurisdiction.
- Exact-match caching. Any difference in the request body is a cache miss. Semantic caching is on the roadmap, not in the product.
- Guardrails cost latency. Around 500 ms, and they buffer streamed responses before returning them.
- No MCP gateway in the product itself. MCP controls live in a separate Cloudflare One product.
Pricing: core features are free. The free plan stores 100,000 logs, Workers Paid raises that to 10 million per gateway, and Logpush beyond 10 million is $0.05 per million.
5. Vercel AI Gateway
Vercel AI Gateway went GA in August 2025 and is the default provider in the AI SDK, though any stack can call it. It offers hundreds of models (its public catalog listed 377 when I checked) through OpenAI- and Anthropic-compatible APIs.
What stands out
- No markup and no platform fee on tokens, including when you bring your own keys on the paid tier.
- Automatic provider failover, provider ordering and team-wide routing rules.
- Budgets per team, project, API key or user, plus zero data retention and no-training controls, and US or EU regional inference.
- OpenTelemetry trace export on Pro and Enterprise.
Best for: teams building on Vercel or the AI SDK that want broad model access, failover and spend controls without running a proxy.
Watch out for
- Managed only, with partial residency. Pinning the inference region does not pin where Vercel processes the request itself.
- Soft budgets. The request that crosses a budget still completes, and bring-your-own-key spend does not count toward it.
- Thin on governance features. No gateway-side semantic cache, no built-in guardrails and no MCP governance; Vercel positions MCP authorization in the tool layer.
- Add-ons add up. Team-wide ZDR and provider allowlists cost $0.10 per 1,000 requests each.
Pricing: provider list price on prepaid credits, with per-request add-ons and invoiced billing on Enterprise.
The hands-on test: overhead and failover
Here is exactly what I ran, so you can repeat it.
A mock OpenAI-compatible provider that answers in about 1 ms, Bifrost 2.2.1 on its defaults, and LiteLLM 1.102.0 with one worker and then eight, all on one 10-core laptop. Each run fired 5,000 chat completions with 50 concurrent connections.
Normal traffic. Median latency and throughput for direct calls, Bifrost, LiteLLM with eight workers and LiteLLM with one worker, against the same mock provider.
The results, rounded across repeat runs:
- Direct to the mock: ~35,000 requests per second, ~1 ms median.
- Bifrost: ~10,000 to 12,000 requests per second, ~3.6 ms median, ~15 ms p99.
- LiteLLM, 8 workers: ~1,200 requests per second, ~28 ms median, ~200 ms p99.
- LiteLLM, 1 worker: ~380 requests per second, ~114 ms median, ~270 ms p99.
Then I broke it on purpose. I pointed each gateway's primary provider at a dead port, kept a healthy backup, and sent 2,000 more requests.
Primary provider down. All three setups kept a 100% success rate, but the median time to answer ranged from about 8 ms (Bifrost) to about 60 ms (LiteLLM with retries off) to 4.7 seconds (LiteLLM defaults).
Every setup answered 100% of requests. The difference was how long each one made the user wait. In Bifrost, the fallback is one field on the request:
curl -X POST http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"openai/gpt-4o-mini","fallbacks":["groq/gpt-4o-mini"],"messages":[{"role":"user","content":"ping"}]}'
In LiteLLM, the fix for the slow failover was two lines of router config:
router_settings:
num_retries: 0
fallbacks: [{"openai/gpt-4o-mini": ["backup"]}]
Two caveats. This is a laptop with the mock, the gateway and the load generator on the same machine, so treat it as direction, not a lab result. And Bifrost was also writing every request to its local log store during the test, work LiteLLM was not doing, which makes the gap more notable, not less.
Comparison at a glance
Based on vendor docs as of September 2026. Ent = Enterprise tier.
Deployment
- Bifrost: Self-hosted; VPC, on-prem and air-gapped (Ent)
- Kong AI Gateway: Self-managed or Konnect SaaS
- LiteLLM: Self-hosted
- Cloudflare AI Gateway: Managed only
- Vercel AI Gateway: Managed only
Open source
- Bifrost: Yes, Apache 2.0
- Kong AI Gateway: Core only; most AI features Ent
- LiteLLM: Yes, MIT core
- Cloudflare AI Gateway: No
- Vercel AI Gateway: No
Keys and budgets
- Bifrost: Yes, free
- Kong AI Gateway: Token budgets (Ent)
- LiteLLM: Yes, free
- Cloudflare AI Gateway: Spend limits
- Vercel AI Gateway: Soft budgets
Failover
- Bifrost: Fallbacks and weighted LB free; adaptive LB (Ent)
- Kong AI Gateway: Advanced LB (Ent)
- LiteLLM: Fallbacks and LB
- Cloudflare AI Gateway: Retries and dynamic routing
- Vercel AI Gateway: Provider failover and routing rules
Caching
- Bifrost: Semantic, free
- Kong AI Gateway: Semantic (Ent)
- LiteLLM: Exact and semantic
- Cloudflare AI Gateway: Exact match only
- Vercel AI Gateway: Provider prompt caching only
Guardrails
- Bifrost: Many integrations (Ent)
- Kong AI Gateway: Several integrations (Ent)
- LiteLLM: Free integrations; per-key scoping (Ent)
- Cloudflare AI Gateway: Llama Guard and DLP
- Vercel AI Gateway: None built in
MCP gateway
- Bifrost: Yes, free
- Kong AI Gateway: Yes (Ent)
- LiteLLM: Yes, free
- Cloudflare AI Gateway: Separate product
- Vercel AI Gateway: No
Pricing
- Bifrost: Open source free; Enterprise custom
- Kong AI Gateway: From $100 per model per month plus gateway fees
- LiteLLM: Open source free; Enterprise quote
- Cloudflare AI Gateway: Core free; 5% fee on Unified Billing credits
- Vercel AI Gateway: No token markup; per-request add-ons
How to choose
Start with the question that removes the most options: can your AI traffic leave your infrastructure?
- It cannot. You are choosing between the self-hosted three. If Kong already runs your APIs, extend it. If you want the fastest gateway with governance and MCP in the open-source tier, pick Bifrost. If you are Python-first and value provider breadth over throughput, LiteLLM works, as long as you scale workers, pin versions and fix the retry defaults.
- It can. Managed is less work. If you already live on Cloudflare, its gateway is nearly free to adopt. If you build on Vercel or the AI SDK, Vercel's gateway is the shortest path.
Then pressure-test your pick on the two things that bite later: what the paid tier costs once you need SSO and audit logs, and what happens during a provider outage.
A decision flow. Must traffic stay in your VPC? If yes, choose Kong for existing Kong shops, Bifrost for performance plus governance and MCP, or LiteLLM for Python-first teams. If no, choose Cloudflare for Cloudflare shops or Vercel for Vercel and AI SDK teams.
FAQ
What is an AI gateway?
An AI gateway is a reverse proxy that sits between your applications and the model providers they call. Your apps talk to one API, and the gateway decides which provider and key to use, fails over when something breaks, enforces budgets and rate limits per team, caches repeated answers, applies guardrails and logs every call. With a gateway like Bifrost, adopting all of that is a base-URL change rather than a rewrite of your application code.
Do I need an AI gateway if I only use one provider?
Usually, yes, once you are in production. Even with a single provider you want per-team budgets, isolated keys you can revoke without touching every service, and a record of who called what. A gateway also keeps your options open: adding a second provider later becomes a configuration change instead of a rewrite. Bifrost's virtual keys give you those budgets and rate limits in the open-source tier.
Is an AI gateway the same as an API gateway?
No. An API gateway manages requests: paths, authentication and request-level rate limits. An AI gateway understands what is inside the request, meaning tokens, models, prompts and cost, so it can cap spend, route between models and cache by meaning. Most enterprises run both, and some take AI governance a step further: Bifrost Edge applies the same gateway policies to the AI apps and coding agents running on employee devices.
Which AI gateway is best for MCP and agents?
Bifrost and LiteLLM ship MCP gateways in their open-source tiers, and Kong offers the deepest MCP policy set in its Enterprise tier. Bifrost's MCP gateway lets you control which tools each virtual key can see, auto-execute approved tools in agent mode, and cut tool-definition tokens with code mode. Cloudflare and Vercel keep MCP outside their AI gateways, in separate products or in the tool layer.
Does an AI gateway add latency?
Some, and how much depends on the gateway's language and architecture. In my test the added median latency ranged from about 2.6 ms (Bifrost) to about 27 ms (LiteLLM with eight workers), and Maxim's published benchmark puts Bifrost at 11 µs of overhead at 5,000 requests per second on server hardware. Provider latency usually dominates the total, so test failover behavior as well as raw overhead: that is where the differences users actually feel show up.
Is LiteLLM safe for production?
Many teams run it in production successfully. Treat it like any critical dependency: scale it with multiple workers, use the official Docker image, pin versions after the March 2026 supply-chain incident, patch quickly, and review the retry settings so a dead provider does not add seconds to every request.
Final Thoughts
There is no single best AI gateway in 2026. There is a best one for where your traffic lives and what you are willing to operate.
If you can go managed, Cloudflare and Vercel remove almost all of the work. If you cannot, the choice is about performance, the open-source line and the ecosystem you already run: Kong for Kong shops, LiteLLM for provider breadth, and Bifrost if you want the fastest gateway with governance and MCP you can run yourself, with Bifrost Edge extending the same policies to employee devices.
Whichever you pick, do what I did before trusting it: send it real load, then kill a provider and watch what your users would have felt.
Resources & References
- Bifrost on GitHub (open source)
- Bifrost documentation
- Bifrost Edge documentation
- Kong AI Gateway 2.0 GA announcement
- LiteLLM documentation
- LiteLLM March 2026 security update
- Cloudflare AI Gateway documentation
- Vercel AI Gateway documentation
Stay in Touch
Short takes and discussions on X
→ x.com/sebuzdugan
Practical AI / ML videos on YouTube
→ youtube.com/@sebuzdugan
Partnerships & collabs
→ sebuzdugan@gmail.com
Originally published on Medium.




Top comments (0)