An AI gateway is easy to justify to engineers and hard to justify to a security review. Engineers ask about throughput and failover. The review asks who can call which model, who approved that, where the prompt went, and whether you can prove any of it six months later.
Those questions are where the five gateways below actually differ, and they are what this comparison is organised around.
TL;DR
- Bifrost: the most specific audit and isolation story of the five: HMAC-signed audit logs, row-level data access control between teams, SCIM 2.0. Open source on GitHub; the governance features are Enterprise.
- LiteLLM: the widest provider coverage (140+) and the largest community, at the cost of a Postgres-and-Redis production topology.
- Kong: one control plane for AI and non-AI traffic, if you already run Kong. Effectively Enterprise-only for enterprise needs.
- Cloudflare: the only one that inspects prompt content in flight (DLP, guardrails). Account-scoped tokens complicate tenant isolation.
- Vercel: traces that never carry prompt content, and ZDR routing. Governance is metered per request.
What "enterprise" changes about the question
For a small team, the free tier is the product. For an enterprise, the free tier is a trial. The tier you deploy is the paid one, and the interesting question is what it buys.
Five requirements come up in almost every procurement review:
Identity. SSO against your IdP, and SCIM so that deprovisioning a leaver in Okta actually removes their gateway access. A shared dashboard password is not an answer.
Authorization. RBAC deep enough to express your org chart, and ideally row-level controls so a developer on one team cannot read another team's keys, prompts, or spend.
Auditability. A record of who changed which policy and when, retained long enough for an audit cycle, and exportable to the SIEM you already run.
Isolation. Multi-tenancy that holds under scrutiny: per-team budgets, per-team credentials, and no path from one tenant's token to another tenant's data.
Deployment control. Where the gateway runs, where prompts go, and whether either can be constrained to a network you operate.
Everything below is assessed against those five. Where the software was self-hostable I ran it; where it is managed, the source is the vendor's documentation as of September 2026.
Quick comparison
| Identity | Authorization | Audit | Deployment | |
|---|---|---|---|---|
| Bifrost | OIDC + SCIM 2.0 (Okta, Entra, Keycloak, Zitadel, Google Workspace, Auth0, generic OIDC) | RBAC with custom roles + row-level data access control | HMAC-signed, 365-day default, S3/GCS archive | Self-host; vendor-supported private network, on-prem, air-gapped |
| LiteLLM | SSO + SCIM, OIDC/JWT | RBAC with org and team admins | Audit logs with retention policies | Self-host; air-gapped and multi-region control plane |
| Kong | OIDC, mTLS, LDAP-advanced; not in the OSS build | RBAC + workspaces; not in the OSS build | Enterprise only, per Kong's pricing page | Self-host (traditional, DB-less, hybrid) or Konnect |
| Cloudflare | Cloudflare Access against your IdP | Account-scoped tokens (see limitation) | Account-level, 18-month retention | Managed only |
| Vercel | SAML SSO (Pro add-on / Enterprise); Directory Sync for SCIM-style provisioning (Enterprise) | Team, project, key, and member scopes | Audit Log Drains to S3/Splunk/Datadog/Panther | Managed only |
Bifrost
Apache 2.0, Go, in front of 10,000+ models across 25+ providers, with 8.4k stars on GitHub.
What stands out: the audit and isolation story is unusually specific.
Most gateways say "audit logs." Bifrost's audit logs record who changed what, when, and which resource was affected, with action type, outcome, initiator, target, request path, IP, and duration. You can configure an HMAC key so entries can be cryptographically verified, retention defaults to 365 days in the database, and archived events go to S3 or GCS under your bucket's lifecycle rules. That combination (signed, retained for an audit cycle, archived independently of the database) is what a compliance reviewer is actually asking for when they ask whether you have audit logging.
The isolation story is equally concrete. User provisioning covers OIDC login and inbound SCIM 2.0 at /scim/v2, with setup guides for Okta, Entra, Keycloak, Zitadel, Google Workspace, Auth0, and any standards-compliant OIDC provider. RBAC supports custom roles when the built-in ones don't match your org, with permissions expressed as resource-plus-operation pairs. On top of that, data access control filters row-level visibility by role, team, and identity: "a developer on Team A cannot see virtual keys, prompts, or routing rules owned by Team B unless their role grants broader scope." Audit logs are scoped the same way. RBAC controls what operations a user can perform; DAC controls which rows they can perform them on. Few gateways separate those two.
Underneath, the virtual key is the organising primitive: model access, budgets nested across key, team and customer, rate limits, routing, and which MCP tools that key may reach all hang off one object. For a multi-tenant product, per-customer budget nesting is what makes cost attribution possible without a separate metering pipeline.
Clustering is specific about its guarantees: six discovery methods (Kubernetes, Consul, etcd, DNS, UDP, mDNS), gossip-based state sync, rolling zero-downtime deploys, and a documented fault tolerance of one node failure at three nodes, two at five. It is eventually consistent; the docs say nodes "converge to the same state within seconds," which is worth knowing if you expect a budget counter to be globally exact at the instant of enforcement.
The open-source build is also more capable than most vendors' free tiers. Telemetry, logging, and governance are auto-loaded built-ins. In the maximhq/bifrost:latest image I pulled on September 2026, the startup log showed seven internal plugins registering with no configuration:
plugin status: logging - active
plugin status: governance - active
plugin status: routing - active
plugin status: telemetry - active
plugin status: compat - active
plugin status: model-catalog-resolver - active
plugin status: prompts - active
/metrics served Prometheus with no exporter configured, and the virtual key, team, budget, and customer endpoints answered immediately against the bundled SQLite store, with no external database needed to create a key with a spend cap and a rate limit. That matters for enterprises mainly as a proof-of-concept path: you can demonstrate governance before procurement approves anything.
Limitations: Bifrost covers well under a fifth of LiteLLM's providers by count and has the smallest community of the five. RBAC, DAC, OIDC provisioning, signed audit logs, clustering, and guardrails are Enterprise, as are vendor-supported private-network, on-premise, and air-gapped deployments; the OSS container itself runs anywhere. The bundled SQLite is single-instance only. The migration docs call it "single-writer and single-host by design," so any HA deployment starts by moving the config and log stores to PostgreSQL.
Best for: Bifrost is tailored for regulated environments requiring signed audit trails, row-level tenant isolation, and a documented HA self-hosting architecture. It deploys air-gapped or on-prem to provide a single governance layer across both LLM calls and MCP tools.
LiteLLM
Python SDK plus proxy, 140+ providers, the widest coverage here, and the largest open-source project of the five at 59.8k GitHub stars.
What stands out: it will already support whatever your teams want to call.
Enterprises rarely standardise on one provider. Between the fine-tuned model one team runs on vLLM, the Bedrock endpoint procurement already approved, and whatever was announced last week, provider coverage stops being a checkbox and becomes the thing that determines whether the gateway can be mandatory. LiteLLM lists 140+ providers on its homepage, OpenAI, Anthropic, Azure OpenAI, Bedrock, Vertex AI, Cohere, Mistral, Groq, Ollama, and vLLM among them, across chat, embedding, image, audio, and rerank endpoints.
Its governance model is correspondingly broad: virtual key auth, a user/team/organization hierarchy, budgets at key, user and team level, RPM and TPM limits, model groups, load balancing across deployments, retries, in-model failover, cross-model fallback, A/B testing, and traffic mirroring. Enterprise adds SSO and SCIM, OIDC/JWT auth, RBAC with org and team admins, audit logs with retention policies, secret-manager integration, key rotation, tag-based spend tracking, a multi-region control plane, and an air-gapped deployment option.
If you already run an observability stack, this is the gateway that feeds it without an adapter: Langfuse, OpenTelemetry, LangSmith, Datadog, Arize, MLflow, S3, GCS, Azure Blob, and Azure Sentinel are all supported destinations. Prometheus metrics are open source, not Enterprise, although several comparisons say otherwise. Enabling the callback on the OSS container (main-latest, pulled 30 Aug 2026) exposed around 150 litellm_* series covering request counts, latency histograms, and spend, labelled by model, team and key, with no license gate.
Limitations: LiteLLM's production topology is the real cost. Virtual keys need PostgreSQL (the docs are specific that it is Postgres, not SQLite), Redis is required "as soon as you run more than one proxy instance," and past "roughly 1000+ requests per second" the production guide recommends a Redis transaction buffer. That is sound engineering, since contended budget counters across replicas need a shared transactional store, but it is a database, a cache, and a tuning guide to stand up before you have governance.
Best for: LiteLLM is the choice when provider breadth is the binding constraint, such as teams running their own inference alongside hosted APIs. It feeds an existing observability stack (Langfuse, Datadog, OpenTelemetry) without an adapter, and assumes a platform team to run the Postgres and Redis it depends on.
Kong AI Gateway
Kong AI Gateway is Kong Gateway plus a set of AI plugins, with an Enterprise tier and the Konnect control plane.
What stands out: one policy layer for AI and everything else.
The other four gateways govern AI traffic. Kong governs AI traffic using the same control plane, the same RBAC model, and the same operational runbooks you already use for your REST and gRPC services. For an enterprise that already runs Kong, that is not a feature. It is the elimination of a second governance system, a second audit surface, and a second on-call rotation. Kong 3.14 (April 2026) extends this to MCP and agent-to-agent traffic in one runtime with per-tool ACLs.
The catch is where the line falls. I listed the plugin directory inside the kong:3.9.3 OSS image and cross-checked it against available_on_server on the running admin API, which reports 45 plugins. The authentication primitives an enterprise needs are mostly not there: openid-connect, mtls-auth, vault-auth, ldap-auth-advanced, jwt-signer, and rbac are all absent from the OSS build. What ships free is the basic set (acl, oauth2, jwt, basic-auth, key-auth, ldap-auth, session) alongside six AI plugins (ai-proxy, ai-prompt-guard, ai-prompt-decorator, ai-prompt-template, ai-request-transformer, ai-response-transformer) and the prometheus and opentelemetry plugins.
So for an enterprise evaluation, Kong is effectively an Enterprise-only product. OIDC against your IdP, mTLS between services, and RBAC are absent from the open-source image; Kong's pricing page lists SSO and audit logging as Enterprise-only; and the AI-specific cost controls (ai-rate-limiting-advanced for token-aware limits, ai-semantic-cache, ai-sanitizer for PII, ai-mcp-proxy) are likewise missing from the OSS build. That is a coherent position; it is just one to price before comparing against the others' free tiers.
Limitations: Kong's open-source and Enterprise lines are not the same codebase at the same version. The newest OSS release is 3.9.3, which is what kong:latest pulls, while the 3.14 line carrying Agent Gateway is Enterprise, so a feature in a 3.14 announcement is not reachable by upgrading the open-source build. Documented self-hosting topologies are traditional, DB-less, and hybrid, plus the Konnect control plane; no air-gapped mode is documented. Adopting Kong solely to proxy LLMs means operating a large general-purpose API platform for a narrow problem.
Best for: Kong is built for enterprises already standardised on Kong Gateway that want AI traffic under the same RBAC, audit, and operational model as their existing APIs. It eliminates a second control plane rather than adding a better one, and requires the Enterprise license to do so.
Cloudflare AI Gateway
Cloudflare AI Gateway is managed, with two dozen providers listed in the docs, plus analytics, logging, caching, rate limiting, retries, dynamic routing, guardrails, and DLP. Nothing to self-host, so everything here is from Cloudflare's documentation rather than something I ran.
What stands out: inspection of the prompt itself.
The other four gateways govern who may call a model. Cloudflare is the one that inspects what they send. DLP applies to prompts and responses in flight, with built-in profiles for financial and identifier data and custom profiles under Cloudflare One, and guardrails flag or block harmful content. If the thing keeping your security team from approving an LLM feature is that an employee might paste a customer record into it, this is the only option here that treats that as a first-class problem rather than leaving it to your application.
Cloudflare Access puts identity in front of the gateway: users authenticate against your IdP, each request carries a verified cf.user_id, and logs, analytics, and spend can be filtered by that identity. Dynamic Routing is configured in a visual editor or JSON: conditional branches on body, headers or metadata, per-user and per-team rate limits, budget limits, model fallback, A/B tests, percentage rollouts, route versioning, and instant rollback. Changing which model serves production traffic requires no application deploy, and someone who does not write code can read the routing graph during a review.
Telemetry is not confined to the dashboard. Spans go to any OTLP-compatible backend in JSON or protobuf via the OpenTelemetry integration, carrying model, provider, token counts, cost estimates, custom metadata, and, unlike Vercel's, the prompts and completions themselves. A GraphQL analytics API and Workers Logpush cover raw logs, and account audit logs record create, update and delete actions with 18-month retention.
Limitations: Cloudflare's tokens are account-scoped. Per the authentication docs, AI Gateway Run permission "cannot be restricted to a single gateway," so a token issued for one gateway reaches every gateway in the account, including those holding stored provider keys; isolating business units means separate accounts or Worker-side bindings. DLP on streaming responses "buffers the full streamed response before scanning," so time-to-first-token grows to the full generation time. Request-only scanning avoids that, at the cost of not inspecting the model's output.
Best for: Cloudflare AI Gateway is designed for security-led adoption where the blocker is what employees might put in a prompt. It inspects request and response content in flight with DLP and guardrails, and fits organisations already inside the Cloudflare One perimeter.
Vercel AI Gateway
Vercel AI Gateway fronts hundreds of models across many providers behind one endpoint, with BYOK on the paid tier.
What stands out: prompts stay out of the telemetry by construction.
Every other gateway here can be configured to keep prompt bodies out of exported traces. Vercel is the one where it is not a setting: trace drains emit an OpenTelemetry trace per request, and the docs state that "traces contain request metadata but not prompt or completion content." Nobody has to remember to turn it off, and nobody can turn it on by mistake, which is a shorter conversation with a compliance reviewer than a configuration flag that someone could flip.
Zero Data Retention extends that upstream. Vercel states that AI Gateway itself "does not retain prompts, outputs, or sensitive data," and ZDR routing restricts traffic to providers Vercel holds retention agreements with. It can be enforced team-wide from the dashboard or per request, and BYOK keys are excluded by default unless you mark them as covered by your own agreement. When no compliant provider exists for a model, the request fails rather than falling back silently, which is the correct behaviour for a compliance control, and worth designing around.
Identity is handled at the Vercel team level rather than inside the gateway. SAML SSO is a paid add-on on Pro and included on Enterprise, and can be enforced so that team members cannot access team resources without an SSO session. Directory Sync, Enterprise only, provisions and deprovisions members from Okta or Google Directory and maps IdP groups to Vercel roles, which is the SCIM-shaped answer to the leaver question. Budgets can be set at team, project, API-key, and team-member level, with member budgets enforced through spend attribution: keys assigned to a user count against both that member's budget and the team's. Team activity reaches your SIEM through Audit Log Drains, with S3, Splunk, Datadog, and Panther as supported destinations. Provider routing is unusually granular (order, only, and sort by cost, TTFT or throughput), which matters when procurement has approved some providers and not others.
Limitations: Vercel prices governance per request rather than per seat. Per the pricing page, team-wide provider allowlisting and team-wide ZDR each cost $0.10 per 1,000 requests, custom reporting is $0.075 per 1,000 writes plus $5 per 1,000 queries, and trace drains are $0.05 per 1,000 traces plus egress; the per-request equivalents carry no surcharge. Model those numbers at your volume before assuming the no-markup token pricing makes it the cheap option. Governance is shaped around Vercel teams and projects rather than an arbitrary org hierarchy, DLP stays in your application, and there is no self-hosted deployment.
Best for: Vercel AI Gateway is suited to teams whose workloads already run on Vercel. It emits telemetry that cannot carry prompt content by design, routes only to providers on an approved list, and enforces zero data retention team-wide or per request.
How to choose
Bifrost: you need signed audit logs, row-level data isolation between teams, and self-hosting with a documented HA story, and you want the governance model to cover MCP tools as well as models. The open-source build is complete enough to prove the design before procurement engages.
LiteLLM: provider coverage is the binding constraint, your teams run their own inference alongside hosted APIs, or you need the gateway to feed an observability stack you already operate. Assumes a platform team to run Postgres, Redis, and the tuning the production guide describes.
Kong: you already run Kong, and a second control plane for AI traffic is the thing you are trying to avoid. Price the Enterprise license first; the OSS build lacks the authentication primitives an enterprise review will ask for.
Cloudflare: the blocker is what employees might put in a prompt rather than which model they call, or you already run Cloudflare One and want AI traffic under the same security perimeter. Check the account-scoped token behaviour against your tenancy model early.
Vercel: your workloads run on Vercel, or you want telemetry that cannot carry prompt content and provider routing constrained to an approved list. Model the per-request governance charges at your volume.
Bottom line
The order these questions get asked in matters. Provider coverage and latency are easy to compare and quick to agree on; identity, isolation, and audit are harder to compare and harder to retrofit. A gateway that cannot deprovision a leaver through SCIM, or cannot show a reviewer who changed a budget policy in March, will not survive the review regardless of how many models it fronts.
Start from the constraint that will actually block the deployment (the compliance framework you are certifying against, the tenancy model you have to enforce, the network your prompts are allowed to cross) and let that eliminate options before comparing features.
Disclosure: this post was made in collaboration with the Bifrost team; the testing and conclusions are mine.





Top comments (0)