<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Swapnoneel Saha</title>
    <description>The latest articles on DEV Community by Swapnoneel Saha (@swapnoneel123).</description>
    <link>https://dev.to/swapnoneel123</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1320336%2Fc601d462-6b3a-47ac-b0ae-889c5fc913d1.png</url>
      <title>DEV Community: Swapnoneel Saha</title>
      <link>https://dev.to/swapnoneel123</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/swapnoneel123"/>
    <language>en</language>
    <item>
      <title>Best Open-Source MCP Gateways for Secure AI Agent Access (2026)</title>
      <dc:creator>Swapnoneel Saha</dc:creator>
      <pubDate>Wed, 30 Sep 2026 10:48:55 +0000</pubDate>
      <link>https://dev.to/swapnoneel123/best-open-source-mcp-gateways-for-secure-ai-agent-access-2026-1i68</link>
      <guid>https://dev.to/swapnoneel123/best-open-source-mcp-gateways-for-secure-ai-agent-access-2026-1i68</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnef0yxcw553xcbl4tg4m.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnef0yxcw553xcbl4tg4m.png" alt="Best Open-Source MCP Gateways for Secure AI Agent Access (2026)" width="800" height="457"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  An open-source MCP gateway is the essential control plane for securing, governing, and observing AI agents that interact with external tools and APIs.&lt;/li&gt;
&lt;li&gt;  Key evaluation criteria for 2026 include performance overhead, governance capabilities (access control, budgets), security features, and deployment flexibility (self-hosted, air-gapped).&lt;/li&gt;
&lt;li&gt;  Bifrost is the best overall open-source MCP gateway for teams requiring microsecond latency, unified LLM and MCP governance, and enterprise-grade security in a self-hostable binary.&lt;/li&gt;
&lt;li&gt;  Other strong open-source options include Docker MCP Gateway for container-native workflows and Obot for teams seeking a comprehensive, self-hosted MCP platform.&lt;/li&gt;
&lt;li&gt;  Self-hosting an MCP gateway provides full control over data, credentials, and audit trails, which is a critical requirement for regulated industries and production AI workloads.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;As AI agents become more autonomous, the need to govern their access to internal systems has become a primary security concern for engineering teams. The &lt;a href="https://www.linuxfoundation.org/press/linux-foundation-and-anthropic-launch-the-model-context-protocol-mcp-to-standardize-how-ai-models-and-agents-interact-with-external-tools-and-data-sources" rel="noopener noreferrer"&gt;Model Context Protocol (MCP)&lt;/a&gt; has emerged as the open standard for this agent-to-tool communication, but the protocol itself doesn't solve for security or governance. This is where an MCP gateway becomes critical infrastructure. &lt;a href="https://www.getmaxim.ai/bifrost" rel="noopener noreferrer"&gt;Bifrost&lt;/a&gt;, an &lt;a href="https://github.com/maximhq/bifrost" rel="noopener noreferrer"&gt;open-source AI gateway&lt;/a&gt; from Maxim AI, is one of several solutions designed to provide this control layer.&lt;/p&gt;

&lt;p&gt;An MCP gateway acts as a centralized proxy between AI clients (like Claude Desktop or custom agents) and the MCP servers that expose tools, APIs, and data. Instead of letting agents connect directly to a sprawling landscape of tools, the gateway enforces access control, manages credentials, and creates a complete audit trail for every tool call. This article compares the best open-source MCP gateways available in 2026 for teams that need to secure agent access while maintaining full control over their infrastructure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key Criteria for Evaluating Open-Source MCP Gateways
&lt;/h2&gt;

&lt;p&gt;When selecting an MCP gateway, the focus has shifted from basic proxying to enterprise-grade governance and performance. A production-ready gateway must provide robust security without becoming a bottleneck.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Criterion&lt;/th&gt;
&lt;th&gt;Why It Matters&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Performance &amp;amp; Overhead&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Every tool call an agent makes passes through the gateway. High latency here directly impacts the end-user experience of agentic applications. Look for gateways with overhead measured in microseconds.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Security &amp;amp; Governance&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The core function of the gateway. This includes role-based access control (RBAC), per-user or per-agent virtual keys, budget and rate-limit enforcement, and the ability to filter which tools are accessible.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Authentication &amp;amp; Identity&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Must support modern authentication protocols like OAuth 2.0 and integrate with enterprise identity providers (IdPs) like Okta or Entra ID to manage user and agent identities.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Observability&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;A centralized gateway is the ideal place to log and trace every tool interaction. Native support for OpenTelemetry (OTLP) and Prometheus metrics is essential for monitoring and auditing agent activity.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Deployment Model&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;For security and compliance, self-hosting is non-negotiable. The gateway must be deployable in any environment, including private clouds (VPC), on-premises data centers, and air-gapped networks.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Unified LLM &amp;amp; MCP Control&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Agents interleave model calls and tool calls. A gateway that can govern both types of traffic through a single control plane simplifies operations, unifies cost tracking, and provides a complete trace of an agent's actions.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh75yhqmsutd4lq7zwpk0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh75yhqmsutd4lq7zwpk0.png" alt="A blueprint schematic of a complex digital lock or vault, representing the granular access controls and security policie" width="800" height="457"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Top 3 Open-Source MCP Gateways in 2026
&lt;/h2&gt;

&lt;p&gt;Based on the criteria above, here is an assessment of the leading open-source options for securing AI agent access.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Bifrost
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://www.getmaxim.ai/bifrost" rel="noopener noreferrer"&gt;Bifrost&lt;/a&gt; is a high-performance, open-source gateway written in Go that unifies LLM routing and MCP governance in a single binary. It is designed for mission-critical workloads where performance and security are paramount.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; Enterprise teams and performance-sensitive applications that need a single, self-hostable control plane for all AI traffic (both model and tool calls) with microsecond latency.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key Features:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Unified Gateway:&lt;/strong&gt; &lt;a href="https://www.getmaxim.ai/bifrost/resources/mcp-gateway" rel="noopener noreferrer"&gt;Bifrost&lt;/a&gt; manages both LLM requests and MCP tool calls, providing a single point of governance. This simplifies infrastructure and provides a holistic view of agent behavior.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;High Performance:&lt;/strong&gt; Adds less than 100 microseconds of overhead at 5,000 requests per second, ensuring that the gateway is never a bottleneck for agentic workflows.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Granular Governance:&lt;/strong&gt; Uses &lt;a href="https://docs.getbifrost.ai/features/governance/virtual-keys" rel="noopener noreferrer"&gt;virtual keys&lt;/a&gt; to enforce per-consumer budgets, rate limits, and access policies. It supports &lt;a href="https://docs.getbifrost.ai/features/governance/mcp-tools" rel="noopener noreferrer"&gt;MCP tool filtering&lt;/a&gt;, allowing administrators to control precisely which tools are available to specific agents or users.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Advanced Security:&lt;/strong&gt; Integrates with enterprise identity providers via OIDC and supports federated authentication for turning existing enterprise APIs into secure MCP tools.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Deployment Flexibility:&lt;/strong&gt; Can be deployed as a single binary, in a Docker container, or in a clustered configuration for high availability within a VPC or on-premise environment.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Code Mode &amp;amp; Agent Mode:&lt;/strong&gt; Features an optimized "&lt;a href="https://docs.getbifrost.ai/mcp/code-mode" rel="noopener noreferrer"&gt;Code Mode&lt;/a&gt;" that can reduce token consumption and latency for complex tool orchestration tasks.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Beyond the gateway, Bifrost's governance and security can be extended to the endpoint. &lt;a href="https://www.getmaxim.ai/bifrost/edge" rel="noopener noreferrer"&gt;Bifrost Edge&lt;/a&gt; ensures that AI traffic from desktop apps and coding agents on employee machines is routed through the gateway, applying the same security policies everywhere and preventing shadow AI.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Docker MCP Gateway
&lt;/h3&gt;

&lt;p&gt;The &lt;a href="https://docs.docker.com/mcp-gateway/" rel="noopener noreferrer"&gt;Docker MCP Gateway&lt;/a&gt; is an open-source solution from Docker that excels at orchestrating MCP servers as containerized workloads. It is a natural fit for development teams already standardized on the Docker ecosystem.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; Development teams and organizations that use Docker for container orchestration and want a secure, isolated environment for running each MCP server.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key Features:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Container-per-Server Isolation:&lt;/strong&gt; Runs each MCP server in its own sandboxed Docker container with restricted network and filesystem access, providing a strong security posture at the infrastructure layer.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Docker Ecosystem Integration:&lt;/strong&gt; Deeply integrated with Docker Desktop and Docker Compose, offering a seamless local development experience.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Credential Management:&lt;/strong&gt; Leverages Docker's built-in secrets management for handling credentials securely.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;OCI-Based Catalog:&lt;/strong&gt; Manages MCP servers as OCI images, enabling versioning and supply-chain security practices.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;While excellent for development and container-native teams, it is less focused on being a unified, multi-user enterprise control plane for both LLM and MCP traffic compared to Bifrost.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Obot
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://www.obot.ai/" rel="noopener noreferrer"&gt;Obot&lt;/a&gt; is an open-source MCP platform that includes a gateway, a server catalog, and agent orchestration tools. It is designed for teams who want to self-host their entire MCP stack and avoid vendor lock-in.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; Organizations looking for a comprehensive, end-to-end open-source MCP platform to run on their own Kubernetes infrastructure.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key Features:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Full Platform:&lt;/strong&gt; Provides more than just a gateway; it's a complete system for managing the lifecycle of MCP servers and agents.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Kubernetes-Native:&lt;/strong&gt; Designed to run on Kubernetes, making it a good fit for teams with existing cloud-native expertise.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Enterprise Identity:&lt;/strong&gt; The enterprise version supports integration with IdPs like Okta and Microsoft Entra for centralized user authentication.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Self-Hosted Control:&lt;/strong&gt; Gives teams complete control over their data, security, and infrastructure without relying on a hosted service.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Obot offers a broader platform scope but requires more operational overhead to manage the full Kubernetes-based deployment compared to a standalone gateway like Bifrost.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5j77v0b9tmdrbv2xok5n.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5j77v0b9tmdrbv2xok5n.png" alt="Three distinct, glowing crystalline structures on a platform. One is a sleek, singular shard representing a unified gate" width="800" height="457"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Recommendation
&lt;/h2&gt;

&lt;p&gt;For the majority of teams seeking a secure, performant, and flexible open-source solution, &lt;strong&gt;Bifrost&lt;/strong&gt; is the leading choice in 2026. Its unique ability to govern both LLM and MCP traffic through a single, low-latency control plane addresses the reality of modern agentic systems. The focus on enterprise-grade governance, comprehensive security features, and the ability to deploy anywhere—from a local machine to an air-gapped cluster—makes it the most versatile and production-ready option.&lt;/p&gt;

&lt;p&gt;While Docker MCP Gateway and Obot are strong contenders in their respective niches (container-native workflows and full-platform self-hosting), Bifrost provides the best balance of performance, security, and operational simplicity for securing AI agent access at scale.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What is an MCP gateway?
&lt;/h3&gt;

&lt;p&gt;An MCP gateway is a centralized control layer that manages how AI agents interact with external tools, APIs, and data sources via the Model Context Protocol (MCP). It acts as a secure intermediary, handling authentication, access control, routing, auditing, and policy enforcement for all tool calls.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why is an open-source MCP gateway important?
&lt;/h3&gt;

&lt;p&gt;Open-source gateways provide transparency, customization, and full control over your infrastructure. For security-conscious organizations, this means data and credentials never have to pass through a third-party service, ensuring data sovereignty and simplifying compliance with regulations like GDPR and HIPAA.&lt;/p&gt;

&lt;h3&gt;
  
  
  How does an MCP gateway improve AI agent security?
&lt;/h3&gt;

&lt;p&gt;It improves security by centralizing control. Instead of agents holding credentials directly, the gateway manages access. It enforces policies on which agents can use which tools, creates immutable audit logs of all activity, and can be integrated with enterprise identity systems to ensure only authorized users and agents can invoke sensitive tools.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can an MCP gateway reduce costs?
&lt;/h3&gt;

&lt;p&gt;Yes, by providing centralized observability and budget controls. A gateway like Bifrost can enforce spending limits per agent, user, or project. By logging every tool and model call, it provides the necessary data to identify and optimize costly agent behaviors.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is the difference between an LLM gateway and an MCP gateway?
&lt;/h3&gt;

&lt;p&gt;An LLM gateway manages requests to large language models (routing, failover, caching, etc.), while an MCP gateway manages requests from AI agents to external tools. Bifrost is unique in that it combines both functions into a single, unified gateway, recognizing that agentic workflows involve a constant mix of both traffic types.&lt;/p&gt;

&lt;h2&gt;
  
  
  Next Steps
&lt;/h2&gt;

&lt;p&gt;Securing agent access is a foundational step in building reliable and enterprise-ready AI applications. Teams evaluating open-source MCP gateways can &lt;a href="https://getmaxim.ai/bifrost/book-a-demo" rel="noopener noreferrer"&gt;request a Bifrost demo&lt;/a&gt; to see how unified governance works in practice or review the &lt;a href="https://github.com/maximhq/bifrost" rel="noopener noreferrer"&gt;open-source repository&lt;/a&gt; on GitHub.&lt;/p&gt;

</description>
      <category>aigateway</category>
      <category>mcp</category>
      <category>llmops</category>
      <category>security</category>
    </item>
    <item>
      <title>Best AI Governance Tools for Responsible AI Usage in 2026</title>
      <dc:creator>Swapnoneel Saha</dc:creator>
      <pubDate>Wed, 30 Sep 2026 10:44:30 +0000</pubDate>
      <link>https://dev.to/swapnoneel123/best-ai-governance-tools-for-responsible-ai-usage-in-2026-1kb8</link>
      <guid>https://dev.to/swapnoneel123/best-ai-governance-tools-for-responsible-ai-usage-in-2026-1kb8</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4khkvjtitewlziszbtv9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4khkvjtitewlziszbtv9.png" alt="Best AI Governance Tools for Responsible AI Usage in 2026" width="800" height="457"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  AI governance tools are now a mandatory component for enterprises, driven by regulations like the EU AI Act and the need to manage risks from shadow AI and autonomous agents.&lt;/li&gt;
&lt;li&gt;  The market is split between two main categories: runtime enforcement platforms that control live AI traffic, and policy management platforms for compliance, risk assessment, and audit.&lt;/li&gt;
&lt;li&gt;  Runtime enforcement tools like &lt;a href="https://www.getmaxim.ai/bifrost" rel="noopener noreferrer"&gt;Bifrost&lt;/a&gt; and Fiddler AI are critical for preventing policy violations in real-time, while platforms like Credo AI and IBM watsonx.governance excel at lifecycle management and documentation.&lt;/li&gt;
&lt;li&gt;  For teams that need a high-performance, open-source, and self-hosted control plane, the Bifrost AI gateway provides the strongest combination of access control, cost management, audit logging, and security.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The adoption of AI is no longer a question for most enterprises; the central challenge has shifted to control. As employees use third-party AI tools and engineering teams deploy autonomous agents, organizations face a growing governance gap. Without a dedicated framework, they risk data leakage, compliance violations, and operational incidents. AI governance platforms are the tools enterprises use to close this gap, providing the structure to enforce policies, manage risk, and ensure AI operates safely and responsibly.&lt;/p&gt;

&lt;p&gt;This has become a strategic necessity, with frameworks like the &lt;a href="https://www.nist.gov/itl/ai-risk-management-framework" rel="noopener noreferrer"&gt;NIST AI Risk Management Framework (AI RMF)&lt;/a&gt; and ISO/IEC 42001 now guiding enterprise strategy. The market for governance tools has matured to address these needs, but it has also split into distinct categories. Choosing the right tool depends on whether the primary goal is documenting policy for compliance or enforcing it on live traffic. This guide compares the top AI governance tools of 2026, with a focus on solutions for responsible enterprise AI usage.&lt;/p&gt;

&lt;h2&gt;
  
  
  Runtime vs. Policy: The Two Halves of AI Governance
&lt;/h2&gt;

&lt;p&gt;AI governance tools generally fall into one of two categories, and many large enterprises use a tool from each.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Policy &amp;amp; Compliance Platforms:&lt;/strong&gt; These are systems of record for an organization's AI governance program. They provide tools for creating an inventory of AI models, conducting risk assessments, mapping controls to regulatory requirements (like the EU AI Act), and generating audit-ready documentation. They help answer the questions: "What AI are we using?" and "Is it compliant with our policies?"&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Runtime Enforcement Platforms:&lt;/strong&gt; These tools sit in the path of live AI requests and enforce governance policies in real time. They act as a control plane for AI traffic, applying rules for access, data security, content safety, and cost before a request reaches a model and before a response reaches a user. They answer the question: "Can we stop a policy violation before it happens?"&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A comprehensive governance strategy needs both. A policy is ineffective without enforcement, and enforcement is arbitrary without a clear policy. This review focuses on the tools that provide the technical controls necessary for a complete governance program.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key Criteria for Evaluating AI Governance Tools
&lt;/h2&gt;

&lt;p&gt;When assessing AI governance platforms, enterprises should look for a core set of capabilities that address the full lifecycle of AI usage.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Criterion&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Access Control &amp;amp; Identity&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Granular control over who (users, teams, services) can access which models, with what budgets, and under which conditions. Support for SSO and RBAC is critical.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Policy Enforcement&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The ability to define and enforce rules at runtime. This includes routing, rate limits, and guardrails for content, data (PII), and security (secrets).&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Audit &amp;amp; Observability&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Immutable, detailed logs of all AI requests, responses, and policy decisions. This is non-negotiable for compliance with SOC 2, HIPAA, or ISO 27001.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Shadow AI &amp;amp; Endpoint Governance&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The ability to discover and govern unmanaged AI usage on employee devices (desktop apps, browser AI) to close compliance gaps.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Deployment &amp;amp; Integration&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Flexibility to deploy in any environment (cloud, VPC, on-prem) and integrate with existing identity providers, observability stacks, and security tools.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Regulatory Alignment&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Features that directly support compliance with major frameworks like the EU AI Act, NIST AI RMF, and ISO 42001.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3o1ol9knopj6hwo3vq9o.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3o1ol9knopj6hwo3vq9o.png" alt="A visual metaphor for evaluation criteria, showing a set of interconnected, glowing gears or mechanisms, each representi" width="800" height="457"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Best AI Governance Tools of 2026 at a Glance
&lt;/h2&gt;

&lt;p&gt;This table compares the leading AI governance tools across the key evaluation criteria.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Primary Focus&lt;/th&gt;
&lt;th&gt;Runtime Enforcement&lt;/th&gt;
&lt;th&gt;Policy Management&lt;/th&gt;
&lt;th&gt;Shadow AI Governance&lt;/th&gt;
&lt;th&gt;Deployment Model&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Bifrost&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Runtime Enforcement&lt;/td&gt;
&lt;td&gt;✅ Yes&lt;/td&gt;
&lt;td&gt;➖ No&lt;/td&gt;
&lt;td&gt;✅ Yes&lt;/td&gt;
&lt;td&gt;Self-Hosted (Cloud/VPC/On-Prem)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Credo AI&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Policy &amp;amp; Compliance&lt;/td&gt;
&lt;td&gt;➖ No&lt;/td&gt;
&lt;td&gt;✅ Yes&lt;/td&gt;
&lt;td&gt;➖ No&lt;/td&gt;
&lt;td&gt;SaaS&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;IBM watsonx.governance&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Model Lifecycle&lt;/td&gt;
&lt;td&gt;✅ Yes&lt;/td&gt;
&lt;td&gt;✅ Yes&lt;/td&gt;
&lt;td&gt;➖ No&lt;/td&gt;
&lt;td&gt;SaaS (IBM Cloud) / Hybrid&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;OneTrust AI Governance&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Privacy &amp;amp; GRC&lt;/td&gt;
&lt;td&gt;➖ No&lt;/td&gt;
&lt;td&gt;✅ Yes&lt;/td&gt;
&lt;td&gt;➖ No&lt;/td&gt;
&lt;td&gt;SaaS&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Microsoft Purview&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Data Governance&lt;/td&gt;
&lt;td&gt;✅ Yes&lt;/td&gt;
&lt;td&gt;✅ Yes&lt;/td&gt;
&lt;td&gt;➖ No&lt;/td&gt;
&lt;td&gt;SaaS (Azure)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Fiddler AI&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;AI Observability &amp;amp; Security&lt;/td&gt;
&lt;td&gt;✅ Yes&lt;/td&gt;
&lt;td&gt;➖ No&lt;/td&gt;
&lt;td&gt;➖ No&lt;/td&gt;
&lt;td&gt;SaaS / VPC&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  A Deeper Look at the Top AI Governance Platforms
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Bifrost
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://www.getmaxim.ai/bifrost" rel="noopener noreferrer"&gt;Bifrost&lt;/a&gt; is a high-performance, &lt;a href="https://github.com/maximhq/bifrost" rel="noopener noreferrer"&gt;open-source AI gateway&lt;/a&gt; from Maxim AI that acts as a centralized runtime enforcement point for all AI traffic. It unifies access to over 1000 models through a single API and applies governance policies to every request. Its position in the network allows it to enforce controls that policy-only platforms cannot.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; Enterprises that need a high-performance, self-hosted, and extensible platform for runtime AI governance, security, and cost control.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key Capabilities:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Unified Access &amp;amp; Cost Control:&lt;/strong&gt; Bifrost uses &lt;a href="https://docs.getbifrost.ai/features/governance/virtual-keys" rel="noopener noreferrer"&gt;virtual keys&lt;/a&gt; to manage access. Teams and services get unique keys with their own budgets, rate limits, and model permissions, providing granular control over spending and usage.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Runtime Policy Enforcement:&lt;/strong&gt; As a gateway, Bifrost enforces &lt;a href="https://docs.getbifrost.ai/features/governance/routing" rel="noopener noreferrer"&gt;routing rules&lt;/a&gt;, provider failover, and load balancing on every request, ensuring reliability and compliance with architectural policies.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Security and Guardrails:&lt;/strong&gt; The platform integrates with security tools like AWS Bedrock Guardrails and Azure Content Safety to apply content moderation, PII redaction, and prompt injection defense at the gateway layer.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Immutable Audit Logs:&lt;/strong&gt; Bifrost generates comprehensive, signed &lt;a href="https://docs.getbifrost.ai/enterprise/audit-logs" rel="noopener noreferrer"&gt;audit logs&lt;/a&gt; for every transaction, providing the evidence required for SOC 2, HIPAA, and ISO 27001 compliance.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Endpoint Governance with Bifrost Edge:&lt;/strong&gt; Beyond the gateway, &lt;a href="https://www.getmaxim.ai/bifrost/edge" rel="noopener noreferrer"&gt;Bifrost Edge&lt;/a&gt; extends the same governance and security controls to AI traffic on employee machines. It provides visibility into and control over desktop apps and browser-based AI, addressing the "shadow AI" problem directly with &lt;a href="https://docs.getbifrost.ai/edge/security" rel="noopener noreferrer"&gt;endpoint enforcement&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  2. Credo AI
&lt;/h3&gt;

&lt;p&gt;Credo AI is a leading AI governance platform focused on policy, risk, and compliance management. It serves as a central registry for an organization's AI systems, helping teams document use cases, assess risks against established frameworks, and generate compliance reports.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; GRC (Governance, Risk, and Compliance) teams and Chief AI Officers who need a system of record for managing AI policies and demonstrating regulatory compliance.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key Capabilities:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;AI Registry:&lt;/strong&gt; Provides a centralized inventory to catalog all AI models, applications, and use cases across the enterprise.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Risk Assessment:&lt;/strong&gt; Offers workflows to evaluate AI systems against fairness, performance, transparency, and security risks.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Policy Packs:&lt;/strong&gt; Translates regulations and standards like the EU AI Act and NIST AI RMF into actionable assessment requirements and report templates.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  3. IBM watsonx.governance
&lt;/h3&gt;

&lt;p&gt;IBM watsonx.governance is an enterprise platform designed to govern the entire AI model lifecycle. It provides tools for tracking models from development to deployment, monitoring for bias and drift, and ensuring transparency and explainability.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; Large enterprises, particularly those in regulated industries like finance and healthcare, that need to manage the lifecycle of a diverse portfolio of AI models.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key Capabilities:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Model Lifecycle Management:&lt;/strong&gt; Tracks model development, validation, deployment, and performance in a central inventory.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Bias and Drift Detection:&lt;/strong&gt; Continuously monitors production models for fairness, bias, and performance degradation.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Explainability:&lt;/strong&gt; Generates explanations for model predictions, helping teams understand and document how AI systems make decisions.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  4. OneTrust AI Governance
&lt;/h3&gt;

&lt;p&gt;OneTrust is a well-established platform in the privacy and trust management space, and its AI Governance solution extends these capabilities to AI systems. It helps organizations inventory AI use cases, assess them against privacy and ethical risks, and connect AI governance to broader compliance programs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; Organizations that already use OneTrust for privacy management and want to integrate AI governance into their existing GRC workflows.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key Capabilities:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;AI Inventory and Discovery:&lt;/strong&gt; Helps organizations create a comprehensive inventory of their AI projects and understand where AI is being used with their data.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Risk Assessment:&lt;/strong&gt; Provides templates and workflows for assessing AI projects against privacy regulations and ethical guidelines.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Regulatory Intelligence:&lt;/strong&gt; Keeps teams updated on evolving AI laws and helps automate compliance workflows.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  5. Microsoft Purview
&lt;/h3&gt;

&lt;p&gt;For organizations heavily invested in the Microsoft ecosystem, Microsoft Purview provides unified data and AI governance. It extends Microsoft's data classification, lineage, and security capabilities to cover AI workloads running in Azure and Microsoft 365.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; Enterprises using Azure Machine Learning, Azure OpenAI, and Microsoft Copilot that need integrated data security and governance.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key Capabilities:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Unified Data Governance:&lt;/strong&gt; Combines data discovery, classification, and policy enforcement across an organization's data estate.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;AI-Specific Controls:&lt;/strong&gt; Applies data security and compliance policies to AI interactions, including retaining prompts and responses for audit purposes.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Deep Microsoft Integration:&lt;/strong&gt; Offers native policy enforcement and metadata synchronization for AI services within the Microsoft ecosystem.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  6. Fiddler AI
&lt;/h3&gt;

&lt;p&gt;Fiddler AI is an AI observability platform with strong capabilities for runtime governance and security. Its Fiddler Guardrails product provides real-time content moderation to detect and block harmful content, PII leaks, and hallucinations before they reach users.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; ML engineering and security teams that need a real-time defense layer for their LLM applications.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key Capabilities:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Real-Time Guardrails:&lt;/strong&gt; Detects and blocks safety risks, PII, and leaked secrets in prompts and responses with low latency.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;AI Observability:&lt;/strong&gt; Provides deep monitoring of model performance, data drift, and explainability for production AI systems.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Hallucination Detection:&lt;/strong&gt; Includes specific checks to identify and flag unfaithful or fabricated responses in RAG applications.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frxiurr6kasmy4yxbvbja.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frxiurr6kasmy4yxbvbja.png" alt="A fleet of various aircraft (representing different AI tools) flying in a precise, orderly formation under the guidance " width="800" height="457"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Recommendation: Choosing the Right Tool for the Job
&lt;/h2&gt;

&lt;p&gt;The best AI governance tool depends entirely on the problem an organization is trying to solve.&lt;/p&gt;

&lt;p&gt;For teams whose primary need is to document policies, create an AI inventory, and generate compliance reports for auditors, a &lt;strong&gt;policy and compliance platform&lt;/strong&gt; like &lt;strong&gt;Credo AI&lt;/strong&gt; or &lt;strong&gt;OneTrust&lt;/strong&gt; is the logical starting point. These tools provide the necessary framework for establishing a formal governance program.&lt;/p&gt;

&lt;p&gt;For teams whose primary need is to &lt;strong&gt;enforce&lt;/strong&gt; those policies on live AI traffic, prevent data leaks, control costs, and create an immutable audit trail, a &lt;strong&gt;runtime enforcement platform&lt;/strong&gt; is essential. In this category, &lt;strong&gt;Bifrost&lt;/strong&gt; stands out as the most powerful and flexible option. Its open-source foundation, high-performance architecture, and comprehensive feature set—from virtual keys for cost control to Bifrost Edge for endpoint governance—provide a complete control plane for enterprise AI. While other tools offer components of runtime governance, Bifrost unifies them in a single, self-hostable gateway that gives organizations full control over their AI infrastructure.&lt;/p&gt;

&lt;p&gt;Ultimately, a mature enterprise AI program needs both layers: a system of record for policy and a system of enforcement for control. Teams evaluating AI governance solutions can &lt;a href="https://getmaxim.ai/bifrost/book-a-demo" rel="noopener noreferrer"&gt;request a Bifrost demo&lt;/a&gt; or review the &lt;a href="https://github.com/maximhq/bifrost" rel="noopener noreferrer"&gt;open-source repository&lt;/a&gt; to see how a runtime gateway can serve as the foundation for a responsible AI strategy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What is AI governance?
&lt;/h3&gt;

&lt;p&gt;AI governance is the framework of policies, processes, and controls an organization uses to ensure its AI systems operate ethically, securely, and in compliance with laws and regulations. It covers the entire AI lifecycle, from development to deployment and monitoring.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why is AI governance important in 2026?
&lt;/h3&gt;

&lt;p&gt;With the rapid adoption of powerful AI and the introduction of binding regulations like the EU AI Act, governance is no longer optional. It is a requirement for managing significant financial, reputational, and operational risks, and for building trust with customers and regulators.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is the difference between AI governance and AI security?
&lt;/h3&gt;

&lt;p&gt;AI governance is a broad discipline focused on oversight, policy, and risk management. AI security is a subset of governance that focuses specifically on protecting AI systems from threats like data leakage, model theft, and prompt injection attacks. Effective governance tools often include strong security features.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do you govern "shadow AI"?
&lt;/h3&gt;

&lt;p&gt;Shadow AI—the unapproved use of AI tools by employees—is a major governance challenge. The most effective way to govern it is with an endpoint agent, like Bifrost Edge, that can discover, monitor, and apply policies to AI applications running on employee devices, regardless of how they are connected.&lt;/p&gt;

&lt;h3&gt;
  
  
  What are the most important features of an AI governance tool?
&lt;/h3&gt;

&lt;p&gt;The most critical features include granular access control (who can use what), real-time policy enforcement (guardrails for content and data), comprehensive audit logging, and the flexibility to deploy in your own environment to maintain data control.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;a href="https://www.nist.gov/itl/ai-risk-management-framework" rel="noopener noreferrer"&gt;NIST AI Risk Management Framework (AI RMF 1.0)&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;  &lt;a href="https://artificialintelligenceact.eu/" rel="noopener noreferrer"&gt;The EU AI Act&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;  &lt;a href="https://www.ibm.com/topics/ai-governance" rel="noopener noreferrer"&gt;IBM, "What is AI Governance?"&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;  &lt;a href="https://www.gartner.com/reviews/market/ai-governance-platforms/vendor/onetrust/product/onetrust-ai-governance" rel="noopener noreferrer"&gt;Gartner Peer Insights, "OneTrust AI Governance"&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;  &lt;a href="https://docs.getbifrost.ai/overview" rel="noopener noreferrer"&gt;Bifrost AI Gateway Documentation&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>aigovernance</category>
      <category>responsibleai</category>
      <category>enterprisearchitecture</category>
      <category>security</category>
    </item>
    <item>
      <title>Top 5 Enterprise AI Gateways for Multi-Model Routing in 2026</title>
      <dc:creator>Swapnoneel Saha</dc:creator>
      <pubDate>Wed, 30 Sep 2026 10:40:22 +0000</pubDate>
      <link>https://dev.to/swapnoneel123/top-5-enterprise-ai-gateways-for-multi-model-routing-in-2026-5bae</link>
      <guid>https://dev.to/swapnoneel123/top-5-enterprise-ai-gateways-for-multi-model-routing-in-2026-5bae</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F06u24cfdvh9kuprp3tmf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F06u24cfdvh9kuprp3tmf.png" alt="Top 5 Enterprise AI Gateways for Multi-Model Routing in 2026" width="800" height="457"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  An AI gateway is essential infrastructure for production AI, providing a unified control plane for routing, failover, governance, and observability across multiple LLM providers.&lt;/li&gt;
&lt;li&gt;  Multi-model routing optimizes for cost, latency, and capability by sending each request to the most appropriate model, rather than relying on a single, expensive one.&lt;/li&gt;
&lt;li&gt;  Bifrost is a top choice for enterprises needing high performance and comprehensive governance in a self-hosted or in-VPC deployment.&lt;/li&gt;
&lt;li&gt;  LiteLLM offers the broadest provider support, making it ideal for teams prioritizing model variety in a self-hosted, open-source environment.&lt;/li&gt;
&lt;li&gt;  Kong, Cloudflare, and Vercel are strong contenders for teams already invested in their respective ecosystems, offering integrated AI routing at the API or edge layer.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;As enterprise AI moves from single-model prototypes to production systems, engineering teams face a common set of challenges: provider outages, unpredictable costs, and the operational complexity of managing dozens of models and API keys. A single-model strategy is no longer viable; no one model is the cheapest, fastest, and most capable for every task. This reality has made the AI gateway a critical piece of infrastructure.&lt;/p&gt;

&lt;p&gt;An AI gateway, or LLM gateway, is a dedicated layer that sits between applications and AI providers. It unifies access to thousands of models through a single, consistent API and centralizes the logic for multi-model routing. Instead of embedding provider-specific SDKs and routing rules in every service, teams point their applications at the gateway and manage policy from one place. This article compares the top five enterprise AI gateways for 2026, focusing on their capabilities for multi-model routing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key Criteria for Evaluating Enterprise AI Gateways
&lt;/h2&gt;

&lt;p&gt;Before comparing platforms, it's important to establish the criteria that matter for enterprise use cases. Beyond a simple unified API, a true enterprise gateway must provide robust controls across several key areas.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Capability&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;th&gt;Why It Matters for Enterprises&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Routing Strategy&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The logic used to select a model for each request. Common strategies include failover, load balancing, cost-based, latency-based, and semantic routing.&lt;/td&gt;
&lt;td&gt;Determines an organization's ability to optimize for cost, user experience, and reliability without application code changes.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Deployment Model&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;How the gateway is hosted: fully managed SaaS, self-hosted in a public cloud, in a VPC, on-premise, or in an air-gapped environment.&lt;/td&gt;
&lt;td&gt;Enterprises in regulated industries often have strict data residency and security requirements that mandate self-hosted or in-VPC deployments.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Performance&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The latency overhead the gateway adds to each request, especially under load.&lt;/td&gt;
&lt;td&gt;For real-time, user-facing AI applications, gateway latency can be a significant bottleneck. Low-overhead gateways are critical for production performance.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Governance&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Features for managing access, cost, and compliance, such as virtual keys, per-user budgets, RBAC, and audit logs.&lt;/td&gt;
&lt;td&gt;Provides the financial and security controls necessary to operate AI safely at scale, preventing budget overruns and unauthorized access.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Observability&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;The ability to log, trace, and monitor every request with detailed information on token usage, cost, latency, and provider choice.&lt;/td&gt;
&lt;td&gt;Essential for debugging, performance tuning, and providing finance and compliance teams with a clear view of AI usage and spend.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;MCP Support&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Native support for the Model Context Protocol (MCP) to manage and govern connections to external tools and agents.&lt;/td&gt;
&lt;td&gt;As AI becomes more agentic, governing tool use becomes as important as governing model use. A gateway with MCP support provides a unified control plane.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F46kqbygm51kip8jvo83h.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F46kqbygm51kip8jvo83h.png" alt="A blueprint of a fortress-like data center, with visible layers for security, routing logic, and performance, representi" width="800" height="457"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Top 5 Platforms Compared
&lt;/h2&gt;

&lt;p&gt;Based on these criteria, here is a detailed comparison of the leading enterprise AI gateways.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Bifrost
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://www.getmaxim.ai/bifrost" rel="noopener noreferrer"&gt;Bifrost&lt;/a&gt; is a high-performance, &lt;a href="https://github.com/maximhq/bifrost" rel="noopener noreferrer"&gt;open-source AI gateway&lt;/a&gt; from Maxim AI, written in Go. It is designed for enterprise-grade workloads that demand low latency and comprehensive governance. Bifrost unifies access to over 1,000 models from more than 25 providers through a single OpenAI-compatible API.&lt;/p&gt;

&lt;p&gt;Its primary differentiator is performance; published benchmarks show it adds only 11 microseconds of overhead per request at a sustained load of 5,000 requests per second. This makes it suitable for latency-sensitive applications like real-time copilots and conversational agents.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Routing Capabilities:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Automatic Failover:&lt;/strong&gt; Bifrost provides robust fallback chains. When a request to a primary provider fails, it automatically retries with exponential backoff before failing over to the next provider or model in a pre-configured sequence.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Load Balancing:&lt;/strong&gt; It can perform weighted load balancing across multiple API keys for the same provider or across different providers to distribute traffic and avoid rate limits.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Routing Rules:&lt;/strong&gt; Administrators can define rules to direct specific types of requests to certain models or providers based on request metadata.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Semantic Caching:&lt;/strong&gt; Reduces costs and improves latency by caching responses based on the semantic meaning of a query, not just an exact string match.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; Enterprises that require a self-hosted, high-performance gateway with a full suite of governance and security features. Its ability to be deployed in a VPC or an air-gapped environment makes it a strong choice for organizations in finance, healthcare, and other regulated industries.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. LiteLLM
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://www.litellm.ai/" rel="noopener noreferrer"&gt;LiteLLM&lt;/a&gt; is a popular and mature open-source AI gateway known for its extensive provider support, covering over 140 providers and 1,800+ models. Built in Python, it has a large community and is often the first gateway to support new and emerging models.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Routing Capabilities:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Flexible Strategies:&lt;/strong&gt; LiteLLM offers multiple routing strategies out of the box, including &lt;code&gt;simple-shuffle&lt;/code&gt;, &lt;code&gt;least-busy&lt;/code&gt;, &lt;code&gt;latency-based&lt;/code&gt;, and &lt;code&gt;cost-based&lt;/code&gt; routing.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Auto Routing:&lt;/strong&gt; A beta feature allows for more intelligent routing by classifying request complexity and routing to different tiers of models accordingly.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Granular Control:&lt;/strong&gt; Routing strategies, fallbacks, and timeouts can be configured on a per-key or per-team basis, allowing for highly customized behavior across an organization.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; Teams that prioritize the breadth of model compatibility and want a flexible, open-source solution they can self-host and customize. Its Python base makes it highly accessible for many development teams.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Kong AI Gateway
&lt;/h3&gt;

&lt;p&gt;The &lt;a href="https://konghq.com/products/kong-ai-gateway" rel="noopener noreferrer"&gt;Kong AI Gateway&lt;/a&gt; extends the widely adopted Kong API Gateway with AI-specific capabilities. For organizations already using Kong to manage their microservices and APIs, this provides a familiar control plane for also managing LLM traffic.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Routing Capabilities:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Model-Based Routing:&lt;/strong&gt; Kong can inspect incoming prompts and route them to different upstream providers based on complexity. A common pattern is routing simple requests to a fast, cheap model and complex reasoning tasks to a more powerful, expensive one.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Plugin Architecture:&lt;/strong&gt; AI features are delivered via plugins, such as &lt;code&gt;ai-proxy-advanced&lt;/code&gt;, which handle credential injection, format translation, and routing logic.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Unified Observability:&lt;/strong&gt; AI traffic metrics are integrated into Kong Konnect's analytics dashboards, alongside standard API traffic.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; Enterprises already invested in the Kong ecosystem. It allows platform teams to apply their existing API management policies (authentication, rate limiting, logging) to AI workloads with minimal friction.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Cloudflare AI Gateway
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://www.cloudflare.com/ai-gateway/" rel="noopener noreferrer"&gt;Cloudflare's AI Gateway&lt;/a&gt; is a managed service that leverages Cloudflare's global edge network to provide observability, caching, and routing for AI applications. It's designed for ease of use and integrates tightly with other Cloudflare products like Workers AI.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Routing Capabilities:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Dynamic Routing:&lt;/strong&gt; Allows configuration of rules to automatically failover between providers when one is unavailable. If a model is available on multiple providers, the gateway can route around an outage without manual intervention.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Model-First Routing:&lt;/strong&gt; Cloudflare is moving towards a model where developers request a specific model (e.g., Llama 3), and the gateway decides the best provider (including Cloudflare's own Workers AI) to serve it from based on availability and capacity.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Caching at the Edge:&lt;/strong&gt; Caching responses directly on Cloudflare's edge network can significantly reduce latency for users globally.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; Teams building applications on the Cloudflare stack or those who prefer a fully managed, zero-ops solution. Its strength lies in its global performance and seamless integration with Workers.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Vercel AI Gateway
&lt;/h3&gt;

&lt;p&gt;Similar to Cloudflare, the &lt;a href="https://vercel.com/ai/gateway" rel="noopener noreferrer"&gt;Vercel AI Gateway&lt;/a&gt; is a managed service tightly integrated into the Vercel ecosystem, primarily serving applications built with Next.js and the Vercel AI SDK. It focuses on providing developers with a simple way to access hundreds of models without managing credentials or complex routing logic.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Routing Capabilities:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Automatic Failover:&lt;/strong&gt; The gateway can be configured with ordered fallbacks, ensuring that if a request to a primary model fails, it transparently retries with an alternative.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Cost-Aware Routing:&lt;/strong&gt; Vercel provides patterns for developers to build their own classifiers that route simple, low-value requests to cheap models and escalate complex ones to more capable frontier models, all through a single gateway endpoint.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Firewall-Style Rules:&lt;/strong&gt; Administrators can set rules to restrict which models the team is allowed to use, enforcing governance at the gateway level.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; Development teams building on Vercel and Next.js. The tight integration with the Vercel AI SDK and platform provides a streamlined developer experience for building AI-powered frontends.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdyqkrdnbzds8ql77ui9d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdyqkrdnbzds8ql77ui9d.png" alt="A side-by-side comparison of five different distinct architectural models on a workbench, each with its own unique desig" width="800" height="457"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Comparison at a Glance
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Feature&lt;/th&gt;
&lt;th&gt;Bifrost&lt;/th&gt;
&lt;th&gt;LiteLLM&lt;/th&gt;
&lt;th&gt;Kong AI Gateway&lt;/th&gt;
&lt;th&gt;Cloudflare AI Gateway&lt;/th&gt;
&lt;th&gt;Vercel AI Gateway&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Deployment&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Self-hosted, In-VPC, Air-gapped&lt;/td&gt;
&lt;td&gt;Self-hosted&lt;/td&gt;
&lt;td&gt;Self-hosted, Managed (Konnect)&lt;/td&gt;
&lt;td&gt;Managed&lt;/td&gt;
&lt;td&gt;Managed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Core Tech&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Go&lt;/td&gt;
&lt;td&gt;Python&lt;/td&gt;
&lt;td&gt;NGINX (Lua)&lt;/td&gt;
&lt;td&gt;Managed Service&lt;/td&gt;
&lt;td&gt;Managed Service&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Performance&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Very High (&amp;lt;15 µs overhead)&lt;/td&gt;
&lt;td&gt;Moderate&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;High (Edge Network)&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Routing&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Failover, Weighted LB, Rules&lt;/td&gt;
&lt;td&gt;Cost, Latency, Usage-based&lt;/td&gt;
&lt;td&gt;Prompt-aware, Plugin-driven&lt;/td&gt;
&lt;td&gt;Dynamic Failover&lt;/td&gt;
&lt;td&gt;Failover, Cost-aware&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Governance&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Virtual Keys, Budgets, RBAC, Audit Logs&lt;/td&gt;
&lt;td&gt;Virtual Keys, Budgets, Teams&lt;/td&gt;
&lt;td&gt;Consumer Auth, RBAC&lt;/td&gt;
&lt;td&gt;Rate Limiting, Logs&lt;/td&gt;
&lt;td&gt;Budgets, Allow-lists&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;MCP Support&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes (Native)&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes (Enterprise)&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Open Source&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes (Apache 2.0)&lt;/td&gt;
&lt;td&gt;Yes (MIT)&lt;/td&gt;
&lt;td&gt;Yes (Core Gateway)&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Choosing an enterprise AI gateway is a foundational architectural decision. The right choice depends less on a raw feature list and more on your organization's specific needs for deployment, performance, and governance.&lt;/p&gt;

&lt;p&gt;For enterprises requiring maximum control, performance, and security within their own infrastructure, a self-hosted solution like &lt;strong&gt;Bifrost&lt;/strong&gt; provides the most comprehensive feature set, from granular governance to native MCP support. For teams prioritizing the widest possible model support in an open-source package, &lt;strong&gt;LiteLLM&lt;/strong&gt; is a strong and flexible choice. For organizations already committed to an existing platform, the AI gateways from &lt;strong&gt;Kong&lt;/strong&gt;, &lt;strong&gt;Cloudflare&lt;/strong&gt;, and &lt;strong&gt;Vercel&lt;/strong&gt; offer a compelling path to unified AI routing and observability with lower operational overhead.&lt;/p&gt;

&lt;p&gt;Ultimately, the goal of a multi-model routing strategy is to build more resilient, cost-effective, and capable AI applications. The gateways listed here provide the essential infrastructure to achieve that goal at enterprise scale.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;a href="https://www.linkedin.com/pulse/model-routing-economics-why-enterprises-quietly-building-vamsi-gunturu-n44xc/" rel="noopener noreferrer"&gt;Model Routing Economics: Why Enterprises Are Quietly Building Multi-LLM Gateways&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;  &lt;a href="https://docs.getbifrost.ai/enterprise/overview" rel="noopener noreferrer"&gt;Bifrost Enterprise AI Gateway Documentation&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;  &lt;a href="https://docs.litellm.ai/docs/routing" rel="noopener noreferrer"&gt;LiteLLM Routing Documentation&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;  &lt;a href="https://developers.cloudflare.com/ai-gateway/" rel="noopener noreferrer"&gt;Cloudflare AI Gateway Documentation&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;  &lt;a href="https://docs.konghq.com/ai-gateway/latest/" rel="noopener noreferrer"&gt;Kong AI Gateway Documentation&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>gateway</category>
      <category>llm</category>
      <category>enterprise</category>
    </item>
    <item>
      <title>Top Enterprise LLM Observability Tools in 2026</title>
      <dc:creator>Swapnoneel Saha</dc:creator>
      <pubDate>Wed, 30 Sep 2026 10:35:33 +0000</pubDate>
      <link>https://dev.to/swapnoneel123/top-enterprise-llm-observability-tools-in-2026-dmp</link>
      <guid>https://dev.to/swapnoneel123/top-enterprise-llm-observability-tools-in-2026-dmp</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0sptlpqpd78ss0noelvm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0sptlpqpd78ss0noelvm.png" alt="Top Enterprise LLM Observability Tools in 2026" width="800" height="457"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  LLM observability platforms are essential for debugging, evaluating, and monitoring production AI applications, which fail in ways traditional APM tools can't detect.&lt;/li&gt;
&lt;li&gt;  The market has three main approaches: traditional APM vendors adding LLM features, AI-native tracing tools, and AI gateways that provide observability at the infrastructure layer.&lt;/li&gt;
&lt;li&gt;  An AI gateway is the most effective point for enterprise observability, as it captures telemetry from every model, agent, and application without requiring code-level instrumentation.&lt;/li&gt;
&lt;li&gt;  Bifrost, an open-source AI gateway, provides comprehensive, low-overhead observability by default, emitting LLM-native metrics, traces, and logs compatible with existing enterprise monitoring stacks.&lt;/li&gt;
&lt;li&gt;  When evaluating tools, key criteria include data ownership, OpenTelemetry support, performance overhead, multi-agent tracing, and integration with governance and security controls.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The use of large language models (LLMs) in production is no longer a question of &lt;em&gt;if&lt;/em&gt; but &lt;em&gt;how&lt;/em&gt;. As enterprises deploy AI agents and copilots, they face a new class of operational challenges. Models hallucinate, agentic workflows get stuck in loops, and a single bad prompt can cause costs to skyrocket. Traditional application performance monitoring (APM) tools, built for a world of deterministic software, are blind to these failures. An API call returning a 200 status code says nothing about whether the response was accurate, safe, or helpful.&lt;/p&gt;

&lt;p&gt;This is the gap LLM observability tools fill. They provide the deep visibility needed to trace, debug, and evaluate the non-deterministic behavior of AI systems. For enterprises, where reliability, security, and cost control are non-negotiable, choosing the right observability strategy is critical. This guide compares the leading enterprise LLM observability tools in 2026, with a focus on where and how telemetry is captured. While several approaches exist, capturing observability signals at the AI gateway layer offers the most comprehensive and scalable solution. &lt;a href="https://www.getmaxim.ai/bifrost" rel="noopener noreferrer"&gt;Bifrost&lt;/a&gt;, an &lt;a href="https://github.com/maximhq/bifrost" rel="noopener noreferrer"&gt;open-source AI gateway&lt;/a&gt; from Maxim AI, exemplifies this approach, providing rich, contextual telemetry for every AI request that passes through it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the Gateway is the Right Place for Observability
&lt;/h2&gt;

&lt;p&gt;LLM observability can be implemented at multiple levels: SDKs, application code, or a centralized gateway. For an enterprise, the gateway is the superior choice for three reasons:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt; &lt;strong&gt;Universal Coverage:&lt;/strong&gt; A gateway sees every request from every application, team, and agent. Instrumenting the gateway once provides visibility into all AI traffic, including "shadow AI" usage from uninstrumented coding agents or desktop apps. This eliminates the need to add and maintain observability code in every single application.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Standardized Telemetry:&lt;/strong&gt; The gateway emits a consistent set of metrics, traces, and logs regardless of the originating application or the destination model provider. This creates a unified source of truth for cost attribution, performance analysis, and error tracking across the entire organization.&lt;/li&gt;
&lt;li&gt; &lt;strong&gt;Zero Application Overhead:&lt;/strong&gt; Observability at the gateway has a negligible impact on application performance. High-performance gateways like Bifrost add mere microseconds of latency, ensuring that monitoring doesn't slow down production systems.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcsr53nx2diqreoni2fgf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcsr53nx2diqreoni2fgf.png" alt="A visual metaphor of a secure, transparent control tower overseeing intersecting pathways of light. The tower itself is " width="800" height="457"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Key Criteria for Evaluating Enterprise LLM Observability Tools
&lt;/h2&gt;

&lt;p&gt;When assessing solutions, enterprises should look beyond dashboards and focus on foundational capabilities:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Deployment &amp;amp; Data Control:&lt;/strong&gt; Can the tool be self-hosted in a VPC or on-premises for maximum data control and to meet compliance requirements?&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Performance Overhead:&lt;/strong&gt; What is the latency impact on production AI requests? Solutions should have minimal, predictable overhead under load.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Open Standards Support:&lt;/strong&gt; Does the platform natively support OpenTelemetry for metrics and traces? This avoids vendor lock-in and ensures compatibility with existing stacks like Prometheus, Grafana, and Datadog.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Agent &amp;amp; Tool (MCP) Tracing:&lt;/strong&gt; Can the tool trace multi-step agent workflows, including tool calls made via the Model Context Protocol (MCP)? A simple prompt-response log is insufficient for modern agents.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Integration with Governance:&lt;/strong&gt; Does the observability data connect directly to governance controls like virtual keys, budgets, rate limits, and security guardrails? Visibility without control is an incomplete solution.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Top LLM Observability Tools for Enterprises in 2026
&lt;/h2&gt;

&lt;p&gt;The market offers several strong contenders, each with a different architectural philosophy.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Bifrost: Gateway-Native Observability
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://www.getmaxim.ai/bifrost" rel="noopener noreferrer"&gt;Bifrost&lt;/a&gt; is a high-performance, open-source AI gateway that provides deep observability as a built-in, first-class feature. Because it sits at the intersection of all AI traffic, it is the natural control point for capturing comprehensive telemetry.&lt;/p&gt;

&lt;p&gt;Bifrost's approach is unique in that observability is not an add-on; it's part of the core infrastructure. It automatically captures detailed metadata for every request and response, including tokens, costs, latency, and provider details, all with negligible performance impact.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key Capabilities:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Built-in Telemetry:&lt;/strong&gt; Bifrost emits native &lt;a href="https://docs.getbifrost.ai/features/observability/prometheus" rel="noopener noreferrer"&gt;Prometheus metrics&lt;/a&gt; and distributed traces via &lt;a href="https://docs.getbifrost.ai/features/observability/otel" rel="noopener noreferrer"&gt;OpenTelemetry (OTLP)&lt;/a&gt;, making it compatible with virtually any modern observability backend, including Grafana, Jaeger, and Datadog.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;AI-Native Signals:&lt;/strong&gt; It generates telemetry that standard tools miss, such as per-request token counts, cost data, provider fallback events, and MCP tool-call spans for agentic workflows.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Asynchronous Logging:&lt;/strong&gt; A powerful logging plugin captures full request/response payloads asynchronously, ensuring that detailed tracing has zero impact on request latency.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Unified Governance:&lt;/strong&gt; Observability is tied directly to Bifrost's governance model. All telemetry is tagged with the corresponding virtual key, allowing for precise cost attribution and usage monitoring by team, project, or user.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; Enterprises that require a scalable, secure, and high-performance solution for unifying observability and governance at the infrastructure layer. Its ability to be self-hosted and its open-standards support make it ideal for regulated industries and organizations with existing monitoring stacks.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Datadog LLM Observability
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://www.datadoghq.com/product/llm-observability/" rel="noopener noreferrer"&gt;Datadog&lt;/a&gt; has extended its market-leading APM platform to include LLM-specific observability. For organizations already invested in the Datadog ecosystem, this provides a single pane of glass for monitoring infrastructure, applications, and AI models.&lt;/p&gt;

&lt;p&gt;Datadog excels at correlating LLM performance with underlying infrastructure metrics. It traces LLM calls from within application code using its libraries and provides dashboards for tracking tokens, latency, and errors. Bifrost also features a native &lt;a href="https://docs.getbifrost.ai/enterprise/datadog-connector" rel="noopener noreferrer"&gt;Datadog connector&lt;/a&gt;, allowing teams to combine the benefits of gateway-level capture with Datadog's analysis and visualization tools.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; Companies already standardized on Datadog for their APM and infrastructure monitoring who want to add LLM visibility within the same platform.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Langfuse
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://langfuse.com/" rel="noopener noreferrer"&gt;Langfuse&lt;/a&gt; is an open-source LLM engineering platform that combines tracing, prompt management, and evaluation capabilities. It is developer-centric and provides detailed, session-based replays that are useful for debugging complex agent conversations.&lt;/p&gt;

&lt;p&gt;Langfuse requires instrumenting application code with its SDK to capture traces. It offers a clean UI for exploring traces and allows teams to create datasets for fine-tuning or evaluation from production data.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; Development teams looking for an open-source, all-in-one platform to debug, evaluate, and manage prompts during the development lifecycle.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Arize AI (Phoenix)
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://www.arize.com/" rel="noopener noreferrer"&gt;Arize AI&lt;/a&gt; focuses on ML monitoring and has a strong open-source offering called Phoenix. Phoenix is particularly effective at detecting model drift and evaluating the quality of RAG (Retrieval-Augmented Generation) pipelines. It provides tools to analyze embeddings and visualize how retrieval quality impacts final responses.&lt;/p&gt;

&lt;p&gt;Like Langfuse, Phoenix generally relies on in-app instrumentation to collect data. Its strength lies in post-production analysis and evaluation rather than real-time gateway control.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; ML engineering teams that need to diagnose and troubleshoot complex RAG systems and monitor for subtle drifts in model quality over time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Comparison at a Glance
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Feature&lt;/th&gt;
&lt;th&gt;Bifrost&lt;/th&gt;
&lt;th&gt;Datadog LLM Observability&lt;/th&gt;
&lt;th&gt;Langfuse&lt;/th&gt;
&lt;th&gt;Arize Phoenix&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Capture Point&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;AI Gateway&lt;/td&gt;
&lt;td&gt;Application SDK&lt;/td&gt;
&lt;td&gt;Application SDK&lt;/td&gt;
&lt;td&gt;Application SDK&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Deployment&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Self-hosted (OSS), Cloud&lt;/td&gt;
&lt;td&gt;SaaS&lt;/td&gt;
&lt;td&gt;Self-hosted (OSS), Cloud&lt;/td&gt;
&lt;td&gt;Self-hosted (OSS), Cloud&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;OpenTelemetry&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Native Support&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Agent Tracing&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Yes (LLM + MCP)&lt;/td&gt;
&lt;td&gt;LLM Tracing&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;RAG Evaluation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Performance&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&amp;lt;15µs overhead&lt;/td&gt;
&lt;td&gt;Low (SDK-dependent)&lt;/td&gt;
&lt;td&gt;Low (SDK-dependent)&lt;/td&gt;
&lt;td&gt;Low (SDK-dependent)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Governance&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Integrated&lt;/td&gt;
&lt;td&gt;Separate&lt;/td&gt;
&lt;td&gt;Separate&lt;/td&gt;
&lt;td&gt;Separate&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fia2l8jtgmaoti8tgv580.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fia2l8jtgmaoti8tgv580.png" alt="A side-by-side comparison visualization. On one side, a complex, tangled web of individual light trails representing unm" width="800" height="457"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Recommendation
&lt;/h2&gt;

&lt;p&gt;For enterprises, LLM observability cannot be an isolated function; it must be an integrated part of the AI infrastructure stack, connected to security, governance, and cost management. While AI-native and APM tools offer valuable insights, they often miss traffic from uninstrumented systems and lack direct control mechanisms.&lt;/p&gt;

&lt;p&gt;An AI gateway provides the most robust and comprehensive foundation for enterprise-grade observability. By capturing standardized telemetry from every AI request at the source, it delivers universal visibility without compromising performance.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bifrost&lt;/strong&gt; stands out as the top choice for enterprises in 2026. Its combination of high performance, open-source flexibility, native OpenTelemetry support, and integrated governance makes it the most effective platform for understanding and controlling production AI systems at scale. Teams can adopt Bifrost to centralize observability and then feed that rich, gateway-level data into downstream systems like Datadog or Grafana, getting the best of both worlds.&lt;/p&gt;

&lt;p&gt;To learn more, teams can review the &lt;a href="https://docs.getbifrost.ai/features/observability/default" rel="noopener noreferrer"&gt;Bifrost documentation&lt;/a&gt; or &lt;a href="https://getmaxim.ai/bifrost/book-a-demo" rel="noopener noreferrer"&gt;request a demo&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sources
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;a href="https://docs.getbifrost.ai/features/observability/default" rel="noopener noreferrer"&gt;Bifrost Docs: Built-in Observability&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;  &lt;a href="https://www.confident-ai.com/blog/llm-observability-tools" rel="noopener noreferrer"&gt;Confident AI: 12 LLM Observability Tools to Evaluate &amp;amp; Monitor AI in 2026&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;  &lt;a href="https://www.paloaltonetworks.com/cyberpedia/what-is-an-ai-gateway" rel="noopener noreferrer"&gt;Palo Alto Networks: What Is an AI Gateway?&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;  &lt;a href="https://signoz.io/blog/llm-observability-tools" rel="noopener noreferrer"&gt;SigNoz: Top LLM Observability Tools&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>llobs</category>
      <category>llmops</category>
      <category>aigateway</category>
      <category>observability</category>
    </item>
    <item>
      <title>Top 5 Enterprise LLM Observability Platforms in 2026</title>
      <dc:creator>Swapnoneel Saha</dc:creator>
      <pubDate>Wed, 30 Sep 2026 10:30:16 +0000</pubDate>
      <link>https://dev.to/swapnoneel123/top-5-enterprise-llm-observability-platforms-in-2026-3oa5</link>
      <guid>https://dev.to/swapnoneel123/top-5-enterprise-llm-observability-platforms-in-2026-3oa5</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  LLM observability provides visibility into the performance, quality, cost, and security of AI applications, which traditional monitoring tools cannot capture.&lt;/li&gt;
&lt;li&gt;  Enterprise platforms are evaluated on tracing depth, evaluation capabilities, security and governance features, cost management, and deployment flexibility.&lt;/li&gt;
&lt;li&gt;  AI Gateways like Bifrost provide observability as a function of centrally routing all AI traffic, capturing complete data without SDK-level changes.&lt;/li&gt;
&lt;li&gt;  APM extensions from vendors like Datadog and New Relic integrate LLM tracing into existing infrastructure monitoring stacks.&lt;/li&gt;
&lt;li&gt;  AI-native platforms like Arize AI and Langfuse offer deep, purpose-built tracing and evaluation features, often with open-source options.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;LLM observability has become a critical component of the enterprise AI stack. Traditional application performance monitoring (APM) can confirm that an API returned a 200 OK status, but it cannot tell you if the model's response was a hallucination, if it leaked sensitive data, or if a minor prompt change just increased costs by 300%. LLM observability platforms are designed to answer these questions by tracking the inputs, outputs, and internal states of AI applications.&lt;/p&gt;

&lt;p&gt;This article compares the top five enterprise LLM observability platforms for 2026, assessing them on the features that matter most for production AI: tracing, evaluation, cost control, security, and integration. The platforms fall into three categories: AI gateways, APM extensions, and AI-native tracing tools.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key Criteria for Evaluating Enterprise LLM Observability
&lt;/h2&gt;

&lt;p&gt;Before comparing the tools, it's important to establish a framework. For enterprise use, an LLM observability solution must go beyond basic logging and provide robust, scalable, and secure insights.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Criteria&lt;/th&gt;
&lt;th&gt;Description&lt;/th&gt;
&lt;th&gt;Why It Matters for Enterprises&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Full Request &amp;amp; Response Tracing&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Captures the entire payload, including prompts, responses, model parameters, tool calls, and retrieval steps.&lt;/td&gt;
&lt;td&gt;Essential for debugging, auditing, and understanding model behavior. Incomplete data leaves blind spots.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Evaluation &amp;amp; Quality Scoring&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Capabilities to assess outputs against defined metrics (e.g., faithfulness, relevance, toxicity) via automated evaluators or human feedback.&lt;/td&gt;
&lt;td&gt;Moves beyond operational metrics (latency, errors) to measure the actual quality and safety of AI responses.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Cost &amp;amp; Token Tracking&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Granular visibility into token consumption and cost per request, aggregated by user, project, or model.&lt;/td&gt;
&lt;td&gt;Enables cost control, budget enforcement, and identification of expensive outliers before they impact the bottom line.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Security &amp;amp; Governance&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Features like PII/secrets redaction, access controls (RBAC), audit logs, and policy enforcement (guardrails).&lt;/td&gt;
&lt;td&gt;Critical for compliance, protecting sensitive data, and managing risk in regulated industries.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Integration &amp;amp; Deployment&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Support for OpenTelemetry, compatibility with existing stacks (APM, data lakes), and flexible deployment options (cloud, VPC, self-hosted).&lt;/td&gt;
&lt;td&gt;Reduces vendor lock-in, integrates with existing workflows, and meets enterprise data residency and security requirements.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  The Top 5 Platforms Compared
&lt;/h2&gt;

&lt;p&gt;This list assesses the leading platforms, highlighting their strengths and ideal use cases for enterprise teams.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Bifrost: Observability Through a Unified AI Gateway
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://www.getmaxim.ai/bifrost" rel="noopener noreferrer"&gt;Bifrost&lt;/a&gt; is a high-performance, &lt;a href="https://github.com/maximhq/bifrost" rel="noopener noreferrer"&gt;open-source AI gateway&lt;/a&gt; that provides observability as a native function of its architecture. By centralizing all LLM traffic through a single point, Bifrost captures complete, structured data for every request and response automatically, without requiring developers to instrument each application with an SDK.&lt;/p&gt;

&lt;p&gt;Bifrost's approach is powerful for enterprises because it guarantees 100% visibility. Since all requests—from any team, application, or model—must pass through the gateway, nothing is missed. This makes it a strong foundation for governance and cost control.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxq9xuf100vs7lsvfmzn4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxq9xuf100vs7lsvfmzn4.png" alt="A close-up of interconnected, illuminated pathways on a circuit board, symbolizing the detailed tracing of an AI request" width="800" height="457"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key Features:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Built-in, Asynchronous Logging&lt;/strong&gt;: Bifrost logs the full request context—inputs, outputs, tokens, cost, latency, provider, and model—with negligible impact on performance.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;OpenTelemetry &amp;amp; Prometheus Exports&lt;/strong&gt;: Natively exports traces and metrics to existing monitoring systems like Datadog, New Relic, Grafana, and Honeycomb, integrating seamlessly into established enterprise stacks.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Integrated Governance and Security&lt;/strong&gt;: Because observability is part of the gateway, the same platform enforces &lt;a href="https://www.getmaxim.ai/bifrost/resources/governance" rel="noopener noreferrer"&gt;governance&lt;/a&gt; controls. Virtual keys, budgets, rate limits, and guardrails are applied to the same traffic being observed.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Enterprise-Grade Guardrails&lt;/strong&gt;: Connects to a wide range of content safety providers, including AWS Bedrock Guardrails, Azure Content Safety, and Google Model Armor, to enforce security policies on live traffic.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Endpoint Governance with Bifrost Edge&lt;/strong&gt;: A key differentiator is &lt;a href="https://www.getmaxim.ai/bifrost/edge" rel="noopener noreferrer"&gt;Bifrost Edge&lt;/a&gt;, which extends the gateway's observability and security policies to AI usage on employee machines, covering desktop apps and browser-based AI to mitigate shadow AI risks.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; Enterprises that need a centralized control plane for all AI traffic, combining comprehensive observability with robust security, governance, and cost management in a single platform. Its high-performance, deployment flexibility (including in-VPC and air-gapped environments) makes it ideal for regulated industries and mission-critical applications.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Datadog Agent Observability
&lt;/h3&gt;

&lt;p&gt;Datadog is a dominant player in the traditional APM and infrastructure monitoring space. Its &lt;a href="https://www.datadoghq.com/product/agent-observability/" rel="noopener noreferrer"&gt;Agent Observability&lt;/a&gt; (formerly LLM Observability) product extends its existing platform to cover AI applications. For companies already invested in the Datadog ecosystem, this offers a path to LLM monitoring with minimal vendor sprawl.&lt;/p&gt;

&lt;p&gt;Datadog's core strength is its ability to correlate LLM traces with the rest of the application stack. Teams can see how a model's latency impacts overall service performance, linking AI behavior to infrastructure metrics, logs, and user experience data in one place.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key Features:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Unified Monitoring&lt;/strong&gt;: Views LLM performance alongside application and infrastructure metrics.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Auto-Instrumentation&lt;/strong&gt;: Provides libraries for popular frameworks to automatically capture trace data.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Built-in Quality &amp;amp; Safety Checks&lt;/strong&gt;: Includes out-of-the-box evaluations for metrics like toxicity and topic relevance.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Cost and Token Tracking&lt;/strong&gt;: Monitors usage to help manage expenses.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; Organizations already standardized on Datadog for their primary monitoring needs. It provides a single pane of glass for teams who want to add LLM visibility to their existing APM workflows without onboarding a new vendor.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. New Relic AI Observability
&lt;/h3&gt;

&lt;p&gt;Similar to Datadog, New Relic has extended its established APM platform to include &lt;a href="https://newrelic.com/platform/ai-observability" rel="noopener noreferrer"&gt;AI observability&lt;/a&gt;. It leverages the OpenTelemetry standard, which gives teams more flexibility and reduces vendor lock-in compared to proprietary instrumentation.&lt;/p&gt;

&lt;p&gt;New Relic focuses on providing a holistic view of the AI stack, from the application layer down to the infrastructure. It offers pre-built dashboards and integrations for popular AI frameworks and providers, allowing teams to get started quickly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key Features:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;OpenTelemetry Native&lt;/strong&gt;: Built on the open standard for telemetry, ensuring data portability.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Full-Stack Visibility&lt;/strong&gt;: Connects AI layer performance with application and infrastructure health.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Real-time Insights&lt;/strong&gt;: Offers dashboards for tracking performance, errors, and cost issues.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Model Context Protocol (MCP) Monitoring&lt;/strong&gt;: Provides visibility into the entire MCP request lifecycle for agentic applications.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; Companies that use New Relic as their APM and want to add LLM monitoring within that ecosystem, especially those who prioritize open standards like OpenTelemetry.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Arize AI
&lt;/h3&gt;

&lt;p&gt;Arize AI is an AI-native observability platform that offers deep capabilities for both LLM evaluation and production monitoring. It is split into two main products: Phoenix, a popular open-source library for tracing and evaluation during development, and AX, the enterprise-scale monitoring platform.&lt;/p&gt;

&lt;p&gt;Arize's strength lies in its comprehensive, model-centric approach. It's designed from the ground up to handle the nuances of AI systems, with strong features for monitoring embedding drift, evaluating RAG pipeline performance, and debugging complex agent workflows.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn0cnpldwrcthkipknrv1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn0cnpldwrcthkipknrv1.png" alt="A botanist in a futuristic greenhouse meticulously inspecting glowing, holographic plants with a digital tablet, a metap" width="800" height="457"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key Features:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Open-Source Foundation&lt;/strong&gt;: Phoenix allows teams to start with a powerful, self-hostable tool for tracing and evaluation.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Enterprise-Scale Monitoring&lt;/strong&gt;: Arize AX is built for high-volume production environments, processing trillions of events per month for its customers.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Deep Evaluation Tooling&lt;/strong&gt;: Offers advanced evaluation capabilities, including support for LLM-as-a-judge and integration with various evaluation frameworks.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;OpenTelemetry-Based&lt;/strong&gt;: Uses the OpenInference standard, which is built on OpenTelemetry, for vendor-agnostic instrumentation.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; ML-focused teams and enterprises that require deep, AI-native tracing and evaluation capabilities across the entire model lifecycle. The open-source entry point makes it accessible, while the enterprise platform provides the scale and features needed for large-scale production deployments.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Langfuse
&lt;/h3&gt;

&lt;p&gt;Langfuse is an open-source LLM engineering platform that combines observability, prompt management, and evaluation in one integrated system. It has gained significant traction for its developer-friendly workflow and comprehensive feature set.&lt;/p&gt;

&lt;p&gt;Langfuse provides detailed, hierarchical traces that capture every step of an LLM application's execution, from model calls to tool use and retrieval steps. Its open-source nature allows for self-hosting, giving enterprises full control over their data—a critical requirement for privacy and compliance.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key Features:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;Open Source and Self-Hostable&lt;/strong&gt;: Gives teams complete data sovereignty and control.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Integrated Platform&lt;/strong&gt;: Combines tracing, prompt management, and evaluation in a single workflow.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Detailed Tracing&lt;/strong&gt;: Offers hierarchical views of complex agent interactions.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;Framework Agnostic&lt;/strong&gt;: Integrates with various LLM frameworks and models.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Best for:&lt;/strong&gt; Teams looking for a powerful, open-source platform that covers the entire LLM development lifecycle. Its self-hosting capability makes it a strong choice for organizations with strict data residency requirements or those who prefer to build on an open-source stack.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What is the difference between LLM observability and traditional APM?
&lt;/h3&gt;

&lt;p&gt;Traditional Application Performance Monitoring (APM) tracks operational metrics like latency, error rates, and resource usage. LLM observability goes further by analyzing the quality and content of AI responses, tracking things like hallucinations, relevance, token costs, and potential data leakage—issues that traditional APM cannot detect.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why is OpenTelemetry important for LLM observability?
&lt;/h3&gt;

&lt;p&gt;OpenTelemetry is an open standard for instrumenting, generating, and collecting telemetry data (traces, metrics, logs). Using an OpenTelemetry-native platform prevents vendor lock-in, as the instrumentation in your code is not tied to a specific vendor's SDK. This allows enterprises to switch observability backends without re-instrumenting their applications.&lt;/p&gt;

&lt;h3&gt;
  
  
  How does an AI gateway provide observability?
&lt;/h3&gt;

&lt;p&gt;An AI gateway acts as a central proxy for all LLM API requests. Because every request and response flows through it, the gateway can log the complete data payload for every transaction automatically. This provides 100% visibility without requiring developers to add observability code to each individual application.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can these platforms detect security issues like prompt injection?
&lt;/h3&gt;

&lt;p&gt;Yes, advanced LLM observability platforms can help detect security threats. By monitoring prompts for anomalous patterns and applying guardrails, they can flag potential prompt injection attacks or attempts to leak sensitive data. Platforms with integrated guardrail systems provide an active defense layer, not just passive monitoring.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do I choose the right platform for my organization?
&lt;/h3&gt;

&lt;p&gt;The right choice depends on your existing infrastructure and primary needs.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;If you need a unified control plane for governance, security, and observability&lt;/strong&gt;, an AI gateway like &lt;strong&gt;Bifrost&lt;/strong&gt; is the most comprehensive solution.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;If you are already heavily invested in an APM platform&lt;/strong&gt;, using the LLM module from &lt;strong&gt;Datadog&lt;/strong&gt; or &lt;strong&gt;New Relic&lt;/strong&gt; is the path of least resistance.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;If your primary need is deep, AI-native evaluation and tracing&lt;/strong&gt;, a purpose-built platform like &lt;strong&gt;Arize AI&lt;/strong&gt; or &lt;strong&gt;Langfuse&lt;/strong&gt; is likely the best fit.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Next Steps
&lt;/h2&gt;

&lt;p&gt;Choosing an enterprise LLM observability platform is a strategic decision that impacts reliability, security, and cost. Gateways provide a holistic control plane, APM extensions offer integration with existing stacks, and AI-native tools deliver specialized depth. Teams evaluating these options can &lt;a href="https://getmaxim.ai/bifrost/book-a-demo" rel="noopener noreferrer"&gt;request a Bifrost demo&lt;/a&gt; to see how gateway-based observability works or explore the &lt;a href="https://github.com/maximhq/bifrost" rel="noopener noreferrer"&gt;open-source repository&lt;/a&gt; to get started.&lt;/p&gt;

</description>
      <category>observability</category>
      <category>ai</category>
      <category>enterprise</category>
      <category>devops</category>
    </item>
    <item>
      <title>5 Best MCP Gateways in 2026</title>
      <dc:creator>Swapnoneel Saha</dc:creator>
      <pubDate>Tue, 15 Sep 2026 11:44:59 +0000</pubDate>
      <link>https://dev.to/swapnoneel123/5-best-mcp-gateways-in-2026-1880</link>
      <guid>https://dev.to/swapnoneel123/5-best-mcp-gateways-in-2026-1880</guid>
      <description>&lt;p&gt;If you’re building AI agents in 2026, the difficult part is no longer making a model call.&lt;/p&gt;

&lt;p&gt;It’s making the model safely use real systems.&lt;/p&gt;

&lt;p&gt;GitHub access. Databases. CRMs. Internal APIs. Cloud infrastructure. File systems. SaaS applications. Every MCP server introduces another transport, credential, tool catalog, session, permission model, and operational surface.&lt;/p&gt;

&lt;p&gt;Among the options below, &lt;a href="https://www.getmaxim.ai/mcp-gateway" rel="noopener noreferrer"&gt;Bifrost by Maxim AI&lt;/a&gt; stands out for pairing an Apache 2.0 open-source core with a production-focused control plane. Its &lt;a href="https://github.com/maximhq/bifrost" rel="noopener noreferrer"&gt;GitHub repository&lt;/a&gt; is available for teams that want to inspect, run, and extend the gateway themselves.&lt;/p&gt;

&lt;p&gt;At small scale, connecting an agent directly to a few MCP servers works fine.&lt;/p&gt;

&lt;p&gt;At production scale, teams start dealing with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Credentials scattered across agent configurations&lt;/li&gt;
&lt;li&gt;Hundreds of tools bloating the model context&lt;/li&gt;
&lt;li&gt;Destructive tools exposed to the wrong users&lt;/li&gt;
&lt;li&gt;No unified audit trail for tool calls&lt;/li&gt;
&lt;li&gt;MCP servers using different transports&lt;/li&gt;
&lt;li&gt;OAuth tokens expiring during long-running workflows&lt;/li&gt;
&lt;li&gt;No reliable way to attribute tool usage to a user, team, or customer&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That’s where MCP gateways become core infrastructure.&lt;/p&gt;

&lt;p&gt;An MCP gateway sits between AI clients and MCP servers, giving teams one place to manage discovery, authentication, authorization, routing, observability, and policy enforcement.&lt;/p&gt;

&lt;p&gt;This article compares the 5 best MCP gateways in 2026 based on real production concerns: security, transport support, governance, reliability, deployment flexibility, developer experience, and cost control.&lt;/p&gt;

&lt;p&gt;We’ll start with a quick comparison table, then examine how each gateway behaves in production.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is an MCP Gateway (And Why Does It Matter in Production)?
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://modelcontextprotocol.io/" rel="noopener noreferrer"&gt;MCP&lt;/a&gt; is an open protocol for connecting AI applications with external data sources and tools.&lt;/p&gt;

&lt;p&gt;An MCP server can expose:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Tools that an AI model can call&lt;/li&gt;
&lt;li&gt;Resources that provide context or data&lt;/li&gt;
&lt;li&gt;Prompts and reusable workflows&lt;/li&gt;
&lt;li&gt;Authentication and session handling&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;An MCP gateway sits in front of one or more MCP servers and presents a controlled interface to AI clients such as Claude, Cursor, ChatGPT, VS Code, or custom agents.&lt;/p&gt;

&lt;p&gt;A production MCP gateway should provide:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;One endpoint for multiple MCP servers&lt;/li&gt;
&lt;li&gt;Centralized authentication and credential management&lt;/li&gt;
&lt;li&gt;Tool-level filtering and access control&lt;/li&gt;
&lt;li&gt;User, team, agent, or tenant-level permissions&lt;/li&gt;
&lt;li&gt;Session-aware routing&lt;/li&gt;
&lt;li&gt;Rate limits and usage budgets&lt;/li&gt;
&lt;li&gt;Per-call logs, metrics, and traces&lt;/li&gt;
&lt;li&gt;Support for remote and local MCP transports&lt;/li&gt;
&lt;li&gt;Safe handling of OAuth tokens and secrets&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Without a gateway, every agent becomes its own integration hub.&lt;/p&gt;

&lt;p&gt;That means each agent must independently manage server URLs, tool schemas, credentials, retries, permissions, and logs. The result is an N×M integration problem: N agents connected directly to M tools.&lt;/p&gt;

&lt;p&gt;An MCP gateway reduces that complexity to a governed control point.&lt;/p&gt;

&lt;h2&gt;
  
  
  How We Evaluated These MCP Gateways
&lt;/h2&gt;

&lt;p&gt;This comparison focuses on production readiness rather than feature-count marketing.&lt;/p&gt;

&lt;p&gt;The main criteria were:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Security and identity: OAuth, SSO, RBAC, delegated access, and secret handling&lt;/li&gt;
&lt;li&gt;Tool governance: allowlists, virtual servers, approvals, scopes, and policy enforcement&lt;/li&gt;
&lt;li&gt;Transport compatibility: STDIO, HTTP, SSE, and Streamable HTTP&lt;/li&gt;
&lt;li&gt;Observability: audit logs, latency, errors, traces, and usage attribution&lt;/li&gt;
&lt;li&gt;Reliability: session handling, retries, scaling, and high availability&lt;/li&gt;
&lt;li&gt;Deployment: SaaS, VPC, on-premises, Kubernetes, and air-gapped options&lt;/li&gt;
&lt;li&gt;Context efficiency: discovery, schema loading, and tool-token reduction&lt;/li&gt;
&lt;li&gt;Pricing and operating model&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The ranking is based on production breadth and architectural fit, not on one synthetic cross-vendor benchmark.&lt;/p&gt;

&lt;h2&gt;
  
  
  Quick Comparison: Top 5 MCP Gateways
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Gateway&lt;/th&gt;
&lt;th&gt;Delivery&lt;/th&gt;
&lt;th&gt;Open Source&lt;/th&gt;
&lt;th&gt;Strengths&lt;/th&gt;
&lt;th&gt;Tradeoffs&lt;/th&gt;
&lt;th&gt;Pricing Model&lt;/th&gt;
&lt;th&gt;Best For&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.getmaxim.ai/mcp-gateway" rel="noopener noreferrer"&gt;Bifrost&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Self-hosted single binary and Enterprise&lt;/td&gt;
&lt;td&gt;✓ Apache 2.0 core&lt;/td&gt;
&lt;td&gt;Unified MCP, LLM, and agent control plane; governance; air-gapped deployment; high performance&lt;/td&gt;
&lt;td&gt;Enterprise features require a paid tier&lt;/td&gt;
&lt;td&gt;Free self-hosted; Enterprise custom&lt;/td&gt;
&lt;td&gt;Production AI systems that need one control plane&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://ibm.github.io/mcp-context-forge/latest/overview/features/" rel="noopener noreferrer"&gt;IBM ContextForge&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Self-hosted Python, Docker, and Kubernetes&lt;/td&gt;
&lt;td&gt;✓ Apache 2.0&lt;/td&gt;
&lt;td&gt;MCP, A2A, REST/gRPC federation; registry; plugins; observability&lt;/td&gt;
&lt;td&gt;Broad architecture requires more operational ownership&lt;/td&gt;
&lt;td&gt;Free software; infrastructure costs apply&lt;/td&gt;
&lt;td&gt;Platform teams that want open-source federation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://developer.konghq.com/ai-gateway/mcp/" rel="noopener noreferrer"&gt;Kong AI Gateway&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Konnect-managed or self-hosted&lt;/td&gt;
&lt;td&gt;Varies&lt;/td&gt;
&lt;td&gt;API-to-MCP mapping; aggregation; authentication; policies; metrics and audit logs&lt;/td&gt;
&lt;td&gt;Heavier API-management setup&lt;/td&gt;
&lt;td&gt;Commercial and enterprise pricing&lt;/td&gt;
&lt;td&gt;Existing Kong and API-platform teams&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://learn.microsoft.com/en-us/azure/api-management/ai-gateway-overview" rel="noopener noreferrer"&gt;Azure API Management AI Gateway&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Managed Azure AI gateway and API Management&lt;/td&gt;
&lt;td&gt;✗&lt;/td&gt;
&lt;td&gt;One governed endpoint for models and MCP; OpenAPI-to-MCP; connectors; runtime keys; policies&lt;/td&gt;
&lt;td&gt;Public preview; regions, limits, and pricing may change&lt;/td&gt;
&lt;td&gt;Preview pricing to be announced; Azure costs apply&lt;/td&gt;
&lt;td&gt;Azure-native platform teams&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://docs.arcade.dev/en/operate/governance/mcp-gateways" rel="noopener noreferrer"&gt;Arcade MCP Gateways&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;Managed Cloud, VPC, and air-gapped Enterprise&lt;/td&gt;
&lt;td&gt;✗&lt;/td&gt;
&lt;td&gt;Tool curation; user-scoped OAuth; remote MCP federation; action-focused integrations&lt;/td&gt;
&lt;td&gt;Remote MCP support is Streamable HTTP and tools-focused&lt;/td&gt;
&lt;td&gt;Free tier; Team from $25/month plus usage; Enterprise custom&lt;/td&gt;
&lt;td&gt;Applications needing user-authorized actions&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;These gateways overlap, but they solve slightly different problems. Some are full AI control planes. Others focus on tool federation, API management, managed connectors, or user-level authorization.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9hi80exk851u8wlpnhx1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9hi80exk851u8wlpnhx1.png" alt=" " width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Bifrost (by Maxim AI)
&lt;/h2&gt;

&lt;p&gt;Many MCP gateways focus on aggregating tools.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.getmaxim.ai/mcp-gateway" rel="noopener noreferrer"&gt;Bifrost&lt;/a&gt; takes a broader approach: MCP traffic is part of the same AI infrastructure as model calls and agent activity.&lt;/p&gt;

&lt;p&gt;According to Maxim AI’s &lt;a href="https://www.getmaxim.ai/llm-info" rel="noopener noreferrer"&gt;official Bifrost reference&lt;/a&gt;, Bifrost is an enterprise AI gateway and control plane for LLMs, MCP tools, and agents. The open-source core is licensed under Apache 2.0 and can run inside a company’s own infrastructure.&lt;/p&gt;

&lt;p&gt;That broader scope is what puts Bifrost at number one.&lt;/p&gt;

&lt;h3&gt;
  
  
  One control plane for MCP, models, and agents
&lt;/h3&gt;

&lt;p&gt;Bifrost can connect to upstream MCP servers over:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;STDIO&lt;/li&gt;
&lt;li&gt;HTTP&lt;/li&gt;
&lt;li&gt;SSE&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It can also expose connected tools through a single MCP endpoint for clients such as Claude Desktop, Cursor, Claude Code, and custom agents.&lt;/p&gt;

&lt;p&gt;The result is one place to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Register MCP servers&lt;/li&gt;
&lt;li&gt;Discover and synchronize tools&lt;/li&gt;
&lt;li&gt;Filter tools by client, request, or virtual key&lt;/li&gt;
&lt;li&gt;Create curated virtual MCP servers&lt;/li&gt;
&lt;li&gt;Apply user, team, customer, or business-unit policies&lt;/li&gt;
&lt;li&gt;Authenticate upstream servers&lt;/li&gt;
&lt;li&gt;Record every tool call&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Bifrost supports server-level and per-user OAuth 2.0, automatic token refresh, shared headers, and token exchange. Credentials remain in the gateway instead of being copied into every agent configuration.&lt;/p&gt;

&lt;h3&gt;
  
  
  Governance is built into the request path
&lt;/h3&gt;

&lt;p&gt;Bifrost combines tool governance with broader &lt;a href="https://www.getmaxim.ai/ai-governance" rel="noopener noreferrer"&gt;AI governance&lt;/a&gt; controls.&lt;/p&gt;

&lt;p&gt;Teams can use:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Role-based access control&lt;/li&gt;
&lt;li&gt;Data access policies&lt;/li&gt;
&lt;li&gt;Virtual keys&lt;/li&gt;
&lt;li&gt;Budgets and rate limits&lt;/li&gt;
&lt;li&gt;Tool groups and access profiles&lt;/li&gt;
&lt;li&gt;Guardrails for requests and responses&lt;/li&gt;
&lt;li&gt;Personal-data and secret detection&lt;/li&gt;
&lt;li&gt;Signed audit logs&lt;/li&gt;
&lt;li&gt;Prometheus metrics and OpenTelemetry tracing&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This matters because MCP tools are not just read-only information sources. They can create tickets, modify infrastructure, send messages, move money, or delete data.&lt;/p&gt;

&lt;p&gt;The gateway needs to understand which tool is being called, who is calling it, and whether that action is allowed.&lt;/p&gt;

&lt;h3&gt;
  
  
  MCP Code Mode reduces context pressure
&lt;/h3&gt;

&lt;p&gt;Large MCP installations can expose hundreds of tools. Sending every tool definition to the model on every request increases token usage and makes tool selection less reliable.&lt;/p&gt;

&lt;p&gt;Bifrost’s MCP Code Mode lets the model orchestrate tools through sandboxed code and load schemas only when needed.&lt;/p&gt;

&lt;p&gt;The official reference reports input-token reductions of up to 92.8% across large tool sets. Maxim AI also reports a published example where a 508-tool workflow reduced token cost from $377 to $29 while maintaining a 100% task pass rate.&lt;/p&gt;

&lt;p&gt;This is one of Bifrost’s most interesting advantages for agent-heavy systems: the gateway is not only routing tool calls; it is also reducing the cost of deciding which tools to call.&lt;/p&gt;

&lt;h3&gt;
  
  
  Performance and deployment
&lt;/h3&gt;

&lt;p&gt;Bifrost is written in Go and is designed for the request path. Its architecture keeps configuration in memory, reuses pooled memory, and maintains provider connections ahead of time.&lt;/p&gt;

&lt;p&gt;Maxim AI’s official reference reports approximately 11 microseconds of gateway overhead at 5,000 requests per second on a single t3.xlarge instance. That is a vendor-published benchmark, so teams should validate results against their own MCP server latency and workload patterns.&lt;/p&gt;

&lt;p&gt;Deployment is flexible:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Virtual machines&lt;/li&gt;
&lt;li&gt;Kubernetes&lt;/li&gt;
&lt;li&gt;Private cloud&lt;/li&gt;
&lt;li&gt;In-VPC environments&lt;/li&gt;
&lt;li&gt;On-premises infrastructure&lt;/li&gt;
&lt;li&gt;Fully disconnected or air-gapped networks&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Bifrost runs inside the customer’s infrastructure, and the official reference states that prompts, responses, files, and telemetry do not leave that environment.&lt;/p&gt;

&lt;p&gt;Teams can also add &lt;a href="https://www.getmaxim.ai/ai-guardrails" rel="noopener noreferrer"&gt;AI guardrails&lt;/a&gt; and route operational data into existing monitoring systems through &lt;a href="https://www.getmaxim.ai/ai-observability" rel="noopener noreferrer"&gt;AI observability&lt;/a&gt; integrations.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pricing
&lt;/h3&gt;

&lt;p&gt;Bifrost’s open-source version is free forever for self-hosted use under Apache 2.0.&lt;/p&gt;

&lt;p&gt;The Enterprise tier is custom priced and adds features such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;High availability&lt;/li&gt;
&lt;li&gt;Cluster mode&lt;/li&gt;
&lt;li&gt;Identity-based governance&lt;/li&gt;
&lt;li&gt;Advanced security&lt;/li&gt;
&lt;li&gt;Compliance controls&lt;/li&gt;
&lt;li&gt;Enterprise support&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Bifrost Enterprise includes a 14-day trial. See &lt;a href="https://www.getmaxim.ai/pricing" rel="noopener noreferrer"&gt;current Bifrost pricing&lt;/a&gt; before making a purchasing decision.&lt;/p&gt;

&lt;h3&gt;
  
  
  Best for
&lt;/h3&gt;

&lt;p&gt;Teams running high-traffic, customer-facing, regulated, or multi-tenant AI systems where MCP tools need to follow the same policies as LLM and agent traffic.&lt;/p&gt;

&lt;p&gt;If you are evaluating it, start with the &lt;a href="https://docs.getbifrost.ai/mcp/overview" rel="noopener noreferrer"&gt;Bifrost MCP Gateway documentation&lt;/a&gt; and the &lt;a href="https://github.com/maximhq/bifrost" rel="noopener noreferrer"&gt;open-source repository&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. IBM ContextForge
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/IBM/mcp-context-forge" rel="noopener noreferrer"&gt;IBM ContextForge&lt;/a&gt; is one of the broadest open-source options in this category.&lt;/p&gt;

&lt;p&gt;It acts as an MCP registry, proxy, and federation layer, but also supports A2A, REST, and gRPC services.&lt;/p&gt;

&lt;p&gt;That makes it useful for organizations whose internal tool ecosystem is not uniformly MCP-based yet.&lt;/p&gt;

&lt;h3&gt;
  
  
  What stands out
&lt;/h3&gt;

&lt;p&gt;ContextForge supports:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Federation across multiple MCP servers&lt;/li&gt;
&lt;li&gt;REST-to-MCP adaptation&lt;/li&gt;
&lt;li&gt;gRPC-to-MCP translation&lt;/li&gt;
&lt;li&gt;HTTP, JSON-RPC, WebSocket, SSE, and Streamable HTTP&lt;/li&gt;
&lt;li&gt;Virtual servers containing selected tools&lt;/li&gt;
&lt;li&gt;JWT, Basic Auth, and custom authentication&lt;/li&gt;
&lt;li&gt;Rate limiting&lt;/li&gt;
&lt;li&gt;OAuth and user-scoped access&lt;/li&gt;
&lt;li&gt;Admin UI and real-time logs&lt;/li&gt;
&lt;li&gt;Prometheus metrics&lt;/li&gt;
&lt;li&gt;OpenTelemetry tracing&lt;/li&gt;
&lt;li&gt;Redis-backed caching and session storage&lt;/li&gt;
&lt;li&gt;Multi-cluster Kubernetes deployments&lt;/li&gt;
&lt;li&gt;Plugin extensibility&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It can wrap legacy REST or gRPC services as MCP-compatible tools, which can reduce the need to rewrite existing internal APIs.&lt;/p&gt;

&lt;h3&gt;
  
  
  Tradeoffs
&lt;/h3&gt;

&lt;p&gt;ContextForge’s breadth is also its main tradeoff.&lt;/p&gt;

&lt;p&gt;This is a platform to operate, not just a lightweight reverse proxy. Production deployments may require PostgreSQL, Redis, Kubernetes configuration, secret management, backups, upgrades, and monitoring.&lt;/p&gt;

&lt;p&gt;Teams also remain responsible for securing the gateway and validating downstream tool behavior. The software provides the control surface, but operating a reliable self-hosted control plane still requires platform engineering effort.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pricing
&lt;/h3&gt;

&lt;p&gt;ContextForge is open source under Apache 2.0. There is no gateway license fee for self-hosting, but teams pay for infrastructure, operations, and any commercial support they choose to add.&lt;/p&gt;

&lt;h3&gt;
  
  
  Best for
&lt;/h3&gt;

&lt;p&gt;Platform teams that want an open-source, self-hosted gateway capable of federating MCP with REST, gRPC, and A2A services.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Kong AI Gateway
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://developer.konghq.com/ai-gateway/mcp/" rel="noopener noreferrer"&gt;Kong AI Gateway&lt;/a&gt; brings MCP into an established API-management platform.&lt;/p&gt;

&lt;p&gt;It supports three traffic types through a unified control plane:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;LLM traffic&lt;/li&gt;
&lt;li&gt;MCP traffic&lt;/li&gt;
&lt;li&gt;Agent-to-agent traffic&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Kong’s MCP layer can map REST API endpoints into MCP-compatible tools, aggregate multiple MCP servers, and expose those tools to clients such as Claude Desktop and Cursor.&lt;/p&gt;

&lt;h3&gt;
  
  
  What stands out
&lt;/h3&gt;

&lt;p&gt;Kong provides:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;AI MCP Server entities&lt;/li&gt;
&lt;li&gt;REST-to-MCP mapping&lt;/li&gt;
&lt;li&gt;Multiple MCP server aggregation&lt;/li&gt;
&lt;li&gt;Key-based and OpenID Connect authentication&lt;/li&gt;
&lt;li&gt;OAuth protected-resource metadata&lt;/li&gt;
&lt;li&gt;ACLs and rate-limiting policies&lt;/li&gt;
&lt;li&gt;MCP audit logs&lt;/li&gt;
&lt;li&gt;Latency, response-size, and error metrics&lt;/li&gt;
&lt;li&gt;OpenTelemetry integrations&lt;/li&gt;
&lt;li&gt;Konnect-based management&lt;/li&gt;
&lt;li&gt;Self-hosted deployment options&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For companies already operating Kong, this can be a natural extension of an existing API governance model.&lt;/p&gt;

&lt;p&gt;MCP tools can be managed alongside other APIs rather than becoming a completely separate infrastructure category.&lt;/p&gt;

&lt;h3&gt;
  
  
  Tradeoffs
&lt;/h3&gt;

&lt;p&gt;Kong is powerful, but it assumes familiarity with API management, control planes, data planes, policies, and Kong configuration.&lt;/p&gt;

&lt;p&gt;It is usually more infrastructure than a small team needs for a handful of MCP servers. The MCP registry functionality is also marked as a technology preview in the current documentation.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pricing
&lt;/h3&gt;

&lt;p&gt;Kong AI Gateway is commercially positioned through Kong Konnect and enterprise deployments. Pricing depends on the selected control-plane, data-plane, support, and deployment model.&lt;/p&gt;

&lt;h3&gt;
  
  
  Best for
&lt;/h3&gt;

&lt;p&gt;Enterprises already standardized on Kong or teams that want MCP governance integrated into an existing API platform.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Azure API Management AI Gateway
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://learn.microsoft.com/en-us/azure/api-management/ai-gateway-overview" rel="noopener noreferrer"&gt;Azure API Management AI Gateway&lt;/a&gt; is Microsoft’s managed gateway for AI models and MCP tools. The AI Gateway tier is currently in public preview.&lt;/p&gt;

&lt;p&gt;It gives platform teams one place to publish, secure, govern, and observe AI traffic.&lt;/p&gt;

&lt;p&gt;A single MCP server can federate backends from:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Remote MCP servers&lt;/li&gt;
&lt;li&gt;OpenAPI specifications&lt;/li&gt;
&lt;li&gt;Built-in SaaS connectors&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That makes it possible to bring existing APIs and MCP services behind one governed endpoint without rewriting every backend.&lt;/p&gt;

&lt;h3&gt;
  
  
  What stands out
&lt;/h3&gt;

&lt;p&gt;Azure AI Gateway provides:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Runtime access keys for client applications&lt;/li&gt;
&lt;li&gt;Backend authentication through API keys, OAuth 2.0, or managed identity&lt;/li&gt;
&lt;li&gt;Content-safety policies&lt;/li&gt;
&lt;li&gt;IP filters&lt;/li&gt;
&lt;li&gt;Request and token rate limits&lt;/li&gt;
&lt;li&gt;Azure Monitor and Application Insights integration&lt;/li&gt;
&lt;li&gt;OpenTelemetry support&lt;/li&gt;
&lt;li&gt;A self-service model and tool catalog&lt;/li&gt;
&lt;li&gt;Private networking options&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The gateway keeps backend credentials away from client applications. Agents authenticate to the gateway, while the gateway handles authentication to the underlying model or tool service.&lt;/p&gt;

&lt;p&gt;Azure API Management can also expose REST API operations as MCP tools, which makes it attractive for enterprises with a large existing API estate.&lt;/p&gt;

&lt;h3&gt;
  
  
  Tradeoffs
&lt;/h3&gt;

&lt;p&gt;The main limitation is preview maturity.&lt;/p&gt;

&lt;p&gt;Microsoft currently lists the AI Gateway tier in East US 2 and Sweden Central, with availability, quotas, telemetry, and APIs subject to change. The preview has no service-level agreement, so teams should use controlled pilots and maintain a rollback plan for critical workloads.&lt;/p&gt;

&lt;p&gt;Azure API Management’s MCP capabilities currently focus on tools rather than MCP resources or prompts.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pricing
&lt;/h3&gt;

&lt;p&gt;Microsoft states that pricing and the business model for the AI Gateway tier will be announced later. Related Azure services, observability resources, networking, and backend providers may still generate costs.&lt;/p&gt;

&lt;h3&gt;
  
  
  Best for
&lt;/h3&gt;

&lt;p&gt;Azure-native organizations that want managed governance for models, APIs, and MCP tools through Microsoft identity, networking, and monitoring.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Arcade MCP Gateways
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://docs.arcade.dev/en/operate/governance/mcp-gateways" rel="noopener noreferrer"&gt;Arcade MCP Gateways&lt;/a&gt; are designed around managed tool access and user-authorized actions.&lt;/p&gt;

&lt;p&gt;A gateway can combine tools from multiple MCP servers into one endpoint. Administrators choose exactly which tools are exposed, and different gateways can provide different tool bundles for different projects or workflows.&lt;/p&gt;

&lt;h3&gt;
  
  
  What stands out
&lt;/h3&gt;

&lt;p&gt;Arcade supports:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Tool federation from multiple MCP servers&lt;/li&gt;
&lt;li&gt;Per-gateway tool selection&lt;/li&gt;
&lt;li&gt;Server instructions for agents&lt;/li&gt;
&lt;li&gt;Remote MCP server registration&lt;/li&gt;
&lt;li&gt;OIDC-based user identity&lt;/li&gt;
&lt;li&gt;OAuth-based authentication&lt;/li&gt;
&lt;li&gt;Project and organization scopes&lt;/li&gt;
&lt;li&gt;Dashboard-based gateway creation&lt;/li&gt;
&lt;li&gt;AI-assisted gateway configuration&lt;/li&gt;
&lt;li&gt;Streamable HTTP connections&lt;/li&gt;
&lt;li&gt;Managed tool and connector catalogs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The user identity model is particularly useful for applications where an agent must act on behalf of an end user rather than using one shared service account.&lt;/p&gt;

&lt;p&gt;Arcade supports Arcade Auth, an external OIDC User Source, or API-key-style headers for clients that cannot complete a browser-based OAuth flow.&lt;/p&gt;

&lt;h3&gt;
  
  
  Tradeoffs
&lt;/h3&gt;

&lt;p&gt;Arcade’s remote MCP support currently focuses on Streamable HTTP servers that are reachable from Arcade.&lt;/p&gt;

&lt;p&gt;Its documentation also states that remote MCP gateways expose tools, while prompts, resources, and sampling are not supported for remote servers today.&lt;/p&gt;

&lt;p&gt;That makes Arcade a strong fit for action-oriented integrations, but it is not the same type of full AI infrastructure control plane as Bifrost or Kong.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pricing
&lt;/h3&gt;

&lt;p&gt;Arcade publishes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Free: $0 per month, including 2,000 auth events and 2,000 tool calls per month&lt;/li&gt;
&lt;li&gt;Team: $25 per month plus usage&lt;/li&gt;
&lt;li&gt;Auth events: $0.10 each&lt;/li&gt;
&lt;li&gt;Tool calls: $0.01 each&lt;/li&gt;
&lt;li&gt;Enterprise: Custom pricing with VPC, air-gapped deployment, SSO, RBAC, audit logs, and private registry access&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;See the &lt;a href="https://www.arcade.dev/pricing/" rel="noopener noreferrer"&gt;Arcade pricing page&lt;/a&gt; for current limits and terms.&lt;/p&gt;

&lt;h3&gt;
  
  
  Best for
&lt;/h3&gt;

&lt;p&gt;Product teams that need user-scoped authorization, managed integrations, and a simple way to expose carefully selected tools to agents or IDE clients.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to Choose the Right MCP Gateway
&lt;/h2&gt;

&lt;p&gt;There is no universal winner for every MCP architecture.&lt;/p&gt;

&lt;p&gt;Ask yourself:&lt;/p&gt;

&lt;h3&gt;
  
  
  Do MCP calls need to follow the same policy as model calls?
&lt;/h3&gt;

&lt;p&gt;If your organization is managing LLMs, agents, and MCP tools as one production system, a unified AI gateway is usually the cleanest model.&lt;/p&gt;

&lt;p&gt;Bifrost is especially strong here because its control plane covers LLM, MCP, and agent traffic, with self-hosted and air-gapped deployment options.&lt;/p&gt;

&lt;p&gt;Kong is a strong alternative for teams that already manage APIs through its platform.&lt;/p&gt;

&lt;h3&gt;
  
  
  Do you need open-source federation?
&lt;/h3&gt;

&lt;p&gt;If you need to connect MCP servers with REST, gRPC, or A2A services while keeping the gateway in your own infrastructure, IBM ContextForge is worth serious consideration.&lt;/p&gt;

&lt;p&gt;Bifrost is also a strong option when performance, unified AI governance, and a simpler deployment model matter more than broad protocol federation.&lt;/p&gt;

&lt;h3&gt;
  
  
  Are you primarily trying to roll out approved connectors?
&lt;/h3&gt;

&lt;p&gt;Azure API Management AI Gateway is a strong fit when your connectors and APIs already live in Azure. It can federate remote MCP servers, expose operations from OpenAPI definitions as tools, and connect to built-in SaaS integrations through one governed endpoint.&lt;/p&gt;

&lt;p&gt;Arcade is compelling when the main requirement is exposing user-authorized actions through curated gateways.&lt;/p&gt;

&lt;h3&gt;
  
  
  Which MCP transports do your servers use?
&lt;/h3&gt;

&lt;p&gt;Transport compatibility matters more than many comparison tables suggest.&lt;/p&gt;

&lt;p&gt;If you depend on local STDIO servers, verify support before choosing a managed gateway. If your servers are already exposed through Streamable HTTP, options such as Arcade and Kong become easier to evaluate.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can you enforce tool-level permissions?
&lt;/h3&gt;

&lt;p&gt;Server-level access is not enough.&lt;/p&gt;

&lt;p&gt;A gateway should let you expose read-only tools while hiding destructive operations. For example, an agent may be allowed to read a GitHub repository but not merge a pull request, query a database but not modify rows, or inspect cloud resources but not delete them.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can you explain every tool call later?
&lt;/h3&gt;

&lt;p&gt;For production systems, logs should answer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which user or agent initiated the call?&lt;/li&gt;
&lt;li&gt;Which MCP server handled it?&lt;/li&gt;
&lt;li&gt;Which tool was executed?&lt;/li&gt;
&lt;li&gt;What authorization policy applied?&lt;/li&gt;
&lt;li&gt;How long did it take?&lt;/li&gt;
&lt;li&gt;Did it fail, retry, or time out?&lt;/li&gt;
&lt;li&gt;What did it cost?&lt;/li&gt;
&lt;li&gt;What model request caused it?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If the gateway only shows server uptime, it is not providing enough operational visibility.&lt;/p&gt;

&lt;h3&gt;
  
  
  How large will your tool catalog become?
&lt;/h3&gt;

&lt;p&gt;A gateway should help agents discover tools without forcing every schema into every prompt.&lt;/p&gt;

&lt;p&gt;For large MCP installations, evaluate lazy discovery, virtual servers, tool groups, schema loading, and approaches such as &lt;a href="https://www.getmaxim.ai/mcp-gateway" rel="noopener noreferrer"&gt;Bifrost MCP Code Mode&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final Thoughts
&lt;/h2&gt;

&lt;p&gt;MCP is becoming the standard interface between AI agents and the systems they need to use.&lt;/p&gt;

&lt;p&gt;That also means MCP gateways are becoming more than connection managers. They are becoming policy boundaries.&lt;/p&gt;

&lt;p&gt;The right gateway determines:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Which tools an agent can see&lt;/li&gt;
&lt;li&gt;Which actions it can execute&lt;/li&gt;
&lt;li&gt;Which identity is attached to the call&lt;/li&gt;
&lt;li&gt;Where credentials are stored&lt;/li&gt;
&lt;li&gt;How sessions are routed&lt;/li&gt;
&lt;li&gt;How failures are handled&lt;/li&gt;
&lt;li&gt;What gets logged&lt;/li&gt;
&lt;li&gt;How tool usage is controlled and optimized&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you want open-source federation across MCP, REST, gRPC, and A2A, IBM ContextForge is a strong choice.&lt;/p&gt;

&lt;p&gt;If you want MCP governance integrated with an established API-management platform, Kong is compelling.&lt;/p&gt;

&lt;p&gt;If you want managed Azure governance for models, APIs, and MCP tools, Azure API Management AI Gateway is worth evaluating, subject to its preview limitations.&lt;/p&gt;

&lt;p&gt;If your priority is user-authorized actions through curated tool bundles, Arcade is a good fit.&lt;/p&gt;

&lt;p&gt;But for teams building high-traffic production AI systems, Bifrost is the strongest overall option because it combines MCP, LLM, and agent governance in one control plane. Its Apache 2.0 core, self-hosted deployment, air-gapped support, tool-level policies, Code Mode, observability, and performance focus give it a particularly strong foundation for serious agent infrastructure.&lt;/p&gt;

&lt;p&gt;Start with the &lt;a href="https://www.getmaxim.ai/mcp-gateway" rel="noopener noreferrer"&gt;Bifrost MCP Gateway&lt;/a&gt;, review the &lt;a href="https://docs.getbifrost.ai/mcp/overview" rel="noopener noreferrer"&gt;documentation&lt;/a&gt;, explore the &lt;a href="https://www.getmaxim.ai/bifrost/resources" rel="noopener noreferrer"&gt;Bifrost resources&lt;/a&gt;, and validate the architecture against your own tools and security requirements.&lt;/p&gt;

&lt;p&gt;Thanks for reading! 🙏&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>ai</category>
      <category>llm</category>
      <category>agents</category>
    </item>
    <item>
      <title>What Is Granular Access Control? RBAC, ABAC, and AI</title>
      <dc:creator>Swapnoneel Saha</dc:creator>
      <pubDate>Wed, 09 Sep 2026 17:25:41 +0000</pubDate>
      <link>https://dev.to/swapnoneel123/what-is-granular-access-control-rbac-abac-and-ai-54eo</link>
      <guid>https://dev.to/swapnoneel123/what-is-granular-access-control-rbac-abac-and-ai-54eo</guid>
      <description>&lt;p&gt;"Can this person use the AI API?" is one question.&lt;/p&gt;

&lt;p&gt;"Can this service call gpt-5, at most 200 times an hour, with a 50 dollar cap this month, with no file system tools attached, and only until Friday?" is a completely different question.&lt;/p&gt;

&lt;p&gt;Granular access control is what turns the first question into the second one. It means defining permissions at the smallest useful unit, so an identity can take exactly one action, on exactly one resource, within exactly one budget, under exactly one set of conditions, and nothing else.&lt;/p&gt;

&lt;p&gt;That idea has a name, the principle of least privilege, and it just means giving each caller the minimum access it needs. Granular access control is how you actually write that down.&lt;/p&gt;

&lt;p&gt;And the reason it stopped being a nice-to-have this year is that most of the things asking for permission are no longer people.&lt;/p&gt;

&lt;h2&gt;
  
  
  Access control is a dial, not a switch
&lt;/h2&gt;

&lt;p&gt;Strip a permission down and it is three things: a &lt;strong&gt;subject&lt;/strong&gt; (who is asking), an &lt;strong&gt;action&lt;/strong&gt; (what they want to do), and a &lt;strong&gt;resource&lt;/strong&gt; (what they want to do it to). Every access control system you have ever used is just a way of writing those three things down and checking them.&lt;/p&gt;

&lt;p&gt;Coarse-grained access control leaves all three parts wide open. "Engineers can use the database" is one permission covering forty people, six verbs, and every table you own. It is one line of config, and it is fast to write, which is exactly why it survives so long.&lt;/p&gt;

&lt;p&gt;Granular access control narrows the same three parts instead. "The reporting service can run SELECT on the &lt;code&gt;orders&lt;/code&gt; table, and nothing else." One caller, one verb, one table.&lt;/p&gt;

&lt;p&gt;Think of a house key against a hotel keycard. The house key opens everything you own, forever, and the only way to take it back is to change the lock. The keycard opens room 402, on floor 4, until Friday at 11 in the morning, and the front desk can kill it from the lobby without touching a single door.&lt;/p&gt;

&lt;p&gt;Both are access control. The keycard is granular.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F01pzq01gmahs1mm3fwa8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F01pzq01gmahs1mm3fwa8.png" alt="Granular access control illustrated as a hotel keycard" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So granularity is not a product you buy or a model you adopt. It is how far down you are willing to turn the dial, and how much extra work you are willing to do in exchange. Turn it too far and you get a permissions table nobody can read. Leave it too coarse and one leaked credential owns the whole system.&lt;/p&gt;

&lt;h2&gt;
  
  
  The five dimensions you can narrow a permission on
&lt;/h2&gt;

&lt;p&gt;Most explanations of granular access control stop at "give people less access", which is true and completely useless as advice. The useful version is that a permission can be narrowed along five separate axes, and you can turn each one independently.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2q78ahupqavcnq2hz7j1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2q78ahupqavcnq2hz7j1.png" alt="The five dimensions of granular access control" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Here they are, from the one everybody already does to the one almost nobody does.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Who is asking
&lt;/h3&gt;

&lt;p&gt;The identity. This is the axis every system starts with, and the one that changed the most recently.&lt;/p&gt;

&lt;p&gt;For a long time a subject meant a person, or a service account somebody created once and forgot about. Now a subject is just as likely to be a coding agent, a CI job, a background worker, or a chain of three agents where the last one has no idea which human started the request.&lt;/p&gt;

&lt;p&gt;These are called non-human identities, and they are the ones quietly holding most of the over-permissioned access. Sonrai's cloud access research, published in May 2026, found that &lt;a href="https://sonraisecurity.com/blog/why-92-percent-cloud-permissions-never-used/" rel="noopener noreferrer"&gt;92% of identities with sensitive permissions&lt;/a&gt; did not use them once in 90 days, and 87% of that group were machine identities. That is a vendor scanning its own customers and not an independent audit, so take the exact number lightly.&lt;/p&gt;

&lt;p&gt;The fix here is easy to say and annoying to do: one identity per caller, never one shared credential per team.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. What action they want to take
&lt;/h3&gt;

&lt;p&gt;Read, write, update, delete. List versus fetch. Call a model versus look at what models exist.&lt;/p&gt;

&lt;p&gt;Coarse systems have two levels here, usually named something like read and admin. Granular systems split up the verbs that can do different amounts of damage. Being able to view a log line and being able to export the whole log table are not the same permission, even though both are "reading".&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Which resource, exactly
&lt;/h3&gt;

&lt;p&gt;This is where granularity gets real, because a resource is rarely one flat thing.&lt;/p&gt;

&lt;p&gt;A database can be scoped at the server, the database, the table, the column, or the row. An AI setup can be scoped at the provider, the model, or a single tool that a model is allowed to invoke. Each step down that ladder means less damage if the key leaks, and one more thing to maintain.&lt;/p&gt;

&lt;p&gt;Row-level scoping is the one worth knowing by name. It answers "which of these records can you see", which is a different question from "can you see this table", and it is the difference between a support agent seeing their own tickets and a support agent seeing everybody's.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. How much they can consume
&lt;/h3&gt;

&lt;p&gt;Here is the axis classic access control mostly ignores, and it is the one that bites hardest with AI.&lt;/p&gt;

&lt;p&gt;Permission has traditionally been a boolean. You can call this endpoint, or you cannot. But when a single call can cost real money and a loop can make ten thousand of them, "yes" without a ceiling is not a permission, it is an open tab.&lt;/p&gt;

&lt;p&gt;So the granular version attaches quantity to the grant. This key may spend 200 dollars a month. This key may burn 10,000 tokens an hour. This key may make 100 requests a minute. Developers have been asking providers for exactly this for years, and OpenAI's own developer forum carries a &lt;a href="https://community.openai.com/t/api-key-specific-rate-and-spending-limits-would-be-good/1374358" rel="noopener noreferrer"&gt;feature request for per-key spending limits&lt;/a&gt; with people piling into it.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Under what conditions
&lt;/h3&gt;

&lt;p&gt;The context around the request. Time of day, source IP, environment, device posture, whether the key has expired yet.&lt;/p&gt;

&lt;p&gt;Expiry is the underrated one. A permission with no end date is a permission you will forget you granted, and every "how did this old key still work" incident starts there.&lt;/p&gt;

&lt;h2&gt;
  
  
  Is granular access control the same as RBAC?
&lt;/h2&gt;

&lt;p&gt;No, and this trips up a lot of people, so it is worth separating properly.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuqjtyhvz59r4ngk3s42q.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuqjtyhvz59r4ngk3s42q.png" alt="RBAC, ABAC, and ReBAC compared" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;RBAC&lt;/strong&gt; (role-based access control) groups permissions into roles and hands roles to identities. It is the most common model in the world because it maps onto how companies actually think. You are a developer, developers get the developer role, done.&lt;/p&gt;

&lt;p&gt;RBAC can be granular or coarse depending entirely on how you write the roles. That is the part people miss. A role called &lt;code&gt;admin&lt;/code&gt; with every permission attached is RBAC and it is not granular at all.&lt;/p&gt;

&lt;p&gt;The trouble shows up when you try to push RBAC down the dial. Every new condition needs its own role, so you get &lt;code&gt;developer-staging&lt;/code&gt;, &lt;code&gt;developer-staging-eu&lt;/code&gt;, &lt;code&gt;developer-staging-eu-readonly&lt;/code&gt;, and pretty soon nobody can tell you what any of them do. The industry name for this is &lt;a href="https://www.osohq.com/learn/rbac-vs-abac" rel="noopener noreferrer"&gt;role explosion&lt;/a&gt;, and it is the standard failure mode of a team that discovered granularity and only had roles to express it with.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ABAC&lt;/strong&gt; (attribute-based access control) fixes that by evaluating attributes at request time instead of pre-baking roles. Department, clearance, resource classification, time, location. It gets you much finer control without the role count exploding, and the price you pay is that when a request gets denied, you have to trace which attribute failed instead of just reading a role name.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ReBAC&lt;/strong&gt; (relationship-based access control) asks a third question: how is this subject related to this resource? You can edit the doc because you created it. You can see the profile because you manage that person. Google Drive works this way, and so does every app where ownership is the real rule.&lt;/p&gt;

&lt;p&gt;Well, which one should you use? Honestly, most real systems end up with RBAC for the broad strokes and one of the other two layered on for the cases roles cannot express. Granularity comes from the rules you write, and not from the model you picked. You can write a terrible coarse ABAC policy. Plenty of people have.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why AI made this urgent
&lt;/h2&gt;

&lt;p&gt;I run Claude Code and GPT Codex daily, and both of them do things a normal API consumer never did. They read files and call tools. And they chain several model calls together off one instruction I typed half-awake.&lt;/p&gt;

&lt;p&gt;That breaks the two assumptions most permission systems were built on: that whoever is asking is a person, and that every request is something a person actually asked for.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft0ckl54bbgbqz57nzaw1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft0ckl54bbgbqz57nzaw1.png" alt="AI agents expand access control risk" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The numbers back this up. 1Password surveyed 1,000 security and engineering staff at large US firms in late May and early June 2026, and found that agents in production reached &lt;a href="https://www.helpnetsecurity.com/2026/07/29/1password-ai-agent-governance/" rel="noopener noreferrer"&gt;roughly twice as much data&lt;/a&gt; as had actually been approved. In the same survey, 33% of developers running agents reported a breach or security incident tied to an over-privileged non-human identity, and 40% said they leave agents holding persistent access to systems and secrets after the task is finished.&lt;/p&gt;

&lt;p&gt;None of that is an AI-specific vulnerability. It is the same over-permissioning problem as always, running at a speed and volume that humans never generated.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where granular access control breaks in practice
&lt;/h2&gt;

&lt;p&gt;Everybody agrees with least privilege. Almost nobody has it. So it is worth being specific about where the plan falls apart, because the reason is usually the setup, and not laziness.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A provider key cannot be narrowed.&lt;/strong&gt; Your OpenAI or Anthropic key is a single credential with your whole account behind it. There is no version of it that means "gpt-4o only, 50 dollars, no tool calls". So the moment more than one service needs model access, you either share one key and lose all attribution, or you mint several and lose all central control.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Enforcement lives in the wrong place.&lt;/strong&gt; If the permission check is inside your application code, then every new service, notebook, cron job, and intern's side script has to re-implement it correctly. One of them will not. And the ones that skip it will not show up in any policy list, because they never registered with the policy system in the first place.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;You cannot enforce what you cannot attribute.&lt;/strong&gt; At Zonko Labs I built an internal tool that captured our AI product's data logs and generated reports on latency and probable slowdowns, and the thing that made it useful was that every log line could be traced back to a specific caller. Without that, a spend spike is just a number going up. You cannot tighten a permission when you do not know which caller needs loosening.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsxibvxloeoljo26wt89q.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsxibvxloeoljo26wt89q.png" alt="Common granular access control failure modes" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So the pattern behind all three is the same. Granular access control needs a place in the request path that sees every call, knows who made it, and can say no before the call leaves your network.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the gateway comes in
&lt;/h2&gt;

&lt;p&gt;Let me be honest about where this actually helps first. If you are one person with one API key on one machine, you do not need any of this, and a gateway will not make your side project safer. It also cannot fix a permission model you have not thought about. A gateway enforces the rules you write, so if you write "allow everything", you get everything, faster.&lt;/p&gt;

&lt;p&gt;But the moment there are several services, several models, and a few agents in the mix, the check has to sit in the traffic path. That is what an AI gateway is: a proxy that every model call goes through. Which makes it the one place you can check a permission once, instead of in every service.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkav5q2ggtprmlizj8nvz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkav5q2ggtprmlizj8nvz.png" alt="AI gateway enforcing permissions on model calls" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.getmaxim.ai/bifrost" rel="noopener noreferrer"&gt;Bifrost&lt;/a&gt; is the open-source one I keep pointing people at, partly because the whole thing sits in &lt;a href="https://github.com/maximhq/bifrost" rel="noopener noreferrer"&gt;a public GitHub repo&lt;/a&gt;, so you can read exactly what it enforces rather than trusting a feature list.&lt;/p&gt;

&lt;p&gt;The thing you hand out is a &lt;strong&gt;virtual key&lt;/strong&gt;. Your real provider credentials sit inside the gateway, and each caller gets its own scoped key instead. And the things you can scope that key on are almost exactly the five dimensions above.&lt;/p&gt;

&lt;p&gt;Take the rule in plain words first:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This key belongs to the support team. It can only use OpenAI, only the &lt;code&gt;gpt-4o&lt;/code&gt; model, and it stops working on the first of next year.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Written out in Bifrost's &lt;a href="https://docs.getbifrost.ai/features/governance" rel="noopener noreferrer"&gt;governance config&lt;/a&gt;, that is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"vk-support-bot"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"is_active"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"allow_all_providers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"expires_at"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2027-01-01T00:00:00Z"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"provider_configs"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"provider"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"openai"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"allowed_models"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"gpt-4o"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"blacklisted_models"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"team_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"team-support"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Reading that back: &lt;code&gt;allow_all_providers&lt;/code&gt; set to false means any provider not listed here is denied, including ones added to the gateway later. &lt;code&gt;allowed_models&lt;/code&gt; narrows dimension three to a single model. &lt;code&gt;expires_at&lt;/code&gt; is dimension five, so the key dies on its own without anybody remembering to revoke it. And &lt;code&gt;team_id&lt;/code&gt; attaches it to a team, which is where the budget lives.&lt;/p&gt;

&lt;p&gt;A few more things worth knowing:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Spending and rate caps sit on the key itself.&lt;/strong&gt; A budget is &lt;code&gt;max_limit&lt;/code&gt; plus a &lt;code&gt;reset_duration&lt;/code&gt; of &lt;code&gt;1m&lt;/code&gt;, &lt;code&gt;1h&lt;/code&gt;, &lt;code&gt;1d&lt;/code&gt;, &lt;code&gt;1w&lt;/code&gt;, &lt;code&gt;1M&lt;/code&gt;, &lt;code&gt;1Q&lt;/code&gt; or &lt;code&gt;1Y&lt;/code&gt;, and rate limits are separate &lt;code&gt;token_max_limit&lt;/code&gt; and &lt;code&gt;request_max_limit&lt;/code&gt; counters. Blowing through them returns real status codes rather than a generic failure: 402 for &lt;code&gt;budget_exceeded&lt;/code&gt;, 429 for &lt;code&gt;token_limited&lt;/code&gt;, 403 for &lt;code&gt;model_blocked&lt;/code&gt;. That distinction sounds small until you are the one reading the error at 2 in the morning.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tool access is deny-by-default.&lt;/strong&gt; If you have wired up MCP servers, a virtual key with no MCP configuration gets no tools at all. Where tools are granted, they are listed explicitly in &lt;code&gt;tools_to_execute&lt;/code&gt;, and the key's list acts as a ceiling that a request cannot widen. That is dimension three pushed all the way down to individual tools, which matters a lot given how much damage one file system tool can do inside an agent loop! (If MCP is new to you, I wrote &lt;a href="https://www.swapnoneel.site/blog/what-is-an-mcp-server" rel="noopener noreferrer"&gt;a beginner's guide to MCP servers&lt;/a&gt; first.)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Seeing the dashboard is a separate grant from calling a model.&lt;/strong&gt; Bifrost ships three system roles, Admin, Developer and Viewer, carrying 42, 27 and 14 permissions, and you can build custom ones by toggling resource and operation pairs. On top of that sits data access control, which decides which rows a user sees at all: &lt;code&gt;own-data&lt;/code&gt;, &lt;code&gt;team-data&lt;/code&gt;, or &lt;code&gt;all-data&lt;/code&gt;. Being allowed to open the logs page and being allowed to see everybody's logs are two different grants.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Write the policy once, not once per person.&lt;/strong&gt; &lt;a href="https://docs.getbifrost.ai/enterprise/access-profiles" rel="noopener noreferrer"&gt;Access profiles&lt;/a&gt; let you define one policy and auto-issue a per-user virtual key from it, each with its own budget counter. Edit the template, propagate, and every key follows. Keys can rotate on a schedule from &lt;code&gt;1h&lt;/code&gt; up to &lt;code&gt;365d&lt;/code&gt;. This is the part that decides whether any of it survives the team getting bigger. If every new engineer needs a hand-written policy, somebody will quietly hand out a shared key by month three.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Audit logs are signed.&lt;/strong&gt; Administrative events can be HMAC-signed and exported as JSON, JSON Lines, or syslog in the RFC 5424 format that SIEM tools ingest. Knowing who did something and being able to prove it later are the same job, and homegrown setups usually only manage the first half.&lt;/p&gt;

&lt;p&gt;Bifrost publishes its own benchmarks (roughly 20 microseconds of added latency at 5,000 requests per second), and those are vendor benchmarks run on a vendor harness, so measure your own. The point holds whichever gateway you use: the check belongs on the path, and not copy-pasted into fifteen codebases.&lt;/p&gt;

&lt;p&gt;If governance at this layer is the actual problem you are solving, I also went through &lt;a href="https://www.swapnoneel.site/blog/top-5-ai-agent-governance-platforms" rel="noopener noreferrer"&gt;the tools in this space&lt;/a&gt; in more detail.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where should you set the dial?
&lt;/h2&gt;

&lt;p&gt;The floor is not RBAC, and it is not ABAC either. The floor is being able to answer one question about your own system right now:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If a credential leaked this minute, what exactly could someone do with it, and how much could they spend before anyone noticed?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If the honest answer is "everything, and I have no idea", the fix is not a bigger access control model. It is a first cut at all five dimensions, in this order: one identity per caller, then a spending cap on each one, then an expiry date, then a model or resource allowlist, then conditions. That order is deliberate, because knowing who called and capping what they spend buy you the most safety for the least work, and conditions buy you the least.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl5be20b0hjp3mg2g397n.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl5be20b0hjp3mg2g397n.png" alt="Recommended order for setting granular access controls" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;You will not get to per-row, per-tool, per-hour granularity, and you should not try. Nobody is running the theoretically correct permission model. The teams doing well are the ones who pushed the dial two notches past a shared key and then actually maintained it.&lt;/p&gt;

&lt;p&gt;Anyway, I am curious where other people drew this line. If you have a permission setup that survived a team getting bigger, or one that collapsed into a shared key by month three, tell me in the comments, I want to hear the failure stories more than the success ones.&lt;/p&gt;

&lt;p&gt;You can find me on &lt;a href="https://x.com/swapnoneel123" rel="noopener noreferrer"&gt;X&lt;/a&gt;, or read the rest of my writing on &lt;a href="https://www.swapnoneel.site" rel="noopener noreferrer"&gt;my site&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>security</category>
      <category>ai</category>
      <category>beginners</category>
    </item>
    <item>
      <title>Three Insights for Building Great Claude Skills</title>
      <dc:creator>Swapnoneel Saha</dc:creator>
      <pubDate>Tue, 08 Sep 2026 18:32:06 +0000</pubDate>
      <link>https://dev.to/swapnoneel123/three-insights-for-building-great-claude-skills-47o1</link>
      <guid>https://dev.to/swapnoneel123/three-insights-for-building-great-claude-skills-47o1</guid>
      <description>&lt;p&gt;Most Claude skills that never fire are not broken. Their description is.&lt;/p&gt;

&lt;p&gt;That one fact is behind almost every "why won't my skill trigger" post you'll find. Three things separate a working skill from one that sits there ignored: a description written like a routing rule, a &lt;code&gt;SKILL.md&lt;/code&gt; kept short on purpose, and pushing fragile steps into scripts instead of long instructions. Get those three right, and everything else gets much easier.&lt;/p&gt;

&lt;h2&gt;
  
  
  First, what a Claude skill actually is
&lt;/h2&gt;

&lt;p&gt;A skill is a folder with a &lt;code&gt;SKILL.md&lt;/code&gt; file in it. At the top of that file, you have a small YAML header with a name and a description, followed by instructions written in markdown. That's the whole concept. Anthropic launched skills in October 2025 and &lt;a href="https://venturebeat.com/ai/anthropic-launches-enterprise-agent-skills-and-opens-the-standard" rel="noopener noreferrer"&gt;made the format an open standard&lt;/a&gt; in December, so Claude Code, the API, and claude.ai all use the exact same folder structure.&lt;/p&gt;

&lt;p&gt;The interesting part is how skills load. Claude doesn't read your entire skill upfront. When it starts up, it reads only the name and description of every installed skill, costing around 100 tokens each. Claude loads the main instructions only when it decides your skill fits what you asked for. And it reads extra helper files only when those instructions specifically point to them. Anthropic calls this &lt;a href="https://platform.claude.com/docs/en/agents-and-tools/agent-skills/overview" rel="noopener noreferrer"&gt;progressive disclosure&lt;/a&gt;. Two of the three insights below come down to working with this design instead of fighting it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgz1vqxln5j6jxssvavdi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgz1vqxln5j6jxssvavdi.png" alt="A folder reveals a skill’s metadata, instructions, and helper script loading in stages." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This means skills are cheap to keep around, costing you tokens only when they actually run. That single design choice shapes what makes a skill good or bad.&lt;/p&gt;

&lt;h2&gt;
  
  
  Insight 1: Your description is a routing rule, not a summary
&lt;/h2&gt;

&lt;p&gt;Here is where people get confused. Your description isn't a friendly summary for humans browsing your repository. It is a routing rule Claude uses to pick your skill out of the pile.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9bmhnhds2rzhe73g61my.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9bmhnhds2rzhe73g61my.png" alt="Prompt paths enter a routing switch; precise triggers pass and vague descriptions stop." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Remember what Claude actually sees at startup: only the name and description of each skill. Nothing else. When you send a prompt, Claude decides whether to open your skill's full text based purely on that description. If the description is vague, Claude never reads the instructions. The best instructions in the world won't help if your skill never triggers.&lt;/p&gt;

&lt;p&gt;That's why a skill that won't trigger is almost never a code bug. It just means your description doesn't explain when to trigger using the words you actually type.&lt;/p&gt;

&lt;p&gt;A good description handles two jobs: what the skill does, and when Claude should run it. Write it in the third person (like "Reviews pull requests" instead of "I review" or "Review this"). Because Claude reads this inside its system prompt, mixing perspectives can confuse the model. You have up to 1,024 characters, so use that space for the exact phrases you would normally type.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;reviewing-pull-requests&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Reviews&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;pull&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;request&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;diffs&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;against&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;the&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;team's&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;checklist&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;flags&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;risky&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;changes.&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Use&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;when&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;the&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;asks&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;to&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;review&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;PR,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;review&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;diff,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;or&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;check&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;staged&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;changes&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;before&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;merging."&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice that the description names the triggers out loud: "review a PR", "review a diff", "check staged changes". Compare that to "Helps with code review", which competes with everything Claude already knows about code review and loses.&lt;/p&gt;

&lt;p&gt;The name matters less than the description, but it has strict rules that catch people off guard. It must use only lowercase letters, numbers, and hyphens, capped at 64 characters. It cannot include the words "anthropic" or "claude". Anthropic recommends names ending in -ing (&lt;code&gt;reviewing-pull-requests&lt;/code&gt;, &lt;code&gt;analyzing-spreadsheets&lt;/code&gt;) to clearly show what action the skill takes. A name like &lt;code&gt;claude-pr-helper&lt;/code&gt; breaks two rules at once, while a generic name like &lt;code&gt;helper&lt;/code&gt; gives no useful information.&lt;/p&gt;

&lt;p&gt;There is real data behind this. One developer ran &lt;a href="https://medium.com/@ivan.seleznov1/why-claude-code-skills-dont-activate-and-how-to-fix-it-86f679409af1" rel="noopener noreferrer"&gt;650 activation trials&lt;/a&gt; on the same skills. They found that passive descriptions triggered far less reliably than direct instructions starting with "Use this skill when." While this isn't an official benchmark from Anthropic, the takeaway matches both the official documentation and real-world use: name the specific situation where Claude should trigger, not just what it can do.&lt;/p&gt;

&lt;p&gt;Also, keep your scope realistic. If it's too broad ("helps with frontend"), Claude will pick its own built-in knowledge instead. If it's too narrow ("fixes React 18 hydration errors in Next.js 14 app router"), it won't run if you phrase your prompt slightly differently. You are describing a general task, not a specific bug ticket.&lt;/p&gt;

&lt;h2&gt;
  
  
  Insight 2: Progressive disclosure is a token budget you spend
&lt;/h2&gt;

&lt;p&gt;Once a skill triggers, its &lt;code&gt;SKILL.md&lt;/code&gt; body loads into the same context window as your conversation, your files, and every other skill's metadata. Anthropic makes a great point in their &lt;a href="https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices" rel="noopener noreferrer"&gt;authoring guide&lt;/a&gt;: the context window is shared space. Every token you waste in your instruction file is memory the model can't use for your actual work.&lt;/p&gt;

&lt;p&gt;This means your main skill file should be smaller than you might expect. The official recommendation is to keep &lt;code&gt;SKILL.md&lt;/code&gt; under 500 lines. Move anything longer into separate files that load only when Claude needs them. Loading happens in stages, and each stage has a different cost:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Level&lt;/th&gt;
&lt;th&gt;When it loads&lt;/th&gt;
&lt;th&gt;Rough cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Name and description&lt;/td&gt;
&lt;td&gt;Always, at startup&lt;/td&gt;
&lt;td&gt;~100 tokens per skill&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SKILL.md body&lt;/td&gt;
&lt;td&gt;When the skill triggers&lt;/td&gt;
&lt;td&gt;Under 5k tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bundled files and scripts&lt;/td&gt;
&lt;td&gt;Only when referenced&lt;/td&gt;
&lt;td&gt;Nothing until read&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F49tkgx6zafxd3thsb00x.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F49tkgx6zafxd3thsb00x.png" alt="Instruction pages narrow through a funnel so only the needed context reaches the working window." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That third row is the real secret. You can include a 2,000-line API guide, datasets, or code examples, and they cost zero tokens until Claude actually needs them. So the best approach is a lean &lt;code&gt;SKILL.md&lt;/code&gt; that acts like a table of contents, pointing Claude to larger files when necessary:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;reviewing-pull-requests/
  SKILL.md          # short: the workflow and when to read what
  checklist.md      # loaded only when reviewing
  examples.md       # loaded only when Claude wants a sample
  scripts/
    diff_stats.py   # executed, never loaded into context
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The documentation highlights two practical rules here. First, keep linked files only one level deep and link them directly inside &lt;code&gt;SKILL.md&lt;/code&gt;. If you bury files inside subfolders, Claude often previews just the first 100 lines and misses the rest. Second, don't explain concepts Claude already understands. Anthropic uses a straightforward example: a good instruction to extract PDF text is about 50 tokens of code. A bad one uses 150 tokens because it starts by explaining what a PDF is. Claude already knows what a PDF is!&lt;/p&gt;

&lt;p&gt;Notice what I'm doing here, too. I'm not explaining what a token is or how context windows work. If you're building skills, you already understand those basics. That's the exact same mindset your skills need: only write what the model doesn't already know.&lt;/p&gt;

&lt;p&gt;I care about this point a lot because of my work building &lt;a href="https://anrl-site.vercel.app" rel="noopener noreferrer"&gt;ANRL&lt;/a&gt;. It's a compact data format created to cut token waste and context fragmentation, and it reduced formatting overhead by more than 40%. When you spend weeks fighting to save 40% on tokens, watching someone waste 3,000 tokens in &lt;code&gt;SKILL.md&lt;/code&gt; re-explaining JSON is painful. Keeping your skills lean isn't just about clean formatting—it protects the model's working memory.&lt;/p&gt;

&lt;h2&gt;
  
  
  Insight 3: Let code do the deterministic work
&lt;/h2&gt;

&lt;p&gt;The last one separates skills that feel reliable from skills that mostly work.&lt;/p&gt;

&lt;p&gt;If a step is deterministic (parse this file, validate this schema, sort these rows), do not write instructions asking Claude to be careful. Write a script and tell Claude to run it. When Claude runs a script through bash, only the output comes back into context, not the code. So a bundled &lt;code&gt;validate.py&lt;/code&gt; is cheaper than asking Claude to generate the same validation every time, and it does the exact same thing on every run instead of a slightly different thing.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg06ndx00o8to85u67nqt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg06ndx00o8to85u67nqt.png" alt="A plan crosses a guarded script bridge, passes validation, and reaches a stable output." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Anthropic uses a great phrase for writing these helper scripts: "solve, don't defer." If a file might be missing, handle that check right inside the script. Don't let the script crash and force Claude to guess a fix, because improvised fixes break reliability. The same rule applies to unexplained settings: setting a 30-second timeout with an explanatory comment is much better than leaving a random number like &lt;code&gt;47&lt;/code&gt; that no one can explain later.&lt;/p&gt;

&lt;p&gt;The bigger idea is giving Claude freedom only when the task allows it. Anthropic's docs share a helpful comparison:&lt;br&gt;
Some tasks are like a narrow bridge with steep drops on both sides, like running database migrations in a strict order. For those, give Claude a single script and tell it not to alter the command. Other tasks are like an open field with no dangers, like reviewing code where the best feedback depends on what Claude finds. For those, give general guidelines and trust the model to choose the right path. Making mistakes here hurts in both directions: adding too many rigid rules in an open field makes the skill inflexible, while giving too little direction on a narrow bridge leads to errors.&lt;/p&gt;

&lt;p&gt;For bulk edits or risky actions, the documentation suggests a three-step pattern: plan, validate, and execute. Have Claude write its planned changes to a file first. Next, run a script that checks that plan for mistakes. Finally, apply the changes only if the check passes. It might sound like extra work, but it's the difference between catching an error in a safe file and catching it after it breaks your database.&lt;/p&gt;

&lt;h2&gt;
  
  
  When one skill becomes twenty
&lt;/h2&gt;

&lt;p&gt;Those three insights help you build a solid individual skill. But developers rarely stop at just one. Once your team starts collecting dozens of skills, two quiet problems start showing up.&lt;/p&gt;

&lt;p&gt;First, skills can connect to external tools through MCP (if you haven't used MCP yet, check out my &lt;a href="https://www.swapnoneel.site/blog/what-is-an-mcp-server" rel="noopener noreferrer"&gt;beginner's guide to MCP servers&lt;/a&gt;). The easiest way to think about the difference: MCP provides the tool, while the skill teaches Claude how to use it for your specific workflow. But pay close attention to Anthropic's security warning: a skill can tell Claude to run tools in ways the author never mentioned. That's why they recommend checking every file inside any skill you didn't write yourself.&lt;/p&gt;

&lt;p&gt;Second, it's best practice to test skills across different models like Haiku, Sonnet, and Opus. A skill is only as reliable as the model running it. In a team environment, that means your skills end up calling different models across different API providers.&lt;/p&gt;

&lt;p&gt;Combine that with several teammates installing community skills from GitHub, and you aren't just managing simple markdown files anymore. You have unmonitored tool calls and model requests running across your team. That's an entirely different problem.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi8hkijfuymdbellxer1x.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi8hkijfuymdbellxer1x.png" alt="A shelf of skills expands into a connected ecosystem inspected for hidden tool risk." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Do you need a gateway for your skills?
&lt;/h2&gt;

&lt;p&gt;If you're running just two skills on your laptop, you don't. Seriously, skip this section and go write your skills instead of setting up infrastructure. A gateway won't make an individual skill better, and it won't fix a bad description.&lt;/p&gt;

&lt;p&gt;A gateway starts to matter when you can no longer answer simple questions: Who is allowed to run which tool? Which skill triggered this request? Whose API budget is paying for it? When you hit that point, routing all that traffic through one place makes life much easier. That place is an AI gateway, and the one I keep returning to is &lt;a href="https://www.getmaxim.ai/bifrost" rel="noopener noreferrer"&gt;Bifrost (by Maxim AI)&lt;/a&gt;, an open-source AI gateway and control plane for LLMs, MCP tools, and agents (&lt;a href="https://github.com/maximhq/bifrost" rel="noopener noreferrer"&gt;repo here&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjcuzh381b5f9rtjrg7ja.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjcuzh381b5f9rtjrg7ja.png" alt="A small setup runs directly while a growing team routes tools and models through one gateway." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Two of its core features line up directly with the scaling problems above.&lt;/p&gt;

&lt;p&gt;To secure tool calls, Bifrost acts as an &lt;a href="https://docs.getbifrost.ai/mcp/filtering" rel="noopener noreferrer"&gt;MCP gateway with deny-by-default tool filtering&lt;/a&gt;. Without explicit permissions, zero tools are accessible. You define what to unlock across client settings, request headers, and virtual key configurations. This means a skill could ask for a filesystem MCP server, but its key can restrict execution strictly to &lt;code&gt;read_file&lt;/code&gt;, completely blocking destructive actions like &lt;code&gt;delete_file&lt;/code&gt;. You don't have to depend on plain text instructions in a prompt to keep your systems safe. (It even has an MCP Code Mode that runs tool workflows inside a sandbox instead of dumping every tool schema into your context, cutting token usage by over 92% across large tool sets.)&lt;/p&gt;

&lt;p&gt;To manage model access, Bifrost provides &lt;a href="https://docs.getbifrost.ai/features/governance" rel="noopener noreferrer"&gt;virtual keys and governance rules&lt;/a&gt;. Each key gets its own list of approved models (across 25+ providers and 10,000+ models), an automatic budget cap with reset periods, and custom rate limits. So when you want your team to test a skill across Haiku, Sonnet, and Opus, you give them one key with an allow-list and a spending cap instead of juggling five API keys in a spreadsheet. Provider fallback and retries are built in, so a skill doesn't fall over if one provider has an outage.&lt;/p&gt;

&lt;p&gt;Performance is where this actually holds up. If a gateway is slow, developers route around it. Bifrost is written in Go with memory pooling and pre-warmed connections, keeping added gateway overhead to about 11 microseconds at 5,000 requests per second on a single t3.xlarge instance. In Maxim's published benchmarks against LiteLLM (tested with 500 concurrent users on AWS), Bifrost delivered 9.5x higher throughput (424 req/s vs 44.84 req/s) and roughly 50x lower P50 latency (804 ms vs 38.65 s). Those are vendor benchmarks, so treat them as such, but the main point holds: the gateway manages your traffic without becoming the slow part.&lt;/p&gt;

&lt;p&gt;None of this is about writing an individual skill. It's just what helps once your skills work well and you start scaling up.&lt;/p&gt;

&lt;h2&gt;
  
  
  So what actually makes a Claude skill great?
&lt;/h2&gt;

&lt;p&gt;It's not about making a giant &lt;code&gt;SKILL.md&lt;/code&gt; file. It's much simpler than that.&lt;/p&gt;

&lt;p&gt;Write your description like a clear routing rule in the third person, using the exact words you'd naturally type in a prompt. Keep the main instructions short and let extra reference files handle the heavy lifting, because context memory isn't free. And move fragile, repetitive tasks into code scripts so they run consistently every single time instead of guessing. Those three habits make the difference between a skill that runs reliably and one that sits forgotten in a folder.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0mqczjiu9kscgq8x4ld4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0mqczjiu9kscgq8x4ld4.png" alt="Three rising habits—routing, lean instructions, and deterministic code—build a reliable skill." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Using an AI gateway matters as well, but only later on. Don't worry about it until you have more skills and tools than you can easily keep track of in your head. You'll know when you cross that line. Until then, setting up a gateway is just extra work you don't need.&lt;/p&gt;

&lt;p&gt;If you've built skills that trigger reliably, I'd love to hear how you wrote your descriptions to make them work. That's still the area where most people spend the least effort. Come share your thoughts on &lt;a href="https://x.com/swapnoneel123" rel="noopener noreferrer"&gt;X&lt;/a&gt;, or check out my other posts at &lt;a href="https://www.swapnoneel.site" rel="noopener noreferrer"&gt;my site&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>programming</category>
      <category>webdev</category>
    </item>
    <item>
      <title>How to Cut Time to First Token (TTFT) in LLM Apps</title>
      <dc:creator>Swapnoneel Saha</dc:creator>
      <pubDate>Fri, 04 Sep 2026 15:49:58 +0000</pubDate>
      <link>https://dev.to/swapnoneel123/how-to-cut-time-to-first-token-ttft-in-llm-apps-b0f</link>
      <guid>https://dev.to/swapnoneel123/how-to-cut-time-to-first-token-ttft-in-llm-apps-b0f</guid>
      <description>&lt;p&gt;Whenever you type something into an AI app like ChatGPT or Gemini, and hit enter, you have to stare at nothing for half a second, one second, and sometimes even more. Then the words start pouring out fast. The wait time that you always experience in between these two events is called time to first token, or TTFT, and it is the single thing your users actually feel.&lt;/p&gt;

&lt;p&gt;Time to first token (TTFT) is the gap between sending a request and the first token coming back. To cut it, shorten and cache your prompt so the model reads less, stream the reply, keep connections warm, and route around slow providers. The biggest lever is prompt length, since most of that wait is the model reading your input.&lt;/p&gt;

&lt;p&gt;Now let me show you where the time actually goes, and how we can fix this. So, let's begin!&lt;/p&gt;

&lt;h2&gt;
  
  
  What Time to First Token Actually Is
&lt;/h2&gt;

&lt;p&gt;When you send a prompt, the model does two very different jobs.&lt;/p&gt;

&lt;p&gt;First it reads everything you sent. Every token of your system prompt, your chat history, your retrieved documents, your question. It runs all of that through the network in one big forward pass and builds an internal memory of it. This step is called prefill, and nothing comes back to you while it runs.&lt;/p&gt;

&lt;p&gt;Then it starts writing, one token at a time. That second phase is where you see text stream in.&lt;/p&gt;

&lt;p&gt;TTFT covers the first phase plus everything around it: the network trip to the provider, any time your request spends waiting in a queue, and the prefill itself. Here is the part that matters. On any prompt that is not tiny, prefill dominates, and the more tokens you send, the more &lt;a href="https://clickhouse.com/resources/engineering/llm-inference-latency" rel="noopener noreferrer"&gt;prefill work&lt;/a&gt; the model does before it can start. A short prompt starts answering quickly. A prompt stuffed with 8,000 tokens of context makes the reader sit and watch the cursor.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwbbzan004icvor2ke7cu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwbbzan004icvor2ke7cu.png" alt="Illustration of the path from a prompt to the first token" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is a Good Time to First Token?
&lt;/h2&gt;

&lt;p&gt;It depends on what the app is for.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://handbook.modular.com/llm-inference-basics/llm-inference-metrics" rel="noopener noreferrer"&gt;Modular LLM Inference Handbook&lt;/a&gt; puts it simply: a chatbot usually needs TTFT under 500 milliseconds to feel responsive, and a code-completion tool needs it below 100 milliseconds or it feels broken. A batch job that emails a summary once a day can take five seconds and nobody cares.&lt;/p&gt;

&lt;p&gt;So before you optimize anything, decide which of these you are. A chat UI and a background report have completely different budgets, and chasing 100ms on a nightly job is wasted effort.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgn4jgcpvibw4m8c9otbc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgn4jgcpvibw4m8c9otbc.png" alt="Illustration comparing TTFT targets for chat, code completion, and batch work" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Fixes, Biggest Wins First
&lt;/h2&gt;

&lt;p&gt;Here is the order I actually work through. The early steps cost you almost nothing and pay back the most. The later ones matter, but only after the cheap wins are done. Do not skip to the bottom.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiueoiqvsei4npgygwcfr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiueoiqvsei4npgygwcfr.png" alt="Illustration of TTFT improvements ranked from measurement to routing" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Measure TTFT on its own, before you touch anything
&lt;/h3&gt;

&lt;p&gt;You cannot cut a number you are not looking at. And most teams are not looking at this one, because their logging records total response time and stops there. Total time hides TTFT completely, since a slow first token and a slow-typing model both show up as one big number.&lt;/p&gt;

&lt;p&gt;So split them. Log the time from request-sent to first-token-received as its own field, separate from the time to finish. Do it at the p50 and the p99, because the average will lie to you and the tail is where users rage-quit.&lt;/p&gt;

&lt;p&gt;This is not theory for me. At Zonko Labs I built the internal tool that captured the data logs for Luffy, an AI co-worker that lived in Slack, and turned them into reports on latency and probable slowdowns. The first thing that tool taught everyone was that our "the bot is slow" complaints were almost always a first-token problem, not a typing-speed problem. You only learn that if you measure the two apart.&lt;/p&gt;

&lt;p&gt;If your traffic runs through an AI gateway, you can track this without building custom timers in your app. For example, &lt;a href="https://www.getmaxim.ai/bifrost" rel="noopener noreferrer"&gt;Bifrost&lt;/a&gt; exports a metric called &lt;a href="https://docs.getbifrost.ai/features/observability/prometheus" rel="noopener noreferrer"&gt;&lt;code&gt;bifrost_stream_first_token_latency_seconds&lt;/code&gt;&lt;/a&gt;. It measures TTFT automatically across OpenAI, Anthropic, or any other provider, so you get clean p50 and p99 charts right out of the box.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Shorten the prompt
&lt;/h3&gt;

&lt;p&gt;This is the biggest lever you fully control, and it is the one no tool can pull for you.&lt;/p&gt;

&lt;p&gt;Prefill scales with input length, so every token you delete is prefill work the model never has to do. Look hard at what you are actually sending. Most apps ship a system prompt that grew over months and never got trimmed, a full chat history when the last four turns would do, and RAG retrieval that dumps ten documents when three answer the question.&lt;/p&gt;

&lt;p&gt;Cut the system prompt to what the model needs, not what makes you feel safe. Truncate or summarize old history instead of resending all of it. Retrieve fewer, better chunks. None of this is glamorous, and all of it moves TTFT more than the fancy stuff.&lt;/p&gt;

&lt;p&gt;Be honest about the tradeoff, though. Cutting context can cost you answer quality, so trim, then check your evals, then trim again. This is the one place where going too far actually hurts, so it earns the care.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Cache the stable part of your prompt
&lt;/h3&gt;

&lt;p&gt;Most of what you send is the same every single call. The system prompt does not change. The tool definitions do not change. Only the user's question at the end changes. Prefill does not know that, so it re-reads the whole thing every time, unless you turn on caching.&lt;/p&gt;

&lt;p&gt;There are two flavors, and they are not the same thing.&lt;/p&gt;

&lt;p&gt;Prompt caching, also called prefix caching, is offered by the model providers themselves. It stores the processed form of a stable prefix so the next call reuses that work instead of redoing it, which is what &lt;a href="https://llm-d.ai/blog/kvcache-wins-you-can-see" rel="noopener noreferrer"&gt;prefix caching&lt;/a&gt; buys you. Put the parts that never change at the very front of your prompt and the changing question at the end, so the cache has the longest possible prefix to hit.&lt;/p&gt;

&lt;p&gt;Semantic caching is the bigger swing. Instead of caching the prompt's prefill, it caches the whole answer, and serves it when a new question means the same thing as an old one. On a hit, you skip the model entirely, so TTFT collapses from "prefill a big prompt" to "look up a vector," which is milliseconds. I wrote a &lt;a href="https://www.swapnoneel.site/blog/what-is-semantic-caching" rel="noopener noreferrer"&gt;full explainer on semantic caching&lt;/a&gt; if you want the mechanics.&lt;/p&gt;

&lt;p&gt;Let me make the caching point concrete with a small test you can predict.&lt;/p&gt;

&lt;p&gt;Call A sends a 3,000-token system prompt that never changes, then a 40-token question at the end. Call B folds the same information into a fresh 3,000-token prompt that gets reworded on every request, then the same 40-token question. Both send roughly the same number of tokens. On the hundredth request, which one has the lower TTFT?&lt;/p&gt;

&lt;p&gt;Call A wins, and it is not close. Call A's 3,000-token prefix is identical each time, so after the first call it lives in the prompt cache and its prefill is skipped. Call B looks different to the cache every time, so it pays the full prefill on all one hundred calls. Same token count, wildly different first-token time, decided entirely by whether the front of your prompt holds still. That is the whole game with caching: keep the front stable.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Stream the response, and know its one limit
&lt;/h3&gt;

&lt;p&gt;Streaming sends each token to the user the instant it is generated, instead of waiting for the full answer. Turn it on. Every serious chat UI streams, and it is a one-line change in most SDKs.&lt;/p&gt;

&lt;p&gt;Now the honest part, because it trips people up. &lt;strong&gt;Does streaming reduce TTFT?&lt;/strong&gt; No. Streaming does not make the first token arrive any sooner, since prefill still has to finish first. What it changes is everything after that first token, and it changes how fast the whole thing &lt;em&gt;feels&lt;/em&gt;. People read a streaming reply as faster than a non-streaming one even when the total time is identical, which is why &lt;a href="https://codeant.ai/blogs/ai-first-token-latency" rel="noopener noreferrer"&gt;faster first tokens&lt;/a&gt; matter more than total speed.&lt;/p&gt;

&lt;p&gt;So streaming is a perception win, not a TTFT win, and the two stack. Streaming only pays off if TTFT is already short, because a three-second blank screen followed by a burst still reads as three seconds of nothing. Fix TTFT first, then stream to make the rest feel quick.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Keep the connection warm and close
&lt;/h3&gt;

&lt;p&gt;Some of your TTFT is pure plumbing, and it hides where nobody looks.&lt;/p&gt;

&lt;p&gt;Every fresh HTTPS connection pays for a TLS handshake before a single byte of your prompt moves. If your app opens a new connection per request, you are paying that tax every time. Reuse connections with keep-alive and connection pooling so the handshake happens once, not on every call.&lt;/p&gt;

&lt;p&gt;Distance costs you too. A request from a server in Mumbai to a model endpoint in Virginia spends real milliseconds just crossing the planet and back, twice, before prefill even starts. Pick a provider region near your users where you can, and cut the number of hops between your app and the model. If you are on a serverless setup, watch for cold starts, since a function spinning up from zero can add more delay than the model does.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. Route around slow and rate-limited providers
&lt;/h3&gt;

&lt;p&gt;Here is the failure mode that ruins TTFT and has nothing to do with your prompt. The provider is having a bad minute.&lt;/p&gt;

&lt;p&gt;When a provider is overloaded, your request sits in its queue, and queue time is part of TTFT. When you hit a rate limit, you get a 429 and your client backs off and retries, which stacks seconds onto the first token before it ever appears. Neither of these is fixed by shortening your prompt, because the model never even started.&lt;/p&gt;

&lt;p&gt;The fix is to not be stuck with one provider. Send the request to the fastest healthy option, and when one starts failing or crawling, fall over to another automatically instead of retrying into the same wall. This is &lt;a href="https://www.swapnoneel.site/blog/what-is-adaptive-load-balancing" rel="noopener noreferrer"&gt;adaptive load balancing&lt;/a&gt; applied to model traffic, and it is the difference between one bad provider-minute taking down your p99 and your users never noticing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where a Gateway Does This For You
&lt;/h2&gt;

&lt;p&gt;Let me be honest about what a gateway cannot do first.&lt;/p&gt;

&lt;p&gt;A gateway cannot shorten your prompt. It cannot make a provider's GPU run faster. If your app only calls one model and your prompt is already lean, you do not need a gateway for TTFT. You will get much bigger gains just by trimming your system prompt.&lt;/p&gt;

&lt;p&gt;Where an AI gateway helps is when you want to handle caching, retries, and monitoring across multiple services without rebuilding them in every backend repo. The gateway sits between your application and your model providers as a reverse proxy, so you wire those fixes once in your infrastructure instead of in your app code.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Floc2kyys6pjispqljjzk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Floc2kyys6pjispqljjzk.png" alt="Illustration of an AI gateway coordinating caching, metrics, and fallbacks" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I have used &lt;a href="https://www.getmaxim.ai/bifrost" rel="noopener noreferrer"&gt;Bifrost&lt;/a&gt;, an open-source AI gateway built in Go, and it handles these levers in one place:&lt;/p&gt;

&lt;p&gt;First, it gives you &lt;strong&gt;TTFT metrics out of the box&lt;/strong&gt;. Measuring first-token latency manually usually means writing custom streaming wrappers in Python or Node.js to timestamp the first chunk. Bifrost measures &lt;code&gt;bifrost_stream_first_token_latency_seconds&lt;/code&gt; at the proxy layer for every request, giving you clean p50 and p99 charts across all your providers without extra client telemetry.&lt;/p&gt;

&lt;p&gt;Second, it provides &lt;a href="https://docs.getbifrost.ai/features/semantic-caching" rel="noopener noreferrer"&gt;semantic caching&lt;/a&gt;. If a user asks a question that matches a previous one, Bifrost serves the answer straight from cache in about 10 to 15 milliseconds. That turns an 800-millisecond prefill wait into almost nothing. Plus, cache writes happen asynchronously in the background, so cache misses do not add any latency to your stream.&lt;/p&gt;

&lt;p&gt;Third, it runs &lt;strong&gt;automatic fallbacks&lt;/strong&gt;. When a provider is overloaded or returns a 429 rate limit, your request sits in a queue and your TTFT spikes. With Bifrost, you can set fallback rules: if your primary model hangs or errors, it instantly shifts the request to a healthy backup provider instead of leaving your user staring at an empty screen.&lt;/p&gt;

&lt;p&gt;Finally, there is the question of &lt;strong&gt;proxy overhead&lt;/strong&gt;. Adding a gateway puts an extra hop in your network path. If that hop adds 30 or 50 milliseconds, it defeats the entire purpose of optimizing latency. Bifrost is written in Go and adds around 20 microseconds of overhead under load, meaning the hop itself is practically invisible.&lt;/p&gt;

&lt;p&gt;You still have to fix your prompts first. But once your prompts are clean, putting a gateway in front of your models lets you handle caching, fallbacks, and latency tracking through simple configuration instead of messy glue code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start Here
&lt;/h2&gt;

&lt;p&gt;If you do nothing else from this whole post, do three things, in this order.&lt;/p&gt;

&lt;p&gt;Measure TTFT as its own metric, so you stop flying blind. Shorten and cache your prompt, because that is where the biggest, cheapest wins live. Then stream, so the rest of the answer feels fast.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd49u0m6wsdsmpjjkqytw.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fd49u0m6wsdsmpjjkqytw.png" alt="Illustration of the three steps to reduce TTFT" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The truth most infrastructure posts bury is that for an app that calls an API, time to first token is mostly about prompt length and caching, not about exotic GPU tricks you will never touch. The GPU tricks are real, but they belong to the people running the models. Your levers are the prompt you send, the cache in front of it, and the route it takes. Those you own completely.&lt;/p&gt;

&lt;p&gt;Start by logging your p50 and p99 first-token time this week. I would bet the number surprises you, and I would bet a fat system prompt is the reason.&lt;/p&gt;

&lt;p&gt;If you have a TTFT horror story, or a fix that worked that I left out, tell me. I am &lt;a href="https://x.com/swapnoneel123" rel="noopener noreferrer"&gt;@swapnoneel123&lt;/a&gt; on X, and I write more of these on &lt;a href="https://www.swapnoneel.site" rel="noopener noreferrer"&gt;my site&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>ai</category>
      <category>machinelearning</category>
      <category>performance</category>
    </item>
    <item>
      <title>What Is an MCP Registry? Discovering &amp; Governing MCP Servers</title>
      <dc:creator>Swapnoneel Saha</dc:creator>
      <pubDate>Thu, 03 Sep 2026 17:01:44 +0000</pubDate>
      <link>https://dev.to/swapnoneel123/what-is-an-mcp-registry-discovering-governing-mcp-servers-2dd3</link>
      <guid>https://dev.to/swapnoneel123/what-is-an-mcp-registry-discovering-governing-mcp-servers-2dd3</guid>
      <description>&lt;p&gt;Most people consider an MCP registry as "npm for MCP servers." It is a nice line to explain the concept, but it is partially true and extremely half-baked.&lt;/p&gt;

&lt;p&gt;An MCP registry is a searchable catalog of MCP servers. It stores metadata, like a server's name, where to install it, and how a client should connect, so an AI agent or a developer can find servers to plug in. But here is the part the npm comparison hides: a registry is a discovery layer, not a trust layer. It tells you a server exists. It does not tell you the server is safe, or that your agents should be allowed to run it.&lt;/p&gt;

&lt;p&gt;That gap between "I found it" and "I trust it" is where teams get burned. Discovering MCP servers and governing MCP servers are two different jobs, and a registry only does the first one. So let's get both straight.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If you want to know more about what an MCP server is and how it works, I wrote a &lt;a href="https://www.swapnoneel.site/blog/what-is-an-mcp-server" rel="noopener noreferrer"&gt;beginner's guide to MCP servers&lt;/a&gt; first. This post picks up where that one left off.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What an MCP Registry Actually Is
&lt;/h2&gt;

&lt;p&gt;An MCP registry is a metadata catalog. It does not hold the server's code. It holds a small record that describes the server and points to wherever the code actually lives.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcefrhqor7m9l0cjfli72.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcefrhqor7m9l0cjfli72.png" alt="A searchable MCP registry catalog with server metadata cards." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This trips people up, so let me separate the two clearly. Your code lives on a package registry: npm, PyPI, Docker Hub. That is where the actual bits get downloaded from. The MCP registry sits one level up and just stores a pointer. For example, a &lt;code&gt;weather-mcp&lt;/code&gt; package can live on npm, and the MCP registry entry maps "weather v1.2.0" to &lt;code&gt;npm:weather-mcp&lt;/code&gt;. The registry is the index card. npm is the shelf.&lt;/p&gt;

&lt;p&gt;Each entry follows a standard shape called &lt;code&gt;server.json&lt;/code&gt;. It records the server's unique name, where to find it (an npm package name, or a remote server URL), how to run it (command-line args, environment variables), and some discovery data like a description and what the server can do. That is it. A registry is a stack of these &lt;code&gt;server.json&lt;/code&gt; cards, plus an API to search them.&lt;/p&gt;

&lt;p&gt;And that is genuinely useful. Before any of this, finding an MCP server meant scrolling GitHub READMEs and random Discord links. A registry turns that into one queryable list. But notice what a stack of index cards can and cannot do for you, because that is where the trouble starts.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Official MCP Registry
&lt;/h2&gt;

&lt;p&gt;The one most people mean by "the MCP registry" is the &lt;a href="https://registry.modelcontextprotocol.io" rel="noopener noreferrer"&gt;official MCP Registry&lt;/a&gt;, which launched in preview on &lt;a href="https://blog.modelcontextprotocol.io/posts/2025-09-08-mcp-registry-preview/" rel="noopener noreferrer"&gt;8 September 2025&lt;/a&gt;. It is open source, and it is backed by a serious group: Anthropic, GitHub, PulseMCP, Microsoft, and others.&lt;/p&gt;

&lt;p&gt;Its job is deliberately narrow. In the maintainers' own framing, it is a centralized metadata repository, and the metadata it stores is intentionally unopinionated. It does not rank servers, and it does not review or rate them. It just holds the &lt;code&gt;server.json&lt;/code&gt; records and serves them over a REST API.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvkj4odc0zcuit2mxa3mi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvkj4odc0zcuit2mxa3mi.png" alt="The official MCP Registry shown as a shared metadata source." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Here is the design choice that matters most, and almost nobody mentions it. The official registry is not really built for your app to read directly. It expects to be consumed by "downstream aggregators", which pull its data on a schedule (say, once an hour) and then layer their own curation, ratings, or search on top. The registry is the wholesale source of truth. The nice front-ends you actually browse are meant to be built by other people.&lt;/p&gt;

&lt;p&gt;And in practice, as of 2026, most MCP clients still do not read the registry directly as a built-in source. So for now you usually meet it through one of those aggregators, not through your agent going and querying it live.&lt;/p&gt;

&lt;p&gt;Two more things worth knowing early. First, the official registry does not host private servers, so an internal &lt;code&gt;mcp.acme-corp.internal&lt;/code&gt; server has no place there. Second, it is in preview, which means data can reset and the shape can still change. So treat it as a fast-moving foundation, not a finished product.&lt;/p&gt;

&lt;h2&gt;
  
  
  MCP Registry vs Marketplace vs Gateway
&lt;/h2&gt;

&lt;p&gt;This is the question I see confused most, so let me draw the lines hard. These are three different things doing three different jobs.&lt;/p&gt;

&lt;p&gt;A &lt;strong&gt;registry&lt;/strong&gt; is discovery in its rawest form. It answers one question: what servers exist, and how do I connect to them? It is a passive list.&lt;/p&gt;

&lt;p&gt;A &lt;strong&gt;marketplace&lt;/strong&gt; is discovery made pleasant. Sites like &lt;a href="https://smithery.ai" rel="noopener noreferrer"&gt;Smithery&lt;/a&gt;, Glama, and mcp.so give you search, categories, install buttons, and community ratings. Increasingly they read from the official registry underneath and add their own polish on top. A marketplace is a registry with a nice storefront and an opinion.&lt;/p&gt;

&lt;p&gt;A &lt;strong&gt;gateway&lt;/strong&gt; is something else entirely. It is not about finding servers, it is about controlling them at runtime. A gateway sits in the traffic path between your agents and your MCP servers, and it enforces the rules every time a call is made: who is allowed to use this tool, how often, at what cost, with what data. A registry is a directory. A gateway is a checkpoint.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjujtawbghe3ufx7n2t90.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjujtawbghe3ufx7n2t90.png" alt="A registry directory, marketplace storefront, and runtime gateway checkpoint." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The one line to remember: a registry without a gateway gives you discovery but no runtime governance. Your agents can find tools, but nothing stops an unauthorized call once they connect. Hold onto that, because it is the seam the rest of this post walks through.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Discovery Actually Works
&lt;/h2&gt;

&lt;p&gt;So how does a client go from "I need a GitHub server" to a running connection? Through the registry API and the &lt;code&gt;server.json&lt;/code&gt; record, and the interesting part is how it proves a server is who it claims to be.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxy6p4mdhykuv6j4v0j2x.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxy6p4mdhykuv6j4v0j2x.png" alt="MCP server discovery from verified namespace to connection." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Server names use a reverse-DNS format: &lt;code&gt;io.github.username/server-name&lt;/code&gt;, or &lt;code&gt;com.example/server&lt;/code&gt;. That name is not decorative. It ties the server to a namespace that the publisher had to prove they own, through GitHub or DNS verification. If you publish under &lt;code&gt;io.github.swpn0neel/*&lt;/code&gt;, the registry made you show you actually control that GitHub account first. So a name is a small ownership claim, verified at publish time.&lt;/p&gt;

&lt;p&gt;Here is the piece that makes the whole thing extensible. The registry ships an &lt;a href="https://modelcontextprotocol.io/registry/about" rel="noopener noreferrer"&gt;open OpenAPI spec&lt;/a&gt;, which means anyone can stand up their own registry that speaks the exact same interface. A public marketplace can implement it. A company can implement a private one for internal servers. Any MCP client that already knows how to talk to the official registry can talk to those too, with no extra work. One shape, many registries.&lt;/p&gt;

&lt;p&gt;That design is genuinely clever, and it is also exactly why governance cannot live at this layer. A standard that lets anyone publish, and lets anyone run a compatible registry, is a standard optimized for reach, not for safety. Openness and vetting pull in opposite directions, and the registry picked openness on purpose.&lt;/p&gt;

&lt;h2&gt;
  
  
  Discovery Is the Easy Half
&lt;/h2&gt;

&lt;p&gt;So that is discovery, more or less solved. You have a standard catalog, verified namespaces, an API, and a way to run your own copy.&lt;/p&gt;

&lt;p&gt;But discovery was always the easy half of the title. The hard half is governing what those servers actually do once your agents start calling them, and that is a problem a catalog structurally cannot touch. Let me show you why with one concrete example.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz4az02iq3frlhjsrlfga.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz4az02iq3frlhjsrlfga.png" alt="A discovery catalog ending before runtime governance begins." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What a Registry Does Not Protect You From
&lt;/h2&gt;

&lt;p&gt;Say you tell your coding agent, "add the Slack MCP server." It searches, finds a package named &lt;code&gt;slack-mcp-server&lt;/code&gt;, sees an install command that looks completely official, and wires it in. Feels routine. Now walk that one action through each layer and watch where it does, and does not, get stopped.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq3ck5pb0l8brtywr3jt6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fq3ck5pb0l8brtywr3jt6.png" alt="A verified registry entry beside a hidden security risk." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Did the registry catch it?&lt;/strong&gt; No, and this is by design. The official registry &lt;a href="https://modelcontextprotocol.io/registry/about" rel="noopener noreferrer"&gt;delegates security scanning&lt;/a&gt; to the underlying package registries and to downstream aggregators. Its own job is namespace authentication and metadata hosting. So namespace verification proves the publisher owns the name they used. It does not prove the code behind that name is safe. A registry entry is a verified business card, not a background check.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;So typosquatting still works.&lt;/strong&gt; This is not hypothetical. Security researchers have already documented &lt;a href="https://www.speakeasy.com/resources/mcp-tool-poisoning" rel="noopener noreferrer"&gt;look-alike PyPI packages&lt;/a&gt; like &lt;code&gt;slack-mcp-server&lt;/code&gt; and &lt;code&gt;slack-mcp-server-v2&lt;/code&gt;, whose install commands look official but are not. A name that reads as legitimate is not the same as a source you vetted, and a catalog optimized for discovery will happily list both.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;And the tool itself can lie.&lt;/strong&gt; The nastier class is tool poisoning: a malicious server hides instructions inside a tool's description, so when the model reads "what can you do?", it also reads "and quietly send recent messages to this URL." The model sees a normal tool. The registry saw a normal &lt;code&gt;server.json&lt;/code&gt;. Neither one inspected behavior, because inspecting behavior is not what either one does.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Then there is the rug pull.&lt;/strong&gt; A server can be perfectly clean the day it is approved and turn malicious three versions later. A registry records the new version. It does not re-judge it. This exact risk is why security researchers keep arguing that tool updates should &lt;a href="https://arxiv.org/pdf/2506.01333" rel="noopener noreferrer"&gt;trigger re-review&lt;/a&gt; instead of being trusted automatically.&lt;/p&gt;

&lt;p&gt;None of this means the registry is broken. It means the registry is doing its one job, discovery, and quietly handing you a job it was never built to do: deciding what to trust, and enforcing that decision on every call. That job has to live somewhere else.&lt;/p&gt;

&lt;h2&gt;
  
  
  Do You Need a Private MCP Registry?
&lt;/h2&gt;

&lt;p&gt;For a lot of teams, the first real answer to the trust problem is a private registry. And it genuinely helps, so let me be fair to it before I tell you where it stops.&lt;/p&gt;

&lt;p&gt;A private, internal registry is an allowlist. Instead of letting agents pull from the open catalog, your platform team curates a list of vetted servers, mixes in your own internal ones, and points every agent at that. Because it implements the same OpenAPI shape, your existing clients keep working. This is the model enterprise registry products (Kong's, Portkey's, and others) are built around, and for regulated or multi-team setups it is close to mandatory, because a community catalog was never meant to carry your compliance requirements.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj8r4azb2ts8ztrotkbki.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj8r4azb2ts8ztrotkbki.png" alt="A private registry as a curated internal MCP server allowlist." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So a private registry moves the trust decision to a place you control. That is a real upgrade over pulling from the open internet.&lt;/p&gt;

&lt;p&gt;But look closely at when it acts. A private registry governs discovery time: what a developer or agent is allowed to install and see. It does not govern runtime: what an approved server actually does on each individual call. Once an agent is connected to a vetted server, the registry is out of the loop. It cannot cap how much that agent spends, cannot say "this key may read files but never delete them," and cannot stop a poisoned tool description that slipped through your review. It curated the shelf. It is not standing at the counter.&lt;/p&gt;

&lt;p&gt;You still need something in the traffic path.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the Gateway Comes In
&lt;/h2&gt;

&lt;p&gt;Here is the honest part first, because it decides whether you should even read this section. If you are running one or two MCP servers you trust, you do not need a gateway. A registry, or even a hardcoded config, is plenty. A gateway earns its place when you have many servers, many agents, and real consequences if one of them misbehaves. And a gateway cannot vet server code for you either, so nothing here removes the need to review what you run.&lt;/p&gt;

&lt;p&gt;For a few months I was the first layer of testing for an AI co-worker that lived inside Slack, catching its mistakes before the real users did. The lesson that stuck with me is simple: you cannot govern what you only hear about after the fact. Governing has to happen at runtime, in the path, while the call is being made.&lt;/p&gt;

&lt;p&gt;What a gateway adds is the one thing a registry structurally cannot: enforcement on every call, at runtime. That is where &lt;a href="https://www.getmaxim.ai/bifrost" rel="noopener noreferrer"&gt;Bifrost&lt;/a&gt;, an &lt;a href="https://github.com/maximhq/bifrost" rel="noopener noreferrer"&gt;open-source AI gateway&lt;/a&gt; built by Maxim AI, fits the picture. It can act as an &lt;a href="https://docs.getbifrost.ai/mcp/gateway" rel="noopener noreferrer"&gt;MCP gateway&lt;/a&gt;: you connect all your MCP servers to Bifrost, and your agents connect to Bifrost instead of to each server directly. One endpoint in front of everything, sitting exactly where the checks need to happen.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fifs86uh0q37jducspypi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fifs86uh0q37jducspypi.png" alt="An MCP gateway controlling calls between agents and servers." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Once you are in that position, three things change.&lt;/p&gt;

&lt;p&gt;First, access stops being all-or-nothing. Bifrost filters tools per virtual key, so a given key only sees the tools you explicitly listed for it. You configure it plainly: a client named &lt;code&gt;filesystem&lt;/code&gt;, with &lt;code&gt;tools_to_execute&lt;/code&gt; set to &lt;code&gt;["read_file", "list_directory"]&lt;/code&gt;, and that key can do exactly those two things and nothing else. It is deny-by-default. The read-only agent literally cannot see a delete tool, so a poisoned or over-eager call has nothing to grab.&lt;/p&gt;

&lt;p&gt;Second, the token problem gets solved instead of tolerated. Remember how a registry lets you connect dozens of servers? Once you do, every request stuffs every tool's description into the model's context, and with 150-plus tools the model burns most of its budget just reading the menu. Bifrost's &lt;a href="https://docs.getbifrost.ai/mcp/code-mode" rel="noopener noreferrer"&gt;Code Mode&lt;/a&gt; fixes this by exposing four generic meta-tools and letting the model write code to discover and run what it needs on demand, instead of pre-loading every definition. In their own benchmark of 508 tools across 16 servers, that cut input tokens by up to 92.8%, from 75.1M tokens down to 5.4M, while keeping a perfect pass rate. That is a vendor benchmark on a vendor harness, so weigh it as one, but the mechanism is sound and the direction is not subtle.&lt;/p&gt;

&lt;p&gt;Third, control lives in one place. All your keys and connections sit behind the gateway instead of scattered across config files, and you can attach &lt;a href="https://docs.getbifrost.ai/features/governance/virtual-keys" rel="noopener noreferrer"&gt;budgets and rate limits&lt;/a&gt; to each virtual key, so an agent that goes haywire hits a spending cap instead of your invoice. This is the runtime governance the registry handed off. The registry told your agents what exists. The gateway decides what is allowed, every single time.&lt;/p&gt;

&lt;p&gt;So the two are not competitors. A registry and a gateway are the two halves the title promised, and they sit at different points in the flow: the registry at discovery, the gateway at runtime.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Do You Actually Need?
&lt;/h2&gt;

&lt;p&gt;Let me put it as plainly as I can, because the whole post comes down to one distinction.&lt;/p&gt;

&lt;p&gt;A registry answers "what exists?" A gateway answers "what is allowed?" You will eventually want both, and the mistake I see most is assuming that a shiny catalog, public or private, has solved the second question. It has not. It was never trying to.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvhiq57fueu4ba0bk6m4d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvhiq57fueu4ba0bk6m4d.png" alt="A practical path from discovery to runtime governance." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So here is the path I would actually take. If you are just exploring, use the official MCP registry or a marketplace to find good servers, and keep the count small. If you are a team standing this up for real, add a private registry so you control what gets discovered in the first place. And the moment agents are making real calls with real credentials and real budgets, put a gateway in the traffic path, because that is the only layer that can enforce a decision at the moment the decision matters.&lt;/p&gt;

&lt;p&gt;Discovery is a catalog problem. Governance is a runtime problem. Solve them in the right places and MCP stops feeling scary. Try to solve governance with a directory, and you will find out the hard way that an index card never stopped anybody from doing anything.&lt;/p&gt;

&lt;p&gt;I run coding agents like Claude Code every day, and my own MCP config has quietly grown from one server to a small pile without me really deciding to. That drift is exactly how this bites you, one harmless-looking addition at a time. So if you have hit the same sprawl, I would genuinely like to hear how you are handling it. Drop a comment, or come find me on &lt;a href="https://x.com/swapnoneel123" rel="noopener noreferrer"&gt;X&lt;/a&gt; and let's compare setups.&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>webdev</category>
      <category>ai</category>
      <category>security</category>
    </item>
    <item>
      <title>What Is an MCP Server? A Beginner's Guide</title>
      <dc:creator>Swapnoneel Saha</dc:creator>
      <pubDate>Wed, 02 Sep 2026 18:48:39 +0000</pubDate>
      <link>https://dev.to/swapnoneel123/what-is-an-mcp-server-a-beginners-guide-305k</link>
      <guid>https://dev.to/swapnoneel123/what-is-an-mcp-server-a-beginners-guide-305k</guid>
      <description>&lt;p&gt;An MCP server is just a small program that gives an AI model access to tools and data it cannot reach on its own.&lt;/p&gt;

&lt;p&gt;In plain words, it is a bridge. Through one shared standard called the Model Context Protocol, it lets an AI read your files, query a database, or call an API without you having to write custom glue code every single time. It can be a tiny script running on your laptop, or a service running in the cloud.&lt;/p&gt;

&lt;p&gt;None of this is as complicated as the acronym makes it sound. And it is worth getting right, because MCP is quickly becoming the default way AI agents interact with the real world. So, let's dive in!&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is an MCP Server
&lt;/h2&gt;

&lt;p&gt;An MCP server is a program that exposes three things to an AI model: tools it can run, data it can read, and prompt templates it can reuse. Simply put, the model connects to the server, asks "what have you got?", and then uses whatever comes back. Strip away the jargon and an MCP server is a standard adapter between one AI and one outside system.&lt;/p&gt;

&lt;p&gt;The "server" word trips people up, so let me kill that confusion first. It does not mean a rack in a data center. A filesystem MCP server can be a tiny script on your own laptop. It is called a server only because it serves requests, the same way your local dev server does when you run &lt;code&gt;npm run dev&lt;/code&gt;. Small program, answers questions, done.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1byxqr8e49h575iwempp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1byxqr8e49h575iwempp.png" alt="A small program running on a laptop opens a compact drawer of tools, readable data, and reusable prompts." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;And that is really the core of what people mean by MCP server meaning: it is the thing on the other end of the connection that knows how to do actual work, like reading a file or querying a database, and knows how to describe that work in a format the model understands.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where MCP Came From and Why It Caught On
&lt;/h2&gt;

&lt;p&gt;The Model Context Protocol is an open standard for connecting AI models to outside tools and data. Anthropic &lt;a href="https://www.anthropic.com/news/model-context-protocol" rel="noopener noreferrer"&gt;released it in November 2024&lt;/a&gt;, and within a year it went from a niche idea to something Claude, ChatGPT, Cursor, and VS Code all speak.&lt;/p&gt;

&lt;p&gt;Here is the problem it fixed. Before MCP, if you wanted an AI to use some tool, you wrote the glue yourself, every single time. A couple of years back at Keploy I built a RAG chatbot over their docs, wiring the vector embeddings and the retrieval together by hand. That was one data source and one integration, fully custom. Now picture doing that again for GitHub, then Slack, then your database, then for a different model that expects a different shape. It does not scale, and everyone was rebuilding the same connectors.&lt;/p&gt;

&lt;p&gt;The official docs describe MCP as &lt;a href="https://modelcontextprotocol.io/introduction" rel="noopener noreferrer"&gt;a USB-C port for AI&lt;/a&gt;. One connector shape, and anything can plug into anything. You build an MCP server once, and every client that speaks the protocol can use it. That is the entire reason it caught on so fast: it turned a pile of one-off integrations into one standard everyone could share.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdojewl5j5fu98gzgh3yb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdojewl5j5fu98gzgh3yb.png" alt="Incompatible custom connectors resolve into a row of identical standard plugs." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How Does an MCP Server Actually Work?
&lt;/h2&gt;

&lt;p&gt;An MCP server works through three roles: a host, a client, and a server. The host is the app you are actually using, like Claude Desktop or Cursor. For each server it wants to talk to, the host spins up one client, and that client keeps a dedicated line open to one MCP server. Messages go back and forth as &lt;a href="https://www.jsonrpc.org/specification" rel="noopener noreferrer"&gt;JSON-RPC 2.0&lt;/a&gt;, which is a plain, boring, well-understood format, and that is a good thing.&lt;/p&gt;

&lt;p&gt;The server can offer three kinds of thing, and it is worth knowing them by name:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Tools&lt;/strong&gt;: actions the model can run, like "read this file", "run this query", or "create this GitHub issue".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Resources&lt;/strong&gt;: data the model can read for context, like a file's contents or a database schema. Read-only, no side effects.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prompts&lt;/strong&gt;: ready-made templates for common tasks, so the model does not start from scratch every time.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The flow itself is simple. The client connects, then asks the server to list what it has (a &lt;code&gt;tools/list&lt;/code&gt; call), then runs one of them (a &lt;code&gt;tools/call&lt;/code&gt;). In JSON it looks about like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;//&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"What can you do?"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"jsonrpc"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2.0"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"method"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"tools/list"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="err"&gt;//&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Okay, do this one."&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"jsonrpc"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2.0"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"method"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"tools/call"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"params"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"get_weather"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"arguments"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"city"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Kolkata"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjmu8tsl3mdvf13g2aavo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjmu8tsl3mdvf13g2aavo.png" alt="A client inside an application exchanges a capability menu, tool call, and result with a server." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;One more thing worth knowing early: a server runs in one of two places. A local server talks over &lt;code&gt;stdio&lt;/code&gt; (standard input and output), which is just two programs on the same machine passing text back and forth, no network involved. A remote server talks over HTTP, so it can live anywhere and serve lots of clients at once. Same protocol either way, which is the point.&lt;/p&gt;

&lt;h2&gt;
  
  
  MCP Server vs API: What's the Difference
&lt;/h2&gt;

&lt;p&gt;This is the question I get most, so let me be direct. An MCP server usually sits on top of a normal API. It is not a replacement for one. The difference is who the thing is built for.&lt;/p&gt;

&lt;p&gt;A regular API is built for a developer who already read the docs and hardcoded the exact calls. It assumes someone knows the endpoints ahead of time. An MCP server is built for a model that shows up knowing nothing, asks what is available, and figures out which tool to use on its own. That runtime discovery is the whole trick, and a plain API does not do it.&lt;/p&gt;

&lt;p&gt;So when people ask about MCP vs API, the honest answer is that most MCP servers are a thin wrapper around an API you could have called yourself. What MCP adds is a standard way to describe the tools, so any model can find them and use them without a human wiring each call. The API does the work. The MCP server makes that work legible to an AI.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff4z8fcn2qfyng9uhr9bg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff4z8fcn2qfyng9uhr9bg.png" alt="An API executes fixed calls; an MCP wrapper makes those same capabilities discoverable to an AI." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What Can You Actually Do With One
&lt;/h2&gt;

&lt;p&gt;Plenty, and you probably already have the pieces. The most common MCP servers connect a model to everyday systems: a filesystem server for reading and writing files, a GitHub server for issues and pull requests, a Postgres server for querying a database, a Slack server for messages. Anthropic keeps a set of &lt;a href="https://github.com/modelcontextprotocol/servers" rel="noopener noreferrer"&gt;reference servers&lt;/a&gt; on GitHub, and most companies now ship an official one for their own product.&lt;/p&gt;

&lt;p&gt;I run Claude Code every day, and that is where most people meet their first MCP server. You bolt one onto a coding agent or a desktop chat app, and suddenly it can touch your files, search the web, or open a pull request without you copy-pasting anything. The MCP tools it offers just show up as things the model can now do.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk5edu92gtx19brfig7so.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk5edu92gtx19brfig7so.png" alt="Four practical capabilities: edit files, work with repositories, query data, and send messages." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The pattern is always the same. Find a server for the system you care about, point your client at it, and the model gains a new skill. One server is genuinely a five-minute setup, and it feels a little bit like magic the first time an agent edits a real file on your disk!&lt;/p&gt;

&lt;h2&gt;
  
  
  Do You Need an MCP Gateway
&lt;/h2&gt;

&lt;p&gt;Here is the honest part first: if you are running one or two MCP servers, you do not need a gateway. It solves a problem you do not have yet, so skip this section and go play with a server instead.&lt;/p&gt;

&lt;p&gt;But things change once the count goes up. Say you have got ten servers wired into your agents. Three problems show up fast. First, every request dumps the full tool list of every server into the model's context, so with 150-plus tools the model burns most of its budget just reading the menu. Second, each server has its own keys and auth, scattered across config files with no single place to rotate them. And third, there is no shared way to say who is allowed to call what, or to cap spending. That is tool sprawl, and it turns a neat setup into a liability.&lt;/p&gt;

&lt;p&gt;An MCP gateway fixes this by sitting right in the middle. Think of it like a power strip. Instead of plugging ten different cables into ten wall sockets, your AI agent connects to the gateway once. The gateway handles the rest.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcl4107emyz3timcqn3ve.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcl4107emyz3timcqn3ve.png" alt="Growing server connections share one controlled gateway with centralized credentials, permissions, and budgets." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is where &lt;a href="https://www.getmaxim.ai/bifrost" rel="noopener noreferrer"&gt;Bifrost&lt;/a&gt;, an &lt;a href="https://github.com/maximhq/bifrost" rel="noopener noreferrer"&gt;open-source AI gateway&lt;/a&gt; built by Maxim AI, comes in. It can &lt;a href="https://docs.getbifrost.ai/mcp/gateway" rel="noopener noreferrer"&gt;act as an MCP gateway&lt;/a&gt;, tying all your servers together so your AI talks to one neat endpoint instead of juggling ten.&lt;/p&gt;

&lt;p&gt;It also fixes that token problem in a really clever way. Remember how 100 tools eat up your context? Normally, the model has to re-read every single tool's description on every prompt just to know what's on the menu. That burns tokens fast.&lt;/p&gt;

&lt;p&gt;Bifrost fixes this using &lt;a href="https://docs.getbifrost.ai/mcp/code-mode" rel="noopener noreferrer"&gt;Code Mode&lt;/a&gt;. Instead of dumping every tool definition into the prompt, it gives the model four generic tools and lets it write code to find and run what it needs on demand. In their benchmarks, that cut input tokens by up to 92.8%. That is not a small difference, that is basically your whole bill! Very similar to the cut-costs-at-the-gateway idea I talked about with &lt;a href="https://www.swapnoneel.site/blog/what-is-semantic-caching" rel="noopener noreferrer"&gt;semantic caching&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;On top of that, you get one central place for safety and control. All your API keys live in one spot instead of being scattered across random config files. You can set &lt;a href="https://docs.getbifrost.ai/features/governance/virtual-keys" rel="noopener noreferrer"&gt;per-key tool filtering and budgets&lt;/a&gt;, meaning you decide exactly which agent can use which tools and how much it can spend. And by default, tools do not run automatically unless you explicitly turn that on. That last part matters a lot, because an agent running tools silently without your permission is a bad day waiting to happen.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where to Start
&lt;/h2&gt;

&lt;p&gt;You do not need to read a massive spec or write complex code to get started with MCP. The whole idea is small once the acronym stops scaring you: it is just a standard plug that lets your AI reach outside tools.&lt;/p&gt;

&lt;p&gt;If you want to try it out today, do not overthink it. Pick an app you already use, like Claude Desktop, Cursor, or Claude Code. Then grab just one pre-built server, like the filesystem or GitHub server. Wiring it into your config file takes about five minutes.&lt;/p&gt;

&lt;p&gt;Once it is running, give the model a real prompt: ask it to inspect a local folder or check a repo. The moment you see an agent reach outside its chat box and touch a real file on your machine, the whole protocol clicks.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvfgkcti6rc3vx7ojy4pz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvfgkcti6rc3vx7ojy4pz.png" alt="Choose an app, connect one filesystem server, and complete one real task before expanding." width="800" height="451"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;And my advice would be to keep it tiny. You do not need a gateway or fancy routing when you are only running one or two servers. Play with one first, see how it feels, and only put a gateway like Bifrost in front when you have a dozen servers and your token bill starts climbing.&lt;/p&gt;

&lt;p&gt;If you set up your first MCP server after reading this, I would genuinely like to hear how it went. Drop it in the comments, or come say hi on &lt;a href="https://x.com/swapnoneel123" rel="noopener noreferrer"&gt;X&lt;/a&gt; and we can chat about it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions (FAQs)
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What is an MCP server in simple terms?&lt;/strong&gt; It is a small program that gives an AI model a set of tools and data through a shared standard, so the model can read a file, query a database, or call an API without custom glue code. Think of it as a standard adapter between one AI and one outside system.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is the difference between an MCP client and an MCP server?&lt;/strong&gt; They are the two ends of one connection. The server offers tools, data, and prompts. The client, created by an app like Claude Desktop, connects to the server and uses what it offers. One app can run many clients, one per server it talks to.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkpair6erwwqme1azp965.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkpair6erwwqme1azp965.png" alt="A compact field guide clarifies the adapter, client/server relationship, permissions, and building a small server." width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Are MCP servers safe to use?&lt;/strong&gt; They are as safe as how you run them. A single local server you trust is low risk. The danger shows up when you run many unvetted servers with their own credentials and let an agent run tools automatically, which is exactly why keeping auto-execution off and routing servers through a gateway with tool filtering is worth doing once you scale.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How do you build your own MCP server?&lt;/strong&gt; You use one of the official SDKs, define a few tools with their inputs, and expose them over &lt;code&gt;stdio&lt;/code&gt; or HTTP. Start tiny, with one or two tools, get it working with a client you already use, and grow from there instead of trying to wrap an entire API on day one.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>beginners</category>
      <category>ai</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Fixing Claude Rate Limit Errors with an AI Gateway</title>
      <dc:creator>Swapnoneel Saha</dc:creator>
      <pubDate>Wed, 26 Aug 2026 18:04:49 +0000</pubDate>
      <link>https://dev.to/swapnoneel123/fixing-claude-rate-limit-errors-with-an-ai-gateway-5dk5</link>
      <guid>https://dev.to/swapnoneel123/fixing-claude-rate-limit-errors-with-an-ai-gateway-5dk5</guid>
      <description>&lt;p&gt;Ever had a &lt;code&gt;429&lt;/code&gt; from Claude that just would not clear, no matter how patiently you backed off?&lt;/p&gt;

&lt;p&gt;Yeah. That one was never a rate limit.&lt;/p&gt;

&lt;p&gt;A Claude rate limit error is an HTTP &lt;code&gt;429&lt;/code&gt; carrying &lt;code&gt;"type": "rate_limit_error"&lt;/code&gt;, and it means your organization crossed a per-minute request ceiling, a per-minute token ceiling, or its monthly spend cap. Which one decides the fix, and the response body tells you which before you write any retry code.&lt;/p&gt;

&lt;p&gt;And the difference matters far more than it sounds, because two of those three conditions get strictly worse the moment you retry them.&lt;/p&gt;

&lt;h2&gt;
  
  
  The five failures behind one Claude rate limit error
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;rate_limit_error&lt;/code&gt; is not one condition. Anthropic returns the same string for several different situations, and the correct response may be to wait a few seconds or to change a setting that retrying cannot fix.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe5joinddh3t7dxjgioei.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fe5joinddh3t7dxjgioei.png" alt="Diagnostic flow for five Claude API failure types" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://platform.claude.com/docs/en/api/errors" rel="noopener noreferrer"&gt;Claude API errors reference&lt;/a&gt; documents them separately, so you can tell them apart from the response alone.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What you see&lt;/th&gt;
&lt;th&gt;What it actually is&lt;/th&gt;
&lt;th&gt;Does retrying help?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;429&lt;/code&gt; + &lt;code&gt;rate_limit_error&lt;/code&gt; + a &lt;code&gt;retry-after&lt;/code&gt; header&lt;/td&gt;
&lt;td&gt;A real per-minute limit. You went too fast.&lt;/td&gt;
&lt;td&gt;Yes, after &lt;code&gt;retry-after&lt;/code&gt; seconds&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;429&lt;/code&gt; + &lt;code&gt;rate_limit_error&lt;/code&gt; + &lt;code&gt;details.error_code: enforced_spend_limit_reached&lt;/code&gt;, and no &lt;code&gt;retry-after&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Your usage tier's monthly spend cap&lt;/td&gt;
&lt;td&gt;No. Access returns at 00:00 UTC on the 1st&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;400&lt;/code&gt; + &lt;code&gt;invalid_request_error&lt;/code&gt;, message starting &lt;code&gt;You have reached your specified API usage limits&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;A spend limit you set yourself in the Console&lt;/td&gt;
&lt;td&gt;No. Raise or remove your own limit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;429&lt;/code&gt; right after a sharp traffic ramp&lt;/td&gt;
&lt;td&gt;An acceleration limit, not your steady-state limit&lt;/td&gt;
&lt;td&gt;Partly. Ramp up gradually instead&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;529&lt;/code&gt; + &lt;code&gt;overloaded_error&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Anthropic is busy. Nothing to do with you.&lt;/td&gt;
&lt;td&gt;Yes, but the queue is global&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That last row is the one that trips people up most. The Claude overloaded error gets searched more than any specific 429 phrasing, and almost every guide lumps it in with rate limits. It isn't yours. A &lt;code&gt;529&lt;/code&gt; means the API is temporarily overloaded across all users, so no amount of tier upgrading or key rotation on your side moves it.&lt;/p&gt;

&lt;p&gt;The second row is easy to misdiagnose. It looks identical to a normal rate limit at the status-code level, the official SDKs will happily auto-retry it, and every retry fails. The tell is the missing &lt;code&gt;retry-after&lt;/code&gt; header plus that &lt;code&gt;error_code&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;So before anything else, log the full error body. Not &lt;code&gt;err.status&lt;/code&gt;. The body.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;anthropic&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;anthropic&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Anthropic&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;msg&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-sonnet-5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;hi&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="n"&gt;anthropic&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;APIStatusError&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status_code&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                          &lt;span class="c1"&gt;# 429
&lt;/span&gt;    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;retry-after&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;  &lt;span class="c1"&gt;# None means spend cap
&lt;/span&gt;    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                                 &lt;span class="c1"&gt;# the error_code lives here
&lt;/span&gt;    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;request_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                           &lt;span class="c1"&gt;# req_018Ee... for support
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;retry-after&lt;/code&gt; check on that second line is doing almost all of the diagnostic work. Present means wait. Absent on a &lt;code&gt;429&lt;/code&gt; means stop retrying and go look at billing.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually counts against your Claude rate limits
&lt;/h2&gt;

&lt;p&gt;Now, the per-minute limits. There are three of them running at once, per model, and you can trip any one.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3tnu0xx83flpiei2xx6f.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3tnu0xx83flpiei2xx6f.png" alt="Claude rate limit token buckets and prompt caching" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;RPM&lt;/strong&gt;, requests per minute.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ITPM&lt;/strong&gt;, input tokens per minute.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OTPM&lt;/strong&gt;, output tokens per minute.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;They are enforced per model class, so Sonnet traffic and Haiku traffic draw from separate buckets and can run at full speed at the same time. Worth knowing: on the current lineup, Sonnet 5 has its own bucket while Sonnet 4.6 and 4.5 share a combined one, and the same split applies to Opus 5 versus the Opus 4.x family. Pinning half your traffic to an older Sonnet does not double your Sonnet 5 headroom.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://platform.claude.com/docs/en/api/rate-limits" rel="noopener noreferrer"&gt;rate limits documentation&lt;/a&gt; also says something that should change how you think about bursts. The limiter is a token bucket, which means capacity refills continuously instead of resetting on a clean minute boundary. A 60 RPM limit behaves much more like 1 request per second than like 60 free requests at the top of every minute. So if you fan out 60 concurrent calls at 12:00:00 you will get rate limited, even though your average for that minute was exactly at the ceiling.&lt;/p&gt;

&lt;p&gt;I hit this exact shape of problem building &lt;a href="https://scholarian.vercel.app" rel="noopener noreferrer"&gt;Scholarian&lt;/a&gt;, which fetched and ranked over 10,000 academic papers. The moment you write a loop that fans out one model call per document, your average throughput looks perfectly fine on a dashboard and your p99 latency (the slowest 1% of calls) is a wall of 429s, because nothing in that loop is pacing itself.&lt;/p&gt;

&lt;h3&gt;
  
  
  The one number most people get wrong
&lt;/h3&gt;

&lt;p&gt;For most Claude models, only uncached input tokens count against ITPM.&lt;/p&gt;

&lt;p&gt;Specifically: &lt;code&gt;input_tokens&lt;/code&gt; and &lt;code&gt;cache_creation_input_tokens&lt;/code&gt; count, and &lt;code&gt;cache_read_input_tokens&lt;/code&gt; does not. So a cache hit is free as far as your rate limit is concerned, and it is billed at a reduced rate on top of that. (Claude Haiku 3.5 is the exception and does count cache reads.)&lt;/p&gt;

&lt;p&gt;Do the arithmetic on that, because it is not a small effect. Say you are on the Start tier with 2,000,000 ITPM on Sonnet 5, and you are running a coding agent that ships roughly 180,000 tokens of context per turn.&lt;/p&gt;

&lt;p&gt;Without caching, 2,000,000 divided by 180,000 is about &lt;strong&gt;11 turns per minute&lt;/strong&gt; for your entire organization.&lt;/p&gt;

&lt;p&gt;With prompt caching where 170,000 of those tokens are cache reads, only 10,000 counts against ITPM. That is &lt;strong&gt;200 turns per minute&lt;/strong&gt; off the exact same limit! Anthropic's own worked example frames it the same way, and it is why they publish your cache hit rate on the Usage page right next to the rate limit charts.&lt;/p&gt;

&lt;p&gt;One nuance people miss: &lt;code&gt;cache_creation_input_tokens&lt;/code&gt; does count. Writing the cache costs you full ITPM once, and reading it is free. So a workload that keeps rebuilding the cache because the prefix keeps shifting gets the worst of both.&lt;/p&gt;

&lt;p&gt;Output is simpler. OTPM is measured on tokens actually produced, in real time, and &lt;code&gt;max_tokens&lt;/code&gt; never factors into it. There is genuinely no rate limit penalty for setting a generous &lt;code&gt;max_tokens&lt;/code&gt;, so stop shaving it down for that reason.&lt;/p&gt;

&lt;h2&gt;
  
  
  The headers you should be reading instead of guessing
&lt;/h2&gt;

&lt;p&gt;Every response, including the successful ones, carries your current standing. And most people never look at them.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffri34c0gmiht98fr2wsa.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffri34c0gmiht98fr2wsa.png" alt="Monitoring Claude rate-limit response headers" width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sS&lt;/span&gt; &lt;span class="nt"&gt;-D&lt;/span&gt; - &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null https://api.anthropic.com/v1/messages &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"x-api-key: &lt;/span&gt;&lt;span class="nv"&gt;$ANTHROPIC_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"anthropic-version: 2023-06-01"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"content-type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"model":"claude-sonnet-5","max_tokens":16,
       "messages":[{"role":"user","content":"hi"}]}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run that and you get back a block of &lt;code&gt;anthropic-ratelimit-*&lt;/code&gt; headers. The ones worth wiring into a metric:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;anthropic-ratelimit-requests-remaining&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;anthropic-ratelimit-input-tokens-remaining&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;anthropic-ratelimit-output-tokens-remaining&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;the matching &lt;code&gt;-reset&lt;/code&gt; fields, which are RFC 3339 timestamps rather than durations&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Two gotchas here. The &lt;code&gt;remaining&lt;/code&gt; values are rounded to the nearest thousand, so treat them as a gauge and not as an accountant. And the generic &lt;code&gt;anthropic-ratelimit-tokens-*&lt;/code&gt; triplet reports whichever limit is currently most restrictive, which means the number can jump between input and output accounting between requests without anything being wrong.&lt;/p&gt;

&lt;p&gt;Identify the failure and read the remaining-capacity headers before changing retry logic. Backoff cannot fix a spend cap, and tuning it first wastes time.&lt;/p&gt;

&lt;p&gt;At one of my previous roles I built an internal tool that pulled the data logs off an AI product and turned them into latency and slowdown reports, and the useful part was never the clever part. It was just having the numbers somewhere you could look at them without reproducing the bug first. Same idea here. Export those three headers as gauges and most rate limit debugging stops being detective work.&lt;/p&gt;

&lt;h2&gt;
  
  
  Retrying correctly, and where retrying stops helping
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgcolh0nckwd3kwh0ef46.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgcolh0nckwd3kwh0ef46.png" alt="Retry backoff with jitter and hard ceilings" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The rule in plain English: if the response gave you a &lt;code&gt;retry-after&lt;/code&gt;, wait exactly that long. If it did not, double your wait each attempt and add randomness so your workers do not all wake up at the same instant.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;random&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;anthropic&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;anthropic&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Anthropic&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;max_retries&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# we handle it ourselves
&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;call_with_retry&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;attempt&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="n"&gt;anthropic&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;APIStatusError&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status_code&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;429&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;502&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;503&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;504&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;529&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
                &lt;span class="k"&gt;raise&lt;/span&gt;
            &lt;span class="n"&gt;hinted&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;retry-after&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status_code&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;429&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;hinted&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="k"&gt;raise&lt;/span&gt;  &lt;span class="c1"&gt;# spend cap, so retrying is pointless
&lt;/span&gt;            &lt;span class="n"&gt;wait&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;float&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;hinted&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;hinted&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt; &lt;span class="n"&gt;attempt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;wait&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;random&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;uniform&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mf"&gt;0.8&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;1.2&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;RuntimeError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;exhausted retries&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The jitter multiplier on the last line is the bit people skip, and it is the bit that matters at concurrency. Without it, 50 workers that all got rate limited at the same moment will all retry at the same moment, and you get a thundering herd that reproduces the original 429 on a fixed schedule forever.&lt;/p&gt;

&lt;p&gt;Also notice &lt;code&gt;max_retries=0&lt;/code&gt; at the top. The official SDKs retry twice by default with their own backoff, honoring &lt;code&gt;retry-after&lt;/code&gt;. That default is fine for a script. It is not fine when you are also running your own retry loop, because the two layers multiply and you quietly end up with 3x the attempts you thought you configured.&lt;/p&gt;

&lt;p&gt;But here is the honest limit of all of this, and it is why the post does not end here.&lt;/p&gt;

&lt;p&gt;Retries are a queueing strategy. They smooth out bursts against a ceiling you are near. They do nothing at all when you are structurally above the ceiling, because every retry is just another request competing for the same bucket. If your steady-state demand is 3,000,000 input tokens per minute and your limit is 2,000,000, no backoff curve in the world fixes that. You need more ceiling.&lt;/p&gt;

&lt;h2&gt;
  
  
  Setup is done, now the part that actually raises the ceiling
&lt;/h2&gt;

&lt;p&gt;Everything above happens inside one API key against one endpoint. Reading the error properly and then caching aggressively will genuinely resolve most Claude rate limit errors, and if you are a solo developer on one key, honestly stop reading here and go turn on prompt caching. You will get more out of that one afternoon than out of any infrastructure.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frzzk1yyuceuqbn0e324l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frzzk1yyuceuqbn0e324l.png" alt="Four independent sources of Claude API capacity" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;But if you are past that point, the interesting fact is that "Claude capacity" is not one pool. It is four, and they are metered by different organizations running different quota systems.&lt;/p&gt;

&lt;h3&gt;
  
  
  Four separate places Claude capacity lives
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;1. The direct Claude API, by usage tier.&lt;/strong&gt; Start, Build, Scale, then Custom. Each tier carries both per-minute limits and a monthly spend cap, and as of mid-2026 those caps sit at 500USD on Start, 1,000USD on Build, and 200,000USD on Scale. Tiers move up automatically as you build usage history, and you can also request an increase from the Console.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Amazon Bedrock.&lt;/strong&gt; Claude on Bedrock runs against &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/quotas.html" rel="noopener noreferrer"&gt;AWS service quotas&lt;/a&gt;, which are per-account and per-region and adjustable through the Service Quotas console. These have no relationship to your Anthropic tier at all. Bedrock even splits its own quotas across two inference endpoints, tracked separately for the same underlying model.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Google's platform.&lt;/strong&gt; Claude models run there through the model garden against Google Cloud quotas, again on a completely separate meter.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Priority Tier.&lt;/strong&gt; Anthropic sells committed capacity with its own &lt;code&gt;anthropic-priority-*&lt;/code&gt; headers and its own bucket, sitting alongside your standard limits rather than replacing them.&lt;/p&gt;

&lt;p&gt;The consequence is the interesting part. An organization pinned at 100% of its direct-API ITPM might have an entirely idle Bedrock quota sitting in two regions. The capacity exists, and it is already paid for. The problem is that your application has one base URL and one key, so it cannot reach any of it.&lt;/p&gt;

&lt;p&gt;One clarification before moving on, because this is the most common mix-up of all. If your error reads &lt;code&gt;API Error: rate limit reached&lt;/code&gt; inside Claude Code on a Pro or Max plan, none of this section applies to you, because that is a subscription meter and not the API. Two different systems, two different clocks, and I wrote about &lt;a href="https://www.swapnoneel.site/blog/claude-pro-usage-limits" rel="noopener noreferrer"&gt;Claude Pro usage limits&lt;/a&gt; separately.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where a gateway changes the math
&lt;/h2&gt;

&lt;p&gt;This is the job an &lt;a href="https://www.getmaxim.ai/bifrost" rel="noopener noreferrer"&gt;AI gateway&lt;/a&gt; does, which you will also see called an LLM gateway or an LLM proxy. It is a proxy that speaks every provider's API shape, sits between your code and the model, and holds the routing decisions your application shouldn't have to.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjgxccsp0dbkrtjsmsbur.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjgxccsp0dbkrtjsmsbur.png" alt="AI gateway routing across keys and providers" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Let me lead with what it cannot do, because that part gets oversold constantly. A gateway does not raise Anthropic's ceiling. Your Start tier is still your Start tier. If you run one Anthropic key through a gateway and change nothing else, you will hit exactly the same 429 at exactly the same token count, plus a few microseconds of hop.&lt;/p&gt;

&lt;p&gt;What it changes is how many buckets a single request can reach, and how fast it gives up on a bad one. &lt;a href="https://docs.getbifrost.ai/overview" rel="noopener noreferrer"&gt;Bifrost&lt;/a&gt; is Maxim AI's high-performance, &lt;a href="https://github.com/maximhq/bifrost" rel="noopener noreferrer"&gt;open-source AI gateway&lt;/a&gt; that unifies access to 20+ providers through a single OpenAI-compatible API. Its &lt;a href="https://docs.getbifrost.ai/features/retries-and-fallbacks" rel="noopener noreferrer"&gt;retry and fallback layer&lt;/a&gt; distinguishes upstream failures from credential failures: it retries &lt;code&gt;5xx&lt;/code&gt; and network errors against the same key, rotates keys for &lt;code&gt;429&lt;/code&gt;, &lt;code&gt;401&lt;/code&gt;, &lt;code&gt;402&lt;/code&gt;, and &lt;code&gt;403&lt;/code&gt; responses, and moves to the next provider only after the current provider's retry budget is exhausted. That behavior maps directly to the different Claude failure modes in this post.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pooling keys so one bucket is not the whole story
&lt;/h3&gt;

&lt;p&gt;Rate limits are set at the organization level, so a second organization with its own key is a second bucket. A gateway can treat several keys as one logical pool.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"providers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"anthropic"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"keys"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"primary"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="nl"&gt;"value"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"env.ANTHROPIC_API_KEY"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;   &lt;/span&gt;&lt;span class="nl"&gt;"models"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"*"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"weight"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.7&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"overflow"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"value"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"env.ANTHROPIC_API_KEY_2"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"models"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"*"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"weight"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.3&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Traffic splits 70/30 by weight. But the part that matters for 429s is what happens on failure. Bifrost rotates keys when the failure is bound to the credential rather than to the request, which it defines as &lt;code&gt;429&lt;/code&gt;, &lt;code&gt;401&lt;/code&gt;, &lt;code&gt;403&lt;/code&gt; and &lt;code&gt;402&lt;/code&gt;. A rate limited key gets marked used for that cycle and the request carries on against another key. Auth and billing failures mark the key dead for the remainder of that request. Once every key has been tried, the set resets and a fresh weighted round begins.&lt;/p&gt;

&lt;p&gt;Key selection is scored rather than round robin, using recent error rates, latency, and observed rate limit hits, with weights recomputed on a short interval and penalties decaying once a key recovers. That is &lt;a href="https://www.swapnoneel.site/blog/what-is-adaptive-load-balancing" rel="noopener noreferrer"&gt;adaptive load balancing&lt;/a&gt; applied to credentials instead of servers, and it matters here because a key that just returned a &lt;code&gt;429&lt;/code&gt; is a bad destination for the next 30 seconds specifically, and not permanently.&lt;/p&gt;

&lt;h3&gt;
  
  
  Falling back across channels, not just keys
&lt;/h3&gt;

&lt;p&gt;Key pooling only helps if you have more keys. Falling back across the four capacity pools helps even when you don't.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"anthropic/claude-sonnet-5"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"messages"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"user"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"fallbacks"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"bedrock/anthropic.claude-sonnet-4-5-20250929-v1:0"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="s2"&gt;"vertex/claude-sonnet-4-5@20250929"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same request, same Anthropic model family, three completely independent quota systems. If the direct API is capped, the call lands on Bedrock's per-region quota, which knows nothing about your Anthropic tier.&lt;/p&gt;

&lt;p&gt;The execution detail worth knowing is that each provider in the chain gets its own full retry budget. A primary at &lt;code&gt;max_retries: 3&lt;/code&gt; with two fallbacks also at 3 means up to 12 attempts on one call, so set your client timeout with that in mind. &lt;a href="https://docs.getbifrost.ai/features/retries-and-fallbacks" rel="noopener noreferrer"&gt;Retries and fallbacks&lt;/a&gt; documents the backoff as &lt;code&gt;min(initial × 2^attempt, max) × jitter(0.8-1.2)&lt;/code&gt;, with defaults of 500ms initial and a 5000ms cap.&lt;/p&gt;

&lt;p&gt;And a real gotcha, stated plainly because it will bite you: &lt;code&gt;max_retries&lt;/code&gt; defaults to 0. Retries are off until you turn them on.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"network_config"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"max_retries"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"retry_backoff_initial"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"retry_backoff_max"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;5000&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Two more things that buy headroom
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Caching at the gateway instead of per service.&lt;/strong&gt; If four services ask the same question, an application-level cache inside each one misses four times. A gateway cache hits three of them, and every hit is a request that never touches your ITPM at all. Bifrost's version only engages when a request carries an &lt;code&gt;x-bf-cache-key&lt;/code&gt; header, defaults to a 5 minute time-to-live and a 0.8 cosine similarity threshold, and can run in exact-match mode with no embedding provider at all. That threshold is the whole product, and I went into why in a post on &lt;a href="https://www.swapnoneel.site/blog/what-is-semantic-caching" rel="noopener noreferrer"&gt;semantic caching&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stopping one tenant from eating the org bucket.&lt;/strong&gt; Your organization limit is shared, so one runaway batch job starves the interactive traffic sitting behind it. Governance limits let you cap consumption per virtual key before the request ever reaches Anthropic, in tokens or in requests, with calendar-aligned reset periods running from &lt;code&gt;1m&lt;/code&gt; up to &lt;code&gt;1Y&lt;/code&gt;. When a cap trips, the gateway reports the breach with its own code instead of a vague 429: &lt;code&gt;token_limited&lt;/code&gt; or &lt;code&gt;request_limited&lt;/code&gt; on a &lt;code&gt;429&lt;/code&gt;, &lt;code&gt;budget_exceeded&lt;/code&gt; on a &lt;code&gt;402&lt;/code&gt;, and &lt;code&gt;403&lt;/code&gt; for a blocked model or provider. The &lt;a href="https://docs.getbifrost.ai/features/governance" rel="noopener noreferrer"&gt;governance docs&lt;/a&gt; carry the full mapping.&lt;/p&gt;

&lt;p&gt;That is a real difference in debuggability. "Your batch job hit its token cap" is a fix you can act on in a minute. "Anthropic said 429" is an afternoon.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pointing Claude Code at it
&lt;/h3&gt;

&lt;p&gt;Since a lot of people arriving at this error are inside a coding agent rather than a Python service, the integration is two environment variables. Bifrost exposes an Anthropic-compatible endpoint, so the client never knows anything changed.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;ANTHROPIC_BASE_URL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;http://localhost:8080/anthropic
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;ANTHROPIC_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;dummy-key
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The path is &lt;code&gt;/anthropic&lt;/code&gt;, and not &lt;code&gt;/v1/anthropic&lt;/code&gt;. And the key can be a placeholder because the real credentials live inside the gateway.&lt;/p&gt;

&lt;p&gt;Gateway traffic is billed per token against API keys. It uses a separate meter from Pro or Max subscriptions, not a larger subscription allowance.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a gateway will not fix
&lt;/h2&gt;

&lt;p&gt;Four things, because a post that only lists the wins is a brochure.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fauxl6crqc7hdkz3hl8ik.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fauxl6crqc7hdkz3hl8ik.png" alt="Limits an AI gateway cannot solve" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It does not help with a &lt;code&gt;529&lt;/code&gt;. An overloaded error is Anthropic-wide capacity, so a fallback to Bedrock might dodge it, but no amount of key pooling on the direct API will.&lt;/p&gt;

&lt;p&gt;It does not make a cross-provider fallback behave identically. Sonnet 5 on the direct API and Sonnet 4.5 on Bedrock are not the same model, and if your prompts are tuned tightly or you lean on a specific tool-calling shape, your fallback path needs its own evals. A fallback that silently produces worse answers is arguably worse than a clean 429.&lt;/p&gt;

&lt;p&gt;It does not fix a workload that is simply too large. If you need 5x your current ceiling permanently, the answer is a tier increase or committed capacity, and routing is the bridge that gets you there without downtime in the meantime.&lt;/p&gt;

&lt;p&gt;And the vendor benchmarks are vendor benchmarks. Bifrost's page claims 20 microseconds of added latency at 5,000 requests per second, and a 50x advantage on that same p99 number over LiteLLM. Those are Maxim's own numbers on Maxim's own harness, and you should treat them the way you treat every published benchmark, which is as a reason to go run your own.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you should actually do, in order
&lt;/h2&gt;

&lt;p&gt;Ordered by what each one costs you, cheapest first. Stop at the first one that fixes it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fap1us1416it1ujunxxrb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fap1us1416it1ujunxxrb.png" alt="Ordered checklist for resolving Claude rate limits" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One.&lt;/strong&gt; Log the full error body and check for &lt;code&gt;retry-after&lt;/code&gt;. If it is missing on a &lt;code&gt;429&lt;/code&gt;, you have a spend cap and no retry strategy will ever help you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Two.&lt;/strong&gt; Export the three &lt;code&gt;anthropic-ratelimit-*-remaining&lt;/code&gt; headers as gauges. You cannot tune what you cannot see.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Three.&lt;/strong&gt; Turn on prompt caching and check your hit rate on the Usage page. This is the single biggest win available and it costs you an afternoon, because cached reads do not count against ITPM on most models.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Four.&lt;/strong&gt; Add &lt;code&gt;retry-after&lt;/code&gt;-aware backoff with jitter, and make sure you are not stacking it on top of the SDK's own retries.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Five.&lt;/strong&gt; Move anything non-interactive to the Message Batches API, which has its own separate rate limits.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Six.&lt;/strong&gt; Request a tier increase. It is free and it is slow, so start it before you need it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Seven.&lt;/strong&gt; Only now, put a gateway in front and pool keys or channels. This is real infrastructure with real operational cost, and it is worth it once you have genuinely run out of ceiling, and mostly not before.&lt;/p&gt;

&lt;p&gt;The reason I would put it dead last rather than first is that most 429s I have run into were not capacity problems at all. They were a cache that never got turned on, or a retry loop with no jitter, or a spend cap that everybody kept retrying into. Fix those and the ceiling stops being the constraint.&lt;/p&gt;

&lt;p&gt;If you have hit a &lt;code&gt;429&lt;/code&gt; shape that doesn't fit any of the five rows in that table, drop it in the comments, I'd genuinely like to see it. And if you want more of this sort of thing, I'm on &lt;a href="https://x.com/swapnoneel123" rel="noopener noreferrer"&gt;X&lt;/a&gt; and I write everything up first on &lt;a href="https://www.swapnoneel.site" rel="noopener noreferrer"&gt;my site&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>programming</category>
      <category>webdev</category>
      <category>ai</category>
      <category>devops</category>
    </item>
  </channel>
</rss>
