<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Artifilog</title>
    <description>The latest articles on DEV Community by Artifilog (@artifilog).</description>
    <link>https://dev.to/artifilog</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4153775%2Fa6e08149-df45-4433-852e-ad100fb3243e.png</url>
      <title>DEV Community: Artifilog</title>
      <link>https://dev.to/artifilog</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/artifilog"/>
    <language>en</language>
    <item>
      <title>How to Build MCP Gateway from Scratch</title>
      <dc:creator>Artifilog</dc:creator>
      <pubDate>Thu, 01 Oct 2026 07:42:28 +0000</pubDate>
      <link>https://dev.to/artifilog/how-to-build-mcp-gateway-from-scratch-4857</link>
      <guid>https://dev.to/artifilog/how-to-build-mcp-gateway-from-scratch-4857</guid>
      <description>&lt;h1&gt;
  
  
  How to Build MCP Gateway from Scratch
&lt;/h1&gt;

&lt;p&gt;This guide covers how to build MCP gateway from scratch - a production middleware layer that routes, authenticates, rate-limits, and observably forwards AI agent tool calls. Covers TypeScript code, architecture decisions, and hardening for multi-tenant deployments.&lt;/p&gt;

&lt;p&gt;An MCP gateway is a middleware layer that sits between AI agents and the external tools, data sources, and services they need. It handles routing, authentication, rate limiting, protocol translation, and observability so agents can call tools without managing connections, credentials, and retries themselves.&lt;/p&gt;

&lt;p&gt;You would build one instead of stitching together ad-hoc clients because production AI systems need consistency, security, and debugging visibility that raw MCP client libraries do not provide. Off-the-shelf MCP server setups assume a single agent, a single tool, a single environment. Real deployments have multiple agents, dozens of tools, per-tenant access rules, and audit requirements. A gateway centralizes all of that.&lt;/p&gt;

&lt;p&gt;This post walks through the architecture, implementation patterns, and production hardening for building an MCP gateway from scratch. For background on Model Context Protocol fundamentals, see &lt;a href="https://modelcontextprotocol.io/" rel="noopener noreferrer"&gt;Anthropic's MCP documentation&lt;/a&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Build an MCP gateway to centralize routing, auth, rate limiting, and observability for AI agent tool calls. Start with TypeScript + MCP SDK, add a routing engine backed by a config store, implement JWT auth with JWKS caching, use Redis for rate limiting, add circuit breakers for resilience, and log everything for debugging.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What Is an MCP Gateway?
&lt;/h2&gt;

&lt;p&gt;The Model Context Protocol lets AI agents talk to external tools over a standardized transport. An MCP server exposes tools; an MCP client connects and calls them. An MCP gateway sits above both, intercepting and managing every request.&lt;/p&gt;

&lt;p&gt;Think of it as the API gateway pattern applied to the MCP layer. Just as Kong or Envoy sit between your services and the internet, an MCP gateway sits between your agents and your MCP servers. It does not replace either - it sits in the middle and adds:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Routing: which agent reaches which server, which tool&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Authentication: verifying the caller and the target&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Rate limiting: preventing one noisy agent from starving others&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Protocol translation: normalizing differences between MCP versions or transports&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Logging and metrics: capturing every tool call for audit and debugging&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Without a gateway, each agent carries its own connection logic, auth tokens, and retry code. When a tool changes auth requirements or throttles, you update every agent. With a gateway, you update one layer.&lt;/p&gt;

&lt;p&gt;If you are routing AI model requests across multiple providers, see &lt;a href="https://www.artifilog.com/posts/route-ai-requests-cheapest-model" rel="noopener noreferrer"&gt;How to Route AI Requests to the Cheapest Model&lt;/a&gt; for the routing patterns that translate directly to MCP tool routing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Core Architecture Components
&lt;/h2&gt;

&lt;p&gt;A production MCP gateway has five essential pieces. Each one can be swapped or scaled independently.&lt;/p&gt;

&lt;h3&gt;
  
  
  Transport Layer
&lt;/h3&gt;

&lt;p&gt;The transport layer accepts incoming MCP connections from agents and outbound connections to MCP servers. Most implementations use stdio or SSE transport per the &lt;a href="https://modelcontextprotocol.io/specification/2025-03-26/transport" rel="noopener noreferrer"&gt;MCP specification&lt;/a&gt;. The gateway needs to handle both simultaneously, managing connection pools for outbound calls.&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;interface TransportConfig {
  inbound: 'stdio' | 'sse' | 'streamable-http';
  outbound: 'stdio' | 'sse';
  maxConnections: number;
  idleTimeoutMs: number;
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;Connection pooling is critical here. Each agent-to-gateway and gateway-to-server connection consumes a file descriptor. Without &lt;code&gt;maxConnections&lt;/code&gt;, you hit OS limits under production load. The MCP SDK's &lt;code&gt;StreamableHttpTransport&lt;/code&gt; and &lt;code&gt;StdioTransport&lt;/code&gt; handle the wire protocol, but you configure the pool boundaries. Per &lt;a href="https://github.com/modelcontextprotocol/typescript-sdk" rel="noopener noreferrer"&gt;Anthropic's MCP SDK documentation&lt;/a&gt;, connection limits depend on your Node.js process ulimit settings.&lt;/p&gt;

&lt;h3&gt;
  
  
  Routing Engine
&lt;/h3&gt;

&lt;p&gt;The routing engine maps an incoming request to the correct MCP server and tool. Routes can be static (agent A always uses server X) or dynamic (based on tool name, tenant, or request metadata).&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;interface Route {
  agentId?: string;
  tenantId?: string;
  toolPattern: string;
  targetServer: string;
  fallbackServer?: string;
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;Route resolution happens on every request. For high-throughput deployments, compile the routing table into a trie or prefix tree instead of scanning an array. The &lt;code&gt;toolPattern&lt;/code&gt; field supports wildcard matching - &lt;code&gt;deploy-*&lt;/code&gt; matches &lt;code&gt;deploy-service&lt;/code&gt;, &lt;code&gt;deploy-config&lt;/code&gt;, etc.&lt;/p&gt;

&lt;h3&gt;
  
  
  Auth Middleware
&lt;/h3&gt;

&lt;p&gt;Auth middleware validates the caller's identity and checks permissions before forwarding the request. JWT verification, OAuth token exchange, and API key validation are common patterns. The gateway should also handle credential rotation without dropping connections.&lt;/p&gt;

&lt;p&gt;For multi-tenant deployments, each tenant may use a different identity provider. The gateway validates against multiple JWKS endpoints and caches the keys. Rotation happens transparently when the JWKS endpoint returns new keys.&lt;/p&gt;

&lt;h3&gt;
  
  
  Rate Limiter &amp;amp; Circuit Breaker
&lt;/h3&gt;

&lt;p&gt;Rate limiting protects backend tools from overload. A token-bucket or sliding-window approach per agent, per tenant, or per tool prevents any single caller from exhausting resources.&lt;/p&gt;

&lt;p&gt;The circuit breaker trips when a downstream server fails repeatedly, failing fast instead of queuing requests that will time out. For an AI agent fleet, this means one misbehaving agent cannot cascade into a full outage.&lt;/p&gt;

&lt;h3&gt;
  
  
  Observability Layer
&lt;/h3&gt;

&lt;p&gt;Logs, metrics, and traces for every tool call. Request duration, success/failure status, error types, and agent identity. This is non-negotiable for debugging why an agent failed in production at 3 AM.&lt;/p&gt;

&lt;p&gt;Structured JSON logs with request IDs let you trace a single agent session across multiple tool calls. OpenTelemetry spans show where latency concentrates - is the tool slow, or is the gateway overhead the bottleneck?&lt;/p&gt;

&lt;p&gt;For more on AI agent infrastructure patterns, see &lt;a href="https://www.artifilog.com/" rel="noopener noreferrer"&gt;Artifilog's coverage of AI development tools&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step-by-Step Build
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Step 1: Project Setup
&lt;/h3&gt;

&lt;p&gt;Start with TypeScript and the MCP SDK. A minimal gateway skeleton:&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;import { McpServer } from '@modelcontextprotocol/sdk/server/mcp.js';
import { StreamableHttpTransport } from '@modelcontextprotocol/sdk/server/streamableHttp.js';
import express from 'express';

const app = express();
const mcpServer = new McpServer({ name: 'gateway', version: '1.0.0' });

app.use('/mcp', async (req, res) =&amp;gt; {
  const transport = new StreamableHttpTransport();
  await mcpServer.connect(transport);
  await transport.handleRequest(req, res);
});

app.listen(3000, () =&amp;gt; console.log('MCP gateway running on :3000'));
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;The &lt;a href="https://github.com/modelcontextprotocol/typescript-sdk" rel="noopener noreferrer"&gt;MCP TypeScript SDK&lt;/a&gt; provides &lt;code&gt;McpServer&lt;/code&gt;, &lt;code&gt;StreamableHttpTransport&lt;/code&gt;, and the protocol message types. Pin to a specific version - the spec is still evolving.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Add the Routing Layer
&lt;/h3&gt;

&lt;p&gt;Define routes that map agent IDs and tool patterns to backend MCP servers. Load from config or a database for dynamic updates.&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;interface Route {
  agentId?: string;
  tenantId?: string;
  toolPattern: string;
  targetServer: string;
  fallbackServer?: string;
}

const routes: Route[] = [
  { agentId: 'agent-ops', toolPattern: 'deploy-*', targetServer: 'http://deploy-svc:8080' },
  { agentId: 'agent-ops', toolPattern: 'db-*', targetServer: 'http://db-tools:8081' },
  { agentId: '*', toolPattern: 'search-*', targetServer: 'http://search-svc:8082' },
];

function resolveRoute(agentId: string, toolName: string): Route | null {
  return routes.find(r =&amp;gt;
    (r.agentId === agentId || r.agentId === '*') &amp;amp;&amp;amp;
    new RegExp(r.toolPattern.replace('*', '.*')).test(toolName)
  ) || null;
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;For production, replace the array with a database-backed resolver. PostgreSQL works well - store routes in a table with priority ordering for overlapping patterns.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: Implement Auth Middleware
&lt;/h3&gt;

&lt;p&gt;Verify JWT tokens on inbound requests and exchange tokens for outbound calls. Cache validated tokens to avoid hitting the auth server every request.&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;import jwt from 'jsonwebtoken';

interface AgentIdentity {
  agentId: string;
  tenantId: string;
  scopes: string[];
}

async function authenticate(req: Request): Promise {
  const authHeader = req.headers.authorization;
  if (!authHeader?.startsWith('Bearer ')) return null;

  const token = authHeader.split(' ')[1];
  try {
    const decoded = jwt.verify(token, process.env.JWT_SECRET!) as AgentIdentity;
    return decoded;
  } catch {
    logger.warn({ token: token.slice(0, 8) + '...' }, 'JWT verification failed');
    return null;
  }
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;Always log auth failures with truncated tokens - never log full credentials. The JWKS cache should refresh periodically (e.g., every few minutes) or on cache miss with stale keys.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 4: Add Rate Limiting
&lt;/h3&gt;

&lt;p&gt;Use Redis for distributed rate limiting. Track token consumption per agent per minute.&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;import Redis from 'ioredis';

const redis = new Redis(process.env.REDIS_URL);
const LIMITS: Record = {
  'deploy-*': 10,
  'db-*': 50,
  'search-*': 100,
  '*': 30,
};

async function checkRateLimit(agentId: string, tool: string): Promise {
  const key = `rl:${agentId}:${tool}`;
  const current = await redis.incr(key);
  if (current === 1) await redis.expire(key, 60);
  return current &amp;lt;= LIMITS[tool] ?? LIMITS['*'];
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;The tiered limit approach protects backend services individually. A search tool gets a higher limit than a deploy tool because search is cheaper and faster. Adjust per your backend capacity.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 5: Wire Observability
&lt;/h3&gt;

&lt;p&gt;Emit structured logs and metrics for every forwarded request.&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;import { metrics, logger } from './observability';

async function forwardRequest(route: Route, payload: unknown) {
  const start = Date.now();
  const requestId = crypto.randomUUID();

  logger.info({ requestId, route, payloadSize: JSON.stringify(payload).length }, 'Forwarding MCP request');

  try {
    const response = await fetch(route.targetServer, {
      method: 'POST',
      headers: { 'Content-Type': 'application/json' },
      body: JSON.stringify(payload),
      signal: AbortSignal.timeout(parseInt(process.env.MCP_TIMEOUT_MS || '30000')),
    });

    const duration = Date.now() - start;
    metrics.histogram('mcp.request_duration_ms', duration, { tool: route.toolPattern });
    metrics.counter('mcp.request_success', 1, { agent: route.agentId });

    logger.info({ requestId, duration, status: response.status }, 'MCP request completed');
    return response;
  } catch (err) {
    metrics.counter('mcp.request_failure', 1, { tool: route.toolPattern, error: err.message });
    logger.error({ requestId, route, error: err.message }, 'MCP request failed');
    throw err;
  }
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;The &lt;code&gt;AbortSignal.timeout&lt;/code&gt; prevents hanging requests from consuming connections indefinitely. Set per-tool timeouts: 30 seconds for fast tools, 5 minutes for heavy ones.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 6: Add Fallback Routing
&lt;/h3&gt;

&lt;p&gt;When a target server is unreachable, retry on the fallback route if one exists.&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;async function callWithFallback(route: Route, payload: unknown) {
  try {
    return await forwardRequest(route, payload);
  } catch (err) {
    if (route.fallbackServer) {
      logger.warn({ route, error: err.message }, 'Primary failed, using fallback');
      return await forwardRequest(
        { ...route, targetServer: route.fallbackServer },
        payload
      );
    }
    throw err;
  }
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;Fallback routing matters for AI agent fleets. If your primary LLM provider goes down, agents should retry on a backup without the agent code needing to know.&lt;/p&gt;

&lt;h2&gt;
  
  
  Production Hardening
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Authentication at Scale
&lt;/h3&gt;

&lt;p&gt;In multi-tenant deployments, each tenant may use a different auth provider. The gateway needs to discover and validate tokens from multiple issuers. Cache JWKS endpoints to avoid fetching on every request. Rotate signing keys without downtime.&lt;/p&gt;

&lt;p&gt;For OAuth2 flows where the gateway exchanges tokens on behalf of agents, implement token refresh with exponential backoff. Store refresh tokens encrypted at rest.&lt;/p&gt;

&lt;h3&gt;
  
  
  Rate Limiting Strategies
&lt;/h3&gt;

&lt;p&gt;Global limits protect the gateway infrastructure. Per-tool limits protect individual backend services. Per-agent limits prevent abuse from noisy agents. Use a three-tier approach:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Soft limit (below cap) - log warning, allow&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Hard limit (at cap) - reject with 429&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Burst allowance (above cap briefly) - absorb short spikes&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Redis sorted sets work well for sliding-window rate limiting with millisecond precision.&lt;/p&gt;

&lt;h3&gt;
  
  
  Audit Logging
&lt;/h3&gt;

&lt;p&gt;Every tool call needs an immutable record: who called what, when, with what parameters, and what the result was. Store in PostgreSQL with partitioned tables by date. Retention requirements vary by compliance framework - common baselines range from 30 to 90 days.&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;CREATE TABLE mcp_audit_log (
  id BIGSERIAL,
  request_id UUID NOT NULL,
  agent_id TEXT NOT NULL,
  tenant_id TEXT NOT NULL,
  tool_name TEXT NOT NULL,
  parameters JSONB,
  status TEXT NOT NULL,
  duration_ms INTEGER,
  created_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
  PRIMARY KEY (id, created_at)
) PARTITION BY RANGE (created_at);
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;Partitioning by date keeps queries fast as the table grows. Drop old partitions for retention compliance without expensive deletes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Health Checks and Circuit Breakers
&lt;/h3&gt;

&lt;p&gt;Health checks on every backend MCP server. The circuit breaker should open after several consecutive failures, half-open after a cooldown period, and close after successful probes. Expose health endpoints for your orchestration layer.&lt;/p&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;interface CircuitBreakerState {&lt;br&gt;
  failures: number;&lt;br&gt;
  lastFailure: number;&lt;br&gt;
  state: 'closed' | 'open' | 'half-open';&lt;br&gt;
  cooldownMs: number;&lt;br&gt;
  successThreshold: number;&lt;br&gt;
}

&lt;p&gt;async function callWithCircuitBreaker(&lt;br&gt;
  route: Route,&lt;br&gt;
  payload: unknown,&lt;br&gt;
  breaker: CircuitBreakerState&lt;br&gt;
): Promise {&lt;br&gt;
  if (breaker.state === 'open') {&lt;br&gt;
    if (Date.now() - breaker.lastFailure &amp;gt; breaker.cooldownMs) {&lt;br&gt;
      breaker.state = 'half-open';&lt;br&gt;
      breaker.failures = 0;&lt;br&gt;
    } else {&lt;br&gt;
      throw new Error(&lt;code&gt;Circuit breaker open for ${route.targetServer}&lt;/code&gt;);&lt;br&gt;
    }&lt;br&gt;
  }&lt;/p&gt;

&lt;p&gt;if (breaker.state === 'half-open') {&lt;br&gt;
    breaker.failures++;&lt;br&gt;
    if (breaker.failures &amp;gt;= breaker.successThreshold) {&lt;br&gt;
      breaker.state = 'closed';&lt;br&gt;
      breaker.failures = 0;&lt;br&gt;
    }&lt;br&gt;
  }&lt;/p&gt;

&lt;p&gt;try {&lt;br&gt;
    const response = await forwardRequest(route, payload);&lt;br&gt;
    if (breaker.state === 'half-open') {&lt;br&gt;
      breaker.failures = 0;&lt;br&gt;
      breaker.state = 'closed';&lt;br&gt;
    }&lt;br&gt;
    return response;&lt;br&gt;
  } catch (err) {&lt;br&gt;
    breaker.failures++;&lt;br&gt;
    breaker.lastFailure = Date.now();&lt;br&gt;
    if (breaker.failures &amp;gt;= breaker.successThreshold) breaker.state = 'open';&lt;br&gt;
    throw err;&lt;br&gt;
  }&lt;br&gt;
}&lt;br&gt;
&lt;/p&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;
&lt;h3&gt;
&lt;br&gt;
  &lt;br&gt;
  &lt;br&gt;
  Configuration Management&lt;br&gt;
&lt;/h3&gt;

&lt;p&gt;Routes, rate limits, and auth settings should be dynamic. Store in PostgreSQL or a config service. Push updates to running gateways without restart via webhooks or polling.&lt;/p&gt;

&lt;p&gt;Hot-reloading config is essential for production. If you need to add a rate limit or update a route, you should not restart the gateway and drop connections. Watch a config table or Redis key for changes and apply them atomically.&lt;/p&gt;

&lt;h2&gt;
  
  
  Common Pitfalls
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Skipping Connection Pooling
&lt;/h3&gt;

&lt;p&gt;Each agent-to-gateway and gateway-to-server connection consumes resources. Without pooling, you hit file descriptor limits under load. Set &lt;code&gt;maxConnections&lt;/code&gt; and reuse connections aggressively.&lt;/p&gt;

&lt;p&gt;The Node.js HTTP agent defaults to unlimited connections. Set &lt;code&gt;agent.maxSockets&lt;/code&gt; explicitly. For outbound connections to MCP servers, use a dedicated agent with a fixed pool size.&lt;/p&gt;

&lt;h3&gt;
  
  
  No Timeout Strategy
&lt;/h3&gt;

&lt;p&gt;MCP tools can hang indefinitely without timeouts. Set per-tool timeouts: 30 seconds for fast tools, 5 minutes for heavy ones. Return structured timeout errors to the agent instead of letting connections stall.&lt;/p&gt;

&lt;p&gt;Agents that receive no response will retry. Without timeouts, you get cascading retries that amplify load. Always set both connection and read timeouts on outbound requests.&lt;/p&gt;

&lt;h3&gt;
  
  
  Treating the Gateway as Stateless
&lt;/h3&gt;

&lt;p&gt;Some tool calls need session state (conversation context, auth tokens, streaming connections). The gateway must preserve state per session or delegate session management to the backend servers consistently.&lt;/p&gt;

&lt;p&gt;If agent A calls tool X with a session header, and the gateway routes the next call to a different server instance, that server needs the session context. Use sticky routing or a shared session store.&lt;/p&gt;

&lt;h3&gt;
  
  
  Ignoring Protocol Version Mismatches
&lt;/h3&gt;

&lt;p&gt;MCP evolves. Backward-incompatible changes between versions break routing. The gateway should detect protocol versions on connection and apply version-specific translation logic.&lt;/p&gt;

&lt;p&gt;The MCP spec uses version negotiation on connection init. Log the negotiated version per connection. When you upgrade the SDK, test against older server versions before rolling out.&lt;/p&gt;

&lt;h3&gt;
  
  
  Forgetting Observability
&lt;/h3&gt;

&lt;p&gt;If you cannot trace a request from agent arrival to tool response and back, debugging is guesswork. Structured JSON logs, OpenTelemetry traces, and Prometheus metrics from day one.&lt;/p&gt;

&lt;p&gt;The cost of adding observability after launch is much higher than building it in from the start. Every &lt;code&gt;forwardRequest&lt;/code&gt; call should emit a trace span. Every auth check should log the agent ID. Every rate-limit hit should increment a counter.&lt;/p&gt;

&lt;h2&gt;
  
  
  When to Build vs Buy
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Requirement&lt;/th&gt;
&lt;th&gt;Custom Gateway&lt;/th&gt;
&lt;th&gt;Off-the-Shelf (e.g., mcporter)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Multi-tenant auth&lt;/td&gt;
&lt;td&gt;Full control over identity providers&lt;/td&gt;
&lt;td&gt;Limited to configured providers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Proprietary routing logic&lt;/td&gt;
&lt;td&gt;Custom rules, database-backed&lt;/td&gt;
&lt;td&gt;Static config only&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Compliance audit trails&lt;/td&gt;
&lt;td&gt;Partitioned PostgreSQL logs&lt;/td&gt;
&lt;td&gt;Limited to built-in logging&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Internal infrastructure integration&lt;/td&gt;
&lt;td&gt;Direct integration&lt;/td&gt;
&lt;td&gt;Requires adapter layer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Single tenant, few tools&lt;/td&gt;
&lt;td&gt;Overkill&lt;/td&gt;
&lt;td&gt;Good fit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No strict compliance requirements&lt;/td&gt;
&lt;td&gt;Overkill&lt;/td&gt;
&lt;td&gt;Good fit&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Build a custom gateway when you have unique requirements: multi-tenant auth, proprietary routing logic, compliance-driven audit trails, or integration with existing internal infrastructure. Building gives you full control over the data path and the ability to evolve with your agent fleet.&lt;/p&gt;

&lt;p&gt;Buy or adopt an existing solution when your needs fit standard patterns: single tenant, few tools, no strict compliance requirements. The &lt;a href="https://github.com/mcporter/mcporter" rel="noopener noreferrer"&gt;mcporter&lt;/a&gt; project and other open-source MCP gateways cover common cases well.&lt;/p&gt;

&lt;p&gt;For teams building serious AI agent infrastructure with multiple agents, tools, and tenants, a custom gateway pays for itself in the first debugging session.&lt;/p&gt;

&lt;p&gt;If you are routing AI requests across multiple providers, see &lt;a href="https://www.artifilog.com/posts/best-free-ai-model-routers" rel="noopener noreferrer"&gt;7 Best Free Zero-Ops AI Model Routers in 2026&lt;/a&gt; for the routing patterns that inform gateway design. For deeper dives into agent infrastructure, see &lt;a href="https://www.artifilog.com/" rel="noopener noreferrer"&gt;Artifilog's coverage of AI development tools&lt;/a&gt; and &lt;a href="https://www.artifilog.com/posts/route-ai-requests-cheapest-model" rel="noopener noreferrer"&gt;How to Route AI Requests to the Cheapest Model&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What is the difference between an MCP gateway and an MCP client?
&lt;/h3&gt;

&lt;p&gt;An MCP client connects to one or more servers on behalf of a single agent. A gateway sits between many agents and many servers, adding routing, auth, rate limiting, and observability that a basic client does not provide.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can an MCP gateway handle streaming responses?
&lt;/h3&gt;

&lt;p&gt;Yes. The transport layer must support streaming (SSE or streamable HTTP), and the gateway needs to pipe chunks through without buffering the entire response. Set the response headers to pass through &lt;code&gt;Content-Type: text/event-stream&lt;/code&gt; and handle backpressure.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do I handle auth token expiry in a gateway?
&lt;/h3&gt;

&lt;p&gt;Cache tokens with refresh logic. When a backend rejects a request with 401, the gateway refreshes the token and retries once before surfacing the error to the agent. Track token expiry proactively - refresh before expiry, not after rejection.&lt;/p&gt;

&lt;h3&gt;
  
  
  What monitoring is essential for an MCP gateway?
&lt;/h3&gt;

&lt;p&gt;Request latency distributions, error rates by tool, rate-limit hit counts, circuit-breaker state changes, and connection pool utilization. These tell you when a backend is degrading before agents start failing.&lt;/p&gt;




&lt;p&gt;&lt;a href="https://dev.to/author/hi3n"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;h&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;a href="https://dev.to/author/hi3n"&gt;hi3n&lt;/a&gt;
&lt;/h3&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.artifilog.com/posts/mcp-gateway-from-scratch" rel="noopener noreferrer"&gt;artifilog.com&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>OpenRouter vs Direct API: Which Is Cheaper in 2026?</title>
      <dc:creator>Artifilog</dc:creator>
      <pubDate>Thu, 01 Oct 2026 07:33:05 +0000</pubDate>
      <link>https://dev.to/artifilog/openrouter-vs-direct-api-which-is-cheaper-in-2026-1of7</link>
      <guid>https://dev.to/artifilog/openrouter-vs-direct-api-which-is-cheaper-in-2026-1of7</guid>
      <description>&lt;h1&gt;
  
  
  OpenRouter vs Direct API: Which Is Cheaper in 2026?
&lt;/h1&gt;

&lt;p&gt;Short answer: compare the same model, provider, and token categories before treating the inference rates as equal. OpenRouter resells at the provider's listed rate, advertising no markup on inference, and charges a fee when you buy credits. Direct access can cost less even at the same inference rate because it avoids OpenRouter’s credit-purchase fee. Discounts, negotiated terms, and integration costs can change the full comparison. Longer answer: per token is the wrong number to optimize, because tokens are rarely where the money or the pain goes. The real comparison is total operating cost, and the answer depends on the setup rather than a measured share of teams.&lt;/p&gt;

&lt;p&gt;This post does the honest math: where each side adds cost, where each side removes it, and a decision rule you can apply in one paragraph. No invented percentages, because fees and provider prices change monthly. Instead, a method you can rerun against the &lt;a href="https://openrouter.ai/models" rel="noopener noreferrer"&gt;live OpenRouter models page&lt;/a&gt; in five minutes whenever you need current numbers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sticker price vs operating cost: the two numbers that matter
&lt;/h2&gt;

&lt;p&gt;Sticker price is dollars per million tokens. Operating cost is sticker price plus everything else: engineering hours managing keys and SDKs, downtime when a provider fails, wasted spend from traffic hitting the wrong model, and the free tier value you get for tests and evals.&lt;/p&gt;

&lt;p&gt;Direct-provider contracts may offer discounts or terms that differ from aggregator billing. OpenRouter also supports BYOK and negotiated enterprise terms, so compare the actual offers rather than assuming a discount cannot be used through a gateway. Compare the same model, provider endpoint, token types, caching, and service tier. OpenRouter advertises no inference markup, but different endpoints or terms can carry different rates. If you run high volume with one provider, compare its direct and committed-use offer with OpenRouter’s matching route or BYOK terms, including all fees.&lt;/p&gt;

&lt;p&gt;OpenRouter competes on the second number. One API key and one request format for hundreds of models can reduce provider-specific integration work, though features and behavior still need testing. &lt;a href="https://openrouter.ai/docs/features/provider-routing" rel="noopener noreferrer"&gt;Provider routing with automatic failover&lt;/a&gt; can route to an eligible fallback when configured; it does not guarantee uninterrupted service or unchanged latency. For a small team running moderate volume across several models, estimate whether the engineering time saved exceeds the fee paid. Price that fee against an oncall page, not against zero.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where OpenRouter adds cost and where it removes cost
&lt;/h2&gt;

&lt;p&gt;The added cost is not a per token markup. OpenRouter passes the underlying provider's pricing through, so compare its listed rate with the matching direct endpoint and service tier. The current OpenRouter pricing page lists a 5.5% platform fee on its Standard pay-as-you-go plan and a 5% fee on BYOK usage above $25,000 of list-price inference per month. Confirm the current figures on the &lt;a href="https://openrouter.ai/pricing" rel="noopener noreferrer"&gt;live OpenRouter pricing page&lt;/a&gt; before you commit. That fee is the number you are deciding about. Everything else in this post is commentary on whether it pays for itself.&lt;/p&gt;

&lt;p&gt;Potentially reduced costs are less visible; estimate them for your team. First, integration: each direct provider means another SDK, another key rotation policy, another billing account to reconcile, another set of rate limit semantics to handle. With five models across three providers, that overhead is a permanent tax on every change.&lt;/p&gt;

&lt;p&gt;Second, failover. Providers go down, degrade, or throttle at the worst moments. Direct setups need hand built fallback logic per provider pair, tested rarely and trusted never. OpenRouter's routing layer does this as its core function: ordered provider lists, fallback chains, latency based selection. Compare the implementation and maintenance effort with the gateway fee; no universal volume threshold establishes the cheaper choice.&lt;/p&gt;

&lt;p&gt;Third, experimentation. Swapping models behind one interface turns evals from a project into a loop. That speed has a dollar value when model prices keep falling and last quarter's optimal pick is this quarter's overpay.&lt;/p&gt;

&lt;h2&gt;
  
  
  The free tier: what it covers and where it stops
&lt;/h2&gt;

&lt;p&gt;OpenRouter's free endpoints change the math for everything before production. Prototypes, evals, demos, and CI smoke tests can run on free tier models at zero marginal cost, but account for the engineering, evaluation, and any paid usage during experimentation as well as production. Direct providers offer trial credits too, but scattered across accounts with separate expiry rules and separate keys.&lt;/p&gt;

&lt;p&gt;Know the caps before you plan around them. Our &lt;a href="https://www.artifilog.com/posts/openrouter-free-tier-2026" rel="noopener noreferrer"&gt;free-tier guide&lt;/a&gt; covers which endpoints are free, and the &lt;a href="https://www.artifilog.com/posts/openrouter-free-rate-limits" rel="noopener noreferrer"&gt;rate limits post&lt;/a&gt; documents the throttles that bound them. Free tier is for learning and testing, not for production traffic. Designs that assume free endpoints scale get exactly one rude awakening, and it arrives during the demo that matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  Compare direct access: volume, model concentration, and terms
&lt;/h2&gt;

&lt;p&gt;Go direct when three conditions hold together. You run serious volume, you have converged on one or two models, and you can use provider pricing instruments (commits, tiers, enterprise terms) whose actual terms improve on the gateway or BYOK offer. At that point the platform fee is pure overhead and your team already owns the integration cost as sunk.&lt;/p&gt;

&lt;p&gt;Direct also wins when you need a required provider-specific feature that is not exposed on your chosen gateway route. Check availability, private deployment, region controls, and fine-tuning for the exact offering; OpenRouter itself documents in-region routing for enterprise customers. Aggregators can trail provider launches and may never expose every knob. If your product depends on a capability only one provider offers, route to that provider directly and skip the middle.&lt;/p&gt;

&lt;p&gt;The operational bar for going direct: automated key rotation, per provider dashboards with spend alerts, tested failover for your top two failure modes, and someone who owns the integration when a provider changes its API. If that list sounds like a team you do not have, the fee is buying you headcount you have not hired.&lt;/p&gt;

&lt;h2&gt;
  
  
  Self-hosted middle path: LiteLLM when the credit fee exceeds ops cost
&lt;/h2&gt;

&lt;p&gt;Between paying OpenRouter's credit fee and hand rolling provider integrations sits &lt;a href="https://github.com/BerriAI/litellm" rel="noopener noreferrer"&gt;LiteLLM&lt;/a&gt;: a self-hosted proxy that normalizes provider APIs behind one OpenAI compatible interface. An eligible open-source LiteLLM deployment avoids OpenRouter’s credit fee, but infrastructure and operations still cost money, and some enterprise features have separate commercial terms. You also own uptime, upgrades, and routing config, so there is ops cost instead.&lt;/p&gt;

&lt;p&gt;LiteLLM wins when your volume makes OpenRouter's credit fee materially larger than the cost of running one more service. That crossover usually needs consistent high throughput against a stable model set: the config rarely changes, the proxy just runs. It loses when your team has no one to page the proxy at night, or when your model set churns weekly and someone must maintain routing tables by hand.&lt;/p&gt;

&lt;p&gt;Think of it as a spectrum, not a rivalry. Prototype on OpenRouter free tier, run production through OpenRouter paid while volume is moderate and models churn, self-host LiteLLM when volume is high and the model set is stable, go fully direct when you converge on one provider with commits. Movement along the spectrum in either direction is normal as products mature.&lt;/p&gt;

&lt;h2&gt;
  
  
  Decision rule in one paragraph
&lt;/h2&gt;

&lt;p&gt;If you run high volume on one model with committed pricing, go direct. If you run moderate volume across several models with a small team, pay OpenRouter's credit fee and spend the saved engineering time on your product. If the fee at your volume exceeds the cost of operating one proxy, self-host LiteLLM. Recompute yearly, because model prices fall, fees move, and today's answer expires. The cheapest setup is the one you recheck, not the one you picked once.&lt;/p&gt;







&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.artifilog.com/posts/openrouter-vs-direct-api-cheaper" rel="noopener noreferrer"&gt;artifilog.com&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
      <category>webdev</category>
    </item>
  </channel>
</rss>
