<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sanghyeok Yoon</title>
    <description>The latest articles on DEV Community by Sanghyeok Yoon (@zerokdevops).</description>
    <link>https://dev.to/zerokdevops</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4161480%2F62e72d21-9aad-4f89-bd57-fdabcaf2c8e9.png</url>
      <title>DEV Community: Sanghyeok Yoon</title>
      <link>https://dev.to/zerokdevops</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/zerokdevops"/>
    <language>en</language>
    <item>
      <title>The LLM Gateway I Put in Production: 4 Decisions That Actually Mattered</title>
      <dc:creator>Sanghyeok Yoon</dc:creator>
      <pubDate>Sun, 04 Oct 2026 11:43:18 +0000</pubDate>
      <link>https://dev.to/zerokdevops/the-llm-gateway-i-put-in-production-4-decisions-that-actually-mattered-5797</link>
      <guid>https://dev.to/zerokdevops/the-llm-gateway-i-put-in-production-4-decisions-that-actually-mattered-5797</guid>
      <description>&lt;p&gt;Your team is wiring services directly to provider APIs. Every app has its own API key, its own retry logic, its own hardcoded model name. Nobody can answer "how much are we spending on LLM this month," and when a model gets deprecated, you find out through a 500 error in production.&lt;/p&gt;

&lt;p&gt;I've been working on LLM infrastructure for a few years now, and the pattern I keep coming back to is an in-house LLM gateway: a stateless proxy and a small control plane. Not an "AI platform" - just a boring, dependable layer. Here are the four decisions that actually mattered when putting one in production.&lt;/p&gt;

&lt;h2&gt;
  
  
  The shape
&lt;/h2&gt;



&lt;pre data-lang="mermaid"&gt;&lt;code&gt;flowchart LR
    subgraph Clients
        A[Service A&amp;lt;br/&amp;gt;team: payments]
        B[Service B&amp;lt;br/&amp;gt;team: growth]
        H[Developer&amp;lt;br/&amp;gt;IDE / CLI]
    end
    subgraph GW["LLM Gateway (stateless)"]
        AUTH[Auth: JWT, RBAC]
        ROUTE[Routing: alias, quotas, failover]
        METER[Metering: usage, cost]
    end
    subgraph CP["Control plane (async)"]
        REG[Model registry: aliases, prices]
        Q[Quotas / budgets]
    end
    subgraph Providers
        PA[Provider A]
        PB[Provider B]
    end
    A --&amp;gt;|short-lived JWT| AUTH
    B --&amp;gt;|short-lived JWT| AUTH
    H --&amp;gt;|user token, OBO| AUTH
    AUTH --&amp;gt; ROUTE --&amp;gt; PA
    ROUTE --&amp;gt; PB
    ROUTE -.-&amp;gt; REG
    PA --&amp;gt;|usage| METER
    PB --&amp;gt;|usage| METER
    METER --&amp;gt; BUS[(message bus)] --&amp;gt; ROLL[(rollups)] --&amp;gt; DASH[Dashboards + alerts]&lt;/code&gt;&lt;/pre&gt;



&lt;p&gt;Two planes, four rules:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;One control plane, one data plane.&lt;/strong&gt; Everything a human changes (registry, roles, budgets, prices) lives in the control plane. The proxy is stateless and only &lt;em&gt;reads&lt;/em&gt; control-plane state.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Everything is metered.&lt;/strong&gt; If a request does not produce a metering event, it did not happen. The metering event is as real as the HTTP response.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No secrets in requests.&lt;/strong&gt; The only credentials on the wire are short-lived identity tokens. Provider credentials never leave the gateway's secret view.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Boring technology.&lt;/strong&gt; One language, one message format, one stream. The gateway's job is to be the least interesting component in your architecture.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The load path: clients → auth → routing → provider → stream back → metering → message bus → rollups → dashboards → alerts. The control plane is never in the hot path; it writes state, and the gateway picks up changes on a short TTL (1-5 seconds is fine).&lt;/p&gt;

&lt;h2&gt;
  
  
  Decision 1: A model registry that syncs itself
&lt;/h2&gt;

&lt;p&gt;Lock-in happens quietly. One day a provider deprecates a model name and your services die. The fix is a model registry with aliases:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Apps talk to aliases (&lt;code&gt;chat-fast&lt;/code&gt;, &lt;code&gt;chat-reasoning&lt;/code&gt;, &lt;code&gt;embeddings-1k&lt;/code&gt;), not provider model IDs.&lt;/li&gt;
&lt;li&gt;A sync job polls each provider's &lt;code&gt;/models&lt;/code&gt; endpoint and updates availability and pricing.&lt;/li&gt;
&lt;li&gt;When a provider model disappears, the alias routes to a fallback model instead of failing.&lt;/li&gt;
&lt;li&gt;Price updates flow through the same sync, so your cost dashboard is never stale.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Two tables, one scheduled job. No microservice forest.&lt;/p&gt;

&lt;h2&gt;
  
  
  Decision 2: Auth that works for services, humans, and agents
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Services&lt;/strong&gt; get short-lived JWTs from your identity platform, verified against a JWKS cache.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Humans&lt;/strong&gt; (developers in IDEs or CLIs) get on-behalf-of tokens, so you can attribute usage to people, not shared service accounts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agents&lt;/strong&gt; get service tokens with narrower scope than the services they call.&lt;/li&gt;
&lt;li&gt;Provider credentials are resolved inside the gateway via managed identities - no API keys in your config files.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Fail closed on unknown alias, fail closed on invalid token. Quota counters are best-effort under high load (over-admit and alert); policies fail closed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Decision 3: Cost tracking where the metering event is the source of truth
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Every request produces a usage object: tokens in/out, model, alias, team, user.&lt;/li&gt;
&lt;li&gt;Usage events go to a durable message bus; rollup jobs aggregate them into per-model / per-tenant / per-user tables.&lt;/li&gt;
&lt;li&gt;Budgets are counters (Redis) plus policies (SQL), with alerts at 80% and 100% of monthly budget per team and per model.&lt;/li&gt;
&lt;li&gt;The rollups DB can lag. It cannot lose. Recompute from the bus.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Once this is in place, "how much does this feature cost?" stops being a product meeting and becomes a dashboard question.&lt;/p&gt;

&lt;h2&gt;
  
  
  Decision 4: MCP as the interface for your AI tools
&lt;/h2&gt;

&lt;p&gt;If agents or tools will talk to models, they should talk through the same gateway, with the Model Context Protocol (MCP) as the stable interface. One endpoint, any provider behind it. The gateway becomes the only place in your org that knows which provider is actually serving your traffic today.&lt;/p&gt;

&lt;h2&gt;
  
  
  What v1 looks like
&lt;/h2&gt;

&lt;p&gt;One stateless service plus scheduled sync jobs. No orchestration framework, no second database. Everything else (portal, per-user budgets, fine-grained failover) can be added later without breaking the seams.&lt;/p&gt;




&lt;p&gt;I wrote everything I learned into a book - reference architecture, configuration walkthroughs, an operations runbook, and a production launch checklist: &lt;strong&gt;The AI Gateway Playbook&lt;/strong&gt; (32,000 words, $19, &lt;a href="https://ko-fi.com/s/5159c92ef2" rel="noopener noreferrer"&gt;PDF + diagrams&lt;/a&gt;). Proceeds keep this kind of writing going.&lt;/p&gt;

&lt;p&gt;If you're building an LLM gateway (or you built the one that's a mess of API keys and hardcoded model names), what's the hardest production problem you hit? I'd love to hear about it.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>architecture</category>
      <category>api</category>
    </item>
  </channel>
</rss>
