Your team is wiring services directly to provider APIs. Every app has its own API key, its own retry logic, its own hardcoded model name. Nobody can answer "how much are we spending on LLM this month," and when a model gets deprecated, you find out through a 500 error in production.
I've been working on LLM infrastructure for a few years now, and the pattern I keep coming back to is an in-house LLM gateway: a stateless proxy and a small control plane. Not an "AI platform" - just a boring, dependable layer. Here are the four decisions that actually mattered when putting one in production.
The shape
flowchart LR
subgraph Clients
A[Service A<br/>team: payments]
B[Service B<br/>team: growth]
H[Developer<br/>IDE / CLI]
end
subgraph GW["LLM Gateway (stateless)"]
AUTH[Auth: JWT, RBAC]
ROUTE[Routing: alias, quotas, failover]
METER[Metering: usage, cost]
end
subgraph CP["Control plane (async)"]
REG[Model registry: aliases, prices]
Q[Quotas / budgets]
end
subgraph Providers
PA[Provider A]
PB[Provider B]
end
A -->|short-lived JWT| AUTH
B -->|short-lived JWT| AUTH
H -->|user token, OBO| AUTH
AUTH --> ROUTE --> PA
ROUTE --> PB
ROUTE -.-> REG
PA -->|usage| METER
PB -->|usage| METER
METER --> BUS[(message bus)] --> ROLL[(rollups)] --> DASH[Dashboards + alerts]
Two planes, four rules:
- One control plane, one data plane. Everything a human changes (registry, roles, budgets, prices) lives in the control plane. The proxy is stateless and only reads control-plane state.
- Everything is metered. If a request does not produce a metering event, it did not happen. The metering event is as real as the HTTP response.
- No secrets in requests. The only credentials on the wire are short-lived identity tokens. Provider credentials never leave the gateway's secret view.
- Boring technology. One language, one message format, one stream. The gateway's job is to be the least interesting component in your architecture.
The load path: clients → auth → routing → provider → stream back → metering → message bus → rollups → dashboards → alerts. The control plane is never in the hot path; it writes state, and the gateway picks up changes on a short TTL (1-5 seconds is fine).
Decision 1: A model registry that syncs itself
Lock-in happens quietly. One day a provider deprecates a model name and your services die. The fix is a model registry with aliases:
- Apps talk to aliases (
chat-fast,chat-reasoning,embeddings-1k), not provider model IDs. - A sync job polls each provider's
/modelsendpoint and updates availability and pricing. - When a provider model disappears, the alias routes to a fallback model instead of failing.
- Price updates flow through the same sync, so your cost dashboard is never stale.
Two tables, one scheduled job. No microservice forest.
Decision 2: Auth that works for services, humans, and agents
- Services get short-lived JWTs from your identity platform, verified against a JWKS cache.
- Humans (developers in IDEs or CLIs) get on-behalf-of tokens, so you can attribute usage to people, not shared service accounts.
- Agents get service tokens with narrower scope than the services they call.
- Provider credentials are resolved inside the gateway via managed identities - no API keys in your config files.
Fail closed on unknown alias, fail closed on invalid token. Quota counters are best-effort under high load (over-admit and alert); policies fail closed.
Decision 3: Cost tracking where the metering event is the source of truth
- Every request produces a usage object: tokens in/out, model, alias, team, user.
- Usage events go to a durable message bus; rollup jobs aggregate them into per-model / per-tenant / per-user tables.
- Budgets are counters (Redis) plus policies (SQL), with alerts at 80% and 100% of monthly budget per team and per model.
- The rollups DB can lag. It cannot lose. Recompute from the bus.
Once this is in place, "how much does this feature cost?" stops being a product meeting and becomes a dashboard question.
Decision 4: MCP as the interface for your AI tools
If agents or tools will talk to models, they should talk through the same gateway, with the Model Context Protocol (MCP) as the stable interface. One endpoint, any provider behind it. The gateway becomes the only place in your org that knows which provider is actually serving your traffic today.
What v1 looks like
One stateless service plus scheduled sync jobs. No orchestration framework, no second database. Everything else (portal, per-user budgets, fine-grained failover) can be added later without breaking the seams.
I wrote everything I learned into a book - reference architecture, configuration walkthroughs, an operations runbook, and a production launch checklist: The AI Gateway Playbook (32,000 words, $19, PDF + diagrams). Proceeds keep this kind of writing going.
If you're building an LLM gateway (or you built the one that's a mess of API keys and hardcoded model names), what's the hardest production problem you hit? I'd love to hear about it.
Top comments (0)