DEV Community

Khaled Hammrouni
Khaled Hammrouni

Posted on

How I Built an Open-Source LLM Key-Pool Gateway (and Why You Might Want One)

I hit a wall that felt familiar to anyone building with LLMs.

The problem: I had three OpenAI keys (two free-tier, one paid), an Anthropic key, and a Gemini key. Every app I built needed its own logic to handle rate limits, and quota exhaustion. When one key hit its limit, the app died - unless I wrote custom failover logic in every single app. Provider keys were scattered across .env files, shared insecurely, and impossible to rotate without breaking something.

The realization: This isn't an app problem. It's an infrastructure problem. Every team building with LLMs solves the same three things badly:

  1. Key pooling - spreading load across multiple keys
  2. Failover - automatically switching when a key fails
  3. Session continuity - keeping the conversation alive across the switch

So I built the thing I wanted to exist: Palimpsest Gateway a self-hosted, OpenAI-compatible gateway that sits in front of all your LLM keys, pools them, fails over automatically, and keeps every conversation going across the switch.

It's open source (MIT), written in Python/FastAPI, and runs with a single docker compose up. This is the story of why I built it, how it works, and the hard parts I didn't expect.


The Problem: Keys Run Out, Apps Break

If you've built anything with LLMs at scale, you know the pain points:

Pain Point What Actually Happens
Rate limits Your free-tier key hits 429. App crashes. You catch it, sleep, retry... but the user already left.
Quota exhaustion Monthly spend cap hit. Or daily free-tier allowance gone. No warning, just 402/429.
Key sprawl 5 apps × 3 providers = 15 keys in 15 .env files. Rotating one is a deployment.
No visibility You don't know which key served which request, how much you spent, or why a request failed.

I wanted one endpoint that:

  • Accepts standard OpenAI client calls (zero code change)
  • Pools all my provider keys behind it
  • Fails over transparently when a key hits a limit or errors
  • Keeps conversation context so the user doesn't notice the switch
  • Runs on my hardware, with my keys encrypted at rest
  • Gives me a dashboard to see everything

Architecture: What Actually Runs

┌─────────────┐     ┌──────────────────────┐     ┌─────────────────┐
│  Your App   │────▶│  Palimpsest Gateway  │────▶│  LLM Providers  │
│ (OpenAI     │     │  • Auth & Admission  │     │  OpenAI         │
│  client)    │     │  • Key Selection     │     │  Anthropic      │
└─────────────┘     │  • Failover Loop     │     │  Google Gemini  │
                    │  • Session Memory    │     │  OpenRouter     │
                    │  • Usage Metering    │     └─────────────────┘
                    │  • Encrypted Storage │
                    └──────────┬───────────┘
                               │
                    ┌──────────▼───────────┐
                    │       Redis          │
                    │  • Key registry      │
                    │  • Session history   │
                    │  • Usage counters    │
                    │  • Audit log         │
                    └──────────────────────┘
Enter fullscreen mode Exit fullscreen mode

The Request Flow

  1. Authenticate - App sends its gateway key (pgw_live_...), never a provider key
  2. Admit - Check user's daily cap + pool capacity
  3. Load context - With X-Session-ID, pull stored history so the app only sends the new message
  4. Choose a key - Pick from active keys serving the requested model, preferring capped keys with room left
  5. Call provider - Forward in provider-native format (OpenAI, Anthropic, Google, OpenRouter)
  6. Fail over if needed - Rate limit, outage, bad key → rest/exhaust/suspend that key, retry on next (same call, context carried)
  7. Answer & record - Return OpenAI-format response, save turn, count usage once

The Hard Parts (And How I Solved Them)

1. Streaming Failover Without Breaking the Client

This was the hardest part. When a provider returns a streamed response, you can't just "retry on another key" - the client is already receiving tokens.

Solution: The gateway buffers the entire provider stream in memory before sending anything to the client. If the provider fails mid-stream, the gateway retries on the next key with full context (system prompt + last N turns + current message). The client only ever sees a complete, valid OpenAI stream.

# Simplified: the failover loop in app/router/exhaustion.py
async def _call_with_failover(chain: list[KeyRecord], payload: dict) -> Response:
    for attempt, key in enumerate(chain):
        try:
            return await provider.call(key, payload)  # buffers full stream
        except ProviderError as e:
            classification = classify(e)
            if classification.state == "transient":
                await registry.mark_cooldown(key.id, model, settings.cooldown_seconds)
                continue  # try next key
            elif classification.state == "exhausted":
                await registry.mark_exhausted(key.id)
                continue
            elif classification.state == "rejected":
                await registry.suspend(key.id)
                continue
            else:  # fatal
                raise
    raise PoolExhaustedError(...)
Enter fullscreen mode Exit fullscreen mode

The key insight: buffer the stream, failover before any bytes leave the gateway. The client never sees a partial response.

2. Context Carry-Over Across Providers

When failover happens, the replacement provider needs the conversation history. But each provider has different context window sizes and message formats.

Solution: The gateway stores every turn in Redis per session. On failover, it reconstructs a "verbatim window" - the last N turns (default 4) sent exactly as they were, plus the system prompt and current message. This fits in any reasonable context window and preserves the conversation thread.

# app/context/rebuild.py - simplified
def build_failover_context(session_history: list[Turn], current_message: Message) -> list[Message]:
    verbatim = session_history[-settings.verbatim_turns_n:]  # last 4 turns
    return [
        {"role": "system", "content": system_prompt},
        *verbatim,
        current_message,
    ]
Enter fullscreen mode Exit fullscreen mode

3. Encrypted Key Storage That Survives Restarts

Provider keys are the crown jewels. They must be encrypted at rest, never logged, never returned after creation - but the gateway needs to decrypt them at request time.

Solution: A master key (Fernet) generated on first boot, persisted in a Docker volume. Every provider key is encrypted with it before hitting Redis. The master key never leaves the container. Rotating it is a destructive operation (make rotate-secrets) - by design.

# app/core/crypto.py
class SecretsManager:
    def __init__(self, master_key: bytes):
        self._fernet = Fernet(master_key)

    def encrypt(self, plaintext: str) -> str:
        return self._fernet.encrypt(plaintext.encode()).decode()

    def decrypt(self, ciphertext: str) -> str:
        return self._fernet.decrypt(ciphertext.encode()).decode()
Enter fullscreen mode Exit fullscreen mode

4. Key Selection That Respects Your Intent

You have free keys (daily caps) and paid keys (uncapped). You want free keys used first, paid keys as backstop.

Solution: Selection algorithm orders keys by:

  1. Keys with a request cap + requests remaining → most remaining first
  2. Keys with no cap → last

Within each group, prefer the provider that served the session's previous turns (locality).

# app/pool/registry.py - routable_keys()
def sort_for_selection(keys: list[KeyRecord]) -> list[KeyRecord]:
    capped = [k for k in keys if k.daily_cap is not None]
    uncapped = [k for k in keys if k.daily_cap is None]
    capped.sort(key=lambda k: k.remaining_requests, reverse=True)
    return capped + uncapped
Enter fullscreen mode Exit fullscreen mode

Quick Start: 60 Seconds to Running

git clone https://github.com/hammrouni/palimpsest-gateway.git
cd palimpsest-gateway
docker compose up -d --wait
Enter fullscreen mode Exit fullscreen mode

That's it. On first boot it generates a master key and operator key, persists them in a Docker volume, and starts the API at http://localhost:8000.

# Add a provider key (one-time, or use the dashboard)
echo 'PALIMPSEST_PROVIDER_API_KEY=sk-your-openai-key' >> .env
docker compose up -d

# Create a user in the dashboard (or via API), get a pgw_live_... key
# Then call it from any OpenAI client:
Enter fullscreen mode Exit fullscreen mode
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="pgw_live_...")
reply = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Hello"}],
)
print(reply.choices[0].message.content)
Enter fullscreen mode Exit fullscreen mode

Send X-Session-ID: my-chat header to keep the conversation going server-side.


What It Looks Like In Practice

The dashboard shows you everything:

  • Pool Keys - health, usage, caps, reset schedules, cooldown status
  • Utilization & Failover - which key served what, failover events, retry chains
  • Users & Sessions - per-account usage, active sessions, cap consumption
  • Costs - spend by model/user/day against prices you enter (nothing bundled)
  • Alerts - low capacity, key problems, error rates, webhook support
  • Audit & Export - every operator action logged, CSV exports for usage/requests/audit

Click to see a sample request log line

{
  "timestamp": "2026-09-04T20:00:00Z",
  "request_id": "req_abc123",
  "account_id": "acc_xyz",
  "session_id": "sess_123",
  "model": "gpt-4o-mini",
  "status_code": 200,
  "latency_ms": 850,
  "outcome": "success",
  "chain": ["key_openai_free_1", "key_openai_paid"],
  "attempts": 2,
  "n": 4
}
Enter fullscreen mode Exit fullscreen mode

First key rate-limited → failover to paid key → context carried → success


Why "Palimpsest"?

A palimpsest is a manuscript page written on, scraped clean, and written on again - the old text still faintly visible underneath. That's what this gateway does: it layers your conversation history across provider switches, so the new provider sees the "old writing" (context) even though the underlying key changed.


What's Next

  • Multi-gateway clustering - run multiple gateway instances behind a load balancer, shared Redis
  • Provider quota sync - read actual provider quota (OpenRouter already supported) instead of manual caps
  • Request transformations - normalize tool calls across providers, inject system prompts
  • Your PR here - good first issues tagged, CONTRIBUTING.md has the dev loop

Try It, Star It, Break It

Repo: github.com/hammrouni/palimpsest-gateway

Docs: User Guide · Development

License: MIT

Docker: docker pull ghcr.io/hammrouni/palimpsest-gateway:latest

If you run it and hit something weird, open an issue. If you fix it, open a PR. If you hate it, tell me why - I've probably already hit that wall too.

Top comments (0)