I hit a wall that felt familiar to anyone building with LLMs.
The problem: I had three OpenAI keys (two free-tier, one paid), an Anthropic key, and a Gemini key. Every app I built needed its own logic to handle rate limits, and quota exhaustion. When one key hit its limit, the app died - unless I wrote custom failover logic in every single app. Provider keys were scattered across .env files, shared insecurely, and impossible to rotate without breaking something.
The realization: This isn't an app problem. It's an infrastructure problem. Every team building with LLMs solves the same three things badly:
- Key pooling - spreading load across multiple keys
- Failover - automatically switching when a key fails
- Session continuity - keeping the conversation alive across the switch
So I built the thing I wanted to exist: Palimpsest Gateway a self-hosted, OpenAI-compatible gateway that sits in front of all your LLM keys, pools them, fails over automatically, and keeps every conversation going across the switch.
It's open source (MIT), written in Python/FastAPI, and runs with a single docker compose up. This is the story of why I built it, how it works, and the hard parts I didn't expect.
The Problem: Keys Run Out, Apps Break
If you've built anything with LLMs at scale, you know the pain points:
| Pain Point | What Actually Happens |
|---|---|
| Rate limits | Your free-tier key hits 429. App crashes. You catch it, sleep, retry... but the user already left. |
| Quota exhaustion | Monthly spend cap hit. Or daily free-tier allowance gone. No warning, just 402/429. |
| Key sprawl | 5 apps × 3 providers = 15 keys in 15 .env files. Rotating one is a deployment. |
| No visibility | You don't know which key served which request, how much you spent, or why a request failed. |
I wanted one endpoint that:
- Accepts standard OpenAI client calls (zero code change)
- Pools all my provider keys behind it
- Fails over transparently when a key hits a limit or errors
- Keeps conversation context so the user doesn't notice the switch
- Runs on my hardware, with my keys encrypted at rest
- Gives me a dashboard to see everything
Architecture: What Actually Runs
┌─────────────┐ ┌──────────────────────┐ ┌─────────────────┐
│ Your App │────▶│ Palimpsest Gateway │────▶│ LLM Providers │
│ (OpenAI │ │ • Auth & Admission │ │ OpenAI │
│ client) │ │ • Key Selection │ │ Anthropic │
└─────────────┘ │ • Failover Loop │ │ Google Gemini │
│ • Session Memory │ │ OpenRouter │
│ • Usage Metering │ └─────────────────┘
│ • Encrypted Storage │
└──────────┬───────────┘
│
┌──────────▼───────────┐
│ Redis │
│ • Key registry │
│ • Session history │
│ • Usage counters │
│ • Audit log │
└──────────────────────┘
The Request Flow
-
Authenticate - App sends its gateway key (
pgw_live_...), never a provider key - Admit - Check user's daily cap + pool capacity
-
Load context - With
X-Session-ID, pull stored history so the app only sends the new message - Choose a key - Pick from active keys serving the requested model, preferring capped keys with room left
- Call provider - Forward in provider-native format (OpenAI, Anthropic, Google, OpenRouter)
- Fail over if needed - Rate limit, outage, bad key → rest/exhaust/suspend that key, retry on next (same call, context carried)
- Answer & record - Return OpenAI-format response, save turn, count usage once
The Hard Parts (And How I Solved Them)
1. Streaming Failover Without Breaking the Client
This was the hardest part. When a provider returns a streamed response, you can't just "retry on another key" - the client is already receiving tokens.
Solution: The gateway buffers the entire provider stream in memory before sending anything to the client. If the provider fails mid-stream, the gateway retries on the next key with full context (system prompt + last N turns + current message). The client only ever sees a complete, valid OpenAI stream.
# Simplified: the failover loop in app/router/exhaustion.py
async def _call_with_failover(chain: list[KeyRecord], payload: dict) -> Response:
for attempt, key in enumerate(chain):
try:
return await provider.call(key, payload) # buffers full stream
except ProviderError as e:
classification = classify(e)
if classification.state == "transient":
await registry.mark_cooldown(key.id, model, settings.cooldown_seconds)
continue # try next key
elif classification.state == "exhausted":
await registry.mark_exhausted(key.id)
continue
elif classification.state == "rejected":
await registry.suspend(key.id)
continue
else: # fatal
raise
raise PoolExhaustedError(...)
The key insight: buffer the stream, failover before any bytes leave the gateway. The client never sees a partial response.
2. Context Carry-Over Across Providers
When failover happens, the replacement provider needs the conversation history. But each provider has different context window sizes and message formats.
Solution: The gateway stores every turn in Redis per session. On failover, it reconstructs a "verbatim window" - the last N turns (default 4) sent exactly as they were, plus the system prompt and current message. This fits in any reasonable context window and preserves the conversation thread.
# app/context/rebuild.py - simplified
def build_failover_context(session_history: list[Turn], current_message: Message) -> list[Message]:
verbatim = session_history[-settings.verbatim_turns_n:] # last 4 turns
return [
{"role": "system", "content": system_prompt},
*verbatim,
current_message,
]
3. Encrypted Key Storage That Survives Restarts
Provider keys are the crown jewels. They must be encrypted at rest, never logged, never returned after creation - but the gateway needs to decrypt them at request time.
Solution: A master key (Fernet) generated on first boot, persisted in a Docker volume. Every provider key is encrypted with it before hitting Redis. The master key never leaves the container. Rotating it is a destructive operation (make rotate-secrets) - by design.
# app/core/crypto.py
class SecretsManager:
def __init__(self, master_key: bytes):
self._fernet = Fernet(master_key)
def encrypt(self, plaintext: str) -> str:
return self._fernet.encrypt(plaintext.encode()).decode()
def decrypt(self, ciphertext: str) -> str:
return self._fernet.decrypt(ciphertext.encode()).decode()
4. Key Selection That Respects Your Intent
You have free keys (daily caps) and paid keys (uncapped). You want free keys used first, paid keys as backstop.
Solution: Selection algorithm orders keys by:
- Keys with a request cap + requests remaining → most remaining first
- Keys with no cap → last
Within each group, prefer the provider that served the session's previous turns (locality).
# app/pool/registry.py - routable_keys()
def sort_for_selection(keys: list[KeyRecord]) -> list[KeyRecord]:
capped = [k for k in keys if k.daily_cap is not None]
uncapped = [k for k in keys if k.daily_cap is None]
capped.sort(key=lambda k: k.remaining_requests, reverse=True)
return capped + uncapped
Quick Start: 60 Seconds to Running
git clone https://github.com/hammrouni/palimpsest-gateway.git
cd palimpsest-gateway
docker compose up -d --wait
That's it. On first boot it generates a master key and operator key, persists them in a Docker volume, and starts the API at http://localhost:8000.
# Add a provider key (one-time, or use the dashboard)
echo 'PALIMPSEST_PROVIDER_API_KEY=sk-your-openai-key' >> .env
docker compose up -d
# Create a user in the dashboard (or via API), get a pgw_live_... key
# Then call it from any OpenAI client:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="pgw_live_...")
reply = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "Hello"}],
)
print(reply.choices[0].message.content)
Send X-Session-ID: my-chat header to keep the conversation going server-side.
What It Looks Like In Practice
The dashboard shows you everything:
- Pool Keys - health, usage, caps, reset schedules, cooldown status
- Utilization & Failover - which key served what, failover events, retry chains
- Users & Sessions - per-account usage, active sessions, cap consumption
- Costs - spend by model/user/day against prices you enter (nothing bundled)
- Alerts - low capacity, key problems, error rates, webhook support
- Audit & Export - every operator action logged, CSV exports for usage/requests/audit
Click to see a sample request log line
{
"timestamp": "2026-09-04T20:00:00Z",
"request_id": "req_abc123",
"account_id": "acc_xyz",
"session_id": "sess_123",
"model": "gpt-4o-mini",
"status_code": 200,
"latency_ms": 850,
"outcome": "success",
"chain": ["key_openai_free_1", "key_openai_paid"],
"attempts": 2,
"n": 4
}
First key rate-limited → failover to paid key → context carried → success
Why "Palimpsest"?
A palimpsest is a manuscript page written on, scraped clean, and written on again - the old text still faintly visible underneath. That's what this gateway does: it layers your conversation history across provider switches, so the new provider sees the "old writing" (context) even though the underlying key changed.
What's Next
- Multi-gateway clustering - run multiple gateway instances behind a load balancer, shared Redis
- Provider quota sync - read actual provider quota (OpenRouter already supported) instead of manual caps
- Request transformations - normalize tool calls across providers, inject system prompts
-
Your PR here - good first issues tagged,
CONTRIBUTING.mdhas the dev loop
Try It, Star It, Break It
Repo: github.com/hammrouni/palimpsest-gateway
Docs: User Guide · Development
License: MIT
Docker: docker pull ghcr.io/hammrouni/palimpsest-gateway:latest
If you run it and hit something weird, open an issue. If you fix it, open a PR. If you hate it, tell me why - I've probably already hit that wall too.
Top comments (0)