The migration is two lines
If your code already uses the OpenAI SDK, a gateway that speaks the OpenAI wire format requires exactly two changes:
from openai import OpenAI
client = OpenAI(
base_url="https://api.liurun.click/v1", # was "https://api.openai.com/v1"
api_key="sk-your-liurun-key", # was your OpenAI key
)
resp = client.chat.completions.create(
model="claude-sonnet-5", # any of 109 models, same call shape
messages=[{"role": "user", "content": "Explain WAL in Postgres"}],
)
print(resp.choices[0].message.content)
That model string is now the only knob between vendors. A fallback chain becomes a list, not a architecture project:
MODELS = ["claude-sonnet-5", "gpt-5.5", "deepseek-chat"] # try in order
Plain curl works too
curl https://api.liurun.click/v1/chat/completions \
-H "Authorization: Bearer sk-your-liurun-key" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-2.5-flash",
"messages": [{"role": "user", "content": "Say hi in 3 words"}],
"stream": true
}'
Streaming is standard SSE — data: lines, data: [DONE] sentinel, nothing exotic.
Anthropic SDK users: you don't have to change either
The gateway also accepts the Claude-format /v1/messages endpoint. So the Anthropic SDK works with a base URL swap:
import anthropic
client = anthropic.Anthropic(
base_url="https://api.liurun.click",
api_key="sk-your-liurun-key",
)
msg = client.messages.create(
model="claude-sonnet-5",
max_tokens=1024,
messages=[{"role": "user", "content": "Hello, Claude"}],
)
The part I actually care about: knowing what things cost
Every request gets logged with its exact token counts and its exact cost. Not "credits" — dollars and cents, computed from a per-model price table that's published in full on a public pricing page:
| Model | Input $/1M | Output $/1M |
|---|---|---|
| Claude Sonnet 5 | 1.45 | 7.27 |
| Claude Opus 4.6 | 3.63 | 18.16 |
| GPT-5.5 | 2.72 | 16.35 |
| Gemini 2.5 Flash | 0.16 | 1.36 |
| DeepSeek Chat | 0.36 | 1.45 |
| Grok 4.5 | 1.09 | 3.27 |
| GPT-6 Luna (budget) | 0.05 | 0.27 |
Prompt-cache reads on Claude models bill at 10% of the input rate — cache-friendly agents stop being a budget gamble. Image models bill per image (GPT-Image-2 ≈ $0.051/image).
Billing is prepaid: you top up a USD balance from $1 and usage draws it down. No subscription, no seats, no minimum. If you stop liking the service, your remaining balance is the only thing at risk — and there's no contract to cancel.
A realistic routing setup
The pattern I use in my own projects now:
def complete(messages, *, tier="smart", **kw):
models = {
"smart": ["claude-sonnet-5", "gpt-5.5"],
"cheap": ["gemini-2.5-flash", "deepseek-chat"],
"budget": ["gpt-6-luna"],
}[tier]
last = None
for m in models:
try:
return client.chat.completions.create(
model=m, messages=messages, **kw)
except Exception as e:
last = e
raise last
Same key, same endpoint, model choice becomes a cost/quality dial instead of an integration decision.
Self-hosted alternative
The gateway runs on New API, which is open source — if you'd rather keep everything in-house, the two-line migration above works against your own deployment too. I self-host the hosted version on a small AWS instance in Tokyo behind CloudFront; a t4g.small handles it comfortably.
Checklist before you switch anything in production
- Verify streaming works through the gateway for your SDK version (SSE quirks are the #1 migration bug)
- Check how cache tokens are metered if you rely on prompt caching
- Make sure per-request logs give you cost per call — not just tokens
- Confirm the model list actually includes the models you call (all 109 are listed on the pricing page)
Links: Sign up · Pricing · User Agreement
(Disclosure: I built and operate LiuRun API. The migration pattern above applies to any OpenAI-compatible gateway, self-hosted included.)
Top comments (0)