DEV Community

Run Liu
Run Liu

Posted on

Stop managing six AI vendor accounts: one key for 109 models

The migration is two lines

If your code already uses the OpenAI SDK, a gateway that speaks the OpenAI wire format requires exactly two changes:

from openai import OpenAI

client = OpenAI(
    base_url="https://api.liurun.click/v1",   # was "https://api.openai.com/v1"
    api_key="sk-your-liurun-key",             # was your OpenAI key
)

resp = client.chat.completions.create(
    model="claude-sonnet-5",                  # any of 109 models, same call shape
    messages=[{"role": "user", "content": "Explain WAL in Postgres"}],
)
print(resp.choices[0].message.content)
Enter fullscreen mode Exit fullscreen mode

That model string is now the only knob between vendors. A fallback chain becomes a list, not a architecture project:

MODELS = ["claude-sonnet-5", "gpt-5.5", "deepseek-chat"]  # try in order
Enter fullscreen mode Exit fullscreen mode

Plain curl works too

curl https://api.liurun.click/v1/chat/completions \
  -H "Authorization: Bearer sk-your-liurun-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemini-2.5-flash",
    "messages": [{"role": "user", "content": "Say hi in 3 words"}],
    "stream": true
  }'
Enter fullscreen mode Exit fullscreen mode

Streaming is standard SSE — data: lines, data: [DONE] sentinel, nothing exotic.

Anthropic SDK users: you don't have to change either

The gateway also accepts the Claude-format /v1/messages endpoint. So the Anthropic SDK works with a base URL swap:

import anthropic

client = anthropic.Anthropic(
    base_url="https://api.liurun.click",
    api_key="sk-your-liurun-key",
)

msg = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Hello, Claude"}],
)
Enter fullscreen mode Exit fullscreen mode

The part I actually care about: knowing what things cost

Every request gets logged with its exact token counts and its exact cost. Not "credits" — dollars and cents, computed from a per-model price table that's published in full on a public pricing page:

Model Input $/1M Output $/1M
Claude Sonnet 5 1.45 7.27
Claude Opus 4.6 3.63 18.16
GPT-5.5 2.72 16.35
Gemini 2.5 Flash 0.16 1.36
DeepSeek Chat 0.36 1.45
Grok 4.5 1.09 3.27
GPT-6 Luna (budget) 0.05 0.27

Prompt-cache reads on Claude models bill at 10% of the input rate — cache-friendly agents stop being a budget gamble. Image models bill per image (GPT-Image-2 ≈ $0.051/image).

Billing is prepaid: you top up a USD balance from $1 and usage draws it down. No subscription, no seats, no minimum. If you stop liking the service, your remaining balance is the only thing at risk — and there's no contract to cancel.

A realistic routing setup

The pattern I use in my own projects now:

def complete(messages, *, tier="smart", **kw):
    models = {
        "smart":  ["claude-sonnet-5", "gpt-5.5"],
        "cheap":  ["gemini-2.5-flash", "deepseek-chat"],
        "budget": ["gpt-6-luna"],
    }[tier]
    last = None
    for m in models:
        try:
            return client.chat.completions.create(
                model=m, messages=messages, **kw)
        except Exception as e:
            last = e
    raise last
Enter fullscreen mode Exit fullscreen mode

Same key, same endpoint, model choice becomes a cost/quality dial instead of an integration decision.

Self-hosted alternative

The gateway runs on New API, which is open source — if you'd rather keep everything in-house, the two-line migration above works against your own deployment too. I self-host the hosted version on a small AWS instance in Tokyo behind CloudFront; a t4g.small handles it comfortably.

Checklist before you switch anything in production

  1. Verify streaming works through the gateway for your SDK version (SSE quirks are the #1 migration bug)
  2. Check how cache tokens are metered if you rely on prompt caching
  3. Make sure per-request logs give you cost per call — not just tokens
  4. Confirm the model list actually includes the models you call (all 109 are listed on the pricing page)

Links: Sign up · Pricing · User Agreement

(Disclosure: I built and operate LiuRun API. The migration pattern above applies to any OpenAI-compatible gateway, self-hosted included.)

Top comments (0)