Every developer eventually hits the same wall: the model you picked in January isn't the model you want in October. Maybe a newer checkpoint shipped, maybe your workload changed, maybe you just want a cheaper route for background jobs. If you integrated a provider's SDK directly, "switching models" quietly becomes a migration project: new client library, new auth flow, new error codes, new billing dashboard.
It doesn't have to be. When your provider speaks the OpenAI protocol, the model name stops being architecture and becomes data — one string in one place. Here is how that works in practice with Yingsuan AI, an OpenAI-compatible gateway for DeepSeek, GLM and Qwen models.
The whole trick: one client, one config
Because the gateway implements the same /v1/chat/completions contract as OpenAI, you write the integration once and treat the model as runtime configuration:
import os
from openai import OpenAI
client = OpenAI(
base_url="https://yingsuan.top/v1",
api_key=os.environ["YINGSUAN_API_KEY"],
)
# The ONLY thing that ever changes when you switch models:
model = os.environ.get("MODEL", "deepseek-chat")
response = client.chat.completions.create(
model=model, # "deepseek-chat" today, "qwen2.5-72b" tomorrow
messages=[
{"role": "system", "content": "You are a concise technical assistant."},
{"role": "user", "content": "Explain vector databases in two sentences."},
],
)
print(response.choices[0].message.content)
Want the Qwen 72B model instead of DeepSeek? Run:
MODEL=qwen2.5-72b python app.py
No new SDK. No new account. No diff in git.
Discover what you can switch to
Hardcoding model names is fine, but you can also ask the gateway what's available on your tier with the standard GET /v1/models endpoint:
models = client.models.list()
print([m.id for m in models.data])
Typical lineup on Yingsuan AI:
| Model ID | Good at | Tier |
|---|---|---|
deepseek-chat |
General-purpose chat and coding | Paid |
deepseek-reasoner |
Math, logic, step-by-step reasoning | Paid |
deepseek-v3 |
Strong all-round generation | Paid |
glm-4-flash |
Fast, cheap everyday tasks | Free forever |
glm-4-air / glm-4-plus
|
Balanced / heavyweight GLM work | Paid |
qwen2.5-7b |
Lightweight multilingual tasks | Free forever |
qwen2.5-72b |
Long-form, multilingual generation | Paid |
Same pattern in JavaScript
Node projects get the identical one-string switch:
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://yingsuan.top/v1",
apiKey: process.env.YINGSUAN_API_KEY,
});
const model = process.env.MODEL ?? "deepseek-chat";
const res = await client.chat.completions.create({
model, // swap via env, not via refactoring
messages: [{ role: "user", content: "Summarize the CAP theorem in 3 bullets." }],
});
console.log(res.choices[0].message.content);
Three routing patterns this unlocks
Once switching is free, you can do things that are painful with direct provider integrations:
-
Cost routing. Send cheap, high-volume traffic (summaries, classifications, autocomplete) to permanently free models like
glm-4-flashorqwen2.5-7b, and reservedeepseek-reasonerfor the hard questions. - Capability routing. Detect the task type and pick the model: a reasoner for math, a 72B model for long-form writing, a small model for formatting.
- Fallback routing. If one model errors out or times out, try the next — in user code, no infrastructure required:
const MODEL_CHAIN = [process.env.MODEL ?? "deepseek-chat", "qwen2.5-72b"];
async function complete(messages) {
for (const model of MODEL_CHAIN) {
try {
const res = await client.chat.completions.create({ model, messages });
return res.choices[0].message.content;
} catch (err) {
console.error(`${model} failed, trying next…`);
}
}
throw new Error("all models failed");
}
That's resilience you'd normally build a whole abstraction layer for — expressed as an array of strings.
What it costs to try
Nothing up front: every new key includes 100 free trial calls, no credit card, and free-tier models stay free. Your existing OpenAI SDK code needs exactly two edits: the base_url and the API key.
Get your key and test it
Grab a free API key at yingsuan.top/api.html, then try running the same script twice with two different MODEL values. You'll switch providers faster than your test suite finishes.
What's your routing rule of choice — cost, latency, or capability? Tell me in the comments.
Top comments (0)