DEV Community

YingSuan AI
YingSuan AI

Posted on

How to Switch Between DeepSeek and Qwen Without Changing Your Code (1 Endpoint, 0 Rewrites)

Every developer eventually hits the same wall: the model you picked in January isn't the model you want in October. Maybe a newer checkpoint shipped, maybe your workload changed, maybe you just want a cheaper route for background jobs. If you integrated a provider's SDK directly, "switching models" quietly becomes a migration project: new client library, new auth flow, new error codes, new billing dashboard.

It doesn't have to be. When your provider speaks the OpenAI protocol, the model name stops being architecture and becomes data — one string in one place. Here is how that works in practice with Yingsuan AI, an OpenAI-compatible gateway for DeepSeek, GLM and Qwen models.

The whole trick: one client, one config

Because the gateway implements the same /v1/chat/completions contract as OpenAI, you write the integration once and treat the model as runtime configuration:

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://yingsuan.top/v1",
    api_key=os.environ["YINGSUAN_API_KEY"],
)

# The ONLY thing that ever changes when you switch models:
model = os.environ.get("MODEL", "deepseek-chat")

response = client.chat.completions.create(
    model=model,   # "deepseek-chat" today, "qwen2.5-72b" tomorrow
    messages=[
        {"role": "system", "content": "You are a concise technical assistant."},
        {"role": "user", "content": "Explain vector databases in two sentences."},
    ],
)
print(response.choices[0].message.content)
Enter fullscreen mode Exit fullscreen mode

Want the Qwen 72B model instead of DeepSeek? Run:

MODEL=qwen2.5-72b python app.py
Enter fullscreen mode Exit fullscreen mode

No new SDK. No new account. No diff in git.

Discover what you can switch to

Hardcoding model names is fine, but you can also ask the gateway what's available on your tier with the standard GET /v1/models endpoint:

models = client.models.list()
print([m.id for m in models.data])
Enter fullscreen mode Exit fullscreen mode

Typical lineup on Yingsuan AI:

Model ID Good at Tier
deepseek-chat General-purpose chat and coding Paid
deepseek-reasoner Math, logic, step-by-step reasoning Paid
deepseek-v3 Strong all-round generation Paid
glm-4-flash Fast, cheap everyday tasks Free forever
glm-4-air / glm-4-plus Balanced / heavyweight GLM work Paid
qwen2.5-7b Lightweight multilingual tasks Free forever
qwen2.5-72b Long-form, multilingual generation Paid

Same pattern in JavaScript

Node projects get the identical one-string switch:

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://yingsuan.top/v1",
  apiKey: process.env.YINGSUAN_API_KEY,
});

const model = process.env.MODEL ?? "deepseek-chat";

const res = await client.chat.completions.create({
  model, // swap via env, not via refactoring
  messages: [{ role: "user", content: "Summarize the CAP theorem in 3 bullets." }],
});
console.log(res.choices[0].message.content);
Enter fullscreen mode Exit fullscreen mode

Three routing patterns this unlocks

Once switching is free, you can do things that are painful with direct provider integrations:

  1. Cost routing. Send cheap, high-volume traffic (summaries, classifications, autocomplete) to permanently free models like glm-4-flash or qwen2.5-7b, and reserve deepseek-reasoner for the hard questions.
  2. Capability routing. Detect the task type and pick the model: a reasoner for math, a 72B model for long-form writing, a small model for formatting.
  3. Fallback routing. If one model errors out or times out, try the next — in user code, no infrastructure required:
const MODEL_CHAIN = [process.env.MODEL ?? "deepseek-chat", "qwen2.5-72b"];

async function complete(messages) {
  for (const model of MODEL_CHAIN) {
    try {
      const res = await client.chat.completions.create({ model, messages });
      return res.choices[0].message.content;
    } catch (err) {
      console.error(`${model} failed, trying next…`);
    }
  }
  throw new Error("all models failed");
}
Enter fullscreen mode Exit fullscreen mode

That's resilience you'd normally build a whole abstraction layer for — expressed as an array of strings.

What it costs to try

Nothing up front: every new key includes 100 free trial calls, no credit card, and free-tier models stay free. Your existing OpenAI SDK code needs exactly two edits: the base_url and the API key.

Get your key and test it

Grab a free API key at yingsuan.top/api.html, then try running the same script twice with two different MODEL values. You'll switch providers faster than your test suite finishes.


What's your routing rule of choice — cost, latency, or capability? Tell me in the comments.

Top comments (0)