DEV Community

zhangjj1988
zhangjj1988

Posted on Edited on

Migrate from OpenAI to DeepSeek/Qwen in 10 Minutes — Zero Code Rewrite

Migrate from OpenAI to DeepSeek/Qwen in 10 Minutes — Zero Code Rewrite

Your app already calls the OpenAI SDK. Maybe it's a support bot, a summarizer, or an agent loop. Now you want to try DeepSeek or Qwen — to cut cost, dodge a card requirement, or just see how a different model handles your prompts. The good news: if your client already speaks the OpenAI Chat Completions format, the migration is two lines. You change base_url and model. Everything else — retries, streaming, tool calls — stays exactly where it is.

This post is a copy-paste walkthrough. No new SDK, no format shim, no refactor. By the end you'll have a working call to a real DeepSeek or Qwen model through one OpenAI-compatible endpoint.

taotok.io is a unified LLM API gateway that puts DeepSeek and Qwen behind a single OpenAI-compatible endpoint, so the only thing your code learns is a new base_url and a new model string.

Why migrate (without dumping on OpenAI)

OpenAI's API is excellent and the SDK is everywhere. The friction most teams hit isn't model quality — it's access and billing:

  • The card wall. Signup asks for a credit card up front, and the billing flow assumes a certain scale. If you're a student, in a region where international cards are scarce, or just experimenting, that wall is real.
  • Regional gaps. Not every model or tier is available in every country, and availability shifts.
  • Cost structure. Pricing is built around a commitment profile that doesn't always match a side project or a thin prototype.

None of that means "switch forever." It means "keep a drop-in alternative for the paths where the friction hurts." An alternative is a gateway that puts DeepSeek and Qwen behind one OpenAI-compatible endpoint, with email signup, no card, and no KYC. You can top up with USDT via crypto.

crypto (USDT) is a payment method only; taotok.io is a developer API service, not a financial service.

That's the whole reason this post exists: a lighter onboarding so you can actually reach the models and benchmark them on your own prompts.

Prerequisites

  1. Register at taotok.io with your email. No card, no KYC, no ID upload.
  2. Create an API key from the dashboard. Store it server-side — never ship it in a client bundle or commit it.
  3. You get 500 free credits on signup, enough to run a few hundred short calls while you evaluate fit.

That's it. No second console to read, no region to pick, no extra account to open.

Before / After: the two-line change

Here is the same function, before and after. The only differences are the two commented lines.

Before — calling OpenAI directly:

from openai import OpenAI

client = OpenAI(api_key="sk-...")  # defaults to https://api.openai.com/v1

resp = client.chat.completions.create(
    model="gpt-4o",
    messages=[
        {"role": "user", "content": "Explain exponential backoff in one sentence."}
    ],
)
print(resp.choices[0].message.content)
Enter fullscreen mode Exit fullscreen mode

After — pointing at DeepSeek or Qwen:

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_KEY",
    base_url="https://api.taotok.io/v1",        # <- line 1: repoint the endpoint
)

resp = client.chat.completions.create(
    model="deepseek-v4-pro",                     # <- line 2: pick a real model
    messages=[
        {"role": "user", "content": "Explain exponential backoff in one sentence."}
    ],
)
print(resp.choices[0].message.content)
Enter fullscreen mode Exit fullscreen mode

That is the migration. Your retry wrapper, your message builder, your response parser — none of it changes, because the request and response shapes follow the OpenAI Chat Completions schema. The base_url is the one fact we verified end to end: /v1/models and /v1/chat/completions both respond (a 401 means the route exists and is auth-gated, exactly as expected).

Streaming works the same way

If your app streams output, you don't rewrite it. stream=True and the same SSE delta shape carry over:

stream = client.chat.completions.create(
    model="qwen-max",
    messages=[
        {"role": "user", "content": "Draft a concise commit message for a login-rate-limit fix."}
    ],
    stream=True,
)

for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")
Enter fullscreen mode Exit fullscreen mode

Same loop you already run against OpenAI. The only thing that moved is the model string.

Model cheat sheet

These are the real model names behind the gateway — use them directly in the model field, with no -class suffix:

Model Family Best for
deepseek-v4-pro DeepSeek Strongest reasoning, math, and agent workloads
deepseek-v4-flash DeepSeek Fast and cheap; high-throughput bulk calls
qwen-max Qwen (Alibaba) Flagship general-purpose model
qwen-plus Qwen (Alibaba) Mid-tier, balanced cost and quality
qwen-turbo Qwen (Alibaba) Fastest and cheapest; light tasks
qwen-max-longcontext Qwen (Alibaba) Flagship with extended context for long inputs
qwen3-235b-a22b Qwen (Alibaba) Open-weight 235B MoE model

Switching between them is a one-string change — no code path differs. Route simple text to deepseek-v4-flash or qwen-turbo, and reserve deepseek-v4-pro or qwen-max for the hard reasoning calls.

Advanced: tool calling still works

Because the endpoint follows the OpenAI schema, function calling is available the same way. Define your tools, send them, and read tool_calls back:

tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Get the current weather for a city",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
        },
    },
}]

resp = client.chat.completions.create(
    model="deepseek-v4-flash",
    messages=[{"role": "user", "content": "What's the weather in Tokyo?"}],
    tools=tools,
)

print(resp.choices[0].message.tool_calls)
Enter fullscreen mode Exit fullscreen mode

Run the tool, append the result as a tool role message, and loop until the model finishes — the identical pattern you use with OpenAI. No adapter layer required.

Common gotchas

A few things to check once you repoint, so a surprise doesn't read as a bug:

  • Billing is metered by the usage object your response returns. Read usage from the API answer and bill or log from there; don't estimate from your own counts, since backends count units their own way.
  • -class tiers are compatible labels, not the official model. On the pricing page you'll see tiers like gpt-4o-class. Those are compatible tiers — routing labels that map to a real DeepSeek or Qwen backend — not the original OpenAI model running underneath. They're a shorthand for "acts like this class." If you want the actual DeepSeek or Qwen model, use the real names from the cheat sheet above.
  • Streaming behavior is consistent. The SSE format, delta fields, and finish_reason follow the OpenAI shape, so your existing streaming parser keeps working.

One more practical note: an extra network hop adds latency versus calling a provider directly, so measure p50/p95 on your real payload sizes before you lean on it in production.

Wrap-up

Migrating off OpenAI doesn't mean rewriting your integration. If your client already speaks the OpenAI Chat Completions format, you change base_url and model, keep your retries and tool logic, and you're calling real DeepSeek and Qwen models through one endpoint. Spend your 500 free credits on a small, representative eval — format adherence, instruction following, reasoning, latency — and decide from your own numbers.

Ready to try it? Sign up by email, grab your 500 free credits, and repoint your client:

👉 https://api.taotok.io/go/devto9.html

Building something and want to compare notes with other developers routing across providers? There is a small Discord where people share setups and war stories: https://discord.gg/eEsTYXpZJn

Top comments (1)

Collapse
 
zhangjj1988 profile image
zhangjj1988 •

Curious — anyone benchmarked deepseek-v4-pro vs gpt-4o on code-gen tasks? Trying to decide where the extra network hop is worth it. What is your p50 like?