Migrate from OpenAI to DeepSeek/Qwen in 10 Minutes — Zero Code Rewrite
Your app already calls the OpenAI SDK. Maybe it's a support bot, a summarizer, or an agent loop. Now you want to try DeepSeek or Qwen — to cut cost, dodge a card requirement, or just see how a different model handles your prompts. The good news: if your client already speaks the OpenAI Chat Completions format, the migration is two lines. You change base_url and model. Everything else — retries, streaming, tool calls — stays exactly where it is.
This post is a copy-paste walkthrough. No new SDK, no format shim, no refactor. By the end you'll have a working call to a real DeepSeek or Qwen model through one OpenAI-compatible endpoint.
taotok.io is a unified LLM API gateway that puts DeepSeek and Qwen behind a single OpenAI-compatible endpoint, so the only thing your code learns is a new base_url and a new model string.
Why migrate (without dumping on OpenAI)
OpenAI's API is excellent and the SDK is everywhere. The friction most teams hit isn't model quality — it's access and billing:
- The card wall. Signup asks for a credit card up front, and the billing flow assumes a certain scale. If you're a student, in a region where international cards are scarce, or just experimenting, that wall is real.
- Regional gaps. Not every model or tier is available in every country, and availability shifts.
- Cost structure. Pricing is built around a commitment profile that doesn't always match a side project or a thin prototype.
None of that means "switch forever." It means "keep a drop-in alternative for the paths where the friction hurts." An alternative is a gateway that puts DeepSeek and Qwen behind one OpenAI-compatible endpoint, with email signup, no card, and no KYC. You can top up with USDT via crypto.
crypto (USDT) is a payment method only; taotok.io is a developer API service, not a financial service.
That's the whole reason this post exists: a lighter onboarding so you can actually reach the models and benchmark them on your own prompts.
Prerequisites
- Register at taotok.io with your email. No card, no KYC, no ID upload.
- Create an API key from the dashboard. Store it server-side — never ship it in a client bundle or commit it.
- You get 500 free credits on signup, enough to run a few hundred short calls while you evaluate fit.
That's it. No second console to read, no region to pick, no extra account to open.
Before / After: the two-line change
Here is the same function, before and after. The only differences are the two commented lines.
Before — calling OpenAI directly:
from openai import OpenAI
client = OpenAI(api_key="sk-...") # defaults to https://api.openai.com/v1
resp = client.chat.completions.create(
model="gpt-4o",
messages=[
{"role": "user", "content": "Explain exponential backoff in one sentence."}
],
)
print(resp.choices[0].message.content)
After — pointing at DeepSeek or Qwen:
from openai import OpenAI
client = OpenAI(
api_key="YOUR_KEY",
base_url="https://api.taotok.io/v1", # <- line 1: repoint the endpoint
)
resp = client.chat.completions.create(
model="deepseek-v4-pro", # <- line 2: pick a real model
messages=[
{"role": "user", "content": "Explain exponential backoff in one sentence."}
],
)
print(resp.choices[0].message.content)
That is the migration. Your retry wrapper, your message builder, your response parser — none of it changes, because the request and response shapes follow the OpenAI Chat Completions schema. The base_url is the one fact we verified end to end: /v1/models and /v1/chat/completions both respond (a 401 means the route exists and is auth-gated, exactly as expected).
Streaming works the same way
If your app streams output, you don't rewrite it. stream=True and the same SSE delta shape carry over:
stream = client.chat.completions.create(
model="qwen-max",
messages=[
{"role": "user", "content": "Draft a concise commit message for a login-rate-limit fix."}
],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")
Same loop you already run against OpenAI. The only thing that moved is the model string.
Model cheat sheet
These are the real model names behind the gateway — use them directly in the model field, with no -class suffix:
| Model | Family | Best for |
|---|---|---|
deepseek-v4-pro |
DeepSeek | Strongest reasoning, math, and agent workloads |
deepseek-v4-flash |
DeepSeek | Fast and cheap; high-throughput bulk calls |
qwen-max |
Qwen (Alibaba) | Flagship general-purpose model |
qwen-plus |
Qwen (Alibaba) | Mid-tier, balanced cost and quality |
qwen-turbo |
Qwen (Alibaba) | Fastest and cheapest; light tasks |
qwen-max-longcontext |
Qwen (Alibaba) | Flagship with extended context for long inputs |
qwen3-235b-a22b |
Qwen (Alibaba) | Open-weight 235B MoE model |
Switching between them is a one-string change — no code path differs. Route simple text to deepseek-v4-flash or qwen-turbo, and reserve deepseek-v4-pro or qwen-max for the hard reasoning calls.
Advanced: tool calling still works
Because the endpoint follows the OpenAI schema, function calling is available the same way. Define your tools, send them, and read tool_calls back:
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}]
resp = client.chat.completions.create(
model="deepseek-v4-flash",
messages=[{"role": "user", "content": "What's the weather in Tokyo?"}],
tools=tools,
)
print(resp.choices[0].message.tool_calls)
Run the tool, append the result as a tool role message, and loop until the model finishes — the identical pattern you use with OpenAI. No adapter layer required.
Common gotchas
A few things to check once you repoint, so a surprise doesn't read as a bug:
-
Billing is metered by the
usageobject your response returns. Readusagefrom the API answer and bill or log from there; don't estimate from your own counts, since backends count units their own way. -
-classtiers are compatible labels, not the official model. On the pricing page you'll see tiers likegpt-4o-class. Those are compatible tiers — routing labels that map to a real DeepSeek or Qwen backend — not the original OpenAI model running underneath. They're a shorthand for "acts like this class." If you want the actual DeepSeek or Qwen model, use the real names from the cheat sheet above. -
Streaming behavior is consistent. The SSE format, delta fields, and
finish_reasonfollow the OpenAI shape, so your existing streaming parser keeps working.
One more practical note: an extra network hop adds latency versus calling a provider directly, so measure p50/p95 on your real payload sizes before you lean on it in production.
Wrap-up
Migrating off OpenAI doesn't mean rewriting your integration. If your client already speaks the OpenAI Chat Completions format, you change base_url and model, keep your retries and tool logic, and you're calling real DeepSeek and Qwen models through one endpoint. Spend your 500 free credits on a small, representative eval — format adherence, instruction following, reasoning, latency — and decide from your own numbers.
Ready to try it? Sign up by email, grab your 500 free credits, and repoint your client:
👉 https://api.taotok.io/go/devto9.html
Building something and want to compare notes with other developers routing across providers? There is a small Discord where people share setups and war stories: https://discord.gg/eEsTYXpZJn
Top comments (1)
Curious — anyone benchmarked deepseek-v4-pro vs gpt-4o on code-gen tasks? Trying to decide where the extra network hop is worth it. What is your p50 like?