DEV Community

XHuoAPI
XHuoAPI

Posted on Originally published at xhuoapi.ai

Running Claude Code on DeepSeek, GLM and Qwen with two environment variables

Claude Code talks to its backend in the Anthropic Messages format. Any endpoint that accepts that format can stand in for Anthropic's API, which means you can run the same agent loop on DeepSeek, GLM, Qwen or Kimi without a proxy or a converter.

I tested this on Claude Code 2.1.284 by asking it to create and write a real file with each model, and re-ran the DeepSeek case on 2.1.296 today. Here is the setup and what I found.

Disclosure: I run XHuoAPI, the gateway used below. The same two variables work with any Anthropic-compatible endpoint.

Setup

Install Claude Code (Node.js 18+):

npm install -g @anthropic-ai/claude-code
Enter fullscreen mode Exit fullscreen mode

Point it at the gateway. Leave /v1 off the base URL; Claude Code appends /v1/messages itself:

export ANTHROPIC_BASE_URL=https://api.xhuoapi.ai
export ANTHROPIC_AUTH_TOKEN=YOUR_KEY
export ANTHROPIC_MODEL=deepseek-v3.2
export ANTHROPIC_DEFAULT_HAIKU_MODEL=deepseek-v3.2
Enter fullscreen mode Exit fullscreen mode

Or put the same values in the env block of ~/.claude/settings.json if you'd rather not touch your shell config.

Check the connection with a one-off task:

claude -p "Reply with the single word: ok"
Enter fullscreen mode Exit fullscreen mode

Switching models

Change ANTHROPIC_MODEL and ANTHROPIC_DEFAULT_HAIKU_MODEL together. The second one matters: Claude Code sends some background work to a "Haiku-class" model, and if that variable still points at a model your endpoint doesn't serve, those background calls fail quietly.

What happened in the test

Same task for every model: create a file and write a small program into it.

Model Result Time
claude-sonnet-4-6 worked first try ~7 s
deepseek-v3.2 worked first try ~9 s
qwen3.8-max worked first try ~24 s
glm-5.3 worked first try ~45 s
kimi-k3 failed once, worked on retry —

Two things worth knowing:

  • Prompt caching does a lot of work here. Claude Code sends a long system prompt with every request. The first request wrote about 36K tokens to the cache; later ones read about 64K from it. Cache reads are billed well below normal input on most models, so long sessions cost less than the raw token count suggests.
  • The unrecognized_model line is expected. Since 2.1.290, Claude Code auto-compacts sessions on model IDs it doesn't know (like deepseek-v3.2) within an assumed context window and prints [claude-code:unrecognized_model]. The task still runs. Set CLAUDE_CODE_DISABLE_UNKNOWN_MODEL_WINDOW_ENFORCEMENT=1 if you want the old behavior.
  • ANTHROPIC_AUTH_TOKEN vs ANTHROPIC_API_KEY. The first is sent as Authorization: Bearer, the second as x-api-key. Use AUTH_TOKEN to avoid the interactive confirmation prompt Claude Code shows the first time it sees an API_KEY.

The full guide with the settings-file version and per-model notes is here: https://xhuoapi.ai/en/guides/claude-code.html?utm_source=devto&utm_medium=article&utm_campaign=claude-code

Top comments (1)

Some comments may only be visible to logged-in visitors. Sign in to view all comments.