DEV Community

1261398983
1261398983

Posted on

Running Claude Code and Codex CLI Behind a Single OpenAI-Compatible Endpoint

If you are using AI coding agents like Claude Code or Codex CLI and you keep hitting the same three walls - regional access, overseas card requirements, and one API key per model - this post is a practical writeup of how I consolidated everything into a single endpoint.

The problem

My setup used to look like this:

  • Claude Code -> Anthropic key (needs an overseas card)
  • Codex CLI -> OpenAI key (needs an overseas card)
  • Local tools (Cherry Studio / LobeChat / NextChat) -> each with its own key

Three wallets, three billing dashboards, three points of failure.

The approach: one base_url, both protocols

The cleanest fix is a gateway that speaks both the OpenAI and the Anthropic protocol on the same host, so every tool points at one base_url.

export OPENAI_BASE_URL="https://api.dshapi.icu/v1"
export OPENAI_API_KEY="sk-..."
Enter fullscreen mode Exit fullscreen mode

For Claude Code:

export ANTHROPIC_BASE_URL="https://api.dshapi.icu"
export ANTHROPIC_AUTH_TOKEN="sk-..."
export ANTHROPIC_MODEL="claude-sonnet-4-5"
Enter fullscreen mode Exit fullscreen mode

That is the whole config. No per-tool proxy, no key juggling.

Endpoint verification (I actually ran these)

Endpoint Protocol Result Latency
GET /v1/models OpenAI 200 1.61s
POST /v1/chat/completions OpenAI 200 2.19s
POST /v1/responses OpenAI Responses 200 2.37s
POST /v1/messages Anthropic 200 3.94s

Requesting /v1/models without a key correctly returns 401, so auth is enforced.

Quick smoke test:

curl https://api.dshapi.icu/v1/models -H "Authorization: Bearer $OPENAI_API_KEY"
curl https://api.dshapi.icu/v1/chat/completions \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"deepseek-v4-flash","messages":[{"role":"user","content":"hi"}]}'
Enter fullscreen mode Exit fullscreen mode




Cost: where the savings actually come from

Cost numbers without a mechanism are just marketing. Here is the mechanism, taken from a real call:

Metric Value
prompt_tokens 1,755
prompt_cache_hit_tokens 1,536 (87.5%)
prompt_cache_miss_tokens 219
completion_tokens 20

Two things compound:

  1. Prompt cache hits - 87.5% of input tokens were served from cache, which bills far below normal input.
  2. Group multiplier - the open-weights group bills at 0.08x of list price.

Aggregated over 30,734 requests and 4.07B tokens: 28.67 CNY paid against 359.51 CNY at list price, i.e. 7.97%.

Models available

deepseek-v4-flash      deepseek-v4.1-flash    deepseek-v4-pro
glm-5.2 glm-5.3 glm-5.3-flash
kimi-k2.8 kimi-k3 minimax-m3
hy3 hy4
Enter fullscreen mode Exit fullscreen mode




Registration friction

Item Requirement
Email Any (QQ mail works)
Overseas credit card Not required
Payment WeChat Pay / Alipay
Billing Pay-as-you-go, balance does not expire

Caveats

  • This is an aggregator, not the official vendor. If you need enterprise SLAs or data-residency guarantees, use the official APIs.
  • I am on the referral program, so the last link pays me roughly 10%. The plain domain works identically if you would rather avoid that.
  • Latency is region dependent; these measurements come from a residential connection in Asia.

Links

If you run a working two-protocol setup on a different gateway, I would like to hear how you handle model-name mapping across providers.

Top comments (0)