DEV Community

1261398983
1261398983

Posted on

I Spent 28.67 CNY on 4.07 Billion Tokens: A Real Cost Breakdown for LLM API Relays

Cost claims about LLM API relays are usually unfalsifiable marketing. Here is a breakdown I can actually defend, because it comes from my own billing dashboard and from request-level data I pulled myself.

The headline numbers

Metric Value
Requests 30,734
Tokens 4.07 billion
Paid 28.67 CNY (~4 USD)
List price equivalent 359.51 CNY (~50 USD)
Effective ratio 7.97%

That is a 92% reduction against list price. The interesting part is not the number, it is why it holds up.

Mechanism 1: prompt cache hits

This is the single most under-discussed lever. Here is a real call from my logs:

Metric Tokens
prompt_tokens 1,755
prompt_cache_hit_tokens 1,536
prompt_cache_miss_tokens 219
completion_tokens 20

87.5% of the input was served from cache. For agentic workflows - Claude Code, Codex CLI, any tool that re-sends a large system prompt and a long file context on every turn - this dominates the bill. If your provider does not expose cache hit/miss split in the response, you are flying blind.

You can verify it yourself. The usage block in the response looks like this:

{
  "usage": {
    "prompt_tokens": 1755,
    "prompt_cache_hit_tokens": 1536,
    "prompt_cache_miss_tokens": 219,
    "completion_tokens": 20
  }
}
Enter fullscreen mode Exit fullscreen mode

If those two cache fields are absent, ask your provider why.

Mechanism 2: group rate multipliers

Many relays expose tiered rate multipliers per model group. In my case the open-weight group bills at 0.08x of list price, and a separate premium group at 0.22x. The multiplier is published in the console, not hidden in a PDF.

Cache hits and multipliers compound: 0.08x applied on top of cache-priced input tokens. That is how you land at 8% of list.

What I actually run through one endpoint

deepseek-v4-flash      deepseek-v4.1-flash    deepseek-v4-pro
glm-5.2                glm-5.3                glm-5.3-flash
kimi-k2.8              kimi-k3                minimax-m3
hy3                    hy4
Enter fullscreen mode Exit fullscreen mode

Both protocols terminate on the same base URL, which matters because it means one credential and one bill:

Endpoint Protocol Status Latency
GET /v1/models OpenAI 200 1.61s
POST /v1/chat/completions OpenAI 200 2.19s
POST /v1/responses OpenAI Responses 200 2.37s
POST /v1/messages Anthropic 200 3.94s

How to sanity-check any relay before you commit

  1. Send a request without an API key and confirm you get 401, not a helpful error page.
  2. Call /v1/models and check the model list matches what the pricing page claims.
  3. Read the usage block and confirm cache hit/miss fields exist.
  4. Send the same prompt twice and verify the second call is cheaper. That is cache working.
  5. Cross-check one model's answer against the official API if you can. This is the only real test of whether you are getting the model you paid for.

Caveats worth stating plainly

  • Aggregators are not the official vendor. No enterprise SLA, no data-residency guarantee. If you need those, pay official prices.
  • Cheap relays can and do disappear. Top up small amounts. I keep a month of runway at most.
  • Latency figures above come from one residential connection in Asia. Your mileage will differ.
  • I am on a referral program, so the last link pays me roughly 10%. The plain domain behaves identically.

Links

If you have run the double-send cache test on a different relay and gotten a different result, I would like to see the numbers.

Top comments (0)