Cost claims about LLM API relays are usually unfalsifiable marketing. Here is a breakdown I can actually defend, because it comes from my own billing dashboard and from request-level data I pulled myself.
The headline numbers
| Metric | Value |
|---|---|
| Requests | 30,734 |
| Tokens | 4.07 billion |
| Paid | 28.67 CNY (~4 USD) |
| List price equivalent | 359.51 CNY (~50 USD) |
| Effective ratio | 7.97% |
That is a 92% reduction against list price. The interesting part is not the number, it is why it holds up.
Mechanism 1: prompt cache hits
This is the single most under-discussed lever. Here is a real call from my logs:
| Metric | Tokens |
|---|---|
| prompt_tokens | 1,755 |
| prompt_cache_hit_tokens | 1,536 |
| prompt_cache_miss_tokens | 219 |
| completion_tokens | 20 |
87.5% of the input was served from cache. For agentic workflows - Claude Code, Codex CLI, any tool that re-sends a large system prompt and a long file context on every turn - this dominates the bill. If your provider does not expose cache hit/miss split in the response, you are flying blind.
You can verify it yourself. The usage block in the response looks like this:
{
"usage": {
"prompt_tokens": 1755,
"prompt_cache_hit_tokens": 1536,
"prompt_cache_miss_tokens": 219,
"completion_tokens": 20
}
}
If those two cache fields are absent, ask your provider why.
Mechanism 2: group rate multipliers
Many relays expose tiered rate multipliers per model group. In my case the open-weight group bills at 0.08x of list price, and a separate premium group at 0.22x. The multiplier is published in the console, not hidden in a PDF.
Cache hits and multipliers compound: 0.08x applied on top of cache-priced input tokens. That is how you land at 8% of list.
What I actually run through one endpoint
deepseek-v4-flash deepseek-v4.1-flash deepseek-v4-pro
glm-5.2 glm-5.3 glm-5.3-flash
kimi-k2.8 kimi-k3 minimax-m3
hy3 hy4
Both protocols terminate on the same base URL, which matters because it means one credential and one bill:
| Endpoint | Protocol | Status | Latency |
|---|---|---|---|
| GET /v1/models | OpenAI | 200 | 1.61s |
| POST /v1/chat/completions | OpenAI | 200 | 2.19s |
| POST /v1/responses | OpenAI Responses | 200 | 2.37s |
| POST /v1/messages | Anthropic | 200 | 3.94s |
How to sanity-check any relay before you commit
- Send a request without an API key and confirm you get 401, not a helpful error page.
- Call /v1/models and check the model list matches what the pricing page claims.
- Read the usage block and confirm cache hit/miss fields exist.
- Send the same prompt twice and verify the second call is cheaper. That is cache working.
- Cross-check one model's answer against the official API if you can. This is the only real test of whether you are getting the model you paid for.
Caveats worth stating plainly
- Aggregators are not the official vendor. No enterprise SLA, no data-residency guarantee. If you need those, pay official prices.
- Cheap relays can and do disappear. Top up small amounts. I keep a month of runway at most.
- Latency figures above come from one residential connection in Asia. Your mileage will differ.
- I am on a referral program, so the last link pays me roughly 10%. The plain domain behaves identically.
Links
- Endpoint: https://api.dshapi.icu/v1
- Getting started: https://1261398983.github.io/ai-api-guide/
- Referral (optional): https://api.dshapi.icu/r/T8KiaeGU
If you have run the double-send cache test on a different relay and gotten a different result, I would like to see the numbers.
Top comments (0)