<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: 1261398983</title>
    <description>The latest articles on DEV Community by 1261398983 (@1261398983).</description>
    <link>https://dev.to/1261398983</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4152787%2F540fabae-1365-4ee2-aff8-3fb5b16d6e98.png</url>
      <title>DEV Community: 1261398983</title>
      <link>https://dev.to/1261398983</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/1261398983"/>
    <language>en</language>
    <item>
      <title>I Spent 28.67 CNY on 4.07 Billion Tokens: A Real Cost Breakdown for LLM API Relays</title>
      <dc:creator>1261398983</dc:creator>
      <pubDate>Wed, 30 Sep 2026 18:12:21 +0000</pubDate>
      <link>https://dev.to/1261398983/i-spent-2867-cny-on-407-billion-tokens-a-real-cost-breakdown-for-llm-api-relays-4dam</link>
      <guid>https://dev.to/1261398983/i-spent-2867-cny-on-407-billion-tokens-a-real-cost-breakdown-for-llm-api-relays-4dam</guid>
      <description>&lt;p&gt;Cost claims about LLM API relays are usually unfalsifiable marketing. Here is a breakdown I can actually defend, because it comes from my own billing dashboard and from request-level data I pulled myself.&lt;/p&gt;

&lt;h2&gt;
  
  
  The headline numbers
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Requests&lt;/td&gt;
&lt;td&gt;30,734&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tokens&lt;/td&gt;
&lt;td&gt;4.07 billion&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Paid&lt;/td&gt;
&lt;td&gt;28.67 CNY (~4 USD)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;List price equivalent&lt;/td&gt;
&lt;td&gt;359.51 CNY (~50 USD)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Effective ratio&lt;/td&gt;
&lt;td&gt;7.97%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That is a 92% reduction against list price. The interesting part is not the number, it is why it holds up.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mechanism 1: prompt cache hits
&lt;/h2&gt;

&lt;p&gt;This is the single most under-discussed lever. Here is a real call from my logs:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Tokens&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;prompt_tokens&lt;/td&gt;
&lt;td&gt;1,755&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;prompt_cache_hit_tokens&lt;/td&gt;
&lt;td&gt;1,536&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;prompt_cache_miss_tokens&lt;/td&gt;
&lt;td&gt;219&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;completion_tokens&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;87.5% of the input was served from cache. For agentic workflows - Claude Code, Codex CLI, any tool that re-sends a large system prompt and a long file context on every turn - this dominates the bill. If your provider does not expose cache hit/miss split in the response, you are flying blind.&lt;/p&gt;

&lt;p&gt;You can verify it yourself. The usage block in the response looks like this:&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;{
  "usage": {
    "prompt_tokens": 1755,
    "prompt_cache_hit_tokens": 1536,
    "prompt_cache_miss_tokens": 219,
    "completion_tokens": 20
  }
}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;If those two cache fields are absent, ask your provider why.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mechanism 2: group rate multipliers
&lt;/h2&gt;

&lt;p&gt;Many relays expose tiered rate multipliers per model group. In my case the open-weight group bills at 0.08x of list price, and a separate premium group at 0.22x. The multiplier is published in the console, not hidden in a PDF.&lt;/p&gt;

&lt;p&gt;Cache hits and multipliers compound: 0.08x applied on top of cache-priced input tokens. That is how you land at 8% of list.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I actually run through one endpoint
&lt;/h2&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;deepseek-v4-flash      deepseek-v4.1-flash    deepseek-v4-pro
glm-5.2                glm-5.3                glm-5.3-flash
kimi-k2.8              kimi-k3                minimax-m3
hy3                    hy4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;Both protocols terminate on the same base URL, which matters because it means one credential and one bill:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Endpoint&lt;/th&gt;
&lt;th&gt;Protocol&lt;/th&gt;
&lt;th&gt;Status&lt;/th&gt;
&lt;th&gt;Latency&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GET /v1/models&lt;/td&gt;
&lt;td&gt;OpenAI&lt;/td&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;td&gt;1.61s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;POST /v1/chat/completions&lt;/td&gt;
&lt;td&gt;OpenAI&lt;/td&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;td&gt;2.19s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;POST /v1/responses&lt;/td&gt;
&lt;td&gt;OpenAI Responses&lt;/td&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;td&gt;2.37s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;POST /v1/messages&lt;/td&gt;
&lt;td&gt;Anthropic&lt;/td&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;td&gt;3.94s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  How to sanity-check any relay before you commit
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Send a request without an API key and confirm you get 401, not a helpful error page.&lt;/li&gt;
&lt;li&gt;Call /v1/models and check the model list matches what the pricing page claims.&lt;/li&gt;
&lt;li&gt;Read the usage block and confirm cache hit/miss fields exist.&lt;/li&gt;
&lt;li&gt;Send the same prompt twice and verify the second call is cheaper. That is cache working.&lt;/li&gt;
&lt;li&gt;Cross-check one model's answer against the official API if you can. This is the only real test of whether you are getting the model you paid for.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Caveats worth stating plainly
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Aggregators are not the official vendor. No enterprise SLA, no data-residency guarantee. If you need those, pay official prices.&lt;/li&gt;
&lt;li&gt;Cheap relays can and do disappear. Top up small amounts. I keep a month of runway at most.&lt;/li&gt;
&lt;li&gt;Latency figures above come from one residential connection in Asia. Your mileage will differ.&lt;/li&gt;
&lt;li&gt;I am on a referral program, so the last link pays me roughly 10%. The plain domain behaves identically.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Endpoint: &lt;a href="https://api.dshapi.icu/v1" rel="noopener noreferrer"&gt;https://api.dshapi.icu/v1&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Getting started: &lt;a href="https://1261398983.github.io/ai-api-guide/" rel="noopener noreferrer"&gt;https://1261398983.github.io/ai-api-guide/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Referral (optional): &lt;a href="https://api.dshapi.icu/r/T8KiaeGU" rel="noopener noreferrer"&gt;https://api.dshapi.icu/r/T8KiaeGU&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you have run the double-send cache test on a different relay and gotten a different result, I would like to see the numbers.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>costoptimization</category>
      <category>llm</category>
    </item>
    <item>
      <title>Running Claude Code and Codex CLI Behind a Single OpenAI-Compatible Endpoint</title>
      <dc:creator>1261398983</dc:creator>
      <pubDate>Wed, 30 Sep 2026 17:19:28 +0000</pubDate>
      <link>https://dev.to/1261398983/running-claude-code-and-codex-cli-behind-a-single-openai-compatible-endpoint-5g6m</link>
      <guid>https://dev.to/1261398983/running-claude-code-and-codex-cli-behind-a-single-openai-compatible-endpoint-5g6m</guid>
      <description>&lt;p&gt;If you are using AI coding agents like Claude Code or Codex CLI and you keep hitting the same three walls - regional access, overseas card requirements, and one API key per model - this post is a practical writeup of how I consolidated everything into a single endpoint.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem
&lt;/h2&gt;

&lt;p&gt;My setup used to look like this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Claude Code -&amp;gt; Anthropic key (needs an overseas card)&lt;/li&gt;
&lt;li&gt;Codex CLI -&amp;gt; OpenAI key (needs an overseas card)&lt;/li&gt;
&lt;li&gt;Local tools (Cherry Studio / LobeChat / NextChat) -&amp;gt; each with its own key&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Three wallets, three billing dashboards, three points of failure.&lt;/p&gt;

&lt;h2&gt;
  
  
  The approach: one base_url, both protocols
&lt;/h2&gt;

&lt;p&gt;The cleanest fix is a gateway that speaks both the OpenAI and the Anthropic protocol on the same host, so every tool points at one base_url.&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;export OPENAI_BASE_URL="https://api.dshapi.icu/v1"
export OPENAI_API_KEY="sk-..."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;For Claude Code:&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;export ANTHROPIC_BASE_URL="https://api.dshapi.icu"
export ANTHROPIC_AUTH_TOKEN="sk-..."
export ANTHROPIC_MODEL="claude-sonnet-4-5"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;That is the whole config. No per-tool proxy, no key juggling.&lt;/p&gt;

&lt;h2&gt;
  
  
  Endpoint verification (I actually ran these)
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Endpoint&lt;/th&gt;
&lt;th&gt;Protocol&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;th&gt;Latency&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GET /v1/models&lt;/td&gt;
&lt;td&gt;OpenAI&lt;/td&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;td&gt;1.61s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;POST /v1/chat/completions&lt;/td&gt;
&lt;td&gt;OpenAI&lt;/td&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;td&gt;2.19s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;POST /v1/responses&lt;/td&gt;
&lt;td&gt;OpenAI Responses&lt;/td&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;td&gt;2.37s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;POST /v1/messages&lt;/td&gt;
&lt;td&gt;Anthropic&lt;/td&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;td&gt;3.94s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Requesting /v1/models without a key correctly returns 401, so auth is enforced.&lt;/p&gt;

&lt;p&gt;Quick smoke test:&lt;/p&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;curl &lt;a href="https://api.dshapi.icu/v1/models" rel="noopener noreferrer"&gt;https://api.dshapi.icu/v1/models&lt;/a&gt; -H "Authorization: Bearer $OPENAI_API_KEY"&lt;br&gt;
curl &lt;a href="https://api.dshapi.icu/v1/chat/completions" rel="noopener noreferrer"&gt;https://api.dshapi.icu/v1/chat/completions&lt;/a&gt; \&lt;br&gt;
  -H "Authorization: Bearer $OPENAI_API_KEY" \&lt;br&gt;
  -H "Content-Type: application/json" \&lt;br&gt;
  -d '{"model":"deepseek-v4-flash","messages":[{"role":"user","content":"hi"}]}'&lt;br&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;
&lt;h2&gt;
&lt;br&gt;
  &lt;br&gt;
  &lt;br&gt;
  Cost: where the savings actually come from&lt;br&gt;
&lt;/h2&gt;

&lt;p&gt;Cost numbers without a mechanism are just marketing. Here is the mechanism, taken from a real call:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;prompt_tokens&lt;/td&gt;
&lt;td&gt;1,755&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;prompt_cache_hit_tokens&lt;/td&gt;
&lt;td&gt;1,536 (87.5%)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;prompt_cache_miss_tokens&lt;/td&gt;
&lt;td&gt;219&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;completion_tokens&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two things compound:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Prompt cache hits - 87.5% of input tokens were served from cache, which bills far below normal input.&lt;/li&gt;
&lt;li&gt;Group multiplier - the open-weights group bills at 0.08x of list price.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Aggregated over 30,734 requests and 4.07B tokens: 28.67 CNY paid against 359.51 CNY at list price, i.e. 7.97%.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models available
&lt;/h2&gt;


&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;deepseek-v4-flash      deepseek-v4.1-flash    deepseek-v4-pro&lt;br&gt;
glm-5.2                glm-5.3                glm-5.3-flash&lt;br&gt;
kimi-k2.8              kimi-k3                minimax-m3&lt;br&gt;
hy3                    hy4&lt;br&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;
&lt;h2&gt;
&lt;br&gt;
  &lt;br&gt;
  &lt;br&gt;
  Registration friction&lt;br&gt;
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;Requirement&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Email&lt;/td&gt;
&lt;td&gt;Any (QQ mail works)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Overseas credit card&lt;/td&gt;
&lt;td&gt;Not required&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Payment&lt;/td&gt;
&lt;td&gt;WeChat Pay / Alipay&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Billing&lt;/td&gt;
&lt;td&gt;Pay-as-you-go, balance does not expire&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Caveats
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;This is an aggregator, not the official vendor. If you need enterprise SLAs or data-residency guarantees, use the official APIs.&lt;/li&gt;
&lt;li&gt;I am on the referral program, so the last link pays me roughly 10%. The plain domain works identically if you would rather avoid that.&lt;/li&gt;
&lt;li&gt;Latency is region dependent; these measurements come from a residential connection in Asia.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Links
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Gateway endpoint: &lt;a href="https://api.dshapi.icu/v1" rel="noopener noreferrer"&gt;https://api.dshapi.icu/v1&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Getting-started page: &lt;a href="https://1261398983.github.io/ai-api-guide/" rel="noopener noreferrer"&gt;https://1261398983.github.io/ai-api-guide/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Referral (optional): &lt;a href="https://api.dshapi.icu/r/T8KiaeGU" rel="noopener noreferrer"&gt;https://api.dshapi.icu/r/T8KiaeGU&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you run a working two-protocol setup on a different gateway, I would like to hear how you handle model-name mapping across providers.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>opensource</category>
      <category>tooling</category>
    </item>
  </channel>
</rss>
