<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Antonne Dillard</title>
    <description>The latest articles on DEV Community by Antonne Dillard (@tony_dillard).</description>
    <link>https://dev.to/tony_dillard</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4050452%2F8ef16b25-cc08-4d78-a208-1ca2f0227101.png</url>
      <title>DEV Community: Antonne Dillard</title>
      <link>https://dev.to/tony_dillard</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/tony_dillard"/>
    <language>en</language>
    <item>
      <title>One Key for Claude Code, Codex, and dsh: The OpenAI-Compatible Pattern</title>
      <dc:creator>Antonne Dillard</dc:creator>
      <pubDate>Tue, 01 Sep 2026 01:26:17 +0000</pubDate>
      <link>https://dev.to/tony_dillard/one-key-for-claude-code-codex-and-dsh-the-openai-compatible-pattern-4ka6</link>
      <guid>https://dev.to/tony_dillard/one-key-for-claude-code-codex-and-dsh-the-openai-compatible-pattern-4ka6</guid>
      <description>

&lt;p&gt;If you run Claude Code for hard problems, Codex for OpenAI-ecosystem work, and DeepSeek Harness (dsh) for cheap high-volume agents, you currently juggle three separate accounts, three API keys, and three billing pipelines. That is the problem the OpenAI-compatible endpoint pattern exists to remove: one base URL, one key, and each tool keeps speaking its own protocol.&lt;/p&gt;

&lt;p&gt;This post shows the pattern concretely — the protocol map, the exact environment variables, and the code samples straight from an integration doc that already supports it. The pattern generalizes to any provider that exposes the same endpoints. I'll use the TeamoRouter endpoint (&lt;code&gt;https://api.teamorouter.cn&lt;/code&gt;, keys &lt;code&gt;sk-teamo-...&lt;/code&gt;) as the concrete example, because that is the documentation being quoted.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pattern: one base URL, many protocols
&lt;/h2&gt;

&lt;p&gt;The key fact is that modern coding agents are not locked to one wire format. Each is configured with a base URL and an API key, and each speaks its own protocol on top:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Claude Code&lt;/strong&gt; speaks the &lt;strong&gt;Anthropic protocol&lt;/strong&gt; (&lt;code&gt;POST /v1/messages&lt;/code&gt;) — the same format the &lt;code&gt;anthropic&lt;/code&gt; SDK uses.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Codex&lt;/strong&gt; speaks the &lt;strong&gt;OpenAI protocol&lt;/strong&gt;. That ecosystem exposes two endpoints: &lt;code&gt;POST /v1/chat/completions&lt;/code&gt; (Chat Completions) and &lt;code&gt;POST /v1/responses&lt;/code&gt; (Responses API). The doc notes &lt;code&gt;/v1/responses&lt;/code&gt; works &lt;strong&gt;for GPT-series models only&lt;/strong&gt; — Claude and Gemini get a 400 there.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;dsh&lt;/strong&gt; also speaks the &lt;strong&gt;OpenAI protocol&lt;/strong&gt; via &lt;code&gt;DEEPSEEK_BASE_URL&lt;/code&gt;, hitting &lt;code&gt;/v1/chat/completions&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gemini CLI&lt;/strong&gt; speaks the &lt;strong&gt;Gemini native protocol&lt;/strong&gt; (&lt;code&gt;/v1beta/models/{model}:generateContent&lt;/code&gt;).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The same holds on the language-SDK side: the &lt;code&gt;openai&lt;/code&gt; SDK points at the OpenAI-compatible base URL; the &lt;code&gt;anthropic&lt;/code&gt; SDK points at the Anthropic base URL; both share the same key.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Protocol&lt;/th&gt;
&lt;th&gt;Endpoint&lt;/th&gt;
&lt;th&gt;Auth header&lt;/th&gt;
&lt;th&gt;Typical tool&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Anthropic&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/v1/messages&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;x-api-key&lt;/code&gt; + &lt;code&gt;anthropic-version: 2023-06-01&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Claude Code, &lt;code&gt;anthropic&lt;/code&gt; SDK&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI Chat Completions&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/v1/chat/completions&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;Authorization: Bearer&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;dsh, GLM, &lt;code&gt;openai&lt;/code&gt; SDK&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI Responses&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;/v1/responses&lt;/code&gt; (GPT only)&lt;/td&gt;
&lt;td&gt;&lt;code&gt;Authorization: Bearer&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Codex-era apps&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini native&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/v1beta/models/{model}:generateContent&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;Authorization: Bearer&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Gemini CLI&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;One key, &lt;code&gt;sk-teamo-&amp;lt;your-key&amp;gt;&lt;/code&gt;, is accepted by every row.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pointing each tool at the same base URL
&lt;/h2&gt;

&lt;p&gt;The integration doc gives the exact environment variables for each agent tool. These are the ones to export. Note the &lt;code&gt;/v1&lt;/code&gt; detail: OpenAI-style SDKs need the &lt;code&gt;/v1&lt;/code&gt; suffix on the base URL, while the Anthropic SDK does not.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Claude Code (Anthropic protocol):&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;ANTHROPIC_BASE_URL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"https://api.teamorouter.cn"&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;ANTHROPIC_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"sk-teamo-&amp;lt;your-key&amp;gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Codex / OpenAI-protocol tools (Chat Completions or Responses):&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;OPENAI_BASE_URL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"https://api.teamorouter.cn/v1"&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;OPENAI_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"sk-teamo-&amp;lt;your-key&amp;gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;dsh (OpenAI protocol):&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;DEEPSEEK_BASE_URL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"https://api.teamorouter.cn/v1"&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;DEEPSEEK_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"sk-teamo-&amp;lt;your-key&amp;gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Gemini CLI (Gemini protocol):&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;GOOGLE_GEMINI_BASE_URL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"https://api.teamorouter.cn"&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;GEMINI_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"sk-teamo-&amp;lt;your-key&amp;gt;"&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;GEMINI_API_KEY_AUTH_MECHANISM&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"bearer"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Drop these into your shell profile (or prefix a command with them) and each tool routes to the same endpoint with the same key.&lt;/p&gt;

&lt;h2&gt;
  
  
  The same pattern in raw HTTP and SDKs
&lt;/h2&gt;

&lt;p&gt;Calling the endpoints directly, the requests are standard. Anthropic protocol:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl https://api.teamorouter.cn/v1/messages &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"x-api-key: sk-teamo-&amp;lt;your-key&amp;gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"anthropic-version: 2023-06-01"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"content-type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"model":"claude-fable-5","max_tokens":1024,"messages":[{"role":"user","content":"Introduce yourself in one sentence."}]}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;OpenAI protocol (Chat Completions):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl https://api.teamorouter.cn/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer sk-teamo-&amp;lt;your-key&amp;gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"content-type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"model":"glm-5.3","max_tokens":1024,"messages":[{"role":"user","content":"Hello"}]}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;OpenAI Responses API (GPT-series):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl https://api.teamorouter.cn/v1/responses &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer sk-teamo-&amp;lt;your-key&amp;gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"content-type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"model":"gpt-5.6-sol","input":"Introduce TeamoRouter in one sentence."}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Gemini native:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="s2"&gt;"https://api.teamorouter.cn/v1beta/models/gemini-3.5-flash:generateContent"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer sk-teamo-&amp;lt;your-key&amp;gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"content-type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"contents":[{"role":"user","parts":[{"text":"Introduce yourself."}]}]}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And in the official SDKs, you only change the constructor. The &lt;code&gt;anthropic&lt;/code&gt; SDK (base URL without &lt;code&gt;/v1&lt;/code&gt;):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;anthropic&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Anthropic&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Anthropic&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sk-teamo-&amp;lt;your-key&amp;gt;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.teamorouter.cn&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-fable-5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Hello&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;openai&lt;/code&gt; SDK (note the &lt;code&gt;/v1&lt;/code&gt;):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sk-teamo-&amp;lt;your-key&amp;gt;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.teamorouter.cn/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gpt-5.6-sol&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Hello&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If your app has already migrated to the Responses API, the &lt;code&gt;openai&lt;/code&gt; SDK's &lt;code&gt;client.responses.create(...)&lt;/code&gt; works against the same base URL.&lt;/p&gt;

&lt;h3&gt;
  
  
  Model discovery and the fast-mode option
&lt;/h3&gt;

&lt;p&gt;Before wiring anything up, hit &lt;code&gt;GET /v1/models&lt;/code&gt; for the live model list — it lets you confirm exact model IDs (all lowercase, mind the &lt;code&gt;-&lt;/code&gt; vs &lt;code&gt;.&lt;/code&gt; distinction) before a tool burns requests on a misspelled name:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl https://api.teamorouter.cn/v1/models &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer sk-teamo-&amp;lt;your-key&amp;gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On the OpenAI side, if your provider supports it, the doc also shows a per-request fast mode. Add &lt;code&gt;"service_tier": "fast"&lt;/code&gt; to a Chat Completions or Responses request; the doc quotes it at up to &lt;strong&gt;2.5x standard speed&lt;/strong&gt; for &lt;code&gt;gpt-5.6-sol&lt;/code&gt;, bills at &lt;strong&gt;2x the standard rate&lt;/strong&gt;, and still accepts the older &lt;code&gt;"priority"&lt;/code&gt; value. That is a per-request speed/cost trade you can now make without switching keys.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters
&lt;/h2&gt;

&lt;p&gt;The win is structural, not cosmetic:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;One key, one billing.&lt;/strong&gt; The source puts it directly: one integration, unified billing — you stop registering and topping up across four platforms. One key switches you between DeepSeek, Claude, GPT, and Gemini.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Route by task, not by tool.&lt;/strong&gt; The comparison post's own workflow: hard problems to Claude Code (flagship Claude), daily high-frequency work to dsh (cheap DeepSeek). Both tools point at the same endpoint and the same key.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Protocol correctness is a cost lever.&lt;/strong&gt; This is the subtle one, and the doc is explicit: calling Claude through the OpenAI-compatible format can &lt;strong&gt;lose prompt cache and thinking capabilities&lt;/strong&gt;, which means higher cost and degraded capability — fine for simple chat, wrong for agent workloads. The pattern's discipline — Anthropic protocol for Claude, OpenAI protocol for OpenAI models — is what preserves the caching that keeps bills down.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Honest risks of third-party endpoints
&lt;/h2&gt;

&lt;p&gt;The pattern concentrates your stack on a third-party gateway, and the docs themselves are candid about the edges:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Rate limits on free tiers.&lt;/strong&gt; The free preview models are capped daily: &lt;code&gt;deepseek-v4-pro-free&lt;/code&gt; at &lt;strong&gt;50 requests per day&lt;/strong&gt;, &lt;code&gt;deepseek-v4-flash-free&lt;/code&gt; and &lt;code&gt;glm-5.3-flash-free&lt;/code&gt; at &lt;strong&gt;200 requests per day&lt;/strong&gt;. When the quota is spent, you must switch to the paid equivalents.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Protocol narrowness.&lt;/strong&gt; &lt;code&gt;/v1/responses&lt;/code&gt; is GPT-series only — Claude and Gemini return a 400. Gemini has its own native endpoint; there is no shortcut around picking the right protocol per model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Latency on big models.&lt;/strong&gt; Opus/Fable-class models can take seconds to tens of seconds to emit a first token during the thinking phase; the doc says that is normal, not a failure, and recommends streaming (&lt;code&gt;"stream": true&lt;/code&gt;) plus client read timeouts up to the server's 600-second cap.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You are adding a dependency.&lt;/strong&gt; A third-party endpoint is another system that can rate-limit, degrade, or go down. The source's own FAQ is even-handed: connecting to the official API is the simplest path; a gateway like this makes sense when you want one key across many models, a free tier to start with, or better connectivity from your region.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Key hygiene.&lt;/strong&gt; Keys prefixed &lt;code&gt;sk-teamo-&lt;/code&gt; belong in environment variables or a secret manager — never hardcoded, never committed to Git, never shipped in a client. If a key leaks, revoke it and replace it.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;The OpenAI-compatible endpoint pattern decouples your agent tools from your model provider: one base URL, one key, and each tool keeps its native protocol — Anthropic for Claude Code, OpenAI for Codex and dsh, Gemini for Gemini CLI. The cost is a new dependency on the gateway, so keep official keys handy for the cases where the third-party route adds more risk than it removes. The payoff is one billing pipeline, unified access to every major model family, and the freedom to route each task to the cheapest model that can actually do it.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>opensource</category>
      <category>coding</category>
    </item>
    <item>
      <title>DeepSeek V4: $0.14/$0.28 per Million Tokens, and Agent Economics</title>
      <dc:creator>Antonne Dillard</dc:creator>
      <pubDate>Mon, 31 Aug 2026 00:39:57 +0000</pubDate>
      <link>https://dev.to/tony_dillard/deepseek-v4-014028-per-million-tokens-and-agent-economics-124a</link>
      <guid>https://dev.to/tony_dillard/deepseek-v4-014028-per-million-tokens-and-agent-economics-124a</guid>
      <description>&lt;p&gt;Your agent loop is a token furnace, and the model you feed it decides whether the monthly bill looks like a coffee habit or a car payment. DeepSeek V4 — the model family that powers the DeepSeek Harness (dsh) agent framework by default — is priced at roughly $0.14 per million input tokens and $0.28 per million output tokens. On a single real task that burns half a million tokens, that works out to about a dime. The same task on a flagship-tier model is a different order of magnitude, and that gap changes how often you are willing to let an agent run at all.&lt;/p&gt;

&lt;p&gt;This article walks through the numbers as the DeepSeek ecosystem’s own documentation presents them, does the arithmetic transparently with the assumptions labeled, and gives you a practical rule for when a cheap engine is the right call and when you should still reach for a flagship.&lt;/p&gt;

&lt;p&gt;The price, as published&lt;br&gt;
The figure appears in two of the source documents about the DeepSeek Harness:&lt;/p&gt;

&lt;p&gt;The comparison post “DeepSeek Harness vs Claude Code / Codex / OpenCode” lists dsh’s cost as starting from $0.14/$0.28 per million tokens, attached to its default model, DeepSeek V4 Flash.&lt;br&gt;
The post “DeepSeek Harness: What It Is” describes DeepSeek V4 pricing as about $0.14 / $0.28 per million tokens and calls it a fraction of flagship-model pricing.&lt;br&gt;
So the working numbers you can rely on: ~$0.14 per M input, ~$0.28 per M output. Everything below is computed from those two rates.&lt;/p&gt;

&lt;p&gt;V4 Flash vs V4 Pro&lt;br&gt;
The DeepSeek V4 family in the ecosystem’s model list has two working tiers:&lt;/p&gt;

&lt;p&gt;deepseek-v4-flash— the default engine for dsh, and the tier the $0.14/$0.28 rate is attached to.&lt;br&gt;
deepseek-v4-pro— the higher tier, listed alongside Flash. The published docs do not give Pro’s per-token rate, so I won’t invent one; treat Flash as the cheap workhorse and Pro as the step-up option.&lt;br&gt;
Both also ship as free preview models — deepseek-v4-flash-free and deepseek-v4-pro-free — which is where the rate limits come in (more in the limitations section).&lt;/p&gt;

&lt;p&gt;Why cheap engines power agent stacks&lt;br&gt;
The agent-loop pattern is what makes price per token the decisive variable. As the comparison post puts it, an agent loop is a token black hole — a real task consuming 500,000 tokens is common. The harness (the loop, tools, file system, shell) decides how many tokens a task spends; the model decides what each of those tokens costs. DeepSeek Harness is explicitly model-agnostic: it defaults to DeepSeek V4 but can be pointed at any OpenAI-compatible endpoint with one environment variable. Its sharpest selling point, though, is unit economics — running agents 30–100x cheaper. The same post states that the identical workload often comes to one-hundredth of the cost of a Claude Code flagship session or a Codex GPT-5.6 session.&lt;/p&gt;

&lt;p&gt;Part of the discipline is architectural, not just about the model. dsh is built on the idea that Agent = Model + Harness: the model is the brain, and the harness is everything else — tools, file system, shell, sub-agent orchestration, context access, and knowing when to stop. Because the harness decides what the model actually sees, it also controls the token bill. Its spill storage, for example, writes oversized tool output to disk and hands the model only a locator, so context is no longer blown up by huge text blobs — the source describes this as the real origin of dsh’s “token-saving” reputation. The everything-is-a-plugin design also means you can trim a capability surface (give the agent terminal and file access, but not the network) or swap the provider with one environment variable, without patching a core.&lt;/p&gt;

&lt;p&gt;The strategic bet, in the source’s own framing, is that agent frameworks have to be cheap enough to actually run at scale. The popularity numbers back that up: dsh crossed 20,000 GitHub stars in roughly an hour after open-sourcing, and the first day carried it into the 28,000–31,000 range.&lt;/p&gt;

&lt;p&gt;The economics math&lt;br&gt;
Let’s do the arithmetic on the source’s own “500K tokens per real task” figure. Agent loops are input-heavy: the bulk of tokens are context, file contents, and tool results, with a smaller share of model-generated output. I’ll show two splits so you can see the sensitivity, and the assumption is right there in the table:&lt;/p&gt;

&lt;p&gt;Split (input / output)  Input @ $0.14/M Output @ $0.28/M    Total per task&lt;br&gt;
80/20 → 400K in, 100K out $0.056  $0.028  ≈ $0.08&lt;br&gt;
50/50 → 250K in, 250K out $0.035  $0.070  ≈ $0.11&lt;br&gt;
At either split, one heavy task is roughly eight to eleven cents on DeepSeek V4 Flash. Run ten of those a day and you are under a dollar and a half. Run fifty and you are still in single-digit dollars per day.&lt;/p&gt;

&lt;p&gt;Now compare against the flagship direction, using only the explicit numbers the sources give. A heavy Claude Code session on default flagship models runs to several to tens of dollars, and the same workload on DeepSeek V4 Flash is described as often one-hundredth of that. Eight cents against eight dollars is exactly that ratio, and the 30–100x unit-economics claim brackets the same range. On the Codex side the comparison post is blunter still: GPT-5.6 is a medium-cost default, and Fast mode bills at 2x the standard rate on top. The point isn’t that every task lands at exactly these cents — it’s that the order-of-magnitude gap is real, and it’s arithmetic, not marketing.&lt;/p&gt;

&lt;p&gt;When cheap makes sense — and when it doesn’t&lt;br&gt;
The same sources are refreshingly honest about the trade-off, and their selection rule is worth reproducing in substance:&lt;/p&gt;

&lt;p&gt;Quality-first → Claude Code, whose reasoning depth and tool reliability remain the benchmark. When one correct answer saves you hours, the flagship premium pays for itself.&lt;br&gt;
Ecosystem-first → Codex, if you are already inside OpenAI/ChatGPT and value Fast mode speed and OAuth login convenience.&lt;br&gt;
Cost-first / high-frequency → dsh + DeepSeek V4 Flash or Pro. When agents run all day, per-token price is the whole game.&lt;br&gt;
The practical pattern most people land on is routing by task: hard problems to a flagship Claude, high-volume daily work to cheap DeepSeek. One harness, two models, and the expensive model only appears when the cheap one’s ceiling is actually the bottleneck.&lt;/p&gt;

&lt;p&gt;Honest limitations&lt;br&gt;
dsh is a developer preview. Version 0.1.0-rc.5, explicitly labeled with breaking changes to come. Bet on the concept, not the API surface.&lt;br&gt;
Cheap does not mean flagship quality. Reasoning depth and tool reliability on the flagship tier are still the benchmark; the V4 family trades those for price. If your agent’s correctness ceiling matters more than its burn rate, that is a real cost — just not a per-token one.&lt;br&gt;
Free preview tiers are rate-limited. deepseek-v4-pro-free allows 50 requests per day and deepseek-v4-flash-free allows 200 per day. When the quota is exhausted, the docs tell you to switch to the paid deepseek-v4-pro / deepseek-v4-flash.&lt;br&gt;
Latency is not free either. Large models in these stacks can take seconds to tens of seconds to emit a first token during the thinking phase — a fact the integration docs flag as normal rather than a failure.&lt;br&gt;
Ignore the “Claude Code killer” framing. The sources explicitly call it media narrative — a 0.1.0-rc.5 preview “replacing” a mature product is premature. What is real is the architectural direction, not the headline.&lt;br&gt;
Conclusion&lt;br&gt;
DeepSeek V4’s ~$0.14/$0.28 per million tokens is not a marginal improvement; it moves an agent-heavy workload from “watch the bill” to “forget the bill exists.” The honest counterweight is equally clear in the sources: flagship models still own reasoning depth and tool reliability, so the winning setup is usually a mix — cheap engine for volume, flagship for the hard calls. Price per token decides how many agents you can afford to run; quality per token decides which tasks you should let them touch.&lt;/p&gt;

</description>
      <category>deepseek</category>
      <category>ai</category>
      <category>llm</category>
      <category>agents</category>
    </item>
    <item>
      <title>Why Is Token Collective Procurement Cheaper? A Breakdown of Cost Structure and Cache Hit Rate</title>
      <dc:creator>Antonne Dillard</dc:creator>
      <pubDate>Tue, 25 Aug 2026 02:51:19 +0000</pubDate>
      <link>https://dev.to/tony_dillard/why-is-token-collective-procurement-cheaper-a-breakdown-of-cost-structure-and-cache-hit-rate-10m0</link>
      <guid>https://dev.to/tony_dillard/why-is-token-collective-procurement-cheaper-a-breakdown-of-cost-structure-and-cache-hit-rate-10m0</guid>
      <description>&lt;h2&gt;
  
  
  The One-Sentence Takeaway
&lt;/h2&gt;

&lt;p&gt;Token collective procurement's price advantage is &lt;strong&gt;not an unverifiable low price but the result of four engineering layers&lt;/strong&gt;: discount coefficients, cache-based billing, model tiering, and channel price comparison. When evaluating a procurement platform, you should compare its &lt;strong&gt;effective cost&lt;/strong&gt;, not its nominal unit price.&lt;/p&gt;

&lt;h2&gt;
  
  
  First, the Real Price Structure: Discount Coefficient × Cache Pricing
&lt;/h2&gt;

&lt;p&gt;Using official list prices and platform discount coefficients as an example (per million tokens):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model tier&lt;/th&gt;
&lt;th&gt;Official output price&lt;/th&gt;
&lt;th&gt;Platform discount coefficient&lt;/th&gt;
&lt;th&gt;Platform output price&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Flagship (GPT-5.6 Sol-class)&lt;/td&gt;
&lt;td&gt;$30&lt;/td&gt;
&lt;td&gt;0.1 (10% of list)&lt;/td&gt;
&lt;td&gt;$3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Value (GPT-5.6 Terra-class)&lt;/td&gt;
&lt;td&gt;$12&lt;/td&gt;
&lt;td&gt;0.1 (10% of list)&lt;/td&gt;
&lt;td&gt;$1.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Batch (DeepSeek V4 Flash-class)&lt;/td&gt;
&lt;td&gt;$1.2&lt;/td&gt;
&lt;td&gt;~Official price&lt;/td&gt;
&lt;td&gt;≈$1.2&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;On the cache side, the portion that hits the cache is billed separately at the &lt;strong&gt;cached input price (cached_input)&lt;/strong&gt;, which can be as low as roughly 10% of the official full price. Nominal discounts solve the "unit price"; cache-based billing solves "repeated content" — &lt;strong&gt;the two are orthogonal cost-reduction dimensions&lt;/strong&gt;, and combined, the effective cost is far lower than either dimension alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Four Layers of Effective Cost
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;effective cost = nominal unit price × volume
                × discount coefficient (supplier negotiation + usage scale)
                − cache savings (cache hit rate × cache price gap)
                − tiering savings (right-sizing high-spec tasks to matched models)
                − price-comparison savings (real-time channel quality monitoring → dispatch to best channel)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Layer 1: Discount Coefficient (Supplier Negotiation + Usage Scale)
&lt;/h3&gt;

&lt;p&gt;Platforms gain bargaining power through first-party supplier aggregation and quality distribution channels, then set discount coefficients per vendor (e.g., 10% of list for the OpenAI family, 20% for Anthropic, 20% for Google). Discounts float with cumulative usage — the larger the scale, the more room for negotiation.&lt;/p&gt;

&lt;h3&gt;
  
  
  Layer 2: Cache-Based Billing (The Most Underrated Cost Saver)
&lt;/h3&gt;

&lt;p&gt;Coding agents' requests have a highly repetitive prefix structure (system prompt, persona settings, context scaffolding). Once they hit the prefix cache, they are billed at the cached_input rate. &lt;strong&gt;Quantitative example&lt;/strong&gt;: official output $12/M, cache read price $1.2/M, monthly usage 80M tokens, 60% hit rate:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Without cache-based billing: 80M × $12 = &lt;strong&gt;$960&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;With cache-based billing (60% hit): 32M × $12 + 48M × $1.2 = &lt;strong&gt;$441.6&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;The cache layer alone cuts costs by roughly &lt;strong&gt;54%&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Layer 3: Model Tiering (Match Specs to Task Complexity)
&lt;/h3&gt;

&lt;p&gt;Flagship-model capability is reserved for difficult reasoning; routine and batch tasks are routed to models matched to the spec. &lt;strong&gt;Quantitative example&lt;/strong&gt;: monthly usage 100M tokens, official output price flagship $30/M, mid-tier $12/M, batch $1.2/M:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;All flagship: 100M × $30 = &lt;strong&gt;$3,000&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Tiered (30% flagship + 40% mid-tier + 30% batch): 30×30 + 40×12 + 30×1.2 = &lt;strong&gt;$1,416&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Model tiering alone cuts costs by roughly &lt;strong&gt;53%&lt;/strong&gt;, and the quality loss is negligible — because high-spec resources are only allocated to high-value tasks.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Layer 4: Channel Price Comparison (Real-Time Dispatch to the Best Channel)
&lt;/h3&gt;

&lt;p&gt;The platform's routing engine monitors each channel's latency (TTFT), stability, and price in real time, dispatching requests to the channel with the best current price-performance ratio; when a single upstream provider raises prices or fails, it switches automatically. This layer addresses the cost and availability risk of &lt;strong&gt;vendor lock-in&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which "Bargains" You Should Not Touch
&lt;/h2&gt;

&lt;p&gt;A low price whose nominal unit price sits significantly below upstream cost and cannot be attributed to any of the engineering layers above usually corresponds to shared quota, abnormal channels, or missing cache capability:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Shared quota&lt;/strong&gt;: abnormal concurrency patterns trigger upstream risk control and mass account bans;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Abnormal channels&lt;/strong&gt;: quota of unknown origin, with no guarantee of availability;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fake low prices that lack cache-based billing&lt;/strong&gt;: the nominal unit price is low, but with no cached_input billing, repeated content is billed at full price — so the effective cost is actually higher.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;The screening formula&lt;/strong&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;explainable low price = discount coefficient (≤50%) + cache billing (≥30–50% hit rate)
                      + model tiering (≥40–50%) + channel price comparison (bounded)
unexplainable low price = none of the above four factors can be attributed,
                          and the nominal unit price is more than 30% below official  →  exclude
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Common Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Q: What determines the cache hit rate?&lt;/strong&gt;&lt;br&gt;
The stability of the request prefix. By keeping the system prompt, persona settings, and few-shot examples fixed as a stable prefix, you can raise the hit rate. The gateway side automatically recognizes and bills at the cached_input rate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: How do discount coefficients relate to official pricing?&lt;/strong&gt;&lt;br&gt;
Discount coefficients are determined by supplier negotiation and usage scale (e.g., 10% of list for the OpenAI family) and float with cumulative usage; a healthy price structure is always explainable and auditable.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Does model tiering affect output quality?&lt;/strong&gt;&lt;br&gt;
It depends on whether the tiering is sensible. Grade by task complexity and value rather than randomly downgrading; high-value tasks are still routed to flagship models.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Q: Can you get cache pricing with a direct official connection?&lt;/strong&gt;&lt;br&gt;
Yes, but you must precisely control the prefix structure and configure caching yourself, and you would lack tiered routing and channel price comparison. The value of a collective procurement platform is in &lt;strong&gt;automating&lt;/strong&gt; cache-based billing, tiering, and price comparison.&lt;/p&gt;

&lt;h2&gt;
  
  
  Summary
&lt;/h2&gt;

&lt;p&gt;Token collective procurement's price advantage = the four-layer engineering of &lt;strong&gt;discount coefficient + cache-based billing + model tiering + channel price comparison&lt;/strong&gt;. The cache and tiering layers combined typically cut effective costs by more than 50%. Evaluate platforms by effective cost, and be wary of low prices that cannot be attributed. &lt;a href="https://teamorouter.com?utm_source=blog&amp;amp;utm_medium=seo&amp;amp;utm_campaign=token-jicai-why-cheaper" rel="noopener noreferrer"&gt;Sign up for TeamoRouter&lt;/a&gt; to verify your effective cost with GPT at 10% of list, Claude at 20% of list, and cached_input cache billing.&lt;/p&gt;

</description>
      <category>apigateway</category>
      <category>teamorouter</category>
      <category>api</category>
    </item>
    <item>
      <title>DeepSeek Harness Architecture: The Ultimate Guide to 'Everything Is a Plugin'</title>
      <dc:creator>Antonne Dillard</dc:creator>
      <pubDate>Tue, 18 Aug 2026 02:09:58 +0000</pubDate>
      <link>https://dev.to/tony_dillard/deepseek-harness-architecture-the-ultimate-guide-to-everything-is-a-plugin-31ak</link>
      <guid>https://dev.to/tony_dillard/deepseek-harness-architecture-the-ultimate-guide-to-everything-is-a-plugin-31ak</guid>
      <description>

&lt;p&gt;"Everything is a plugin" is the kind of line that sounds like marketing. In DeepSeek Harness (&lt;code&gt;dsh&lt;/code&gt;) it's literal — and it's the source of every useful capability the framework has.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is DeepSeek Harness (dsh)?
&lt;/h2&gt;

&lt;p&gt;DeepSeek Harness is &lt;strong&gt;DeepSeek AI's open-source agent harness&lt;/strong&gt; — the framework that turns a language model into a worker that can edit code and run commands. It follows the formula &lt;strong&gt;&lt;code&gt;Agent = Model + Harness&lt;/code&gt;&lt;/strong&gt;: the model (usually DeepSeek V4) is the brain; the harness is everything else — tools, filesystem, shell, sub-agents. Released under &lt;strong&gt;MIT&lt;/strong&gt;, powered by the &lt;strong&gt;Cordis&lt;/strong&gt; runtime, and in &lt;strong&gt;developer preview&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The key architectural claim is this: the filesystem, the shell, the model adapter, the Web UI, and sub-agents are &lt;strong&gt;not hardcoded&lt;/strong&gt;. They're swappable plugins, orchestrated by a runtime called &lt;a href="https://github.com/cordiverse/cordis" rel="noopener noreferrer"&gt;Cordis&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The two ideas behind Cordis
&lt;/h2&gt;

&lt;p&gt;Cordis is the MIT plugin runtime that drives dsh. Two ideas capture it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Registrations are reversible effects&lt;/strong&gt; — tools, providers, and listeners install via &lt;code&gt;ctx.effect()&lt;/code&gt; / &lt;code&gt;ctx.on()&lt;/code&gt;, so reload and teardown unwind them cleanly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Spatiotemporal composability&lt;/strong&gt; — plugins compose in one context by declaration order; capabilities stack and unstack.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You don't need to read the source. Just remember: &lt;strong&gt;dsh's capability = the combination of plugins you load.&lt;/strong&gt; Add a plugin to gain ability, remove one to lose it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Seams: the "three-role" swappable capability
&lt;/h2&gt;

&lt;p&gt;One word recurs in dsh's docs — &lt;strong&gt;seam&lt;/strong&gt;, a swappable capability. It's the mechanism behind "swap one provider and change the whole product." A seam has three roles:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Role&lt;/th&gt;
&lt;th&gt;Responsibility&lt;/th&gt;
&lt;th&gt;Example (shell)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Service Definition&lt;/td&gt;
&lt;td&gt;Declares the interface&lt;/td&gt;
&lt;td&gt;&lt;code&gt;dsh-shell&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Service Provider&lt;/td&gt;
&lt;td&gt;Implements it&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;dsh-bash-local&lt;/code&gt; / &lt;code&gt;dsh-bash-sandbox&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Consumer&lt;/td&gt;
&lt;td&gt;The model-facing tool&lt;/td&gt;
&lt;td&gt;&lt;code&gt;dsh-tool-bash&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Because the three are decoupled, changing an implementation touches only the &lt;strong&gt;Provider&lt;/strong&gt;. Move the shell from local to a remote sandbox and the &lt;code&gt;bash&lt;/code&gt; tool's model-visible schema doesn't change — invisible to the agent.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three practical payoffs
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Trimmable capability surface.&lt;/strong&gt; Not every agent needs network. Mount only &lt;code&gt;bash&lt;/code&gt; + file ops, omit web search → a more controllable, cheaper agent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Swappable providers.&lt;/strong&gt; The model backend is itself a seam — this is exactly how dsh plugs into &lt;strong&gt;DeepSeek V4 Pro&lt;/strong&gt; (1M context, flagship reasoning) or &lt;strong&gt;V4 Flash&lt;/strong&gt; (fast + cheap). Point &lt;code&gt;DEEPSEEK_BASE_URL&lt;/code&gt; from the official API to any endpoint — one env var, nothing else changes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;DEEPSEEK_BASE_URL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"https://api.teamorouter.com/v1"&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;DEEPSEEK_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"sk-teamo-your-key"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;3. Custom tools.&lt;/strong&gt; Want a tool that queries your internal ticketing system? Write a Cordis plugin, register it on &lt;code&gt;ctx.tools&lt;/code&gt;, done — faster than waiting for official support.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mental model: a slot board
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[ bash ] [ files ] [ subagent ] [ web-search ] [ model-adapter ] ...
        └────────── all mounted on the Cordis runtime ──────────┘
                              │
                   the agent's actual capability
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You decide what's plugged in. That model explains most of dsh's design — why the default preset bundles bash + file ops, and why headless mode needs only env vars.&lt;/p&gt;

&lt;p&gt;The direct value is two things: &lt;strong&gt;cheaper&lt;/strong&gt; (plug a cheap DeepSeek V4 in and the same loop costs an order of magnitude less) and &lt;strong&gt;more control&lt;/strong&gt; (you decide exactly what the agent can touch).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Try the architecture&lt;/strong&gt; → &lt;a href="https://teamorouter.com?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=dsh-architecture" rel="noopener noreferrer"&gt;Get a key on TeamoRouter&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>The Ultimate Guide to DeepSeek Harness (DSH): What It Is and How to Install</title>
      <dc:creator>Antonne Dillard</dc:creator>
      <pubDate>Mon, 17 Aug 2026 02:40:04 +0000</pubDate>
      <link>https://dev.to/tony_dillard/the-ultimate-guide-to-deepseek-harness-dsh-what-it-is-and-how-to-install-1f3g</link>
      <guid>https://dev.to/tony_dillard/the-ultimate-guide-to-deepseek-harness-dsh-what-it-is-and-how-to-install-1f3g</guid>
      <description>

&lt;p&gt;DeepSeek just shipped an open-source &lt;strong&gt;agent harness&lt;/strong&gt; — and you can run it with one command:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx @deepseek-ai/dsh web
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That pulls the package and starts a Web UI at &lt;code&gt;http://127.0.0.1:3080&lt;/code&gt;. No clone, no build.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is DeepSeek Harness (dsh)?
&lt;/h2&gt;

&lt;p&gt;DeepSeek Harness (&lt;code&gt;dsh&lt;/code&gt;) is &lt;strong&gt;DeepSeek AI's open-source agent harness&lt;/strong&gt; — the framework that turns a language model into a worker that can actually edit code and run commands. Built on an &lt;strong&gt;"everything is a plugin"&lt;/strong&gt; architecture, powered by the &lt;a href="https://github.com/cordiverse/cordis" rel="noopener noreferrer"&gt;Cordis&lt;/a&gt; runtime, and released under the &lt;strong&gt;MIT&lt;/strong&gt; license. It's in &lt;strong&gt;developer preview&lt;/strong&gt;, so interfaces will still shift — but for solo devs and cheap agent experiments it's fully usable today.&lt;/p&gt;

&lt;p&gt;The industry has converged on a clean formula:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Agent = Model + Harness
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;strong&gt;model&lt;/strong&gt; is the brain — the weights that predict tokens. The &lt;strong&gt;harness&lt;/strong&gt; is everything else: the tools the agent can call, the filesystem and shell it can touch, how sub-agents pass context, and when execution stops. A bare model is bad at memory and tool use on its own; the harness turns a chat model into a worker.&lt;/p&gt;

&lt;h2&gt;
  
  
  What model does it run?
&lt;/h2&gt;

&lt;p&gt;dsh is model-agnostic, but its natural fit is the &lt;strong&gt;DeepSeek V4 family&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;DeepSeek V4 Pro&lt;/strong&gt; — the flagship reasoning model, up to 1M context, multiple reasoning-effort levels.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;DeepSeek V4 Flash&lt;/strong&gt; — the fast, cheap sibling for high-frequency work.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;To route dsh through a free tier, set two env vars:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;DEEPSEEK_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"sk-teamo-your-key"&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;DEEPSEEK_BASE_URL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"https://api.teamorouter.com/v1"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then &lt;code&gt;npx @deepseek-ai/dsh web&lt;/code&gt; again. That's the whole setup — dsh speaks the OpenAI protocol, so any OpenAI-compatible endpoint works.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's in the Web UI
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;http://127.0.0.1:3080&lt;/code&gt; is the front-end of an agent runtime:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Sessions&lt;/strong&gt; — create, switch, rename, run several in parallel.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model selection&lt;/strong&gt; — swap &lt;code&gt;deepseek-v4-pro&lt;/code&gt; vs &lt;code&gt;deepseek-v4-flash&lt;/code&gt; per session.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool panel&lt;/strong&gt; — the tools the agent can call: bash, file read/write, sub-agents, web search.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Goals&lt;/strong&gt; — break long tasks into stateful goals, pause / resume.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Don't treat it as "another chat window" — you give a goal, the agent loops over tools until done, and the UI is your observation deck.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prove it works
&lt;/h2&gt;

&lt;p&gt;Give the agent this task:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;List the current directory, find the README, summarize its first
paragraph &lt;span class="k"&gt;in &lt;/span&gt;one sentence, and write it to /tmp/summary.txt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;dsh calls &lt;code&gt;bash&lt;/code&gt; and the file tools in sequence, then reports back. That's the whole idea — &lt;strong&gt;the harness turns a model into a worker&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Run your first agent loop&lt;/strong&gt; → &lt;a href="https://teamorouter.com?utm_source=devto&amp;amp;utm_medium=social&amp;amp;utm_campaign=dsh-install" rel="noopener noreferrer"&gt;Get a free key on TeamoRouter&lt;/a&gt;&lt;/p&gt;

</description>
      <category>deepseek</category>
      <category>ai</category>
      <category>agents</category>
      <category>devtools</category>
    </item>
    <item>
      <title>What Is Kimi K3? Complete 2026 Guide to Moonshot AI's Open Source Model</title>
      <dc:creator>Antonne Dillard</dc:creator>
      <pubDate>Tue, 28 Jul 2026 04:11:30 +0000</pubDate>
      <link>https://dev.to/tony_dillard/what-is-kimi-k3-complete-2026-guide-to-moonshot-ais-open-source-model-565j</link>
      <guid>https://dev.to/tony_dillard/what-is-kimi-k3-complete-2026-guide-to-moonshot-ais-open-source-model-565j</guid>
      <description>

&lt;h2&gt;
  
  
  Quick Answer
&lt;/h2&gt;

&lt;p&gt;Kimi K3 is a frontier-class, open-source large language model built by Moonshot AI (月之暗面). Released in July 2026, it packs 2.8 trillion parameters, a 1-million-token context window, and a novel hybrid architecture combining linear and full attention mechanisms. It is the first open-source model to beat Claude and GPT in frontend coding benchmarks, and it costs significantly less per solved task. This guide covers everything you need to know: what K3 is, how it works under the hood, how it compares to the competition, and how to start using it today.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is Kimi K3?
&lt;/h2&gt;

&lt;p&gt;Kimi K3 is the third generation of Moonshot AI's flagship model family, succeeding the well-regarded Kimi K2 and K2.5. Unlike many Chinese AI models that prioritize benchmark scores over real-world usability, K3 was designed from the ground up for long-horizon agentic tasks — the kind where an AI codes a full feature, navigates a large codebase, or maintains coherent reasoning across hundreds of thousands of tokens.&lt;/p&gt;

&lt;p&gt;Moonshot AI announced K3 on July 16, 2026, and released the open weights on July 27, 2026 — a deliberate five-day gap that let the inference ecosystem (vLLM, NVIDIA, AMD) prepare day-zero support. The model is available through the Kimi web app, the Kimi API, the Kimi Code terminal agent, and third-party API gateways.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key Specifications at a Glance
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Spec&lt;/th&gt;
&lt;th&gt;Detail&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Parameters&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;2.8 trillion (Mixture-of-Experts)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Context window&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1,048,576 tokens (1M)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Architecture&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Hybrid: KDA linear attention + MLA full attention + sparse MoE&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Experts&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;896 routed experts, 16 active per token, plus shared experts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Depth&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;93 layers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Quantization&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;MXFP4 weights in release configuration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Multimodality&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Native vision support with dedicated vision tower&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Activation&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;SiTU (in MoE path)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;License&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Open weights (specific license at release)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  The Architecture That Makes K3 Different
&lt;/h2&gt;

&lt;p&gt;K3 is not simply a scaled-up Kimi K2. It introduces several architectural innovations that change how the model handles long contexts and reduces serving costs.&lt;/p&gt;

&lt;h3&gt;
  
  
  KDA (Kimi Delta Attention): The Core Innovation
&lt;/h3&gt;

&lt;p&gt;Traditional Transformer models use full attention — every token attends to every other token, storing per-token key-value pairs in a KV cache. This works well but scales poorly: a 1M-token context with a model this large would consume enormous GPU memory just for the cache.&lt;/p&gt;

&lt;p&gt;KDA replaces most attention layers with a &lt;strong&gt;recurrent linear attention mechanism&lt;/strong&gt;. Instead of storing per-token KV pairs, it maintains a fixed-size matrix state plus a short convolution state. The recurrent state updates incrementally as new tokens arrive — like an RNN, but with far better parallelization during training. This is how K3 achieves its 1M-token context without requiring a data center's worth of VRAM per request.&lt;/p&gt;

&lt;h3&gt;
  
  
  Attention Residuals (AttnRes): Cross-Layer Memory
&lt;/h3&gt;

&lt;p&gt;In a standard Transformer, information flows through layers sequentially via a single residual stream. K3 adds &lt;strong&gt;Attention Residuals&lt;/strong&gt; — shortcuts that let deeper layers retrieve representations from much earlier layer blocks. Think of it as the model being able to "look back" across dozens of layers instead of relying on whatever survived the journey through intermediate layers. The vLLM team describes this as "creating cross-layer memory traffic" that improves the model's ability to maintain context over extremely long sequences.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Hybrid Design
&lt;/h3&gt;

&lt;p&gt;K3 alternates between KDA layers (efficient, recurrent) and MLA layers (Multi-head Latent Attention, a compressed full-attention variant) every four layers. This gives the model the efficiency of linear attention for most of the sequence while periodically re-establishing global context with full attention. The 896 MoE experts (16 active per token) handle the feed-forward computation, with shared experts always active for common patterns.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Kimi K3 Compares to GPT and Claude
&lt;/h2&gt;

&lt;p&gt;K3's benchmark performance tells a clear story: it is broadly competitive with the best closed-source models, and in some domains it leads outright.&lt;/p&gt;

&lt;h3&gt;
  
  
  DeepSWE (Software Engineering)
&lt;/h3&gt;

&lt;p&gt;On the DeepSWE benchmark, which tests real-world software engineering tasks:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;pass@1&lt;/th&gt;
&lt;th&gt;pass@4&lt;/th&gt;
&lt;th&gt;Cost/Rollout&lt;/th&gt;
&lt;th&gt;Tasks Solved per $100&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Kimi K3&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;68.5%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;89.4%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$4.65&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;14.7&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Fable 5&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;69.9%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;88.5%&lt;/td&gt;
&lt;td&gt;$13.41&lt;/td&gt;
&lt;td&gt;5.3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Sol&lt;/td&gt;
&lt;td&gt;72.7%&lt;/td&gt;
&lt;td&gt;85.8%&lt;/td&gt;
&lt;td&gt;$8.37&lt;/td&gt;
&lt;td&gt;8.2&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;K3 trails slightly on single-attempt accuracy but pulls ahead with multiple attempts — and solves roughly &lt;strong&gt;2.8x more tasks per dollar&lt;/strong&gt; than either competitor.&lt;/p&gt;

&lt;h3&gt;
  
  
  Frontend Code Arena
&lt;/h3&gt;

&lt;p&gt;This is where K3 really shines. In anonymous blind tests on the Frontend Code Arena benchmark, K3 scored &lt;strong&gt;1,679 points — ranking #1 overall&lt;/strong&gt;, ahead of Claude Fable 5 (1,631) and GPT-5.6 Sol (1,618). It took first place in 6 of 7 domains: Brand &amp;amp; Marketing, Reference-Based Design, Data &amp;amp; Analytics, Consumer Product, Simulations, and Content Creation Tools. This was a 17-position leap from Kimi K2.6's #18 ranking.&lt;/p&gt;

&lt;h3&gt;
  
  
  SWE Marathon
&lt;/h3&gt;

&lt;p&gt;For ultra-long software engineering tasks, K3 ranked #1 on the SWE Marathon benchmark. It also scored 91.2 on BrowseComp and came in #2 on Terminal Bench 2.1 and FrontierSWE.&lt;/p&gt;

&lt;h3&gt;
  
  
  Language-Specific Strengths
&lt;/h3&gt;

&lt;p&gt;K3 shows particular strength in Go (79% vs Fable 5's 71%), while Fable 5 leads in Python (74% vs 68%), JavaScript (70% vs 65%), TypeScript (64% vs 60%), and Rust (75% vs 65%). If your stack is Go-heavy, K3 is arguably the best model available.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Kimi K3 Matters
&lt;/h2&gt;

&lt;p&gt;K3 represents several converging trends in AI that make it significant beyond raw benchmark numbers:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Open-source catching up to closed-source.&lt;/strong&gt; K3 matches or exceeds GPT-5.6 Sol and Claude Fable 5 in key coding benchmarks while being fully open-weight. This means you can self-host it, fine-tune it, and inspect its architecture — none of which you can do with GPT or Claude.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Cost efficiency as a competitive moat.&lt;/strong&gt; K3's aggressive input caching (90% discount on cache hits, with observed hit rates above 90% in coding scenarios) means the effective per-task cost is dramatically lower than sticker prices suggest. Moonshot's Mooncake serving architecture is purpose-built for this.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The return of architectural innovation.&lt;/strong&gt; Rather than scaling a standard Transformer, Moonshot invested in novel attention mechanisms (KDA, AttnRes) that genuinely change the compute/capability tradeoff curve.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  How to Access Kimi K3
&lt;/h2&gt;

&lt;p&gt;You have several options, each suited to different use cases:&lt;/p&gt;

&lt;h3&gt;
  
  
  Kimi Web App (kimi.com)
&lt;/h3&gt;

&lt;p&gt;The easiest way to try K3. Sign up at kimi.com and start chatting. Note that free-tier users face context window and rate limitations. Paid subscriptions unlock the full 1M-token context.&lt;/p&gt;

&lt;h3&gt;
  
  
  Kimi API
&lt;/h3&gt;

&lt;p&gt;For programmatic access, the Kimi API is OpenAI-compatible. Get an API key from platform.moonshot.ai (international) or platform.moonshot.cn (China), then use the model ID &lt;code&gt;kimi-k3&lt;/code&gt; with the base URL &lt;code&gt;https://api.moonshot.ai/v1&lt;/code&gt;. API pricing is $3.00/M input tokens (cache miss), $0.30/M input tokens (cache hit), and $15.00/M output tokens. A typical single API call costs about $0.007.&lt;/p&gt;

&lt;h3&gt;
  
  
  API Gateways (Recommended for Production)
&lt;/h3&gt;

&lt;p&gt;For production workloads, running directly against the Moonshot API introduces single-provider risk. What happens during an outage or when rate limits hit? This is where API gateways like &lt;a href="https://teamorouter.com?utm_source=blog&amp;amp;utm_medium=seo&amp;amp;utm_campaign=kimi-k3" rel="noopener noreferrer"&gt;TeamoRouter&lt;/a&gt; come in.&lt;/p&gt;

&lt;p&gt;TeamoRouter acts as a stable API gateway that sits between your application and the Kimi K3 API. It provides automatic failover, intelligent routing, and unified billing across multiple providers — so if Moonshot's API experiences issues, your requests seamlessly fall back to alternative endpoints without any code changes. For teams building on K3, this eliminates the single biggest operational risk of depending on one API provider.&lt;/p&gt;

&lt;h3&gt;
  
  
  Kimi Code (Terminal Agent)
&lt;/h3&gt;

&lt;p&gt;Moonshot's native coding agent, available via &lt;code&gt;npm i @moonshot-ai/kimi-code&lt;/code&gt;. It requires a paid subscription and provides K3-powered code generation, debugging, and repository navigation directly in your terminal.&lt;/p&gt;

&lt;h3&gt;
  
  
  Self-Hosting
&lt;/h3&gt;

&lt;p&gt;With open weights available, you can deploy K3 on your own infrastructure. Be aware that this requires substantial GPU resources — multi-node clusters with high-VRAM GPUs. The vLLM project provides day-zero inference support with validated NVIDIA and AMD launch recipes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limitations to Know
&lt;/h2&gt;

&lt;p&gt;K3 is impressive but not without tradeoffs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Inference speed:&lt;/strong&gt; K3 generates ~33-35 tokens per second on the standard tier, notably slower than GPT-5.6 Sol (~80+ t/s) and Claude Fable 5 (~60+ t/s). The "Kimi K3 Fast" variant improves this to ~117 t/s.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Always-on reasoning:&lt;/strong&gt; At launch, only &lt;code&gt;reasoning_effort="max"&lt;/code&gt; was available. You cannot dial down thinking to save costs on simple queries. Lighter modes are promised for later releases.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Parameter lock:&lt;/strong&gt; Temperature, top_p, and penalty parameters are fixed. Developers must omit these from API requests.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Image handling:&lt;/strong&gt; Public image URLs are not supported through the API at launch. Use base64 encoding or uploaded files for vision inputs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reliability gap:&lt;/strong&gt; While K3 casts a wide net (high pass@k), both Fable 5 and Sol solve more tasks consistently across all attempts. K3 is better for exploration and iteration than for one-shot perfection.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Launch Day Story
&lt;/h2&gt;

&lt;p&gt;K3's launch was dramatic. Within 48 hours of release, Moonshot AI paused new consumer subscriptions — the existing GPU cluster could not handle the exponential call volume. Elon Musk commented "Impressive" on independent benchmark results. Chinese tech media called it a potential "DeepSeek moment" — a reference to when DeepSeek-R1 stunned the world by matching frontier models at a fraction of the cost. Whether K3 achieves the same mindshare remains to be seen, but the technical foundation is there.&lt;/p&gt;

&lt;h2&gt;
  
  
  Getting Started with Kimi K3
&lt;/h2&gt;

&lt;p&gt;The fastest way to start building with K3:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Try it free&lt;/strong&gt; at kimi.com to understand its capabilities.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Get an API key&lt;/strong&gt; from platform.moonshot.ai and make your first API call in under 5 minutes (it's OpenAI-compatible — just change the base URL and model name).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;For production&lt;/strong&gt;, consider routing through &lt;a href="https://teamorouter.com?utm_source=blog&amp;amp;utm_medium=seo&amp;amp;utm_campaign=kimi-k3" rel="noopener noreferrer"&gt;TeamoRouter&lt;/a&gt; for automatic failover, load balancing, and stable access across multiple providers. TeamoRouter gives you a single endpoint that handles provider selection, health monitoring, and failover — so you focus on building, not on API reliability.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Kimi K3 is a genuine milestone: the first open-source model that competes head-to-head with the best closed-source systems in the areas that matter most to developers. Its hybrid architecture points toward a future where long-context AI is not a premium feature but the default. And with open weights now available, the ecosystem around K3 is only going to grow.&lt;/p&gt;

</description>
      <category>kimik3</category>
      <category>moonshotai</category>
      <category>opensourcellm</category>
      <category>guide</category>
    </item>
  </channel>
</rss>
