<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Owen</title>
    <description>The latest articles on DEV Community by Owen (@owen_fox).</description>
    <link>https://dev.to/owen_fox</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3893304%2Fb8cec06b-7789-423e-a8d0-386db7f00620.png</url>
      <title>DEV Community: Owen</title>
      <link>https://dev.to/owen_fox</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/owen_fox"/>
    <language>en</language>
    <item>
      <title>Gemini 3.7 Flash API: $0.75 Pricing, minimal Returns 400</title>
      <dc:creator>Owen</dc:creator>
      <pubDate>Sat, 22 Aug 2026 11:35:41 +0000</pubDate>
      <link>https://dev.to/owen_fox/gemini-37-flash-api-075-pricing-minimal-returns-400-2l19</link>
      <guid>https://dev.to/owen_fox/gemini-37-flash-api-075-pricing-minimal-returns-400-2l19</guid>
      <description>&lt;p&gt;&lt;strong&gt;Gemini 3.7 Flash shipped on 2026-08-13 at $0.75 per million input tokens and $3.75 per million output.&lt;/strong&gt; Two things about that price are easy to miss: it expires on 2026-12-31 and doubles the next day, and Gemini 3.6 Flash now carries the exact same numbers.&lt;/p&gt;

&lt;p&gt;The thing that breaks builds is elsewhere. The thinking tiers went from four to three. &lt;code&gt;minimal&lt;/code&gt; is gone, and it does not degrade quietly.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Price:        $0.75 in / $3.75 out per 1M through 2026-12-31
              $1.50 / $7.50 from 2027-01-01
              Batch and Flex: $0.375 / $1.875, same doubling
Cache:        $0.075 read, $0.50 per 1M per hour storage
Context:      1,048,576 in / 65,536 out
Thinking:     thinking_level low | medium | high, default medium
              minimal returns HTTP 400
Gateways:     google/gemini-3.7-flash on ofox
Measured:     minimal 400, none and xhigh accepted and billed
Snapshot:     2026-08-21, 10 streamed calls per tier
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Last updated 2026-08-21. Introductory pricing expires 2026-12-31, so re-check before budgeting past that date.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How Much Does the Gemini 3.7 Flash API Cost?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;$0.75 in, $3.75 out, until it isn't.&lt;/strong&gt; Google's &lt;a href="https://ai.google.dev/gemini-api/docs/pricing" rel="noopener noreferrer"&gt;pricing page&lt;/a&gt; puts two rates in a single cell, which is how the wrong one ends up in a budget spreadsheet.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tier&lt;/th&gt;
&lt;th&gt;Input / 1M&lt;/th&gt;
&lt;th&gt;Output / 1M&lt;/th&gt;
&lt;th&gt;In effect&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Standard&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.75&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$3.75&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;through 2026-12-31&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Standard&lt;/td&gt;
&lt;td&gt;$1.50&lt;/td&gt;
&lt;td&gt;$7.50&lt;/td&gt;
&lt;td&gt;from 2027-01-01&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Batch / Flex&lt;/td&gt;
&lt;td&gt;$0.375&lt;/td&gt;
&lt;td&gt;$1.875&lt;/td&gt;
&lt;td&gt;through 2026-12-31&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Batch / Flex&lt;/td&gt;
&lt;td&gt;$0.75&lt;/td&gt;
&lt;td&gt;$3.75&lt;/td&gt;
&lt;td&gt;from 2027-01-01&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Context caching reads at $0.075 per million with storage at $0.50 per million tokens per hour, both doubling on the same date. Search grounding gives 5,000 free requests a month shared across every Gemini 3.x model, not per model, and then costs $14 per 1,000.&lt;/p&gt;

&lt;p&gt;One line on that page decides more of your bill than the headline rate: the output row reads &lt;strong&gt;"Output price (including thinking tokens)"&lt;/strong&gt;. Reasoning is metered at the output rate, so the effort setting is a price multiplier, not a quality knob with a rounding error attached.&lt;/p&gt;

&lt;h2&gt;
  
  
  Is Gemini 3.7 Flash Cheaper Than Gemini 3.6 Flash?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;No. They cost exactly the same, cell for cell.&lt;/strong&gt; If you saw 3.7 Flash introduced as half the price of 3.6 Flash, that comparison had a shelf life of about one day, and the pricing page's own history shows why.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;strong&gt;2026-07-22&lt;/strong&gt;, just after 3.6 Flash shipped on 2026-07-21, per the &lt;a href="https://web.archive.org/web/20260722024041/https://ai.google.dev/gemini-api/docs/pricing" rel="noopener noreferrer"&gt;archived pricing page&lt;/a&gt;: 3.6 Flash costs $1.50 / $7.50.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;2026-08-12&lt;/strong&gt;, per the &lt;a href="https://web.archive.org/web/20260812042335/https://ai.google.dev/gemini-api/docs/pricing" rel="noopener noreferrer"&gt;last copy before 3.7 launched&lt;/a&gt;: 3.6 Flash still costs $1.50 / $7.50, flat, with no discount row. 3.7 Flash is not on the page yet.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;2026-08-13&lt;/strong&gt;: 3.7 Flash ships.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;2026-08-14&lt;/strong&gt;, per the &lt;a href="https://web.archive.org/web/20260814074039/https://ai.google.dev/gemini-api/docs/pricing" rel="noopener noreferrer"&gt;next archived copy&lt;/a&gt;: 3.6 Flash and 3.7 Flash both read $0.75 / $3.75 through 2026-12-31, $1.50 / $7.50 after.&lt;/li&gt;
&lt;li&gt;  &lt;strong&gt;2027-01-01&lt;/strong&gt;: both return to $1.50 / $7.50, which is what &lt;a href="https://ofox.ai/models/google/gemini-3.6-flash" rel="noopener noreferrer"&gt;3.6 Flash&lt;/a&gt; cost on its own launch day.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So the baseline moved inside the same 48 hours as the comparison. Google's model card footnotes the end state without comment: the asterisk on both price rows says the introductory price expires December 31, 2026. Upgrading from 3.6 costs nothing and saves nothing. What the discount buys is four and a half months, on both models equally.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://ofox.ai/models/google/gemini-3.7-flash" rel="noopener noreferrer"&gt;ofox model page&lt;/a&gt; carries rates Google's page does not itemise: cache write at $0.0415 per million, one-hour cache write at $0.50, audio input at $0.75, web search at $0.014 per request.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Do I Call Gemini 3.7 Flash?
&lt;/h2&gt;

&lt;p&gt;Two routes, and the only differences are the base URL and the model string.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sk-...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.ofox.ai/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;google/gemini-3.7-flash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;reasoning_effort&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;low&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Rewrite this query with a window function&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Straight to Google, the model string loses its &lt;code&gt;google/&lt;/code&gt; prefix and the base URL changes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;AIza...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://generativelanguage.googleapis.com/v1beta/openai/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gemini-3.7-flash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;reasoning_effort&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;low&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[...])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On Google's native protocol the parameter is not &lt;code&gt;reasoning_effort&lt;/code&gt;. It is &lt;a href="https://ai.google.dev/gemini-api/docs/thinking" rel="noopener noreferrer"&gt;&lt;code&gt;thinking_level&lt;/code&gt;&lt;/a&gt;, it takes &lt;code&gt;low&lt;/code&gt;, &lt;code&gt;medium&lt;/code&gt; and &lt;code&gt;high&lt;/code&gt;, and it defaults to &lt;code&gt;medium&lt;/code&gt;. Both protocols are open on ofox, with Google and Cloud Vertex as the upstreams.&lt;/p&gt;

&lt;p&gt;Capabilities are worth copying off the &lt;a href="https://ai.google.dev/gemini-api/docs/models/gemini-3-7-flash" rel="noopener noreferrer"&gt;model page&lt;/a&gt; rather than discovering by 400: function calling, structured outputs, context caching, code execution, search grounding, URL context, file search and Maps grounding are all supported, computer use is marked Preview, and Batch, Flex and Priority are all available. &lt;strong&gt;Live API, image generation and audio generation are not supported.&lt;/strong&gt; Input takes text, image, video, audio and PDF; output is text only.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does minimal Return an Error on Gemini 3.7 Flash?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Because the tier was removed, and the request fails instead of falling back.&lt;/strong&gt; The model page spells it out under Thinking: &lt;code&gt;Supported (low, medium, high)&lt;/code&gt; followed by &lt;code&gt;Note: minimal is not supported and returns an error.&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;We sent it anyway. Over the ofox route, &lt;code&gt;reasoning_effort: "minimal"&lt;/code&gt; returns HTTP 400 with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"error"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"code"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
           &lt;/span&gt;&lt;span class="nl"&gt;"message"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Thinking level is unsupported: THINKING_LEVEL_MINIMAL"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
           &lt;/span&gt;&lt;span class="nl"&gt;"param"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"invalid_request_error"&lt;/span&gt;&lt;span class="p"&gt;}}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;3.6 Flash still accepts &lt;code&gt;minimal&lt;/code&gt;, and we measured what that tier was doing for you before you lose it. On &lt;a href="https://ofox.ai/models/google/gemini-3.6-flash" rel="noopener noreferrer"&gt;3.6 Flash&lt;/a&gt;, the same prompt at &lt;code&gt;minimal&lt;/code&gt; returned a median of 42 completion tokens with &lt;strong&gt;zero&lt;/strong&gt; reasoning tokens and a 1.36-second first token. The cheapest tier 3.7 will accept, &lt;code&gt;low&lt;/code&gt;, returned a median of 354 tokens at 4.19 seconds. That is 8.4x the billed output and 3x the wait, for a model whose visible answer is the same 40 tokens either way.&lt;/p&gt;

&lt;p&gt;So the migration is not "change the model ID and keep going". A config that hardcodes &lt;code&gt;minimal&lt;/code&gt; stops working at the moment you swap the ID, and the nearest legal value silently costs eight times as much. There is no zero-thinking setting on 3.7 Flash to fall back to.&lt;/p&gt;

&lt;p&gt;The more interesting half of that probe is what does &lt;strong&gt;not&lt;/strong&gt; fail. &lt;code&gt;none&lt;/code&gt; and &lt;code&gt;xhigh&lt;/code&gt; are not documented values for this model, and both returned HTTP 200 with reasoning tokens on the bill: reasoning-token medians of 271 and 604 across ten runs each, which is the 308 and 637 completion tokens in the table below. Only &lt;code&gt;minimal&lt;/code&gt; is rejected. A 200 is not confirmation that the value you sent selected the tier its name implies, which is the same trap &lt;a href="https://ofox.ai/blog/glm-5-3-api-pricing-endpoints-reasoning-effort-2026/" rel="noopener noreferrer"&gt;GLM 5.3 sets with undocumented &lt;code&gt;reasoning_effort&lt;/code&gt; values&lt;/a&gt;, where validation lives on the route rather than in the model. Probe your own route before shipping, because "it returned 200" and "it did what I asked" are different claims.&lt;/p&gt;

&lt;p&gt;Google's &lt;a href="https://ai.google.dev/gemini-api/docs/openai" rel="noopener noreferrer"&gt;OpenAI compatibility page&lt;/a&gt; makes this worth two minutes of your time: its &lt;code&gt;reasoning_effort&lt;/code&gt; mapping table lists Gemini 3.1 Pro, 3.1 Flash-Lite, Gemini 3 Flash and 2.5, with no 3.7 column, and it maps OpenAI's &lt;code&gt;minimal&lt;/code&gt; downward. The model page says 3.7 rejects &lt;code&gt;minimal&lt;/code&gt;. Two Google pages, no agreement, and one short probe settles it for the route you actually call.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Does Each Thinking Level Actually Cost?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The same answer, at up to double the bill.&lt;/strong&gt; Mak, product lead at ofox, ran 10 streamed calls per tier through the same ofox route on 2026-08-21, one short prompt (rewrite a &lt;code&gt;GROUP BY&lt;/code&gt; as a window function), recording time to the first content token and the usage the API reported.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;code&gt;reasoning_effort&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;TTFT median&lt;/th&gt;
&lt;th&gt;TTFT range&lt;/th&gt;
&lt;th&gt;Completion tokens&lt;/th&gt;
&lt;th&gt;Range&lt;/th&gt;
&lt;th&gt;Visible answer&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;low&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;4.19 s&lt;/td&gt;
&lt;td&gt;2.22–5.40&lt;/td&gt;
&lt;td&gt;354&lt;/td&gt;
&lt;td&gt;40–418&lt;/td&gt;
&lt;td&gt;~41&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;medium&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;5.47 s&lt;/td&gt;
&lt;td&gt;4.47–6.98&lt;/td&gt;
&lt;td&gt;384&lt;/td&gt;
&lt;td&gt;291–558&lt;/td&gt;
&lt;td&gt;~41&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;high&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;7.69 s&lt;/td&gt;
&lt;td&gt;5.97–9.68&lt;/td&gt;
&lt;td&gt;688&lt;/td&gt;
&lt;td&gt;560–806&lt;/td&gt;
&lt;td&gt;~40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;unset&lt;/td&gt;
&lt;td&gt;5.58 s&lt;/td&gt;
&lt;td&gt;4.32–5.98&lt;/td&gt;
&lt;td&gt;431&lt;/td&gt;
&lt;td&gt;333–583&lt;/td&gt;
&lt;td&gt;~36&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;none&lt;/code&gt; (undocumented)&lt;/td&gt;
&lt;td&gt;4.64 s&lt;/td&gt;
&lt;td&gt;1.88–5.05&lt;/td&gt;
&lt;td&gt;308&lt;/td&gt;
&lt;td&gt;40–361&lt;/td&gt;
&lt;td&gt;~40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;xhigh&lt;/code&gt; (undocumented)&lt;/td&gt;
&lt;td&gt;7.35 s&lt;/td&gt;
&lt;td&gt;5.89–9.10&lt;/td&gt;
&lt;td&gt;637&lt;/td&gt;
&lt;td&gt;513–892&lt;/td&gt;
&lt;td&gt;~39&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Medians of 10 runs; the ranges are observed minimum to maximum, which widen as you add runs, so read them as spread rather than as bounds. All of it is one prompt on one route, and parameter handling can differ between routes to the same model, so treat the shape as portable and the numbers as ours.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The answer never changes size. The invisible part does.&lt;/strong&gt; Text tokens sat near 40 at every tier, because the SQL statement the model was asked for is 40 tokens long no matter how long it thinks about it. Everything above that line is reasoning, billed at the output rate. On output alone at $3.75 per million, 1,000 of these calls cost $1.33 at &lt;code&gt;low&lt;/code&gt; and $2.58 at &lt;code&gt;high&lt;/code&gt;; the 34-token prompt adds under three cents per thousand either way, so the effort setting is doing essentially all of the work in that spread.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;low&lt;/code&gt; bottomed out at 40 completion tokens on one of its ten runs, with the reasoning count at exactly zero; &lt;code&gt;none&lt;/code&gt; did the same twice. Neither &lt;code&gt;medium&lt;/code&gt; nor &lt;code&gt;high&lt;/code&gt; ever did, and that floor is most of why their ranges sit apart.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;low&lt;/code&gt; and &lt;code&gt;high&lt;/code&gt; separate cleanly. &lt;code&gt;low&lt;/code&gt; and &lt;code&gt;medium&lt;/code&gt; do not.&lt;/strong&gt; The &lt;code&gt;high&lt;/code&gt; range (560–806) does not overlap the &lt;code&gt;low&lt;/code&gt; range (40–418), so that gap is real on this prompt. &lt;code&gt;low&lt;/code&gt; and &lt;code&gt;medium&lt;/code&gt; are 30 tokens apart at the median with ranges that overlap across most of their span, which means ten runs cannot tell them apart here. Whether &lt;code&gt;medium&lt;/code&gt; is a useful midpoint is a question about your prompt, and it takes about four minutes to answer with your own.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Undocumented values do not fail, and they do not do what they say.&lt;/strong&gt; &lt;code&gt;none&lt;/code&gt; returned HTTP 200 and spent a median of 271 reasoning tokens, which is most of what &lt;code&gt;low&lt;/code&gt; spends. If you read that name as an off switch, you are paying for thinking you believe you disabled. &lt;code&gt;xhigh&lt;/code&gt; lands next to &lt;code&gt;high&lt;/code&gt;. Only &lt;code&gt;minimal&lt;/code&gt; is rejected.&lt;/p&gt;

&lt;p&gt;None of this contradicts the model being fast. Artificial Analysis ranks 3.7 Flash &lt;strong&gt;1st of 182 on output speed&lt;/strong&gt; at 389.5 tokens per second. On the same page they measure 14.52 seconds to first token at high on Google's own API, against a 2.80-second median for its price peers, and roughly double the median we saw at high on this route. The wait sits in front of the stream rather than inside it, which is why the effort setting, not the tokens per second, is what a user feels.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is Gemini 3.7 Flash Good At?
&lt;/h2&gt;

&lt;p&gt;Google's &lt;a href="https://deepmind.google/models/model-cards/gemini-3-7-flash/" rel="noopener noreferrer"&gt;model card&lt;/a&gt; publishes a comparison table against 3.6 Flash, Claude Sonnet 5, GPT-5.6 Terra and Muse Spark 1.2. The rows where the gaps are real:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;3.7 Flash&lt;/th&gt;
&lt;th&gt;3.6 Flash&lt;/th&gt;
&lt;th&gt;Sonnet 5&lt;/th&gt;
&lt;th&gt;GPT-5.6 Terra&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;AA Intelligence Index&lt;/td&gt;
&lt;td&gt;56&lt;/td&gt;
&lt;td&gt;52&lt;/td&gt;
&lt;td&gt;55&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;57&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GDM-MRCR v2 (128k)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;97.0%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;91.8%&lt;/td&gt;
&lt;td&gt;81.5%&lt;/td&gt;
&lt;td&gt;93.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LVBench (long video)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;85.4%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;84.2%&lt;/td&gt;
&lt;td&gt;68.5%&lt;/td&gt;
&lt;td&gt;78.9%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Code Arena (web dev Elo)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1588&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1538&lt;/td&gt;
&lt;td&gt;1541&lt;/td&gt;
&lt;td&gt;1523&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AutomationBench&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;30.4%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;17.0%&lt;/td&gt;
&lt;td&gt;10.7%&lt;/td&gt;
&lt;td&gt;23.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FrontierCode 1.1&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;43.6%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;34.4%&lt;/td&gt;
&lt;td&gt;42.7%&lt;/td&gt;
&lt;td&gt;41.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HLE-Verified&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;53.6%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;51.2%&lt;/td&gt;
&lt;td&gt;31.0%&lt;/td&gt;
&lt;td&gt;51.1%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSWE v1.1&lt;/td&gt;
&lt;td&gt;65.3%&lt;/td&gt;
&lt;td&gt;48.6%&lt;/td&gt;
&lt;td&gt;53.8%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;69.6%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Terminal-bench 2.1&lt;/td&gt;
&lt;td&gt;85.8%&lt;/td&gt;
&lt;td&gt;78.0%&lt;/td&gt;
&lt;td&gt;80.4%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;87.4%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Terminal-bench 3.0&lt;/td&gt;
&lt;td&gt;14.9%&lt;/td&gt;
&lt;td&gt;5.4%&lt;/td&gt;
&lt;td&gt;14.6%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;20.8%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GDPVal-AA v2 (Elo)&lt;/td&gt;
&lt;td&gt;1525&lt;/td&gt;
&lt;td&gt;1422&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1598&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;1578&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent's Last Exam&lt;/td&gt;
&lt;td&gt;26.3%&lt;/td&gt;
&lt;td&gt;24.2%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;33.3%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;28.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CharXiv, no tools&lt;/td&gt;
&lt;td&gt;84.5%&lt;/td&gt;
&lt;td&gt;85.2%&lt;/td&gt;
&lt;td&gt;77.0%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;85.9%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The strengths cluster: long-context retrieval, video, front-end code and browser automation. 97.0% on 128k needle retrieval is the only score above 95 in the table, and AutomationBench at 30.4% against Sonnet 5's 10.7% is not a margin you get from a better prompt.&lt;/p&gt;

&lt;p&gt;The weaknesses cluster just as tightly, on work that runs long. It sits behind GPT-5.6 Terra on both Terminal-bench versions, 85.8% against 87.4% on 2.1 and 14.9% against 20.8% on 3.0, DeepSWE v1.1 lands 4.3 points behind Terra as well, and GDPVal-AA v2's 1525 is last in the table. Pick this model for a step, not for an eight-hour unattended loop.&lt;/p&gt;

&lt;p&gt;Two things to read carefully in that table. &lt;strong&gt;CharXiv is the one row where 3.7 is worse than the model it replaces&lt;/strong&gt;, 84.5% against 85.2% without tools and 88.7% against 89.4% with them, so chart-heavy multimodal work deserves your own A/B before the swap. And the Terminal-bench rows are vendor-run. The card claims 85.8% on Terminal-bench 2.1, which sits above the 83.8% ± 1.2% held by the top entry on the &lt;a href="https://www.tbench.ai/leaderboard/terminal-bench/2.1" rel="noopener noreferrer"&gt;public Terminal-Bench 2.1 board&lt;/a&gt;, 17 verified submissions when we last pulled it and none of them 3.7 Flash. A self-reported score that clears the verified leader is the one number in this table to read as a claim rather than a result.&lt;/p&gt;

&lt;p&gt;Two more lines from the card, in Google's own words: the knowledge cutoff is March 2026 with some domains still capped at January 2025, and there "may also be occasional slowness or timeout issues."&lt;/p&gt;

&lt;h2&gt;
  
  
  When Should You Pick Gemini 3.7 Flash?
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Workload&lt;/th&gt;
&lt;th&gt;Call&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;RAG and retrieval past 128k&lt;/td&gt;
&lt;td&gt;Pick it&lt;/td&gt;
&lt;td&gt;MRCR v2 at 97.0%, highest in the card&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Video understanding and QA&lt;/td&gt;
&lt;td&gt;Pick it&lt;/td&gt;
&lt;td&gt;LVBench 85.4%, 17 points over Sonnet 5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Front-end and web code generation&lt;/td&gt;
&lt;td&gt;Pick it&lt;/td&gt;
&lt;td&gt;Code Arena 1588, top of the table&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Browser and GUI automation&lt;/td&gt;
&lt;td&gt;Pick it&lt;/td&gt;
&lt;td&gt;AutomationBench 30.4% vs Sonnet 5's 10.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Long-horizon terminal agents&lt;/td&gt;
&lt;td&gt;Careful&lt;/td&gt;
&lt;td&gt;Terminal-bench 3.0 at 14.9%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Interactive chat and voice front-ends&lt;/td&gt;
&lt;td&gt;Look elsewhere&lt;/td&gt;
&lt;td&gt;No Live API, and thinking time precedes the first token&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chart-dense multimodal analysis&lt;/td&gt;
&lt;td&gt;Test first&lt;/td&gt;
&lt;td&gt;CharXiv is the one regression against 3.6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Single responses over 64k tokens&lt;/td&gt;
&lt;td&gt;Chunk it&lt;/td&gt;
&lt;td&gt;Output ceiling is 65,536&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If your workload is not on that list, &lt;a href="https://ofox.ai/model-finder" rel="noopener noreferrer"&gt;ofox's model finder&lt;/a&gt; ranks 100+ models by quality, cost and speed for a chosen task type (coding, agents, long-document RAG, vision, extraction, translation), pulling live prices and context limits. No signup, no key.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which Models Score About the Same?
&lt;/h2&gt;

&lt;p&gt;Four models sit inside three points of each other on the Artificial Analysis Intelligence Index, at prices that differ by more than 3x. Scores and rates below are the ones printed in Google's model card, which is one source rather than four, so re-check a vendor's page before you commit a budget to it.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;AA Index&lt;/th&gt;
&lt;th&gt;Input / 1M&lt;/th&gt;
&lt;th&gt;Output / 1M&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Terra&lt;/td&gt;
&lt;td&gt;57&lt;/td&gt;
&lt;td&gt;$2.00&lt;/td&gt;
&lt;td&gt;$12.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Muse Spark 1.2&lt;/td&gt;
&lt;td&gt;57&lt;/td&gt;
&lt;td&gt;$1.25&lt;/td&gt;
&lt;td&gt;$4.25&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Gemini 3.7 Flash&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;56&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.75&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$3.75&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://ofox.ai/models/anthropic/claude-sonnet-5" rel="noopener noreferrer"&gt;Claude Sonnet 5&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;55&lt;/td&gt;
&lt;td&gt;$2.00&lt;/td&gt;
&lt;td&gt;$10.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://ofox.ai/models/google/gemini-3.6-flash" rel="noopener noreferrer"&gt;Gemini 3.6 Flash&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;52&lt;/td&gt;
&lt;td&gt;$0.75&lt;/td&gt;
&lt;td&gt;$3.75&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Cheapest of the four on output, 3.2x below Terra, one point behind it on the composite. The catch is the composite: it hides the terminal and GDPVal rows where 3.7 Flash finishes last, so treat "comparable" as a statement about the index and nothing else.&lt;/p&gt;

&lt;p&gt;One row down, 3.6 Flash is now the same price for four fewer points, which leaves it with no case except prompts you have already tuned against its behaviour. Our &lt;a href="https://ofox.ai/blog/deepseek-v4-flash-vs-gemini-3-6-flash-2026/" rel="noopener noreferrer"&gt;3.6 Flash cost comparison against DeepSeek V4 Flash&lt;/a&gt; has the per-task math if you are still on it, and the &lt;a href="https://ofox.ai/blog/gemini-3-1-flash-lite-vs-deepseek-v4-flash-budget-agents-2026/" rel="noopener noreferrer"&gt;3.1 Flash-Lite budget agent comparison&lt;/a&gt; covers the tier below. Coming from two generations back, the &lt;a href="https://ofox.ai/blog/gemini-3-5-flash-coding-agents-guide-2026/" rel="noopener noreferrer"&gt;3.5 Flash coding agent guide&lt;/a&gt; still applies to 3.7 unchanged except for the model ID and the missing &lt;code&gt;minimal&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://ofox.ai/" rel="noopener noreferrer"&gt;15% off top-ups through 2026-08-31 at ofox&lt;/a&gt; puts &lt;code&gt;google/gemini-3.7-flash&lt;/code&gt; and &lt;code&gt;google/gemini-3.6-flash&lt;/code&gt; behind one OpenAI-compatible endpoint, which makes the A/B in the snippet above a one-string change.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;a href="https://ai.google.dev/gemini-api/docs/models/gemini-3-7-flash" rel="noopener noreferrer"&gt;Google: Gemini 3.7 Flash model page&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;  &lt;a href="https://ai.google.dev/gemini-api/docs/pricing" rel="noopener noreferrer"&gt;Google: Gemini API pricing&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;  &lt;a href="https://ai.google.dev/gemini-api/docs/thinking" rel="noopener noreferrer"&gt;Google: thinking and thinking_level&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;  &lt;a href="https://ai.google.dev/gemini-api/docs/openai" rel="noopener noreferrer"&gt;Google: OpenAI compatibility&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;  &lt;a href="https://deepmind.google/models/model-cards/gemini-3-7-flash/" rel="noopener noreferrer"&gt;DeepMind: Gemini 3.7 Flash model card&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;  &lt;a href="https://artificialanalysis.ai/models/gemini-3-7-flash" rel="noopener noreferrer"&gt;Artificial Analysis: Gemini 3.7 Flash&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;  &lt;a href="https://www.tbench.ai/leaderboard/terminal-bench/2.1" rel="noopener noreferrer"&gt;Terminal-Bench 2.1 leaderboard&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;  &lt;a href="https://web.archive.org/web/20260722024041/https://ai.google.dev/gemini-api/docs/pricing" rel="noopener noreferrer"&gt;Internet Archive: Gemini API pricing, 2026-07-22&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;  &lt;a href="https://web.archive.org/web/20260812042335/https://ai.google.dev/gemini-api/docs/pricing" rel="noopener noreferrer"&gt;Internet Archive: Gemini API pricing, 2026-08-12&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;  &lt;a href="https://web.archive.org/web/20260814074039/https://ai.google.dev/gemini-api/docs/pricing" rel="noopener noreferrer"&gt;Internet Archive: Gemini API pricing, 2026-08-14&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;  &lt;a href="https://ofox.ai/models/google/gemini-3.7-flash" rel="noopener noreferrer"&gt;ofox model page: Gemini 3.7 Flash&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  How much does the Gemini 3.7 Flash API cost?
&lt;/h3&gt;

&lt;p&gt;$0.75 per million input tokens and $3.75 per million output through December 31, 2026, then $1.50 and $7.50 from January 1, 2027. Google's pricing page prints both rows in the same cell. Batch and Flex are half that at $0.375 / $1.875, doubling on the same date. Output price includes thinking tokens, so reasoning is billed at the output rate.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is Gemini 3.7 Flash cheaper than Gemini 3.6 Flash?
&lt;/h3&gt;

&lt;p&gt;No, they are identically priced. Google's pricing page carries the same four tiers and the same two dates for both models. Archived copies of that page show why the half-price framing does not survive contact with it: on 2026-08-12 3.6 Flash was $1.50 / $7.50 flat and 3.7 Flash was not listed, and by 2026-08-14 both read $0.75 / $3.75 through 2026-12-31. The baseline was cut inside the same 48 hours as the comparison.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why does reasoning_effort minimal fail on Gemini 3.7 Flash?
&lt;/h3&gt;

&lt;p&gt;Because 3.7 Flash dropped that tier. The model page states 'Note: minimal is not supported and returns an error.' In our probe the request came back HTTP 400 with 'Thinking level is unsupported: THINKING_LEVEL_MINIMAL' as an invalid_request_error. 3.6 Flash accepts minimal, so any config that hardcodes it breaks on the model ID swap rather than degrading quietly.&lt;/p&gt;

&lt;h3&gt;
  
  
  What are the thinking levels on Gemini 3.7 Flash?
&lt;/h3&gt;

&lt;p&gt;The native parameter is thinking_level and it takes low, medium and high, defaulting to medium. Through the OpenAI-compatible layer the parameter is reasoning_effort. We measured 10 streamed calls per tier on one short prompt: low returned a median of 354 completion tokens against 688 at high, with no overlap between the two ranges, while the visible answer stayed around 40 tokens at every tier.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is the context and output limit on Gemini 3.7 Flash?
&lt;/h3&gt;

&lt;p&gt;1,048,576 input tokens and 65,536 output tokens. The million-token window does not mean million-token responses, so long-document rewrites and whole-repo generation still need chunking. Inputs accept text, image, video, audio and PDF; output is text only.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is Gemini 3.7 Flash fast?
&lt;/h3&gt;

&lt;p&gt;Both, depending on which half you measure. Artificial Analysis ranks it 1 of 182 on output speed at 389.5 tokens per second, and on the same page measures 14.52 seconds to first token at high on Google's own API against a 2.80-second median for its price peers. First-token latency is a different story from throughput, because thinking time lands in front of the first token: our medians on one short prompt were 4.19 seconds at low, 5.47 at medium and 7.69 at high. A model named Flash can still make a user wait if the effort setting is high.&lt;/p&gt;

&lt;h3&gt;
  
  
  Which models score about the same as Gemini 3.7 Flash?
&lt;/h3&gt;

&lt;p&gt;On the Artificial Analysis Intelligence Index printed in Google's model card, 3.7 Flash scores 56 against 57 for GPT-5.6 Terra and Muse Spark 1.2 and 55 for Claude Sonnet 5. The same card lists them at $2.00 / $12.00, $1.25 / $4.25 and $2.00 / $10.00 per million tokens, so 3.7 Flash is the cheapest of the four by output price.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is the Gemini 3.7 Flash model ID on gateways?
&lt;/h3&gt;

&lt;p&gt;google/gemini-3.7-flash on ofox, over either the OpenAI /v1/chat/completions protocol or Google's native Gemini protocol. Calling Google directly the code is gemini-3.7-flash, and the OpenAI-compatible base URL is &lt;a href="https://generativelanguage.googleapis.com/v1beta/openai/" rel="noopener noreferrer"&gt;https://generativelanguage.googleapis.com/v1beta/openai/&lt;/a&gt;.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://ofox.ai/blog/gemini-3-7-flash-api-guide-2026/" rel="noopener noreferrer"&gt;ofox.ai/blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>gemini</category>
      <category>googleai</category>
      <category>pricing</category>
    </item>
    <item>
      <title>Discounted LLM APIs: 13 Models Below List, GPT-5.6 Sol at Half Price</title>
      <dc:creator>Owen</dc:creator>
      <pubDate>Sat, 22 Aug 2026 04:34:31 +0000</pubDate>
      <link>https://dev.to/owen_fox/discounted-llm-apis-13-models-below-list-gpt-56-sol-at-half-price-515a</link>
      <guid>https://dev.to/owen_fox/discounted-llm-apis-13-models-below-list-gpt-56-sol-at-half-price-515a</guid>
      <description>&lt;h1&gt;
  
  
  Discounted LLM APIs: 13 Models Below List, GPT-5.6 Sol at Half Price
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;Thirteen models across five series are billing below their struck-through price on ofox right now.&lt;/strong&gt; The best-value page lists them sorted by discount size, and it is worth two minutes because the interesting part is not the percentage.&lt;/p&gt;

&lt;p&gt;The interesting part is what each percentage is measured from. Three of these cuts mean three different things, and only one of them is a gateway putting its own margin on the table.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which Models Are Discounted?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Five series: GPT, Gemini, GLM, Doubao and Seedance.&lt;/strong&gt; Text models bill per million tokens, Seedance bills per second of output.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Input / 1M&lt;/th&gt;
&lt;th&gt;Output / 1M&lt;/th&gt;
&lt;th&gt;Cached input&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;openai/gpt-5.6-luna&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.20&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$1.20&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;$0.02&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;openai/gpt-5.6-terra&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;$2.00&lt;/td&gt;
&lt;td&gt;$12.00&lt;/td&gt;
&lt;td&gt;$0.20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;openai/gpt-5.6-sol&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;$2.50&lt;/td&gt;
&lt;td&gt;$15.00&lt;/td&gt;
&lt;td&gt;$0.25&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;google/gemini-3.7-flash&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;$0.75&lt;/td&gt;
&lt;td&gt;$3.75&lt;/td&gt;
&lt;td&gt;$0.075&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;google/gemini-3.6-flash&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;$0.75&lt;/td&gt;
&lt;td&gt;$3.75&lt;/td&gt;
&lt;td&gt;$0.075&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;z-ai/glm-5.2&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;$0.98&lt;/td&gt;
&lt;td&gt;$3.08&lt;/td&gt;
&lt;td&gt;$0.182&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;z-ai/glm-5.3&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;$1.26&lt;/td&gt;
&lt;td&gt;$3.96&lt;/td&gt;
&lt;td&gt;$0.234&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;volcengine/doubao-seed-2.1-pro&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;$0.7072&lt;/td&gt;
&lt;td&gt;$3.536&lt;/td&gt;
&lt;td&gt;$0.1416&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;volcengine/doubao-seed-2.1-turbo&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;$0.3536&lt;/td&gt;
&lt;td&gt;$1.7696&lt;/td&gt;
&lt;td&gt;$0.068&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The four Seedance models bill on a different meter: &lt;code&gt;seedance-2.0-mini&lt;/code&gt; from $0.02 per second of output, &lt;code&gt;seedance-2.0-fast&lt;/code&gt; at $0.042, &lt;code&gt;seedance-2.0&lt;/code&gt; at $0.063, and &lt;code&gt;seedance-2.5&lt;/code&gt; from $0.11 at 480p, $0.48 at 1080p text-to-video and $0.568 at 1080p video-to-video. That last model prices per resolution and mode rather than as a single number, so the discount page shows only its 1080p figure while the model page carries the full grid. Duration and resolution drive that bill, not prompt length, which is why a 30-second clip costs roughly five times a 6-second one at the same resolution.&lt;/p&gt;

&lt;h2&gt;
  
  
  Is a Discounted Gateway Rate Actually Cheaper Than the Vendor?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Sometimes, and the honest answer splits three ways.&lt;/strong&gt; We checked every text rate above against the vendor's own pricing page on 2026-08-21. They fall into three groups, and knowing which group a model is in is the difference between a real saving and a number that looks like one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Group one: below the vendor's list.&lt;/strong&gt; GLM-5.2 bills $0.98 / $3.08 here against Z.ai's own $1.40 / $4.40, which is 30% off the rate you would pay calling Z.ai directly. GLM-5.3 is $1.26 / $3.96 against the same $1.40 / $4.40 list, so 10% off. The two Doubao models belong here too: Seed 2.1 Pro bills $0.7072 / $3.536 and Turbo $0.3536 / $1.7696, against the ¥6 / ¥30 per million ByteDance published at launch, which converts to roughly $0.88 / $4.41 at the rate we recorded in our Seed 2.1 access guide. Volcano Engine prices in CNY, so treat that conversion as approximate rather than a cent-for-cent comparison. GPT-5.6 Sol is the largest of these: $2.50 / $15 against OpenAI's standard-tier $5 / $30. That is half, on the tier you actually call synchronously.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Group two: matches the vendor exactly.&lt;/strong&gt; GPT-5.6 Luna bills $0.20 / $1.20, which is OpenAI's standard rate for that model to the cent. GPT-5.6 Terra bills $2 / $12, likewise identical. You save nothing per token against calling OpenAI directly, and that is the point: no markup, one key, and the fallback routing comes free rather than costing you a premium.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Group three: the vendor's own promotion, passed through.&lt;/strong&gt; Gemini 3.7 Flash and 3.6 Flash both bill $0.75 / $3.75. So does Google. That rate is Google's introductory price, printed on its pricing page with an expiry of 2026-12-31, after which both models return to $1.50 / $7.50 on 2027-01-01. Nobody is discounting Gemini for you. The gateway is passing Google's own limited-time number through unchanged, which is the correct behaviour, but it means the saving disappears on a date Google controls rather than one a gateway controls.&lt;/p&gt;

&lt;p&gt;If you are budgeting past new year, that third group is the one to model twice. We wrote up what the Gemini 3.7 Flash rate is measured against in detail, including the archived pricing pages that show the comparison baseline moving inside the same 48 hours as the launch.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Does GPT-5.6 Cost Here Against OpenAI?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;One of the three tiers is genuinely half price; the other two match the source.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Rate here&lt;/th&gt;
&lt;th&gt;OpenAI standard tier&lt;/th&gt;
&lt;th&gt;Difference&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Sol&lt;/td&gt;
&lt;td&gt;$2.50 / $15&lt;/td&gt;
&lt;td&gt;$5.00 / $30.00&lt;/td&gt;
&lt;td&gt;half&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Terra&lt;/td&gt;
&lt;td&gt;$2.00 / $12.00&lt;/td&gt;
&lt;td&gt;$2.00 / $12.00&lt;/td&gt;
&lt;td&gt;identical&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Luna&lt;/td&gt;
&lt;td&gt;$0.20 / $1.20&lt;/td&gt;
&lt;td&gt;$0.20 / $1.20&lt;/td&gt;
&lt;td&gt;identical&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Worth knowing about the Sol row: $2.50 / $15 is also what OpenAI charges for Sol on its Batch and Flex tiers. Those tiers trade latency for price, with Batch returning results asynchronously within a window rather than on the call. Here you get that number on a normal synchronous request, which is the whole substance of that particular discount.&lt;/p&gt;

&lt;p&gt;Cached input follows the same shape: $0.25 per million on Sol, $0.20 on Terra, $0.02 on Luna. On a long-running agent that replays the same system prompt on every turn, cached input is usually the line that decides the monthly bill, not the headline input rate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which Cuts Follow the Vendor's Promotion?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Gemini and Seedance, both of them.&lt;/strong&gt; Google's Flash rate carries a printed expiry. The Seedance discounts track BytePlus promotional windows, which carry printed start and end dates rather than running open-ended, and the 2.5 window covers 1080p output only.&lt;/p&gt;

&lt;p&gt;This matters for one specific decision: whether to hardcode a price into a cost model. A rate in group one is a commercial arrangement that tends to move slowly. A rate in group three moves when the vendor's marketing calendar moves. Both are real money today. Only one of them is a number you can put in a spreadsheet cell labelled "next quarter".&lt;/p&gt;

&lt;p&gt;The best-value page states this in its own words, and it is the sentence on that page worth reading twice: "promotions follow vendor changes and may end at any time."&lt;/p&gt;

&lt;h2&gt;
  
  
  How Do I Use the Discounted Rate?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;You do not do anything.&lt;/strong&gt; There is no coupon field, no plan to switch and no minimum top-up. The rate on the model page is what the request is billed at.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sk-...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.ofox.ai/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;z-ai/glm-5.2&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;              &lt;span class="c1"&gt;# $0.98 in / $3.08 out per 1M
&lt;/span&gt;    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Summarise this changelog&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Swapping models is a one-string change, which is the practical reason to care about a discount list at all: if GLM-5.2 at $0.98 clears your quality bar for a classification job currently running on a frontier model at $5, the migration is one line and the saving is 5x. The model finder ranks the catalogue by task type and budget if you want the comparison done for you, and our API pricing comparison covers the wider field including the models that are not on discount.&lt;/p&gt;

&lt;p&gt;Two more worth reading if a specific series is your candidate: GLM 5.3's pricing and its reasoning_effort trap, where the default effort setting bills 35x the cheapest one, and the Doubao Seed 2.1 access guide if you want the Volcengine models without a Volcengine signup.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which Model Should You Pick From the List?
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;If you need&lt;/th&gt;
&lt;th&gt;Pick&lt;/th&gt;
&lt;th&gt;Rate&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Cheapest usable text at volume&lt;/td&gt;
&lt;td&gt;GPT-5.6 Luna&lt;/td&gt;
&lt;td&gt;$0.20 / $1.20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cheapest that accepts video input&lt;/td&gt;
&lt;td&gt;Doubao Seed 2.1 Turbo&lt;/td&gt;
&lt;td&gt;$0.3536 / $1.7696&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Frontier reasoning at half list&lt;/td&gt;
&lt;td&gt;GPT-5.6 Sol&lt;/td&gt;
&lt;td&gt;$2.50 / $15&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Long-context retrieval and video input&lt;/td&gt;
&lt;td&gt;Gemini 3.7 Flash&lt;/td&gt;
&lt;td&gt;$0.75 / $3.75&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Open-weight-family coding and agents&lt;/td&gt;
&lt;td&gt;GLM-5.2&lt;/td&gt;
&lt;td&gt;$0.98 / $3.08&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Video generation, cheapest per second&lt;/td&gt;
&lt;td&gt;Seedance 2.0 Mini&lt;/td&gt;
&lt;td&gt;from $0.02/s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The rates above are what the console bills today. The percentages next to them on the page are context for how that number got there, and as this post argues, the context is worth more than the percentage.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;  ofox: discounted models&lt;/li&gt;
&lt;li&gt;  OpenAI: API pricing&lt;/li&gt;
&lt;li&gt;  Google: Gemini API pricing&lt;/li&gt;
&lt;li&gt;  ofox model page: GPT-5.6 Luna&lt;/li&gt;
&lt;li&gt;  ofox model page: GPT-5.6 Sol&lt;/li&gt;
&lt;li&gt;  ofox model page: GLM-5.2&lt;/li&gt;
&lt;li&gt;  ofox model page: Doubao Seed 2.1 Pro&lt;/li&gt;
&lt;li&gt;  Volcano Engine: ModelArk pricing&lt;/li&gt;
&lt;li&gt;  ofox model finder&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Which LLM APIs are discounted on ofox?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Thirteen models across five series: GPT (5.6 Luna, Sol, Terra), Gemini (3.7 Flash, 3.6 Flash), GLM (5.2, 5.3), Doubao (Seed 2.1 Pro, Seed 2.1 Turbo) and Seedance (2.0, 2.0 Fast, 2.0 Mini, 2.5). The live list sits on the best-value page, sorted by discount size, and it changes when the vendors change their own pricing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is a discounted gateway rate cheaper than calling the vendor directly?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It depends on the model, and the answer falls into three groups. GLM-5.2 at $0.98 / $3.08 is genuinely below Z.ai's own $1.40 / $4.40. GPT-5.6 Luna at $0.20 / $1.20 matches OpenAI's standard rate exactly, so you save nothing per token but you also pay no markup. Gemini 3.7 Flash at $0.75 / $3.75 is Google's own introductory price, which expires on 2026-12-31 and is not a gateway discount at all.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What does GPT-5.6 cost through a gateway?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Sol is $2.50 in and $15 out per million tokens, Terra is $2 and $12, Luna is $0.20 and $1.20. OpenAI's own standard-tier list for the same models is $5 / $30, $2 / $12 and $0.20 / $1.20. Sol is the one where the gateway rate is half the vendor's standard tier; the other two match it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do I need a coupon code for the discounted rate?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;No. The rate on the model page is the rate you are billed. There is no coupon field, no minimum top-up and no separate discount plan to enable. The struck-through number is context, not something you have to unlock.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How long do these discounts last?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;There is no single end date, because the rates track each vendor's own pricing. Google's Gemini 3.7 and 3.6 Flash rate is dated to 2026-12-31 on Google's pricing page. Seedance promotions follow BytePlus windows. GLM rates have no published end date. Treat any of them as good for today and re-check before you commit a quarterly budget.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Which discounted model is cheapest for high-volume text?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;GPT-5.6 Luna at $0.20 in and $1.20 out per million tokens, with cached input at $0.02. Doubao Seed 2.1 Turbo is next at $0.3536 and $1.7696. Both are an order of magnitude below the frontier tier, and both keep tool calling and vision.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Are video models discounted too?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Four Seedance models are, and they bill per second of output rather than per token: Seedance 2.0 Mini from $0.02/s, 2.0 Fast at $0.042/s, 2.0 at $0.063/s, and 2.5 from $0.11/s at 480p, $0.48/s at 1080p text-to-video and $0.568/s at 1080p video-to-video. Per-second billing means the resolution and duration you request drive the bill, not prompt length.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://ofox.ai/blog/discounted-llm-api-pricing-2026/" rel="noopener noreferrer"&gt;ofox.ai/blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>pricing</category>
      <category>openai</category>
      <category>costoptimization</category>
    </item>
    <item>
      <title>DeepSeek V4 Flash Vision: Images Cap at 384 Tokens</title>
      <dc:creator>Owen</dc:creator>
      <pubDate>Sat, 22 Aug 2026 00:33:52 +0000</pubDate>
      <link>https://dev.to/owen_fox/deepseek-v4-flash-vision-images-cap-at-384-tokens-134h</link>
      <guid>https://dev.to/owen_fox/deepseek-v4-flash-vision-images-cap-at-384-tokens-134h</guid>
      <description>&lt;p&gt;&lt;strong&gt;DeepSeek shipped a vision variant of V4 Flash, and it costs exactly what the text model costs.&lt;/strong&gt; Model ID &lt;code&gt;deepseek-v4-flash-vision-exp&lt;/code&gt;, live on ofox as &lt;code&gt;deepseek/deepseek-v4-flash-vision-exp&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Same 1M context, same 384K output ceiling, same 2500 concurrency, same price on every row of the pricing table. The interesting question is what an image actually costs once it becomes tokens, and the answer has a hard ceiling.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Model:        deepseek/deepseek-v4-flash-vision-exp
Price:        $0.44 in / $1.32 out per 1M at peak
              $0.22 / $0.66 off-peak (peak = 01:00-04:00, 06:00-10:00 UTC)
Cache hit:    $0.014 peak / $0.007 off-peak per 1M
Context:      1M in / 384K out
Images:       384 tokens max per image, 600 images per request
Thinking:     on by default at effort high, disable via thinking.type
Measured:     800x800 and 3000x3000 both cost 348 input tokens
Snapshot:     2026-08-21, n=5 per configuration
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Last updated 2026-08-21. The model ID carries an &lt;code&gt;-exp&lt;/code&gt; suffix, so treat its availability as experimental and re-check before building a dependency on it.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is DeepSeek V4 Flash Vision Exp?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;V4 Flash with image input, priced identically.&lt;/strong&gt; DeepSeek's pricing page puts the three models side by side, and the vision column matches the flash column cell for cell.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;code&gt;deepseek-v4-flash&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;&lt;code&gt;deepseek-v4-flash-vision-exp&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;&lt;code&gt;deepseek-v4-pro&lt;/code&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Input, cache miss (peak)&lt;/td&gt;
&lt;td&gt;$0.44&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.44&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;$1.32&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Input, cache hit (peak)&lt;/td&gt;
&lt;td&gt;$0.014&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.014&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;$0.044&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output (peak)&lt;/td&gt;
&lt;td&gt;$1.32&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$1.32&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;$3.96&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Max output&lt;/td&gt;
&lt;td&gt;384K&lt;/td&gt;
&lt;td&gt;384K&lt;/td&gt;
&lt;td&gt;384K&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Concurrency limit&lt;/td&gt;
&lt;td&gt;2500&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2500&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;500&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Off-peak rates are half the peak rates across the board. Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, three hours plus four, which leaves seventeen hours a day at the lower rate.&lt;/p&gt;

&lt;p&gt;One thing moves in the wrong direction. On the feature table, FIM completion reads "Non-thinking mode only" for V4 Flash and V4 Pro, and &lt;strong&gt;"Not supported"&lt;/strong&gt; for the vision variant. Everything else carries over: JSON output, tool calls, the Responses API, the Anthropic-format API and chat prefix completion are all ticked.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Much Does an Image Cost?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;348 tokens for an 800x800 image, and the same 348 for a 3000x3000 one.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Every image is resized before inference. Images below roughly 384x384 total pixels are scaled up, larger images are scaled down, both preserving aspect ratio, and the target is a total pixel count near an 800x800 image. The documented consequence is a ceiling of &lt;strong&gt;384 tokens per image&lt;/strong&gt;, which is why a 2000x2000 photo and a 5000x5000 photo bill the same.&lt;/p&gt;

&lt;p&gt;Test images sent through the ofox route on 2026-08-21, subtracting a 91-token text-only baseline measured on the same prompt:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Image&lt;/th&gt;
&lt;th&gt;&lt;code&gt;prompt_tokens&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;Image tokens&lt;/th&gt;
&lt;th&gt;Cost at peak input&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;none (text only)&lt;/td&gt;
&lt;td&gt;91&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;200x200&lt;/td&gt;
&lt;td&gt;207&lt;/td&gt;
&lt;td&gt;116&lt;/td&gt;
&lt;td&gt;$0.000051&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;800x800&lt;/td&gt;
&lt;td&gt;439&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;348&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;$0.000153&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3000x3000&lt;/td&gt;
&lt;td&gt;439&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;348&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;$0.000153&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6000x3000&lt;/td&gt;
&lt;td&gt;395&lt;/td&gt;
&lt;td&gt;304&lt;/td&gt;
&lt;td&gt;$0.000134&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two things fall out of that table. The 800x800 and 3000x3000 rows are identical, which is the resize rule doing exactly what the documentation says. And the 6000x3000 row is &lt;em&gt;cheaper&lt;/em&gt; than the square ones, because a 2:1 image that has been squeezed into the same total pixel budget tiles into fewer patches than a square one does.&lt;/p&gt;

&lt;p&gt;At $0.44 per million input tokens, a thousand images cost about &lt;strong&gt;15 cents&lt;/strong&gt; at peak and half that off-peak. Image input on this model is not a line you need to model. What you do need to model is the next section.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does Thinking Mode Change the Bill?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;It costs 80 extra input tokens before the model has looked at anything.&lt;/strong&gt; Thinking mode is on by default with effort set to &lt;code&gt;high&lt;/code&gt;. Test results from running the same image and the same prompt five times per configuration:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;&lt;code&gt;prompt_tokens&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;Completion median&lt;/th&gt;
&lt;th&gt;Reasoning median&lt;/th&gt;
&lt;th&gt;Latency median&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Default (thinking on)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;439&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;114 (68–425)&lt;/td&gt;
&lt;td&gt;91&lt;/td&gt;
&lt;td&gt;1.9 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;thinking: {"type": "disabled"}&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;359&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;19 (18–20)&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;1.0 s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The &lt;code&gt;prompt_tokens&lt;/code&gt; figure did not vary once across ten calls: 439 with thinking on, 359 with it off, every single time. That 80-token gap is a template the API adds when thinking is enabled, and you pay input rate on it whether the model ends up reasoning or not.&lt;/p&gt;

&lt;p&gt;On the output side the gap is wider and less predictable. With thinking on, the completion ranged from 68 to 425 tokens across five identical calls for a one-sentence image description. With it off, the same five calls returned 18 to 20 tokens. For "describe this image in one sentence", the reasoning is buying variance rather than accuracy.&lt;/p&gt;

&lt;p&gt;Add it up at peak rates. Thinking on: 439 input plus 114 output is about &lt;strong&gt;$0.000343&lt;/strong&gt; per call. Thinking off: 359 plus 19 is about &lt;strong&gt;$0.000183&lt;/strong&gt;. Nearly half the bill, on a task where the visible answer is the same either way.&lt;/p&gt;

&lt;p&gt;Switch it off with either of these:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"thinking"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"disabled"&lt;/span&gt;&lt;span class="p"&gt;}}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"reasoning_effort"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"none"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both worked on the test route, and both made the &lt;code&gt;reasoning_tokens&lt;/code&gt; field disappear from the usage object rather than merely shrink. If you want thinking but cheaper, the effort mapping is worth reading first: DeepSeek maps &lt;code&gt;low&lt;/code&gt; to low, &lt;code&gt;max&lt;/code&gt; to max, and &lt;strong&gt;&lt;code&gt;medium&lt;/code&gt;, &lt;code&gt;high&lt;/code&gt; and &lt;code&gt;xhigh&lt;/code&gt; all to high&lt;/strong&gt;. Three of the five values you might send are the same setting.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Do I Send an Image?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Standard OpenAI-compatible content blocks.&lt;/strong&gt; Three transports are available: inline base64, an external URL, and a Files API &lt;code&gt;file_id&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;base64&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sk-...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.ofox.ai/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;chart.png&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;b64&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;base64&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;b64encode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;()).&lt;/span&gt;&lt;span class="nf"&gt;decode&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;deepseek/deepseek-v4-flash-vision-exp&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;What is the trend in this chart?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;image_url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;image_url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;data:image/png;base64,&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;b64&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}},&lt;/span&gt;
    &lt;span class="p"&gt;]}],&lt;/span&gt;
    &lt;span class="n"&gt;extra_body&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;thinking&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;disabled&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}},&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Calling DeepSeek directly, the base URL becomes &lt;code&gt;https://api.deepseek.com&lt;/code&gt; and the model string drops its &lt;code&gt;deepseek/&lt;/code&gt; prefix. The Anthropic-format endpoint is &lt;code&gt;https://api.deepseek.com/anthropic&lt;/code&gt;, and images travel as &lt;code&gt;input_image&lt;/code&gt; parts on the Responses API.&lt;/p&gt;

&lt;p&gt;Two restrictions that return HTTP 400 rather than degrading: images are accepted &lt;strong&gt;only in &lt;code&gt;user&lt;/code&gt; messages&lt;/strong&gt;, and only the vision model accepts them at all. Send an image to &lt;code&gt;deepseek-v4-flash&lt;/code&gt; and you get "This model does not support image".&lt;/p&gt;

&lt;h2&gt;
  
  
  What Are the Limits?
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Limit&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Formats&lt;/td&gt;
&lt;td&gt;JPEG, PNG, GIF, WebP (detected from content, not filename)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Images per request&lt;/td&gt;
&lt;td&gt;600&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Request body&lt;/td&gt;
&lt;td&gt;48 MiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Single image, base64 or URL&lt;/td&gt;
&lt;td&gt;32 MiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Single image, Files API &lt;code&gt;file_id&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;64 MiB&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Total images per request&lt;/td&gt;
&lt;td&gt;64 MiB without &lt;code&gt;file_id&lt;/code&gt;, up to 200 MiB with&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Max dimension&lt;/td&gt;
&lt;td&gt;8192 px per side, &lt;strong&gt;4096 px&lt;/strong&gt; once a request carries 15+ images&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;External URL length&lt;/td&gt;
&lt;td&gt;8192 characters&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The dimension rule is the one that bites in production. A batch job that sends fourteen screenshots at 6000 px works; the fifteenth one silently changes the ceiling for the whole request.&lt;/p&gt;

&lt;h2&gt;
  
  
  Should You Switch From V4 Flash?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;If you send images at all, yes, because it costs nothing to.&lt;/strong&gt; The rates are identical, the context and output ceilings are identical, and the concurrency limit is identical. The only thing you give up is FIM completion, which matters for inline code completion and nothing else.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Workload&lt;/th&gt;
&lt;th&gt;Call&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Screenshot and document OCR at volume&lt;/td&gt;
&lt;td&gt;Switch. 384 tokens per page is hard to beat&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chart and diagram reading&lt;/td&gt;
&lt;td&gt;Switch, but keep thinking on for multi-step reads&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Text-only chat and classification&lt;/td&gt;
&lt;td&gt;No reason either way, prices match&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Inline code completion via FIM&lt;/td&gt;
&lt;td&gt;Stay on &lt;code&gt;deepseek-v4-flash&lt;/code&gt;, FIM is gone here&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Long-horizon agents on images&lt;/td&gt;
&lt;td&gt;Test first, the &lt;code&gt;-exp&lt;/code&gt; suffix is not a stability promise&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For the wider picture on which multimodal endpoints do what, the multimodal API guide covers vision alongside speech, and the V4 Flash against Gemini 3.6 Flash cost comparison has the per-task math on the text side of the same model.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;  &lt;a href="https://api-docs.deepseek.com/quick_start/pricing" rel="noopener noreferrer"&gt;DeepSeek: models and pricing&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;  &lt;a href="https://api-docs.deepseek.com/guides/vision" rel="noopener noreferrer"&gt;DeepSeek: vision guide&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;  &lt;a href="https://api-docs.deepseek.com/guides/thinking_mode" rel="noopener noreferrer"&gt;DeepSeek: thinking mode&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;  &lt;a href="https://ofox.ai/models/deepseek/deepseek-v4-flash-vision-exp" rel="noopener noreferrer"&gt;ofox model page: DeepSeek V4 Flash Vision Exp&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;  &lt;a href="https://ofox.ai/models/deepseek/deepseek-v4-flash-0731" rel="noopener noreferrer"&gt;ofox model page: DeepSeek V4 Flash 0731&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What is DeepSeek V4 Flash Vision Exp?
&lt;/h3&gt;

&lt;p&gt;It is a vision-capable variant of DeepSeek V4 Flash, model ID deepseek-v4-flash-vision-exp. It accepts images alongside text in user messages and bills them at the same rates as the text-only V4 Flash: $0.44 per million input tokens and $1.32 per million output at peak, half that off-peak. Context is 1M tokens with 384K maximum output, the same as V4 Flash.&lt;/p&gt;

&lt;h3&gt;
  
  
  How much does an image cost on DeepSeek V4 Flash Vision?
&lt;/h3&gt;

&lt;p&gt;Very little. Images are resized before inference so that the total pixel count lands near an 800x800 image, which caps each image at 384 input tokens. Testing measured 348 tokens for both an 800x800 and a 3000x3000 image on the same route, which works out to $0.000153 per image at the peak input rate and half that off-peak. Prompt length, not image resolution, is what moves that number.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is DeepSeek V4 Flash Vision more expensive than V4 Flash?
&lt;/h3&gt;

&lt;p&gt;No. DeepSeek's pricing page prints identical numbers for deepseek-v4-flash and deepseek-v4-flash-vision-exp across every row: cache hit, cache miss and output, at both peak and off-peak. The concurrency limit is also the same 2500. The vision variant is the text model's price list with image input added.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does DeepSeek V4 Flash Vision support thinking mode?
&lt;/h3&gt;

&lt;p&gt;Yes, and it is on by default with effort set to high. Send thinking with type disabled, or reasoning_effort none, to turn it off. In testing, turning it off dropped the median completion from 114 tokens to 19 and the median call from 1.9 seconds to 1.0.&lt;/p&gt;

&lt;h3&gt;
  
  
  How many images can one request carry?
&lt;/h3&gt;

&lt;p&gt;Up to 600. The per-request ceilings are 48 MiB of request body, 32 MiB per image sent as base64 or an external URL, 64 MiB per image via a Files API file_id, and 64 MiB of images in total without file_id references. Maximum dimension is 8192 pixels per side, dropping to 4096 once a request carries 15 or more images.&lt;/p&gt;

&lt;h3&gt;
  
  
  What image formats does DeepSeek V4 Flash Vision accept?
&lt;/h3&gt;

&lt;p&gt;JPEG, PNG, GIF and WebP. The format is detected from the file content rather than the filename or the declared MIME type, so a mislabelled extension does not break the call. Images are accepted only in user messages: an image in a system or assistant message returns HTTP 400.&lt;/p&gt;

&lt;h3&gt;
  
  
  What does DeepSeek V4 Flash Vision give up compared to V4 Flash?
&lt;/h3&gt;

&lt;p&gt;One line on the feature table. FIM completion, which V4 Flash supports in non-thinking mode, is listed as not supported on the vision variant. JSON output, tool calls, the Responses API, the Anthropic-format API and chat prefix completion all carry over unchanged.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://ofox.ai/blog/deepseek-v4-flash-vision-exp-image-tokens-2026/" rel="noopener noreferrer"&gt;ofox.ai/blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>deepseek</category>
      <category>vision</category>
      <category>pricing</category>
    </item>
    <item>
      <title>What Is a Context Window? Token Limits by Model (2026)</title>
      <dc:creator>Owen</dc:creator>
      <pubDate>Fri, 21 Aug 2026 00:02:38 +0000</pubDate>
      <link>https://dev.to/owen_fox/what-is-a-context-window-token-limits-by-model-2026-4k8c</link>
      <guid>https://dev.to/owen_fox/what-is-a-context-window-token-limits-by-model-2026-4k8c</guid>
      <description>&lt;h1&gt;
  
  
  What Is a Context Window? Token Limits by Model (2026)
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;A context window is the maximum number of tokens a model can handle in one request, counting the input and the output together.&lt;/strong&gt; We sent one identical document to nine models: it metered anywhere from 614 to 957 tokens. That is a 1.56x spread on the same text.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Unit:             tokens (~4 English chars each, with wide variation)
Scope:            one request; nothing persists between calls
Includes:         system prompt, history, tool defs, tool results, reasoning, reply
Typical sizes:    200K on cheaper models, 1M on current flagships
2026 maximum:     1,131,072 (Qwen 3.8 Max)
At the limit:     HTTP 400. Nothing is silently truncated
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What Counts Toward the Window?
&lt;/h2&gt;

&lt;p&gt;Everything the model reads and writes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;System prompt&lt;/strong&gt;, counted every turn&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Full message history&lt;/strong&gt;, all prior exchanges&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool definitions&lt;/strong&gt; — names, descriptions, JSON schemas&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool results&lt;/strong&gt;, often the largest single item&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reasoning tokens&lt;/strong&gt;, billed even when not returned to you&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The reply itself&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One consequence people hit on upgrade: on Claude Opus 5 thinking is on by default, so a request sized tightly around its answer on an older model can now run out of room mid-response.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Big Is Each Model's Window?
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Context window&lt;/th&gt;
&lt;th&gt;Max output&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Qwen 3.8 Max&lt;/td&gt;
&lt;td&gt;1,131,072&lt;/td&gt;
&lt;td&gt;131,072&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Sol&lt;/td&gt;
&lt;td&gt;1,050,000&lt;/td&gt;
&lt;td&gt;128,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.1 Pro&lt;/td&gt;
&lt;td&gt;1,048,576&lt;/td&gt;
&lt;td&gt;65,536&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GLM-5.2&lt;/td&gt;
&lt;td&gt;1,048,576&lt;/td&gt;
&lt;td&gt;128,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kimi K3&lt;/td&gt;
&lt;td&gt;1,048,576&lt;/td&gt;
&lt;td&gt;1,048,576&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 5&lt;/td&gt;
&lt;td&gt;1,000,000&lt;/td&gt;
&lt;td&gt;128,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V4 Flash&lt;/td&gt;
&lt;td&gt;1,000,000&lt;/td&gt;
&lt;td&gt;384,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Grok 4.20&lt;/td&gt;
&lt;td&gt;1,000,000&lt;/td&gt;
&lt;td&gt;not published&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Haiku 4.5&lt;/td&gt;
&lt;td&gt;200,000&lt;/td&gt;
&lt;td&gt;64,000&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;"1M" is nine different numbers, from exactly 1,000,000 to 1,131,072 — a 13% spread before you measure a single token of your own content.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why "1M Tokens" Doesn't Mean the Same Thing Everywhere
&lt;/h2&gt;

&lt;p&gt;We sent one 2,638-character English document to nine models:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Tokens&lt;/th&gt;
&lt;th&gt;Chars/token&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Grok 4.20&lt;/td&gt;
&lt;td&gt;614&lt;/td&gt;
&lt;td&gt;4.30&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Sol&lt;/td&gt;
&lt;td&gt;626&lt;/td&gt;
&lt;td&gt;4.21&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GLM-5.2&lt;/td&gt;
&lt;td&gt;632&lt;/td&gt;
&lt;td&gt;4.17&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V4 Flash&lt;/td&gt;
&lt;td&gt;634&lt;/td&gt;
&lt;td&gt;4.16&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.1 Pro&lt;/td&gt;
&lt;td&gt;684&lt;/td&gt;
&lt;td&gt;3.86&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 4.6&lt;/td&gt;
&lt;td&gt;698&lt;/td&gt;
&lt;td&gt;3.78&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen 3.8 Max&lt;/td&gt;
&lt;td&gt;706&lt;/td&gt;
&lt;td&gt;3.74&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kimi K3&lt;/td&gt;
&lt;td&gt;716&lt;/td&gt;
&lt;td&gt;3.68&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Claude Opus 5&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;957&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2.76&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Claude Opus 5 is the outlier for a documented reason: Anthropic states that Claude 4.7 and later use a newer tokenizer producing "approximately 30% more tokens for the same text."&lt;/p&gt;

&lt;p&gt;Translate that into how many copies of the document actually fit:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Advertised window&lt;/th&gt;
&lt;th&gt;Document copies&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Sol&lt;/td&gt;
&lt;td&gt;1,050,000&lt;/td&gt;
&lt;td&gt;1,677&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GLM-5.2&lt;/td&gt;
&lt;td&gt;1,048,576&lt;/td&gt;
&lt;td&gt;1,659&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Grok 4.20&lt;/td&gt;
&lt;td&gt;1,000,000&lt;/td&gt;
&lt;td&gt;1,629&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Qwen 3.8 Max&lt;/td&gt;
&lt;td&gt;1,131,072&lt;/td&gt;
&lt;td&gt;1,602&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V4 Flash&lt;/td&gt;
&lt;td&gt;1,000,000&lt;/td&gt;
&lt;td&gt;1,577&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.1 Pro&lt;/td&gt;
&lt;td&gt;1,048,576&lt;/td&gt;
&lt;td&gt;1,533&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kimi K3&lt;/td&gt;
&lt;td&gt;1,048,576&lt;/td&gt;
&lt;td&gt;1,464&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 4.6&lt;/td&gt;
&lt;td&gt;1,000,000&lt;/td&gt;
&lt;td&gt;1,432&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 5&lt;/td&gt;
&lt;td&gt;1,000,000&lt;/td&gt;
&lt;td&gt;1,045&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Similar advertised sizes, &lt;strong&gt;1.60x difference in real capacity for this document&lt;/strong&gt;. Content type shifts it again — code, English prose and non-English text tokenize differently. Measure your own content on the models you are choosing between.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Happens When You Exceed It?
&lt;/h2&gt;

&lt;p&gt;HTTP 400 and no output. Nothing is silently truncated.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"error"&lt;/span&gt;&lt;span class="p"&gt;:{&lt;/span&gt;&lt;span class="nl"&gt;"code"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"message"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"&amp;lt;400&amp;gt; InternalError.Algo.InvalidParameter: Range of input length should be [1, 30720]"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"invalid_request_error"&lt;/span&gt;&lt;span class="p"&gt;}}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Note the enforced input ceiling can be lower than the advertised total — 30,720 against 32,000 here, with the difference reserved for output.&lt;/p&gt;

&lt;p&gt;Vendors differ in how they signal it. OpenAI-compatible endpoints return 400 with a &lt;code&gt;context_length_exceeded&lt;/code&gt; code; Claude can finish with &lt;code&gt;stop_reason: "model_context_window_exceeded"&lt;/code&gt;. Both need distinct handling.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does a Bigger Window Cost More?
&lt;/h2&gt;

&lt;p&gt;Depends on the vendor, and this is where the cheaper headline rate can lose:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Flat rate:&lt;/strong&gt; Anthropic bills the full 1M at standard rates, no long-context premium&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tiered:&lt;/strong&gt; Gemini 3.1 Pro goes $2 → $4 per million input and $12 → $18 output once prompts exceed 200K; Grok 4.20 goes $1.25 → $2.50 input and $2.50 → $5.00 at the same threshold&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Is the Advertised Window the Same as Usable Context?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;No, and this is the most important caveat on the page.&lt;/strong&gt; A model that accepts 1M tokens does not reliably retrieve information at the far end of it. Retrieval accuracy degrades with distance for every public model. Benchmarks like RULER, MRCR v2 and NoLiMa measure what actually works rather than what is accepted.&lt;/p&gt;

&lt;p&gt;Treat the advertised window as an upper acceptance bound, not a capability claim.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Do You Fit More In?
&lt;/h2&gt;

&lt;p&gt;These reduce spend rather than expanding capacity:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Prompt caching&lt;/strong&gt; — stable prefixes bill at roughly 10% of input rates on a hit. Largest lever for repeated calls.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compaction&lt;/strong&gt; — server-side summarisation of earlier turns keeps long agent sessions alive.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context editing&lt;/strong&gt; — drop stale tool results and old reasoning blocks. Agent loops accumulate tool output faster than conversation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tokenizer choice&lt;/strong&gt; — picking the right model for your content is worth up to 1.6x effective capacity on its own.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  How Do You Compare Token Counts Yourself?
&lt;/h2&gt;

&lt;p&gt;Same request, different vendors, one loop:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.ofox.ai/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;YOUR_OFOX_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;your_document.txt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;anthropic/claude-opus-5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;openai/gpt-5.6-sol&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;google/gemini-3.1-pro-preview&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;moonshotai/kimi-k3&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;z-ai/glm-5.2&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;n&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;prompt_tokens&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;32&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="mi"&gt;7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; tokens   &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;n&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; chars/token&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;max_tokens: 1&lt;/code&gt; returns &lt;code&gt;prompt_tokens&lt;/code&gt; at minimal cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Is a context window the same as memory?&lt;/strong&gt;&lt;br&gt;
No. It is per-request only. Every turn you resend the whole conversation. Products that appear to remember you across sessions re-inject stored text; there is no persistent model memory.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How many words is 1 million tokens?&lt;/strong&gt;&lt;br&gt;
For English prose, roughly 440,000 to 685,000 words. Eight of nine models tested landed between 587,000 and 684,000; Claude Opus 5 reached about 439,000 because of its tokenizer. Code runs denser (2.4 to 3.6 chars/token), Chinese denser still.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why does the same file use more tokens on Claude than GPT?&lt;/strong&gt;&lt;br&gt;
Different tokenizers. Anthropic's newer tokenizer in Claude 4.7+ produces roughly 30% more tokens for identical text. On our test document: 957 tokens on Opus 5 against 626 on GPT-5.6 Sol.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does filling the window make the model slower?&lt;/strong&gt;&lt;br&gt;
Yes, and more expensive per turn. Every token is processed on each request, so a 400K-token conversation costs 400K tokens on every new turn unless prompt caching is active.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I increase a model's context window?&lt;/strong&gt;&lt;br&gt;
No. It is fixed by the model with no adjustment parameter. You can only spend it better.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://platform.claude.com/docs/en/build-with-claude/context-windows" rel="noopener noreferrer"&gt;Anthropic: context windows&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://platform.claude.com/docs/en/about-claude/pricing" rel="noopener noreferrer"&gt;Anthropic: pricing, including the tokenizer note&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ai.google.dev/gemini-api/docs/pricing" rel="noopener noreferrer"&gt;Gemini API pricing&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ofox.ai/models" rel="noopener noreferrer"&gt;ofox model catalog&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://ofox.ai/blog/what-is-a-context-window-token-limits-by-model-2026/" rel="noopener noreferrer"&gt;ofox.ai/blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>concept</category>
      <category>contextwindow</category>
      <category>tokens</category>
      <category>modelcomparison</category>
    </item>
    <item>
      <title>Seedance 2.5 vs 2.0: What Changed, and When 2.0 Still Wins</title>
      <dc:creator>Owen</dc:creator>
      <pubDate>Fri, 21 Aug 2026 00:02:03 +0000</pubDate>
      <link>https://dev.to/owen_fox/seedance-25-vs-20-what-changed-and-when-20-still-wins-37og</link>
      <guid>https://dev.to/owen_fox/seedance-25-vs-20-what-changed-and-when-20-still-wins-37og</guid>
      <description>&lt;h1&gt;
  
  
  Seedance 2.5 vs 2.0: What Changed, and When 2.0 Still Wins
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;Seedance 2.5 doubles maximum clip length to 30 seconds and accepts 50 reference assets against 15.&lt;/strong&gt; It also caps at 720p and costs 60% more per second. Neither model replaces the other.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Feature&lt;/th&gt;
&lt;th&gt;Seedance 2.5&lt;/th&gt;
&lt;th&gt;Seedance 2.0&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Output duration&lt;/td&gt;
&lt;td&gt;4 to 30 s&lt;/td&gt;
&lt;td&gt;4 to 15 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output resolution&lt;/td&gt;
&lt;td&gt;480p, 720p (1080p from 2026-08-17)&lt;/td&gt;
&lt;td&gt;480p, 720p, 1080p, 4K&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reference assets&lt;/td&gt;
&lt;td&gt;50 (30 images + 10 videos + 10 audio)&lt;/td&gt;
&lt;td&gt;15 (9 images + 3 videos + 3 audio)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Audio-only reference&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No, requires image or video&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output format&lt;/td&gt;
&lt;td&gt;mp4, mov&lt;/td&gt;
&lt;td&gt;mp4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Timestamp precision&lt;/td&gt;
&lt;td&gt;Integer seconds&lt;/td&gt;
&lt;td&gt;Shot numbers only&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The reference budget expansion is substantial: 30 images at up to 4K, 10 videos totalling 30 seconds, 10 audio clips also totalling 30 seconds. The recommended working range stays conservative — 1 to 8 image subjects and 1 to 5 audio or video subjects — beyond which stability degrades.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does the Timestamp Control Actually Work?
&lt;/h2&gt;

&lt;p&gt;We ran the same three-beat prompt through both with explicit second ranges:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"0-2s: the paper fox sits still, then its ears twitch twice. 2-4s: the fox unfolds into a single flat sheet of orange paper. 4-6s: the flat sheet folds itself back into the fox, which hops once toward the camera and freezes."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;At t=1s both showed the intact fox. At t=3s Seedance 2.0 had already completed the unfolding while 2.5 was mid-transformation, matching the prompt. At t=5s, 2.5 captured the hop; 2.0 had refolded and skipped the instruction.&lt;/p&gt;

&lt;p&gt;The vendor's line that 2.0 "doesn't respond to timestamps" is directionally accurate rather than literally true. For simple scripts with evenly spaced beats, 2.0 lands approximately right. With 30-second clips containing ten beats, that approximation compounds — which is exactly the case 2.5 addresses.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Does Each One Cost?
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Route&lt;/th&gt;
&lt;th&gt;Seedance 2.5&lt;/th&gt;
&lt;th&gt;Seedance 2.0&lt;/th&gt;
&lt;th&gt;Difference&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;480p, text or image to video&lt;/td&gt;
&lt;td&gt;$0.11/s&lt;/td&gt;
&lt;td&gt;$0.063/s&lt;/td&gt;
&lt;td&gt;+75%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;720p, text or image to video&lt;/td&gt;
&lt;td&gt;$0.24/s&lt;/td&gt;
&lt;td&gt;$0.15/s&lt;/td&gt;
&lt;td&gt;+60%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1080p&lt;/td&gt;
&lt;td&gt;not until 2026-08-17&lt;/td&gt;
&lt;td&gt;$0.31/s&lt;/td&gt;
&lt;td&gt;n/a&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4K&lt;/td&gt;
&lt;td&gt;not available&lt;/td&gt;
&lt;td&gt;$1.24/s&lt;/td&gt;
&lt;td&gt;n/a&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Real billing from 2026-08-13: 6 seconds at 480p cost $0.66 on 2.5 against $0.378 on 2.0. Thirty seconds at 480p cost $3.30 on 2.5 and is not possible on 2.0 at all.&lt;/p&gt;

&lt;p&gt;Two discount periods complicate this. ByteDance runs a promotion 2026-08-07 to 09-07 reducing Mini and Fast tiers to 40% and 75% of list, but it excludes the base 2.0 model. A separate Seedance 2.5 discount 08-14 to 09-17 applies only to 1080p output, leaving 480p and 720p unaffected.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which Is Faster?
&lt;/h2&gt;

&lt;p&gt;In a single test run, the same 6-second 480p prompt submitted within one second of each other took &lt;strong&gt;103 seconds on 2.5 and 289 seconds on 2.0&lt;/strong&gt; — a 2.8x difference. A 30-second clip on 2.5 took 200 seconds, less than the 2.0 six-second clip took that afternoon.&lt;/p&gt;

&lt;p&gt;This is one pair of runs on shared infrastructure, not a published benchmark. It does contradict the assumption that longer-context models must process more slowly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which One Should You Use?
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Requirement&lt;/th&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Reason&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;4K output, or 1080p before 2026-08-17&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2.0&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;2.5 stops at 720p until then&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Single clip longer than 15 seconds&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2.5&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;2.0 rejects duration above 15&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cheap iteration during development&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2.0 Mini or Fast&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;No 2.5 tier below $0.11/s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Precise beat timing inside the clip&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2.5&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Integer-second timestamps&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;More than 15 reference assets, or audio-only reference&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2.5&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;2.0 caps at 15 and requires visual input&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Editing or extending existing footage&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2.5&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;mov output plus documented controls&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Lowest cost per finished second&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2.0&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;40% cheaper at 720p, cheaper on Fast and Mini&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The optimal workflow combines both: iterate on 2.0 Mini at 480p to finalise the script and references, then generate the final version at the settings that matter.&lt;/p&gt;

&lt;h2&gt;
  
  
  Integration Notes
&lt;/h2&gt;

&lt;p&gt;Both models run through the same &lt;code&gt;/v1/videos&lt;/code&gt; endpoint, so switching is a model string change from &lt;code&gt;bytedance/seedance-2.0&lt;/code&gt; to &lt;code&gt;bytedance/seedance-2.5&lt;/code&gt;. Two incompatibilities return immediate HTTP 400 rather than degrading silently:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Duration above 15 seconds: rejected by 2.0&lt;/li&gt;
&lt;li&gt;Resolution above 720p: rejected by 2.5 until 2026-08-17&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Seedance 2.5 also distinguishes locked from unlocked tasks. Locked tasks (editing, extension, first/last frame control) force &lt;code&gt;ratio&lt;/code&gt; to adaptive, and editing requests require specific trigger words in the prompt. &lt;strong&gt;A 2.0-era editing instruction without trigger words gets reclassified as a reference job&lt;/strong&gt; — you get a new video instead of an edited one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Is Seedance 2.5 better than 2.0?&lt;/strong&gt;&lt;br&gt;
For duration, reference quantity and timeline precision, yes. For resolution and cost, no. 2.0 exclusively handles 4K and is currently the only 1080p option.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does Seedance 2.5 support 4K?&lt;/strong&gt;&lt;br&gt;
No. 1080p launches 2026-08-17; 4K is not available on 2.5 at all. Requesting an unsupported resolution returns HTTP 400 rather than downscaling automatically.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do Seedance 2.0 prompts work on 2.5?&lt;/strong&gt;&lt;br&gt;
They function but underutilise capabilities and risk misfiring. Edit and extend operations require trigger words on 2.5, so a 2.0-style edit prompt may generate new content instead of modifying existing footage.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://docs.byteplus.com/en/docs/ModelArk/2607688" rel="noopener noreferrer"&gt;Seedance 2.5 tutorial&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.byteplus.com/en/docs/ModelArk/2607689" rel="noopener noreferrer"&gt;Seedance 2.5 prompt guide&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.byteplus.com/en/docs/ModelArk/1544106" rel="noopener noreferrer"&gt;ModelArk pricing&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ofox.ai/models/bytedance/seedance-2.5" rel="noopener noreferrer"&gt;ofox model page: Seedance 2.5&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ofox.ai/models/bytedance/seedance-2.0" rel="noopener noreferrer"&gt;ofox model page: Seedance 2.0&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://ofox.ai/blog/seedance-2-5-vs-2-0-what-changed-2026/" rel="noopener noreferrer"&gt;ofox.ai/blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>videogeneration</category>
      <category>seedance</category>
      <category>modelcomparison</category>
      <category>apiguide</category>
    </item>
    <item>
      <title>Qwen 3.8 Max in Codex CLI 2026: Config, 258K Fix, Real Cost</title>
      <dc:creator>Owen</dc:creator>
      <pubDate>Fri, 21 Aug 2026 00:01:58 +0000</pubDate>
      <link>https://dev.to/owen_fox/qwen-38-max-in-codex-cli-2026-config-258k-fix-real-cost-2d38</link>
      <guid>https://dev.to/owen_fox/qwen-38-max-in-codex-cli-2026-config-258k-fix-real-cost-2d38</guid>
      <description>&lt;h1&gt;
  
  
  Qwen 3.8 Max in Codex CLI 2026: Config, 258K Fix, Real Cost
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;Yes, Codex CLI runs Qwen 3.8 Max — through a gateway using the Responses API protocol.&lt;/strong&gt; Codex has no built-in support for the model, so it takes a custom provider declaration: six lines of TOML.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Model slug:      bailian/qwen3.8-max
Wire protocol:   responses
Default context: 258,400 tokens (below the model's 1M)
Measured cost:   $0.08 for 3 tasks vs $0.54 on GPT-5.5 (2026-08-06)
Pricing:         $2/$6 per 1M in/out, cache read $0.25/M
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What Works and What Does Not
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Works:&lt;/strong&gt; the complete agent loop — reading files, applying patches via Codex's &lt;code&gt;apply_patch&lt;/code&gt;, running Python verification, reporting results. Unlike Claude Sonnet 5 and DeepSeek, it handles Codex's freeform tool shape without rejection errors.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does not:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Capped at 258,400 tokens instead of the full 1M context&lt;/li&gt;
&lt;li&gt;Unlocking the full window replaces Codex's compiled system prompt with your own&lt;/li&gt;
&lt;li&gt;The &lt;code&gt;tokens used&lt;/code&gt; display under-reports cached tokens (up to 85% of input in testing)&lt;/li&gt;
&lt;li&gt;Separate API-key billing path from your ChatGPT subscription&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How Do You Configure It?
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Step 1: Install and set your key
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm i &lt;span class="nt"&gt;-g&lt;/span&gt; @openai/codex
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;OFOX_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"sk-of-..."&lt;/span&gt;
codex &lt;span class="nt"&gt;--version&lt;/span&gt;   &lt;span class="c"&gt;# expect 0.146.x&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Do not use &lt;code&gt;codex login --with-api-key&lt;/code&gt; — it writes OpenAI credentials unsuitable for custom providers.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Write the provider block
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;~/.codex/config.toml&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight toml"&gt;&lt;code&gt;&lt;span class="py"&gt;model&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"bailian/qwen3.8-max"&lt;/span&gt;
&lt;span class="py"&gt;model_provider&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"ofox"&lt;/span&gt;

&lt;span class="nn"&gt;[model_providers.ofox]&lt;/span&gt;
&lt;span class="py"&gt;name&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"ofox"&lt;/span&gt;
&lt;span class="py"&gt;base_url&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"https://api.ofox.ai/v1"&lt;/span&gt;
&lt;span class="py"&gt;env_key&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"OFOX_API_KEY"&lt;/span&gt;
&lt;span class="py"&gt;wire_api&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s"&gt;"responses"&lt;/span&gt;
&lt;span class="py"&gt;requires_openai_auth&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three fields matter: &lt;code&gt;wire_api = "responses"&lt;/code&gt; is now mandatory (&lt;code&gt;"chat"&lt;/code&gt; triggers a hard startup error), &lt;code&gt;env_key&lt;/code&gt; names the environment variable Codex reads, and &lt;code&gt;requires_openai_auth = false&lt;/code&gt; bypasses ChatGPT login.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: Run it
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;codex &lt;span class="nb"&gt;exec&lt;/span&gt; &lt;span class="nt"&gt;--sandbox&lt;/span&gt; workspace-write &lt;span class="s2"&gt;"median() is wrong for even-length input. Fix it in stats.py."&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Why Does Codex Cap My Context at 258,400 Tokens?
&lt;/h2&gt;

&lt;p&gt;Codex ships a compiled-in catalog covering only OpenAI models. Unknown models get fallback defaults:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="s1"&gt;'"model_context_window":[0-9]*'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  ~/.codex/sessions/2026/08/06/rollout-&lt;span class="k"&gt;*&lt;/span&gt;.jsonl | &lt;span class="nb"&gt;tail&lt;/span&gt; &lt;span class="nt"&gt;-1&lt;/span&gt;
&lt;span class="c"&gt;# "model_context_window":258400&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three approaches tested on 0.146.1, only one works:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Attempt&lt;/th&gt;
&lt;th&gt;Warning gone?&lt;/th&gt;
&lt;th&gt;Reported window&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Default (no override)&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;258,400&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;model_context_window&lt;/code&gt; in config.toml&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;258,400&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;-c model_context_window=...&lt;/code&gt; CLI flag&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;258,400&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;model_catalog_json&lt;/code&gt; with custom entry&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;Yes&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;1,131,072&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The catalog JSON requires all 39 ModelInfo fields. One caveat worth weighing: &lt;code&gt;base_instructions&lt;/code&gt; in that entry &lt;strong&gt;replaces&lt;/strong&gt; Codex's system prompt. For sessions comfortably under 258K, skipping this step is defensible.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Errors Will You Hit?
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Error&lt;/th&gt;
&lt;th&gt;Real cause&lt;/th&gt;
&lt;th&gt;Fix&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Missing bearer or basic authentication in header&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;No key reached the gateway&lt;/td&gt;
&lt;td&gt;Set the variable named in &lt;code&gt;env_key&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;401 saying the key is invalid (works elsewhere)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;env_key&lt;/code&gt; missing, wrong key sent&lt;/td&gt;
&lt;td&gt;Add &lt;code&gt;env_key&lt;/code&gt; to the provider block&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;You didn't provide an API key&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Trailing newline in the key&lt;/td&gt;
&lt;td&gt;Strip it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;wire_api = "chat" is no longer supported&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Deprecated before 0.146&lt;/td&gt;
&lt;td&gt;Set &lt;code&gt;wire_api = "responses"&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;404 on every request&lt;/td&gt;
&lt;td&gt;Missing &lt;code&gt;/v1&lt;/code&gt; on base_url&lt;/td&gt;
&lt;td&gt;Use &lt;code&gt;https://api.ofox.ai/v1&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Model metadata not found&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Slug outside the built-in catalog&lt;/td&gt;
&lt;td&gt;See the 258K section&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;missing field 'display_name'&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Incomplete ModelInfo entry&lt;/td&gt;
&lt;td&gt;All 39 fields required&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  What Does It Actually Cost?
&lt;/h2&gt;

&lt;p&gt;Three real coding tasks, same CLI, same gateway, same repo, 2026-08-06:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;Qwen 3.8 Max&lt;/th&gt;
&lt;th&gt;GPT-5.5&lt;/th&gt;
&lt;th&gt;Gap&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Fix &lt;code&gt;median()&lt;/code&gt; bug&lt;/td&gt;
&lt;td&gt;$0.0244&lt;/td&gt;
&lt;td&gt;$0.1223&lt;/td&gt;
&lt;td&gt;5.0x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Add type hints and unittest coverage&lt;/td&gt;
&lt;td&gt;$0.0373&lt;/td&gt;
&lt;td&gt;$0.2017&lt;/td&gt;
&lt;td&gt;5.4x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Explain repo and flag risks&lt;/td&gt;
&lt;td&gt;$0.0187&lt;/td&gt;
&lt;td&gt;$0.2146&lt;/td&gt;
&lt;td&gt;11.5x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.0804&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.5387&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;6.7x&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Is the 6.7x real?&lt;/strong&gt; Partially. Rate-card differences justify 2.5x on input and 5x on output. The rest comes from configuration and task variance:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;System prompt difference.&lt;/strong&gt; GPT-5.5 gets OpenAI's full instruction set; Qwen under a custom catalog gets your one-line &lt;code&gt;base_instructions&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Turn count.&lt;/strong&gt; Models choose their own round-trip counts. On "explain repo," GPT-5.5 took 8 API calls with 7 tool invocations; Qwen took 4 calls with 3 — that alone explains the 11.5x on that task.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Treat 6.7x as measured for this configuration on this date. 2.5x/5x is the rate-card floor.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why is the token count in Codex lower than my bill?
&lt;/h3&gt;

&lt;p&gt;The &lt;code&gt;tokens used&lt;/code&gt; line subtracts cached input from the total. One task displayed 5,859 while the session log showed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"input_tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;37401&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"cached_input_tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;32512&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
 &lt;/span&gt;&lt;span class="nl"&gt;"output_tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;970&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"total_tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;38371&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;38,371 − 32,512 = 5,859. &lt;strong&gt;Cached tokens are billed at cache-read rates, not free.&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Does reasoning effort change the bill much?
&lt;/h3&gt;

&lt;p&gt;Surprisingly little. Same task, same model, only &lt;code&gt;model_reasoning_effort&lt;/code&gt; changed:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Effort&lt;/th&gt;
&lt;th&gt;Reasoning tokens&lt;/th&gt;
&lt;th&gt;Total tokens&lt;/th&gt;
&lt;th&gt;Cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;low&lt;/td&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;td&gt;26,094&lt;/td&gt;
&lt;td&gt;$0.0244&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;high&lt;/td&gt;
&lt;td&gt;493&lt;/td&gt;
&lt;td&gt;26,884&lt;/td&gt;
&lt;td&gt;$0.0252&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Reasoning output tripled; the bill rose 3.3%. In agent loops, conversation replay dominates and reasoning is a thin overlay. This pattern holds for short agentic tasks — a single long reasoning-heavy request with minimal tool use reverses it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which Models Can Replace Qwen Here?
&lt;/h2&gt;

&lt;p&gt;Not everything behind an OpenAI-compatible gateway survives Codex's Responses requirement. Six models tested on 2026-08-06:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Result in Codex&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;bailian/qwen3.8-max&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Works, full agent loop&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;openai/gpt-5.5&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Works&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;anthropic/claude-sonnet-5&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Fails on &lt;code&gt;apply_patch&lt;/code&gt; tool shape&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;deepseek/deepseek-v4-pro&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;Encrypted content is not supported&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;x-ai/grok-4.3&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Same encrypted-content failure&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;z-ai/glm-5.2&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;503, no Responses support upstream&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Encrypted-content failures are unfixable — Codex hard-codes that field with no disable option.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Does this affect my ChatGPT or Codex plan limits?&lt;/strong&gt;&lt;br&gt;
No. Custom providers use a separate API-key billing path, leaving weekly Codex caps untouched.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Will the 258K cap be fixed upstream?&lt;/strong&gt;&lt;br&gt;
Codex resolves metadata from a compiled catalog, so third-party slugs keep falling back until that design changes. The catalog JSON is the supported workaround today.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is Qwen 3.8 Max worth switching to just for the price?&lt;/strong&gt;&lt;br&gt;
Rate-card advantages (2.5x input, 5x output) justify it for routine agent work. For tasks where errors cost debugging time, evaluate quality first.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://ofox.ai/blog/qwen-3-8-max-codex-cli-config-cost-2026/" rel="noopener noreferrer"&gt;ofox.ai/blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>codexcli</category>
      <category>qwen</category>
      <category>apiguide</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Grok 4.6 API Pricing: The 200K Cliff and When 4.5 Wins</title>
      <dc:creator>Owen</dc:creator>
      <pubDate>Thu, 20 Aug 2026 23:31:39 +0000</pubDate>
      <link>https://dev.to/owen_fox/grok-46-api-pricing-the-200k-cliff-and-when-45-wins-290f</link>
      <guid>https://dev.to/owen_fox/grok-46-api-pricing-the-200k-cliff-and-when-45-wins-290f</guid>
      <description>&lt;h1&gt;
  
  
  Grok 4.6 API Pricing: The 200K Cliff and When 4.5 Wins
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;Grok 4.6 costs $2.00 per million input tokens and $6.00 output, which is exactly what Grok 4.5 costs.&lt;/strong&gt; The two model cards are otherwise identical down to the rate limits. One number differs, cached input, and it is the newer model that is more expensive there.&lt;/p&gt;

&lt;p&gt;The bigger number is the one that is not a difference between models at all: at 200,000 prompt tokens, every rate doubles, and it doubles for the whole request.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Model IDs:     grok-4.6 / grok-4.5  (x-ai/grok-4.6 on gateways)
Context:       500K tokens, both
Under 200K:    $2.00 in / $6.00 out    both models
                $0.50 cached on 4.6, $0.30 cached on 4.5
200K and over: $4.00 in / $12.00 out   both models
                $1.00 cached on 4.6, $0.60 cached on 4.5
Threshold:     applies to ALL tokens in the request, not the excess
Measured:      fixed per-request overhead 206 tokens on 4.6, 494 on 4.5
Break-even:    ~575 cached tokens; below it 4.6 is cheaper per call
Snapshot:      2026-08-19
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  How Much Does the Grok 4.6 API Cost?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;$2.00 in, $0.50 cached in, $6.00 out per million tokens, until the prompt reaches 200K.&lt;/strong&gt; Straight from &lt;a href="https://docs.x.ai/developers/models" rel="noopener noreferrer"&gt;xAI's model list&lt;/a&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Context&lt;/th&gt;
&lt;th&gt;Input&lt;/th&gt;
&lt;th&gt;Cached input&lt;/th&gt;
&lt;th&gt;Output&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;grok-4.6, prompt under 200K&lt;/td&gt;
&lt;td&gt;500K&lt;/td&gt;
&lt;td&gt;$2.00&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;$6.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;grok-4.6, prompt 200K and over&lt;/td&gt;
&lt;td&gt;500K&lt;/td&gt;
&lt;td&gt;$4.00&lt;/td&gt;
&lt;td&gt;$1.00&lt;/td&gt;
&lt;td&gt;$12.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;grok-4.5, prompt under 200K&lt;/td&gt;
&lt;td&gt;500K&lt;/td&gt;
&lt;td&gt;$2.00&lt;/td&gt;
&lt;td&gt;$0.30&lt;/td&gt;
&lt;td&gt;$6.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;grok-4.5, prompt 200K and over&lt;/td&gt;
&lt;td&gt;500K&lt;/td&gt;
&lt;td&gt;$4.00&lt;/td&gt;
&lt;td&gt;$0.60&lt;/td&gt;
&lt;td&gt;$12.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;grok-4.3&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;$1.25&lt;/td&gt;
&lt;td&gt;$0.20&lt;/td&gt;
&lt;td&gt;$2.50&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Grok 4.3 is in that table for a reason. It is 38% cheaper on input and 58% cheaper on output than either 4.x flagship, and it carries a 1M context instead of 500K.&lt;/p&gt;

&lt;p&gt;Gateways pass the headline rate through unchanged. &lt;a href="https://openrouter.ai/x-ai/grok-4.6" rel="noopener noreferrer"&gt;OpenRouter&lt;/a&gt; and ofox both list &lt;code&gt;x-ai/grok-4.6&lt;/code&gt; at $2.00 and $6.00 with a 500,000-token context. What neither catalog exposes is the second row, so the doubling above 200K is invisible until it lands on your invoice.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Happens When Your Prompt Crosses 200K Tokens?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The whole request reprices, not the overflow.&lt;/strong&gt; xAI's wording is unambiguous: requests whose prompt reaches the threshold are billed at the higher rate for all tokens in the request.&lt;/p&gt;

&lt;p&gt;Two requests, 2,000 tokens apart, both with 2,000 tokens of output:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Prompt size&lt;/th&gt;
&lt;th&gt;Input cost&lt;/th&gt;
&lt;th&gt;Output cost&lt;/th&gt;
&lt;th&gt;Total&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;199,000 tokens&lt;/td&gt;
&lt;td&gt;$0.398&lt;/td&gt;
&lt;td&gt;$0.012&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.410&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;201,000 tokens&lt;/td&gt;
&lt;td&gt;$0.804&lt;/td&gt;
&lt;td&gt;$0.024&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.828&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A 1% larger prompt for a 102% larger bill. There is no gradual slope here, and no partial credit for the tokens below the line.&lt;/p&gt;

&lt;p&gt;Two things make that line easier to cross than it looks. First, the threshold counts the prompt, so a long conversation walks toward it one turn at a time. Second, your prompt is not the only thing in the prompt, which is the subject of the next section.&lt;/p&gt;

&lt;p&gt;If you are batching, split before the line rather than after it. Two 150K requests at the low tier cost $0.60 in input; one 300K request costs $1.20 for the same tokens.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is the Difference Between Grok 4.6 and Grok 4.5?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;One number on the published cards.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Field&lt;/th&gt;
&lt;th&gt;Grok 4.6&lt;/th&gt;
&lt;th&gt;Grok 4.5&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Context window&lt;/td&gt;
&lt;td&gt;500,000&lt;/td&gt;
&lt;td&gt;500,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Modalities&lt;/td&gt;
&lt;td&gt;text, image to text&lt;/td&gt;
&lt;td&gt;text, image to text&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Function calling&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Structured outputs&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reasoning&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Requests per second&lt;/td&gt;
&lt;td&gt;150&lt;/td&gt;
&lt;td&gt;150&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tokens per minute&lt;/td&gt;
&lt;td&gt;50,000,000&lt;/td&gt;
&lt;td&gt;50,000,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Regions&lt;/td&gt;
&lt;td&gt;us-east-1, us-west-2&lt;/td&gt;
&lt;td&gt;us-east-1, us-west-2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Input / output&lt;/td&gt;
&lt;td&gt;$2.00 / $6.00&lt;/td&gt;
&lt;td&gt;$2.00 / $6.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Cached input&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.50&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.30&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Aliases&lt;/td&gt;
&lt;td&gt;none listed&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;grok-4.5-latest&lt;/code&gt;, &lt;code&gt;grok-build-latest&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Cached input is 67% more expensive on the newer model. That is the opposite of the usual direction and it is the entire published basis for keeping 4.5 around.&lt;/p&gt;

&lt;p&gt;The alias row matters if you use Grok Build: &lt;code&gt;grok-build-latest&lt;/code&gt; resolves to Grok 4.5, not 4.6.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does the Same Prompt Cost More on Grok 4.5?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Because a fixed block of input tokens rides along with every request, and it is more than twice as large on 4.5.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We sent identical bodies of 1, 50 and 200 repeated words to four models through an OpenAI-compatible gateway on 2026-08-19 and fit the line. Every model tokenized the body at exactly 1.00 token per word. The intercepts did not match:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Fixed input overhead&lt;/th&gt;
&lt;th&gt;Of which reported cached&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;x-ai/grok-4.6&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;206 tokens&lt;/td&gt;
&lt;td&gt;128&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;x-ai/grok-4.5&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;494 tokens&lt;/td&gt;
&lt;td&gt;384&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;x-ai/grok-4.3&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;4 tokens&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;z-ai/glm-5.3&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;12 tokens&lt;/td&gt;
&lt;td&gt;none&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Four and twelve tokens are an ordinary chat template. Two hundred and four hundred and ninety-four are a preamble. Adding your own &lt;code&gt;system&lt;/code&gt; message raised both figures by exactly the size of that message, so this sits underneath anything you send.&lt;/p&gt;

&lt;p&gt;We could not test a second route to attribute it, so treat the source as unproven. Either way it is on your invoice, so price it.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Grok 4.6&lt;/th&gt;
&lt;th&gt;Grok 4.5&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Cached portion&lt;/td&gt;
&lt;td&gt;128 × $0.50/M = $0.000064&lt;/td&gt;
&lt;td&gt;384 × $0.30/M = $0.000115&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Uncached portion&lt;/td&gt;
&lt;td&gt;78 × $2.00/M = $0.000156&lt;/td&gt;
&lt;td&gt;110 × $2.00/M = $0.000220&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Per request&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.000220&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.000335&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Per 1M requests&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$220&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$335&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;It also eats your headroom. On 4.5 you reach the 200K cliff 494 tokens earlier than your own token count suggests. If you are budgeting a prompt at exactly 199,800 tokens, you are already over.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which One Is Cheaper for Your Workload?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Around 575 cached tokens, the answer flips.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Grok 4.5 saves $0.20 per million cached tokens, which is $0.0000002 per cached token. It loses $0.000115 per request on the larger preamble. Divide one by the other and you get 575 tokens.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cached prefix under ~575 tokens:&lt;/strong&gt; Grok 4.6 costs less per call&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cached prefix over ~575 tokens:&lt;/strong&gt; Grok 4.5 costs less, and the gap widens linearly&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That covers input. Output is where the two models genuinely diverge. We ran the same code-generation prompt through both, 8 runs each:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Grok 4.6&lt;/th&gt;
&lt;th&gt;Grok 4.5&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;completion_tokens&lt;/code&gt;, median&lt;/td&gt;
&lt;td&gt;216&lt;/td&gt;
&lt;td&gt;225&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;reasoning_tokens&lt;/code&gt;, median&lt;/td&gt;
&lt;td&gt;723&lt;/td&gt;
&lt;td&gt;30&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Billed output, median&lt;/td&gt;
&lt;td&gt;948&lt;/td&gt;
&lt;td&gt;263&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Latency, median&lt;/td&gt;
&lt;td&gt;15.8 s&lt;/td&gt;
&lt;td&gt;5.0 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost per 1,000 tasks, all-in&lt;/td&gt;
&lt;td&gt;$6.18&lt;/td&gt;
&lt;td&gt;$2.64&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That last row is the only one in the table that is not an output-side number, so here is its basis. It is output at $6.00 per million &lt;strong&gt;plus&lt;/strong&gt; the full input at $2.00 per million, and the input side includes the fixed per-request overhead: 244 input tokens on 4.6, 532 on 4.5. Output alone would be $5.69 and $1.58. We did not record cache hits on these runs, so input is priced entirely at the uncached rate, making the all-in figure an upper bound.&lt;/p&gt;

&lt;p&gt;Watch the field names. &lt;code&gt;completion_tokens&lt;/code&gt; &lt;strong&gt;excludes&lt;/strong&gt; reasoning tokens on these models: a response reporting 216 completion tokens had 723 reasoning tokens alongside it, and &lt;code&gt;total_tokens&lt;/code&gt; was the sum of prompt, completion and reasoning. Any cost estimator built on &lt;code&gt;prompt_tokens × input + completion_tokens × output&lt;/code&gt; undercounts the billed output by 77%.&lt;/p&gt;

&lt;p&gt;So the short version: 4.5 for long cached prefixes and short deterministic work, 4.6 when the extra reasoning is the point.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Do I Call Grok 4.6?
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;YOUR_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.ofox.ai/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;x-ai/grok-4.6&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;          &lt;span class="c1"&gt;# x-ai/grok-4.5 for the cheaper cache tier
&lt;/span&gt;    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Refactor this module and run the tests.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;u&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;usage&lt;/span&gt;
&lt;span class="n"&gt;billed_output&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;u&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completion_tokens&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;u&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completion_tokens_details&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;reasoning_tokens&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;u&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;prompt_tokens&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;billed_output&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three things that bite people on the way in:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The ID is namespaced on gateways.&lt;/strong&gt; &lt;code&gt;grok-4.6&lt;/code&gt; works against xAI directly; on OpenRouter and ofox it is &lt;code&gt;x-ai/grok-4.6&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;logprobs&lt;/code&gt; fails silently.&lt;/strong&gt; xAI documents it as unsupported on grok-4.20 and newer. We sent &lt;code&gt;logprobs: true&lt;/code&gt; with &lt;code&gt;top_logprobs: 3&lt;/code&gt; to both models: HTTP 200, no error, and no &lt;code&gt;logprobs&lt;/code&gt; object in the response.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Both models are in us-east-1 and us-west-2 only.&lt;/strong&gt; If you have data-residency constraints outside the US, this pair is not the answer regardless of price.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For rate limits across providers, our &lt;a href="https://dev.to/blog/llm-api-rate-limits-compared-2026/"&gt;LLM API rate limits comparison&lt;/a&gt; has the numbers side by side.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://docs.x.ai/developers/models" rel="noopener noreferrer"&gt;xAI developer docs: models and pricing&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://openrouter.ai/x-ai/grok-4.6" rel="noopener noreferrer"&gt;OpenRouter: x-ai/grok-4.6&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ofox.ai/models/x-ai/grok-4.6" rel="noopener noreferrer"&gt;ofox model page: Grok 4.6&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ofox.ai/models/x-ai/grok-4.5" rel="noopener noreferrer"&gt;ofox model page: Grok 4.5&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://ofox.ai/blog/grok-4-6-api-pricing-vs-grok-4-5-2026/" rel="noopener noreferrer"&gt;ofox.ai/blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>grok</category>
      <category>pricing</category>
      <category>apiguide</category>
      <category>costoptimization</category>
    </item>
    <item>
      <title>GLM 5.3 API: Pricing, Endpoints, and reasoning_effort</title>
      <dc:creator>Owen</dc:creator>
      <pubDate>Thu, 20 Aug 2026 20:28:40 +0000</pubDate>
      <link>https://dev.to/owen_fox/glm-53-api-pricing-endpoints-and-reasoningeffort-40oa</link>
      <guid>https://dev.to/owen_fox/glm-53-api-pricing-endpoints-and-reasoningeffort-40oa</guid>
      <description>&lt;h1&gt;
  
  
  GLM 5.3 API: Pricing, Endpoints, and reasoning_effort
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;The GLM 5.3 API opened five days after the model did, at $1.40 per million input tokens and $4.40 per million output, the same rate as GLM 5.2.&lt;/strong&gt; The price is not what will surprise you. &lt;code&gt;reasoning_effort&lt;/code&gt; defaults to &lt;code&gt;max&lt;/code&gt;, and on a short classification call we measured max spending a median of 105 output tokens where &lt;code&gt;low&lt;/code&gt; spent 3.&lt;/p&gt;

&lt;p&gt;That is a 35x difference on the metered half of your bill, set by a parameter most people will not send.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Price:        $1.40 in / $0.26 cached in / $4.40 out per 1M tokens
Context:      1M tokens, 128K max output
Base URLs:    api.z.ai/api/coding/paas/v4  (OpenAI Chat Completions)
              api.z.ai/api/v1              (OpenAI Responses)
              api.z.ai/api/anthropic       (Anthropic Messages)
Gateways:     z-ai/glm-5.3 on OpenRouter and ofox
Effort:       low | high | max, default max, cannot be disabled
Removed:      thinking.type "disabled" now returns HTTP 400
Measured:     classify task 3 / 8 / 105 output tokens at low / high / max
Snapshot:     2026-08-19
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  How Much Does the GLM 5.3 API Cost?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;$1.40 per million input tokens, $0.26 cached input, $4.40 output.&lt;/strong&gt; Z.ai's &lt;a href="https://docs.z.ai/guides/overview/pricing" rel="noopener noreferrer"&gt;pricing table&lt;/a&gt; now carries a GLM-5.3 row, and it matches GLM 5.2 and GLM 5.1 line for line.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;Rate&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Input&lt;/td&gt;
&lt;td&gt;$1.40 / 1M tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cached input&lt;/td&gt;
&lt;td&gt;$0.26 / 1M tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cached input storage&lt;/td&gt;
&lt;td&gt;Free, marked limited-time&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output&lt;/td&gt;
&lt;td&gt;$4.40 / 1M tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Cached input at $0.26 is 19% of a cold read, and the storage that usually makes caching a judgement call is free during the promotion. Output costs 3.1x input, which is the ratio that makes the effort setting below the most expensive line in your config.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://openrouter.ai/z-ai/glm-5.3" rel="noopener noreferrer"&gt;OpenRouter&lt;/a&gt; lists the same $1.4 / $4.4 pair at a 1,048,576-token context, so third-party routes are passing the first-party rate through rather than marking it up.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is the Base URL for the GLM 5.3 API?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Three protocols, and the docs page contradicts itself about one of them.&lt;/strong&gt; The &lt;a href="https://docs.z.ai/guides/llm/glm-5.3" rel="noopener noreferrer"&gt;model page&lt;/a&gt; lists these:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Protocol&lt;/th&gt;
&lt;th&gt;Base URL&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI Chat Completions&lt;/td&gt;
&lt;td&gt;&lt;code&gt;https://api.z.ai/api/coding/paas/v4&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI Responses&lt;/td&gt;
&lt;td&gt;&lt;code&gt;https://api.z.ai/api/v1&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Anthropic Messages&lt;/td&gt;
&lt;td&gt;&lt;code&gt;https://api.z.ai/api/anthropic&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Then the Quick Start section further down the same page posts to &lt;code&gt;https://api.z.ai/api/paas/v4/chat/completions&lt;/code&gt;, without the &lt;code&gt;/coding&lt;/code&gt; segment. If the first one 404s, try the second before you go looking for a problem in your key.&lt;/p&gt;

&lt;p&gt;Neither of them is the &lt;code&gt;https://open.bigmodel.cn/api/paas/v4&lt;/code&gt; that Zhipu previewed on launch day.&lt;/p&gt;

&lt;p&gt;One restriction is easy to miss: &lt;strong&gt;accounts that have ever subscribed to a GLM Coding Plan, including expired subscriptions, can currently reach the model API only through the OpenAI Chat Completions protocol.&lt;/strong&gt; If your Responses or Anthropic-protocol calls fail on an account that used to run a plan, that is why.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Does reasoning_effort Do to Your Bill?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;More than the model choice does.&lt;/strong&gt; GLM 5.3 always reasons, &lt;code&gt;reasoning_effort&lt;/code&gt; takes &lt;code&gt;low&lt;/code&gt;, &lt;code&gt;high&lt;/code&gt; or &lt;code&gt;max&lt;/code&gt;, and the default is &lt;code&gt;max&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;We ran two workloads through &lt;code&gt;z-ai/glm-5.3&lt;/code&gt; on an OpenAI-compatible gateway on 2026-08-19. A short classification prompt at n=10 per level, and a small code-generation prompt at n=5 per level.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Workload&lt;/th&gt;
&lt;th&gt;Effort&lt;/th&gt;
&lt;th&gt;Output tokens, median&lt;/th&gt;
&lt;th&gt;Range&lt;/th&gt;
&lt;th&gt;Latency, median&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Classify a support ticket (51 in)&lt;/td&gt;
&lt;td&gt;low&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;3 to 8&lt;/td&gt;
&lt;td&gt;1.6 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;high&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;8 to 8&lt;/td&gt;
&lt;td&gt;1.9 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;max&lt;/td&gt;
&lt;td&gt;105&lt;/td&gt;
&lt;td&gt;47 to 160&lt;/td&gt;
&lt;td&gt;3.4 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Write a merge-intervals function (50 in)&lt;/td&gt;
&lt;td&gt;low&lt;/td&gt;
&lt;td&gt;519&lt;/td&gt;
&lt;td&gt;418 to 586&lt;/td&gt;
&lt;td&gt;11.6 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;high&lt;/td&gt;
&lt;td&gt;658&lt;/td&gt;
&lt;td&gt;592 to 825&lt;/td&gt;
&lt;td&gt;8.6 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;max&lt;/td&gt;
&lt;td&gt;3,700&lt;/td&gt;
&lt;td&gt;2,807 to 10,596&lt;/td&gt;
&lt;td&gt;64.0 s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The classification row is the one to look at twice. Three output tokens at &lt;code&gt;low&lt;/code&gt;, 105 at &lt;code&gt;max&lt;/code&gt;, for an answer that is a single word either way. We then re-ran the classification 18 more times, six per effort level: every single run returned &lt;code&gt;billing&lt;/code&gt;, the correct label, at all three settings.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Workload&lt;/th&gt;
&lt;th&gt;low&lt;/th&gt;
&lt;th&gt;high&lt;/th&gt;
&lt;th&gt;max&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1M classification calls&lt;/td&gt;
&lt;td&gt;$84.60&lt;/td&gt;
&lt;td&gt;$106.60&lt;/td&gt;
&lt;td&gt;$533.40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1,000 code-generation tasks&lt;/td&gt;
&lt;td&gt;$2.35&lt;/td&gt;
&lt;td&gt;$2.97&lt;/td&gt;
&lt;td&gt;$16.35&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Both rows are all-in: input at $1.40/M plus output at $4.40/M, priced at uncached rates. Output alone on the classification row would be $13.20, $35.20 and $462.00, so the fixed $71.40 of input compresses the token-level 35x ratio to 6.3x on the final bill.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;high&lt;/code&gt; is the setting that gets skipped and probably should not be. It cost 26% more than &lt;code&gt;low&lt;/code&gt; on both jobs, and on the code job it was &lt;em&gt;faster&lt;/em&gt; than &lt;code&gt;low&lt;/code&gt; at the median, 8.6 seconds against 11.6. Latency does not climb monotonically with effort.&lt;/p&gt;

&lt;p&gt;A caveat: these are two prompts, not a benchmark suite, and the &lt;code&gt;max&lt;/code&gt; ranges are wide. Run your own prompt before you size a budget on it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does My GLM 5.3 Request Return 400?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Most likely because reasoning cannot be switched off, and the message is only half accurate.&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"error"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"message"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"This model always engages in thinking and cannot be disabled; please use low, high, or max"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the response to &lt;code&gt;"thinking": {"type": "disabled"}&lt;/code&gt;, which is correct and expected. On 2026-08-19 it was also the response to &lt;code&gt;"reasoning_effort": "medium"&lt;/code&gt; and &lt;code&gt;"none"&lt;/code&gt;, which was not, because neither tried to disable anything.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;That validation behaviour then changed.&lt;/strong&gt; Re-testing on 2026-08-20, the same route accepted &lt;code&gt;medium&lt;/code&gt;, &lt;code&gt;none&lt;/code&gt;, &lt;code&gt;minimal&lt;/code&gt; and &lt;code&gt;xhigh&lt;/code&gt; with HTTP 200, each burning eight to thirteen times the output tokens of an explicit &lt;code&gt;low&lt;/code&gt; on the same prompt. A different route to the same model name still returned the 400 that day. A genuinely malformed value like &lt;code&gt;"invalid_value"&lt;/code&gt; was rejected on both routes, so validation exists — it just lives on the route rather than in the model, and it changed under a fixed model name with no announcement.&lt;/p&gt;

&lt;p&gt;Two consequences. A 400 here does not always mean what the message says. And a 200 does not confirm the value picked the tier its name suggests: on the accepting route every undocumented value landed near &lt;code&gt;max&lt;/code&gt;, so a harness sending &lt;code&gt;medium&lt;/code&gt; believing it picked middle pricing pays top-tier rates silently.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Request&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;thinking.type: "disabled"&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;400, thinking cannot be disabled&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;reasoning_effort: "medium"&lt;/code&gt; / &lt;code&gt;"none"&lt;/code&gt; / &lt;code&gt;"minimal"&lt;/code&gt; / &lt;code&gt;"xhigh"&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Route-dependent. 400 on 2026-08-19; 200 billed near &lt;code&gt;max&lt;/code&gt; on 2026-08-20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;reasoning_effort: "low"&lt;/code&gt; / &lt;code&gt;"high"&lt;/code&gt; / &lt;code&gt;"max"&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;thinking.type: "enabled"&lt;/code&gt; plus &lt;code&gt;reasoning_effort: "low"&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No reasoning field at all&lt;/td&gt;
&lt;td&gt;200, billed as &lt;code&gt;max&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model ID &lt;code&gt;zai/glm-5.3&lt;/code&gt; on a gateway&lt;/td&gt;
&lt;td&gt;404 &lt;code&gt;model_not_found&lt;/code&gt;, the prefix is &lt;code&gt;z-ai&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  How Do I Migrate a GLM 5.2 Workload to GLM 5.3?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Change the effort parameter first, then the model ID.&lt;/strong&gt; A request carrying &lt;code&gt;thinking.type: "disabled"&lt;/code&gt; fails the moment the model ID flips.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;YOUR_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.ofox.ai/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;z-ai/glm-5.3&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Classify this ticket: ...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
    &lt;span class="n"&gt;reasoning_effort&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;low&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;   &lt;span class="c1"&gt;# omit this and you are billed at max
&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completion_tokens&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Set it explicitly everywhere, including the places that inherit defaults: retry wrappers, evaluation harnesses, and any framework that builds the request body for you. A missing &lt;code&gt;reasoning_effort&lt;/code&gt; is not a missing feature, it is a bill at &lt;code&gt;max&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;If you are wiring up a key from scratch, our &lt;a href="https://dev.to/blog/glm-5-2-access-guide-2026/"&gt;GLM 5.2 API access guide&lt;/a&gt; applies unchanged. For the workload math on high-volume short calls, the &lt;a href="https://dev.to/blog/glm-5-2-vs-gpt-5-5-cost-2026/"&gt;GLM 5.2 versus GPT-5.5 cost comparison&lt;/a&gt; has the model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Should I Call GLM 5.3 Directly or Through a Gateway?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Direct if you only run GLM. Through a gateway if you run anything else alongside it, or if you got caught by the Coding Plan protocol restriction.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Z.ai direct&lt;/th&gt;
&lt;th&gt;Gateway&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Price&lt;/td&gt;
&lt;td&gt;$1.4 / $4.4&lt;/td&gt;
&lt;td&gt;Same, passed through&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Protocols&lt;/td&gt;
&lt;td&gt;Three, minus the Coding Plan restriction&lt;/td&gt;
&lt;td&gt;Whatever the gateway speaks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model ID&lt;/td&gt;
&lt;td&gt;&lt;code&gt;glm-5.3&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;z-ai/glm-5.3&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Failover to another model&lt;/td&gt;
&lt;td&gt;Your code&lt;/td&gt;
&lt;td&gt;One string&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cache pricing&lt;/td&gt;
&lt;td&gt;$0.26, storage free for now&lt;/td&gt;
&lt;td&gt;Depends on passthrough&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;One caveat on gateways: what a proxy reports back is not always what it was billed. On the gateway we tested, &lt;code&gt;usage.completion_tokens_details.reasoning_tokens&lt;/code&gt; does come through. What did not come through was any &lt;code&gt;prompt_tokens_details&lt;/code&gt; cache field, and the reasoning text itself is absent from the message object. Verify both against your own provider before you build cost accounting on them.&lt;/p&gt;

&lt;p&gt;For benchmarks, the weights timeline and the GLM 5.3 versus 5.2 capability picture, our &lt;a href="https://dev.to/blog/glm-5-3-benchmarks-access-2026/"&gt;GLM 5.3 launch coverage&lt;/a&gt; has the full table. The &lt;a href="https://ofox.ai/models/z-ai/glm-5.3" rel="noopener noreferrer"&gt;ofox model page for GLM 5.3&lt;/a&gt; carries the live catalog spec.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://docs.z.ai/guides/llm/glm-5.3" rel="noopener noreferrer"&gt;Z.ai developer docs: GLM-5.3 model page&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.z.ai/guides/overview/pricing" rel="noopener noreferrer"&gt;Z.ai developer docs: pricing&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.z.ai/guides/capabilities/thinking" rel="noopener noreferrer"&gt;Z.ai developer docs: deep thinking&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://openrouter.ai/z-ai/glm-5.3" rel="noopener noreferrer"&gt;OpenRouter: z-ai/glm-5.3&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ofox.ai/models/z-ai/glm-5.3" rel="noopener noreferrer"&gt;ofox model page: GLM 5.3&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://ofox.ai/blog/glm-5-3-api-pricing-endpoints-reasoning-effort-2026/" rel="noopener noreferrer"&gt;ofox.ai/blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>glm</category>
      <category>apiguide</category>
      <category>pricing</category>
      <category>costoptimization</category>
    </item>
    <item>
      <title>Opus 5 vs Grok 4.6 Cost: 3.6x the Bill for 1.75x the Code</title>
      <dc:creator>Owen</dc:creator>
      <pubDate>Thu, 20 Aug 2026 17:24:42 +0000</pubDate>
      <link>https://dev.to/owen_fox/opus-5-vs-grok-46-cost-36x-the-bill-for-175x-the-code-1k2h</link>
      <guid>https://dev.to/owen_fox/opus-5-vs-grok-46-cost-36x-the-bill-for-175x-the-code-1k2h</guid>
      <description>&lt;h1&gt;
  
  
  Opus 5 vs Grok 4.6 Cost: 3.6x the Bill for 1.75x the Code
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Opus 5 bills 3.6x what Grok 4.6 bills on the same task, while returning 1.75x as much code in half the wall clock. Underneath that, the two providers disagree about what counts as output, so one cost formula gives a correct answer on Opus 5 and an answer 78% too low on Grok.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Much More Does Opus 5 Cost Than Grok 4.6?
&lt;/h2&gt;

&lt;p&gt;Both prices read off the ofox model pages on 2026-08-20.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Claude Opus 5&lt;/th&gt;
&lt;th&gt;Grok 4.6&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Input / 1M&lt;/td&gt;
&lt;td&gt;$5.00&lt;/td&gt;
&lt;td&gt;$2.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output / 1M&lt;/td&gt;
&lt;td&gt;$25.00&lt;/td&gt;
&lt;td&gt;$6.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cache read / 1M&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cache write / 1M&lt;/td&gt;
&lt;td&gt;$6.25 (5 min), $10 (1 hr)&lt;/td&gt;
&lt;td&gt;not charged&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context window&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;500K&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Max output&lt;/td&gt;
&lt;td&gt;128K&lt;/td&gt;
&lt;td&gt;66K&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Released&lt;/td&gt;
&lt;td&gt;2026-07-25&lt;/td&gt;
&lt;td&gt;2026-08-12&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two rows deserve more attention than the headline rates.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cache writes are asymmetric.&lt;/strong&gt; Anthropic bills you to put a prompt into the cache and again, at a lower rate, to read it back. xAI bills only the read. If your agent rewrites a long system prompt or a large tool schema on every session, that column decides more of your invoice than the $5-versus-$2 input rate does.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Grok 4.6 has a second price tier that the catalog does not expose.&lt;/strong&gt; Once a prompt reaches 200K tokens, every token in that request bills at double ($4 in, $12 out), not just the tokens past the line. Nothing below crosses it; all the runs here sit at a few hundred prompt tokens.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Does One Real Task Actually Bill?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;$0.1417 on Opus 5, $0.0397 on Grok 4.6. That is 3.6x.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The task: rewrite a Python CSV-parsing module to stream instead of buffering, detect its header reliably, surface malformed rows instead of dropping them, and keep money as &lt;code&gt;Decimal&lt;/code&gt;. Four runs per model, same prompt, same OpenAI-compatible endpoint, non-streaming, &lt;code&gt;max_tokens&lt;/code&gt; 8000, on 2026-08-20.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Opus 5&lt;/th&gt;
&lt;th&gt;Grok 4.6&lt;/th&gt;
&lt;th&gt;Ratio&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Prompt tokens&lt;/td&gt;
&lt;td&gt;270&lt;/td&gt;
&lt;td&gt;381&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Visible output tokens (median)&lt;/td&gt;
&lt;td&gt;5,616&lt;/td&gt;
&lt;td&gt;1,308&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reasoning tokens (median)&lt;/td&gt;
&lt;td&gt;not reported separately&lt;/td&gt;
&lt;td&gt;5,229&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Total tokens (median)&lt;/td&gt;
&lt;td&gt;5,886&lt;/td&gt;
&lt;td&gt;6,865&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Wall clock (median)&lt;/td&gt;
&lt;td&gt;59.0s&lt;/td&gt;
&lt;td&gt;109.4s&lt;/td&gt;
&lt;td&gt;Grok 1.85x slower&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output characters (median)&lt;/td&gt;
&lt;td&gt;9,044&lt;/td&gt;
&lt;td&gt;5,160&lt;/td&gt;
&lt;td&gt;Opus 1.75x more&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Bill per run (median)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.1417&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.0397&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;3.6x&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Cost per 1,000 output chars&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.01567&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.00769&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;2.0x&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The bill row is all-in: input at the listed rate plus everything the provider counts as output, priced as &lt;code&gt;total_tokens - prompt_tokens&lt;/code&gt;. On Grok that deliberately includes the reasoning tokens.&lt;/p&gt;

&lt;p&gt;So the headline holds and then stops holding. Opus 5 costs 3.6x as much per run. Normalise by what actually came back and it costs 2.0x as much. Still more expensive, but roughly double rather than roughly quadruple.&lt;/p&gt;

&lt;p&gt;Both models produced a working module every time. All eight runs returned &lt;code&gt;finish_reason: stop&lt;/code&gt;. The difference is thoroughness, not correctness: Opus 5 wrote more error branches, more docstring text, and in one run a small usage example.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why Is Grok Slower If It Writes Less?
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Because most of what it produces, you never see.&lt;/strong&gt; Median reasoning was 5,229 tokens against 1,308 tokens of visible answer, four reasoning tokens for every token in the file it hands you. Opus 5 does think on this prompt too, but it does not report the split.&lt;/p&gt;

&lt;p&gt;You can see Opus 5's hidden portion indirectly. In the first pass, without an explicit &lt;code&gt;max_tokens&lt;/code&gt;, one Opus 5 run reported 4,096 completion tokens and returned 845 characters of text. Four thousand tokens do not produce 845 characters of Python. The rest was thinking that was billed and not returned.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Is Your Cost Estimate Wrong for Exactly One of These Models?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Because &lt;code&gt;completion_tokens&lt;/code&gt; means different things on the two APIs.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Almost every cost snippet on the internet computes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;cost&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;prompt_tokens&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;in_rate&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completion_tokens&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;out_rate&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mf"&gt;1e6&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run that against the same four Grok 4.6 responses and it returns a median of &lt;strong&gt;$0.0086&lt;/strong&gt;. The real median is &lt;strong&gt;$0.0397&lt;/strong&gt;. The formula understates the bill by &lt;strong&gt;78%&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The reason is one field:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Field&lt;/th&gt;
&lt;th&gt;Opus 5&lt;/th&gt;
&lt;th&gt;Grok 4.6&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;completion_tokens&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;includes thinking&lt;/td&gt;
&lt;td&gt;excludes reasoning&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;completion_tokens_details.reasoning_tokens&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;not present&lt;/td&gt;
&lt;td&gt;present, and large&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;total_tokens&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;prompt + completion&lt;/td&gt;
&lt;td&gt;prompt + completion + reasoning&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Verified on a short streamed request to both: Grok returned &lt;code&gt;prompt 227 + completion 189 + reasoning 444 = total 860&lt;/code&gt;, and 227 + 189 alone is 416. Opus 5 returned &lt;code&gt;prompt 40 + completion 891 = total 931&lt;/code&gt;, with no reasoning field at all.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The portable formula is &lt;code&gt;total_tokens - prompt_tokens&lt;/code&gt; for the output side.&lt;/strong&gt; It is correct on both, and it survives a provider adding a reasoning field later.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;out_tokens&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;total_tokens&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;prompt_tokens&lt;/span&gt;
&lt;span class="n"&gt;cost&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;prompt_tokens&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;in_rate&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;out_tokens&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;out_rate&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mf"&gt;1e6&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Does the Same Prompt Cost the Same Tokens on Both?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;No. Opus 5 charges fewer tokens for identical English text.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The 1,130-character prompt metered at 270 tokens on Opus 5 and 381 on Grok 4.6. Grok 4.6 carries a fixed per-request overhead of roughly 206 tokens that is not your text. Subtract it and your 1,130 characters cost about 175 tokens on Grok against 270 on Opus 5, or 6.5 versus 4.2 characters per token.&lt;/p&gt;

&lt;p&gt;Two practical consequences:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;At a few hundred characters per call, 206 tokens of preamble is most of your input bill on Grok.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Input is the small half here anyway.&lt;/strong&gt; These runs produced 15-25x more output tokens than input tokens. Tokenizer differences move the total by single-digit percent.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What Do the Benchmarks Actually Say?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Neither model is on the Terminal-Bench 2.1 leaderboard.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A widely circulated summary this month put Opus 5 at 86.7% on Terminal-Bench 2.1. Pulling the official board on 2026-08-20:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Rank&lt;/th&gt;
&lt;th&gt;Agent&lt;/th&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Accuracy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;Claude Code&lt;/td&gt;
&lt;td&gt;Fable 5&lt;/td&gt;
&lt;td&gt;83.8% ± 1.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Codex&lt;/td&gt;
&lt;td&gt;GPT-5.5&lt;/td&gt;
&lt;td&gt;83.1% ± 1.1%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Terminus 2&lt;/td&gt;
&lt;td&gt;Fable 5&lt;/td&gt;
&lt;td&gt;80.4% ± 1.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;Cursor CLI&lt;/td&gt;
&lt;td&gt;Grok 4.5&lt;/td&gt;
&lt;td&gt;79.3% ± 1.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;Claude Code&lt;/td&gt;
&lt;td&gt;Opus 4.8&lt;/td&gt;
&lt;td&gt;78.9% ± 1.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Seventeen entries in total, all verified by a Terminal-Bench team member, most recent dated 2026-07-11. No Opus 5. No Grok 4.6. The top score is 83.8%, so an 86.7% would sit above the top of a board it is not on.&lt;/p&gt;

&lt;p&gt;That does not make 86.7% fabricated. Vendors run these suites internally and publish before submitting. It does mean the number is vendor-reported rather than a verified board entry.&lt;/p&gt;

&lt;p&gt;The one thing the official board does say about the family is a detail nobody quotes: the Grok 4.5 entry at rank 4 carries a &lt;strong&gt;-9.0% hack rate&lt;/strong&gt;, the largest on the board by a factor of ten, meaning the graders found that share of its passes came from gaming the test. Rank 5, Opus 4.8, is at -0.0%.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Breaks When You Switch?
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Symptom&lt;/th&gt;
&lt;th&gt;Cause&lt;/th&gt;
&lt;th&gt;Fix&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Output stops mid-file at exactly 4,096 tokens&lt;/td&gt;
&lt;td&gt;No &lt;code&gt;max_tokens&lt;/code&gt; set; that is the default&lt;/td&gt;
&lt;td&gt;Set it explicitly. Opus 5 allows 128K, Grok 4.6 allows 66K&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Bill is ~4x your estimate on Grok&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;completion_tokens&lt;/code&gt; excludes &lt;code&gt;reasoning_tokens&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Use &lt;code&gt;total_tokens - prompt_tokens&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No reasoning text in the response on Opus 5&lt;/td&gt;
&lt;td&gt;Thinking is billed inside &lt;code&gt;completion_tokens&lt;/code&gt; and not returned in non-streaming responses&lt;/td&gt;
&lt;td&gt;Stream if you need it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prompt over 200K suddenly doubles on Grok&lt;/td&gt;
&lt;td&gt;Second price tier applies to the whole request&lt;/td&gt;
&lt;td&gt;Keep prompts under 200K or budget for $4/$12&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The switch itself is a string change if you are already on an OpenAI-compatible endpoint:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;
&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.ofox.ai/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;OFOX_KEY&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;tuple&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;8000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;u&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;usage&lt;/span&gt;
    &lt;span class="n"&gt;rates&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;anthropic/claude-opus-5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;25&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;x-ai/grok-4.6&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt;
    &lt;span class="n"&gt;ri&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ro&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;rates&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;out&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;u&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;total_tokens&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;u&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;prompt_tokens&lt;/span&gt;        &lt;span class="c1"&gt;# correct on both APIs
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;u&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;prompt_tokens&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;ri&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;out&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;ro&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mf"&gt;1e6&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Which One Should You Actually Pick?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Split by task, and let output length be the deciding variable rather than the price.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Grok 4.6 for:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Routine agent passes where a shorter, more focused answer is fine: test scaffolding, mechanical refactors, code explanation&lt;/li&gt;
&lt;li&gt;Anything with a long cached system prompt, where the absent cache-write charge compounds every session&lt;/li&gt;
&lt;li&gt;Batch and offline jobs where 109 seconds versus 59 seconds does not matter to anyone&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Opus 5 for:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Passes where completeness is the product: the run that has to enumerate all the error branches, not most of them&lt;/li&gt;
&lt;li&gt;Interactive work where a human is waiting. Half the wall-clock on this task&lt;/li&gt;
&lt;li&gt;Anything above 500K of context, which Grok 4.6 cannot hold at all&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Neither, if:&lt;/strong&gt; you are picking on a leaderboard number. Neither model is on the board that number came from.&lt;/p&gt;

&lt;p&gt;Four runs on one task, summarised honestly: the expensive model is genuinely more thorough, the cheap model is genuinely cheap, and the ratio between those two facts is 2.0x, not 3.6x. Measure your own workload before you commit to either.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Is Grok 4.6 cheaper than Claude Opus 5?&lt;/strong&gt;&lt;br&gt;
Yes, substantially. On list price Grok 4.6 is 26.7% of the combined rate. Measured on the same task, the median bill came out at 28.0% of Opus 5. But Opus 5 returned 1.75x as much code, so per thousand characters of output the gap narrows to 2.0x.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why is my Grok 4.6 cost estimate too low?&lt;/strong&gt;&lt;br&gt;
Because &lt;code&gt;completion_tokens&lt;/code&gt; does not include &lt;code&gt;reasoning_tokens&lt;/code&gt; on Grok, while &lt;code&gt;total_tokens&lt;/code&gt; does. The standard formula understates the bill by 78%.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does Grok 4.6 or Opus 5 answer faster?&lt;/strong&gt;&lt;br&gt;
Opus 5, on this workload. Median wall-clock was 59.0 seconds for Opus 5 and 109.4 seconds for Grok 4.6 across four runs each.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does Terminal-Bench 2.1 show Opus 5 ahead of Grok 4.6?&lt;/strong&gt;&lt;br&gt;
Neither model is on the leaderboard. Any 86-88% figure attributed to this benchmark for either model is vendor-reported, not a verified board entry.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Do both models charge for cache writes?&lt;/strong&gt;&lt;br&gt;
No. Anthropic charges separately to write a prompt into the cache ($6.25/M for the 5-minute TTL, $10/M for the 1-hour TTL) on top of $0.5/M cache reads. Grok 4.6 lists cache reads at $0.5/M and no write charge.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://ofox.ai/models/anthropic/claude-opus-5" rel="noopener noreferrer"&gt;ofox model page: Claude Opus 5&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://ofox.ai/models/x-ai/grok-4.6" rel="noopener noreferrer"&gt;ofox model page: Grok 4.6&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.tbench.ai/leaderboard/terminal-bench/2.1" rel="noopener noreferrer"&gt;Terminal-Bench 2.1 leaderboard&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.x.ai/developers/models" rel="noopener noreferrer"&gt;xAI developer docs: models and pricing&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://ofox.ai/blog/claude-opus-5-vs-grok-4-6-cost-2026/" rel="noopener noreferrer"&gt;ofox.ai/blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>grok</category>
      <category>claude</category>
      <category>pricing</category>
      <category>comparison</category>
    </item>
    <item>
      <title>Pi Coding Agent Custom Provider: Setup, Cost, and 3 Fixes</title>
      <dc:creator>Owen</dc:creator>
      <pubDate>Wed, 19 Aug 2026 04:34:57 +0000</pubDate>
      <link>https://dev.to/owen_fox/pi-coding-agent-custom-provider-setup-cost-and-3-fixes-c7f</link>
      <guid>https://dev.to/owen_fox/pi-coding-agent-custom-provider-setup-cost-and-3-fixes-c7f</guid>
      <description>&lt;h1&gt;
  
  
  Pi Coding Agent Custom Provider: Setup, Cost, and 3 Fixes
&lt;/h1&gt;

&lt;p&gt;Pi ships with 36 API-key providers and six subscription logins already wired up, and there is still a decent chance yours is not among them. Pointing it somewhere else takes one JSON file and about five minutes, which is the easy part. The interesting part is the three things that go wrong afterwards, none of which announce themselves as configuration problems.&lt;/p&gt;

&lt;p&gt;Everything below was run on 2026-08-18 against pi 0.84.1 on macOS, with a gateway that serves 131 models over an OpenAI-compatible endpoint.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Config file:        ~/.pi/agent/models.json
Provider fields:    baseUrl, api, apiKey, models[]
Protocols:          openai-completions, openai-responses,
                    anthropic-messages, google-generative-ai
Key from env:       "$OFOX_API_KEY"
Key from keychain:  "!security find-generic-password -ws ofox"
Unlisted model:     runs, with a warning, at 128K context
Cost accounting:    $0.00 until you add a cost block
Reasoning + Claude: 500 until you add supportsDeveloperRole: false
Built-in override:  baseUrl only, stop before /v1
Reload:             automatic, every time /model opens
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What Can You Do After This Setup, and What Can't You?
&lt;/h2&gt;

&lt;p&gt;You get every model your gateway sells, in Pi's own loop, billed to one key. You do not get a model list fetched from the gateway, working cost numbers, or any warning when a model quietly runs at an eighth of its real context window.&lt;/p&gt;

&lt;p&gt;Works right away:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Any endpoint that speaks OpenAI chat completions, OpenAI Responses, Anthropic Messages or Google Generative AI.&lt;/li&gt;
&lt;li&gt;Model switching mid-session through &lt;code&gt;/model&lt;/code&gt;, including across providers, because the file is re-read each time the picker opens.&lt;/li&gt;
&lt;li&gt;Keys pulled from the environment or from a shell command, so nothing secret has to sit in the JSON.&lt;/li&gt;
&lt;li&gt;Thinking levels, image input and tool calling, as long as you declare them per model.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Does not work right away:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No catalogue discovery.&lt;/strong&gt; Pointed at a local server that logs every request, Pi made four POSTs to &lt;code&gt;/v1/chat/completions&lt;/code&gt; per run and not a single GET to &lt;code&gt;/v1/models&lt;/code&gt;. What you type is what it knows.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No cost numbers.&lt;/strong&gt; Token counts are exact, money is zero, until you write the rates yourself.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No protocol sniffing.&lt;/strong&gt; Override a built-in provider with the wrong base path and the error you get back describes auth, not protocol.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Should You Point Pi at a Gateway or Just Use Its Built-In Providers?
&lt;/h2&gt;

&lt;p&gt;Use the built-ins if you already pay Anthropic, OpenAI or Google directly. Add a custom provider when the model you want is not in that set, or when one key across every tool matters more than per-vendor dashboards.&lt;/p&gt;

&lt;p&gt;When a custom provider earns its keep:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You run models that no first-party provider carries, which in practice means most Chinese open-weight flagships and anything hosted rather than official.&lt;/li&gt;
&lt;li&gt;You already route Claude Code or Codex CLI through a gateway and want one key and one bill rather than four.&lt;/li&gt;
&lt;li&gt;You want to A/B a cheap default against an expensive escalation model without opening a second account for the second model.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When it is not worth the file:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You use one vendor and one plan. &lt;code&gt;/login&lt;/code&gt; covers six subscriptions directly, including ChatGPT Plus and Pro, Claude Pro and Max, GitHub Copilot, xAI and OpenRouter, without any of this.&lt;/li&gt;
&lt;li&gt;You are on a local runtime. Ollama, vLLM and llama.cpp are the documented cases and need &lt;code&gt;baseUrl&lt;/code&gt; plus a model id, nothing else in this post.&lt;/li&gt;
&lt;li&gt;You only wanted to change the key on a built-in provider. That is a one-line override, covered near the end.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Stop rule:&lt;/strong&gt; if &lt;code&gt;pi --list-models&lt;/code&gt; already shows the model you intend to run, close this tab. Everything here exists to add models Pi does not know about.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Do You Need Before You Start?
&lt;/h2&gt;

&lt;p&gt;Node 22 or newer, a key, and a base URL you have already curled once.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Requirement&lt;/th&gt;
&lt;th&gt;What we used&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Node.js&lt;/td&gt;
&lt;td&gt;24.14.1&lt;/td&gt;
&lt;td&gt;Package declares &lt;code&gt;engines: node &amp;gt;=22.19.0&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pi&lt;/td&gt;
&lt;td&gt;0.84.1 (latest 0.84.2)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;@earendil-works/pi-coding-agent&lt;/code&gt;, MIT&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Endpoint&lt;/td&gt;
&lt;td&gt;&lt;code&gt;https://api.ofox.ai/v1&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Must answer &lt;code&gt;/chat/completions&lt;/code&gt;, not just &lt;code&gt;/models&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Key&lt;/td&gt;
&lt;td&gt;one gateway key&lt;/td&gt;
&lt;td&gt;Held in &lt;code&gt;$OFOX_API_KEY&lt;/code&gt;, never inline&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model ids&lt;/td&gt;
&lt;td&gt;exact strings&lt;/td&gt;
&lt;td&gt;Gateway ids, not vendor ids&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Install, if you have not:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; @earendil-works/pi-coding-agent
pi &lt;span class="nt"&gt;--version&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One thing worth deciding before you write the file: the provider name you pick becomes part of every &lt;code&gt;--provider&lt;/code&gt; flag and every session record. Rename it later and old sessions point at a provider that no longer exists.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Do You Add a Custom Provider to Pi?
&lt;/h2&gt;

&lt;p&gt;Four fields in one file, then one command to prove it works.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 1: Write the provider block
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;~/.pi/agent/models.json&lt;/code&gt; holds everything. The minimum viable entry:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"providers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"ofox"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"baseUrl"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://api.ofox.ai/v1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"api"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"openai-completions"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"apiKey"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"$OFOX_API_KEY"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"models"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"deepseek/deepseek-v4-flash"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"contextWindow"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1000000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"maxTokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;384000&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;openai-completions&lt;/code&gt; is the one to reach for first. It is the most widely implemented shape, and on our gateway &lt;code&gt;openai-responses&lt;/code&gt; also worked for the same model, which is not something you can assume elsewhere.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Put the key in the environment, not the file
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;OFOX_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;sk-...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;apiKey&lt;/code&gt; resolves three ways: a literal string, &lt;code&gt;$VAR&lt;/code&gt; or &lt;code&gt;${VAR}&lt;/code&gt; interpolation, and &lt;code&gt;!command&lt;/code&gt;, which runs a shell command and uses stdout. The third form is the one to use on a shared machine:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"apiKey"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"!security find-generic-password -ws ofox"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Step 3: Confirm Pi sees the models
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;pi --list-models ofox

provider  model                       context  max-out  thinking  images
ofox      deepseek/deepseek-v4-flash  1M       384K     no        no
ofox      moonshotai/kimi-k3          1M       1M       no        no
ofox      z-ai/glm-5.2                1M       128K     yes       no
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Those columns come from your file, not from the gateway. Delete &lt;code&gt;contextWindow&lt;/code&gt; and &lt;code&gt;maxTokens&lt;/code&gt; from an entry and the same command prints &lt;code&gt;128K&lt;/code&gt; and &lt;code&gt;16.4K&lt;/code&gt; for it, which are Pi's documented defaults. If &lt;code&gt;thinking&lt;/code&gt; says no on a model that reasons, that is your declaration missing, not the endpoint refusing.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 4: Run something that touches the disk
&lt;/h3&gt;

&lt;p&gt;Print mode is the fastest proof, because it exercises the tool loop rather than just the completion endpoint:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pi &lt;span class="nt"&gt;--provider&lt;/span&gt; ofox &lt;span class="nt"&gt;--model&lt;/span&gt; deepseek/deepseek-v4-flash &lt;span class="nt"&gt;-p&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s2"&gt;"Read buggy.py, run it, and state the one-line bug. Do not edit files."&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In a scratch directory holding a two-line file with &lt;code&gt;return a - b&lt;/code&gt; in an &lt;code&gt;add&lt;/code&gt; function, DeepSeek V4 Flash read the file, ran the interpreter through the bash tool, and answered correctly on the first try. That is the whole integration test: file read, shell execution, answer.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does &lt;code&gt;openai-responses&lt;/code&gt; Work Too?
&lt;/h3&gt;

&lt;p&gt;On our gateway, yes, for the same model, with the protocol name as the only change. Swapping &lt;code&gt;"api": "openai-completions"&lt;/code&gt; for &lt;code&gt;"api": "openai-responses"&lt;/code&gt; and re-running the same prompt returned the same answer.&lt;/p&gt;

&lt;p&gt;Do not generalise that. Responses support is decided per model by whoever hosts it, not per gateway, so an endpoint that answers &lt;code&gt;/v1/responses&lt;/code&gt; for one model can have no Responses route at all for the next one. &lt;code&gt;openai-completions&lt;/code&gt; is the shape with the widest coverage, and there is no advantage to picking anything else unless a specific model needs it. Codex CLI is the tool that forces the question, because it speaks Responses and nothing else.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why Does &lt;code&gt;pi auth check&lt;/code&gt; Say Ready With a Key That Does Not Work?
&lt;/h3&gt;

&lt;p&gt;Because it checks that a key is present, not that it is valid. We pointed the provider at a deliberately invalid key and asked:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pi auth check &lt;span class="nt"&gt;--provider&lt;/span&gt; ofox
&lt;span class="c"&gt;# ready&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The same config, one request later:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="mi"&gt;401&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"message"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"Invalid or expired API key"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"invalid_api_key"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"code"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;401&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;ready&lt;/code&gt; means Pi resolved something into the &lt;code&gt;apiKey&lt;/code&gt; slot. Treat it as a spelling check on your environment variable and nothing more. The real readiness test is Step 4.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does Pi Report $0 for Every Session?
&lt;/h2&gt;

&lt;p&gt;Because a custom provider has no price list, and Pi will not invent one. Token accounting is exact. Here is the usage record Pi wrote for the two turns of that first run, straight from the session file under &lt;code&gt;~/.pi/agent/sessions/&lt;/code&gt;:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Turn&lt;/th&gt;
&lt;th&gt;input&lt;/th&gt;
&lt;th&gt;output&lt;/th&gt;
&lt;th&gt;cacheRead&lt;/th&gt;
&lt;th&gt;reasoning&lt;/th&gt;
&lt;th&gt;total&lt;/th&gt;
&lt;th&gt;cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;2,840&lt;/td&gt;
&lt;td&gt;122&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;2,962&lt;/td&gt;
&lt;td&gt;$0.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;52&lt;/td&gt;
&lt;td&gt;56&lt;/td&gt;
&lt;td&gt;2,944&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;3,052&lt;/td&gt;
&lt;td&gt;$0.00&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two things in that table are worth separating. The &lt;code&gt;cacheRead&lt;/code&gt; figure is real: the gateway returned &lt;code&gt;prompt_tokens_details.cached_tokens&lt;/code&gt; on the second call, and Pi recorded it. The &lt;code&gt;cost&lt;/code&gt; column is not real, it is absent. Every field under &lt;code&gt;cost&lt;/code&gt; sits at zero because the model entry never declared rates.&lt;/p&gt;

&lt;p&gt;Add them and the arithmetic starts working:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"anthropic/claude-sonnet-5"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"contextWindow"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1000000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"maxTokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;128000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"cost"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"input"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"output"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"cacheRead"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"cacheWrite"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;2.5&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Rates are per million tokens, taken from the provider's own pricing page rather than from the vendor's. The next run against Claude Sonnet 5 recorded 4,111 input and 4 output tokens and priced them at &lt;code&gt;$0.008222&lt;/code&gt; plus &lt;code&gt;$0.00004&lt;/code&gt;, total &lt;code&gt;$0.008262&lt;/code&gt;, which is those counts multiplied by $2 and $10 per million. Pi does that arithmetic locally, so the number is only as honest as the rates you typed. Type the gateway's rates, not the model maker's, and re-check them when the page changes.&lt;/p&gt;

&lt;p&gt;The same applies to &lt;code&gt;contextWindow&lt;/code&gt;. Sonnet 5 is a 1M-context model on this gateway, and the entry has to say so, otherwise the 128K default quietly takes over.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Does a Model Fail With "unsupported message role: developer"?
&lt;/h2&gt;

&lt;p&gt;Because &lt;code&gt;reasoning: true&lt;/code&gt; makes Pi send the system prompt as a &lt;code&gt;developer&lt;/code&gt; role message, and not every upstream accepts that role. The failure is loud and looks like a server fault:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"code"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"message"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"Request error: failed to convert messages:
unsupported message role: developer"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"param"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"api_error"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nothing in that string points at your config, which is why it is worth isolating properly. Three runs, same gateway, same model, one field at a time:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model entry&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;reasoning: true&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;500, &lt;code&gt;unsupported message role: developer&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;reasoning: true&lt;/code&gt; plus &lt;code&gt;compat: { supportsDeveloperRole: false }&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Works&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;reasoning: false&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Works&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;So the trigger is the &lt;code&gt;developer&lt;/code&gt; role, and there are two fixes with different costs. Running the same three configs against a local server that logs request bodies shows exactly what changes:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model entry&lt;/th&gt;
&lt;th&gt;&lt;code&gt;messages[].role&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;
&lt;code&gt;reasoning_effort&lt;/code&gt; sent&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;reasoning: true&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;["developer", "user"]&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;"medium"&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;plus &lt;code&gt;supportsDeveloperRole: false&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;&lt;code&gt;["system", "user"]&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;"medium"&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;reasoning: false&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;["system", "user"]&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;absent&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The compat switch moves the system prompt to a &lt;code&gt;system&lt;/code&gt; message and leaves &lt;code&gt;reasoning_effort&lt;/code&gt; in place, so thinking survives. Setting &lt;code&gt;reasoning: false&lt;/code&gt; also clears the error, by dropping &lt;code&gt;reasoning_effort&lt;/code&gt; from the request entirely, which is usually the wrong trade.&lt;/p&gt;

&lt;p&gt;The same capture answers a question people ask about &lt;code&gt;maxTokensField&lt;/code&gt;: on &lt;code&gt;openai-completions&lt;/code&gt; Pi sends &lt;code&gt;max_completion_tokens&lt;/code&gt;, not &lt;code&gt;max_tokens&lt;/code&gt;. If your endpoint only understands the older field, that is the switch to flip.&lt;/p&gt;

&lt;p&gt;The role is only rejected on part of the catalogue. Same gateway, same &lt;code&gt;reasoning: true&lt;/code&gt;, three model families:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;
&lt;code&gt;reasoning: true&lt;/code&gt; result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V4 Flash&lt;/td&gt;
&lt;td&gt;Works&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GLM 5.2&lt;/td&gt;
&lt;td&gt;Works&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 5&lt;/td&gt;
&lt;td&gt;500 until &lt;code&gt;supportsDeveloperRole: false&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The pattern is upstream shape, not gateway policy. Anthropic's API has no &lt;code&gt;developer&lt;/code&gt; role, so a gateway translating OpenAI-shaped requests into Messages has nothing to map it to. OpenAI-shaped upstreams take it and move on. This is the same class of problem as Codex CLI emitting an empty tool description that some upstreams validate and others ignore, which we hit while testing nine harnesses against one gateway. The lesson repeats: when a client and an endpoint disagree, read the body before you change settings.&lt;/p&gt;

&lt;p&gt;Pi's docs list two more switches in the same family, &lt;code&gt;supportsReasoningEffort&lt;/code&gt; for servers that reject reasoning parameters and &lt;code&gt;maxTokensField&lt;/code&gt; for servers that want &lt;code&gt;max_completion_tokens&lt;/code&gt; instead of &lt;code&gt;max_tokens&lt;/code&gt;. If a model 400s the moment thinking is enabled, those are the next two to try.&lt;/p&gt;

&lt;h2&gt;
  
  
  Do You Have to List Every Model?
&lt;/h2&gt;

&lt;p&gt;No, and on a gateway with 131 of them you should not try. An id Pi has never seen still runs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pi &lt;span class="nt"&gt;--provider&lt;/span&gt; ofox &lt;span class="nt"&gt;--model&lt;/span&gt; z-ai/glm-5.2 &lt;span class="nt"&gt;-p&lt;/span&gt; &lt;span class="s2"&gt;"say ok"&lt;/span&gt;
&lt;span class="c"&gt;# Warning: Model "z-ai/glm-5.2" not found for provider "ofox". Using custom model id.&lt;/span&gt;
&lt;span class="c"&gt;# ok&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That fallback is the difference between a five-line config and a five-hundred-line one. It also hides a cost. An unlisted model inherits Pi's defaults, documented as 128,000 context and 16,384 max output, and &lt;code&gt;pi --list-models&lt;/code&gt; prints exactly those two figures for any entry that omits them. Auto-compaction then fires at &lt;code&gt;contextTokens &amp;gt; contextWindow - reserveTokens&lt;/code&gt;, with &lt;code&gt;reserveTokens&lt;/code&gt; defaulting to 16,384, so a 1M-context model summarises itself somewhere around 111,600 tokens instead of near a million. Nothing in the output says why. People do report compaction arriving sooner than expected in the project's community; this is at least one mechanism that produces it, and it is cheap to rule out before blaming the model.&lt;/p&gt;

&lt;p&gt;The practical split: let unlisted ids cover exploration, then write a real entry for the two or three models you run daily, with &lt;code&gt;contextWindow&lt;/code&gt;, &lt;code&gt;maxTokens&lt;/code&gt;, &lt;code&gt;reasoning&lt;/code&gt;, &lt;code&gt;input&lt;/code&gt; and &lt;code&gt;cost&lt;/code&gt; filled in. Everything Pi displays about a model, including whether it will accept an image, comes from that entry rather than from the endpoint.&lt;/p&gt;

&lt;h2&gt;
  
  
  Can You Point Pi's Built-In Anthropic Provider at a Gateway?
&lt;/h2&gt;

&lt;p&gt;Yes, and it is the better route for Claude models, as long as you give it the Anthropic base path rather than the OpenAI one. The override is one line and no model list of your own:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"providers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"anthropic"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"baseUrl"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://api.ofox.ai/anthropic"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"apiKey"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"$OFOX_API_KEY"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pi &lt;span class="nt"&gt;--provider&lt;/span&gt; anthropic &lt;span class="nt"&gt;--model&lt;/span&gt; claude-sonnet-5 &lt;span class="nt"&gt;-p&lt;/span&gt; &lt;span class="s2"&gt;"Reply with exactly: ok"&lt;/span&gt;
&lt;span class="c"&gt;# ok&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Pi keeps its entire built-in Claude catalogue, with the windows already correct, Claude Fable 5 at 1M and the Opus and Haiku entries at 200K. Nothing to declare, nothing to keep in sync, and no &lt;code&gt;developer&lt;/code&gt; role in sight because Messages is the native shape here. The vendor id works as-is, no gateway prefix.&lt;/p&gt;

&lt;p&gt;Getting the base URL wrong produces two errors that both describe the wrong problem. Point it at the OpenAI path:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="mi"&gt;401&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"error"&lt;/span&gt;&lt;span class="p"&gt;:{&lt;/span&gt;&lt;span class="nl"&gt;"message"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"You didn't provide an API key. You need to provide your API key
in an Authorization header using Bearer auth ..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"invalid_request_error"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"code"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;401&lt;/span&gt;&lt;span class="p"&gt;}}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is no auth bug there. Pi is speaking Messages, so it sends &lt;code&gt;x-api-key&lt;/code&gt;, and the OpenAI path only accepts &lt;code&gt;Authorization: Bearer&lt;/code&gt;. Add a Bearer header through the provider's &lt;code&gt;headers&lt;/code&gt; field and the honest answer appears: &lt;code&gt;404 Unsupported OpenAI API endpoint&lt;/code&gt;. The 401 was a protocol mismatch wearing an auth costume.&lt;/p&gt;

&lt;p&gt;The other way to get it wrong is doubling the version segment:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="mi"&gt;404&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"error"&lt;/span&gt;&lt;span class="p"&gt;:{&lt;/span&gt;&lt;span class="nl"&gt;"message"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"Unsupported Anthropic API endpoint. ..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"code"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;404&lt;/span&gt;&lt;span class="p"&gt;}}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is &lt;code&gt;.../anthropic/v1&lt;/code&gt; in the config. Pi appends &lt;code&gt;/v1/messages&lt;/code&gt; itself, so the base URL stops at &lt;code&gt;/anthropic&lt;/code&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;code&gt;baseUrl&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;https://api.ofox.ai/anthropic&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Works, full built-in Claude catalogue&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;https://api.ofox.ai/anthropic/v1&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;404, &lt;code&gt;Unsupported Anthropic API endpoint&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;https://api.ofox.ai/v1&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;401 that is really a 404 on the wrong protocol&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;So Claude has two routes through the same key: the built-in override above, or a custom &lt;code&gt;openai-completions&lt;/code&gt; entry with the gateway's own ids, which is what the cost example used. The override is less typing and avoids the &lt;code&gt;developer&lt;/code&gt; role entirely. The custom entry is the one to use when you want per-model &lt;code&gt;cost&lt;/code&gt; and &lt;code&gt;contextWindow&lt;/code&gt; values Pi does not already know.&lt;/p&gt;

&lt;h3&gt;
  
  
  Should You Run Claude Models in Pi at All?
&lt;/h3&gt;

&lt;p&gt;They work, and Pi's own author has documented a schema problem on the newest ones. Writing on 2026-07-04, Armin Ronacher reported that "newer Claude models sometimes call Pi's edit tool with extra, invented fields in the nested &lt;code&gt;edits[]&lt;/code&gt; array", with the result that "the model invents made-up keys and Pi thus rejects the tool call and asks to try again". His summary of the trend is the uncomfortable part: "this is getting worse with newer Anthropic models as both Opus 4.8 and Sonnet 5 show it but none of the older models."&lt;/p&gt;

&lt;p&gt;That is a training-and-tooling mismatch, not something a base URL can fix, and it costs a retry each time it fires. It does not make Claude unusable in Pi. It does mean that if you are picking a default model for a harness whose edit tool is its own, the newest Claude is not automatically the safest choice, and it is worth watching your session log for repeated tool calls on the same edit.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Do You See What Pi Actually Sends?
&lt;/h2&gt;

&lt;p&gt;Reproduce the call with curl, then compare it to what Pi recorded. Two files and one command cover most of what you need, and neither requires a proxy.&lt;/p&gt;

&lt;p&gt;The session log is the first stop. Every run writes a JSON Lines file under &lt;code&gt;~/.pi/agent/sessions/&amp;lt;project&amp;gt;/&lt;/code&gt;, one record per event, including a &lt;code&gt;model_change&lt;/code&gt; line naming the provider and model id Pi resolved and an assistant message carrying the &lt;code&gt;usage&lt;/code&gt; block. If the model id in that file is not the one you meant to run, the problem is your flag or your fallback, and no amount of provider tuning will fix it.&lt;/p&gt;

&lt;p&gt;The endpoint is the second. Send the same shape yourself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; https://api.ofox.ai/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$OFOX_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Content-Type: application/json'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"model":"deepseek/deepseek-v4-flash","messages":[{"role":"user","content":"hi"}],"max_tokens":8}'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  | python3 &lt;span class="nt"&gt;-c&lt;/span&gt; &lt;span class="s1"&gt;'import sys,json; print(json.load(sys.stdin)["usage"])'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two things in &lt;code&gt;usage&lt;/code&gt; are worth reading closely. &lt;code&gt;prompt_tokens_details.cached_tokens&lt;/code&gt; is what Pi maps to its &lt;code&gt;cacheRead&lt;/code&gt; column, so if that key never appears, the cache figures in your session log will stay at zero no matter how stable your prefix is. And &lt;code&gt;completion_tokens_details.reasoning_tokens&lt;/code&gt; tells you whether thinking is actually running, which is a faster check than reading output and guessing.&lt;/p&gt;

&lt;p&gt;Doing this before you edit config is the whole lesson from every harness integration we have written up. An error string is written by whoever raised it, and that is frequently not the component that is wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Breaks During Setup, and How Do You Fix It?
&lt;/h2&gt;

&lt;p&gt;Six failures, five of them reproduced against a live endpoint on 0.84.1, and one demonstrated with Pi's own model table.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Symptom&lt;/th&gt;
&lt;th&gt;Cause&lt;/th&gt;
&lt;th&gt;Fix&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;401: {"message":"Invalid or expired API key"...}&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Key resolved but wrong, or the variable is empty in this shell&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;echo $OFOX_API_KEY&lt;/code&gt; before blaming the file; &lt;code&gt;pi auth check&lt;/code&gt; will not catch this&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;404 404 page not found&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;baseUrl&lt;/code&gt; missing the version segment&lt;/td&gt;
&lt;td&gt;Use &lt;code&gt;https://host/v1&lt;/code&gt;, not &lt;code&gt;https://host&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;500 ... unsupported message role: developer&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;reasoning: true&lt;/code&gt; on a model whose upstream has no &lt;code&gt;developer&lt;/code&gt; role&lt;/td&gt;
&lt;td&gt;Add &lt;code&gt;compat: { supportsDeveloperRole: false }&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;401 ... provide your API key ... using Bearer auth&lt;/code&gt; on the &lt;code&gt;anthropic&lt;/code&gt; provider&lt;/td&gt;
&lt;td&gt;Built-in override pointed at the OpenAI path, so Pi sends &lt;code&gt;x-api-key&lt;/code&gt; where only Bearer is accepted&lt;/td&gt;
&lt;td&gt;Use the gateway's Anthropic base path, ending at &lt;code&gt;/anthropic&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;404: {"message":"Model 'openai/gpt-5.6' not found","type":"model_not_found"}&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The id is well formed but not in this gateway's catalogue&lt;/td&gt;
&lt;td&gt;Read the id off the gateway's own model page; a vendor announcing a model does not put it in every catalogue&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Session compacts far earlier than the model's real window&lt;/td&gt;
&lt;td&gt;Unlisted model fell back to the 128K default&lt;/td&gt;
&lt;td&gt;Declare &lt;code&gt;contextWindow&lt;/code&gt; and &lt;code&gt;maxTokens&lt;/code&gt; for that model&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The fourth and fifth rows are the ones that cost the most time, because both error messages describe something other than the actual problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Do Teams Share a Pi Provider Config?
&lt;/h2&gt;

&lt;p&gt;Share the file, never the key. &lt;code&gt;models.json&lt;/code&gt; holds no secret when every &lt;code&gt;apiKey&lt;/code&gt; is a &lt;code&gt;$VAR&lt;/code&gt; or a &lt;code&gt;!command&lt;/code&gt;, which makes it safe to commit into a dotfiles repo or a bootstrap script.&lt;/p&gt;

&lt;p&gt;A split that survives contact with more than one developer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Commit&lt;/strong&gt; the provider block: base URL, protocol, and full model entries with &lt;code&gt;contextWindow&lt;/code&gt;, &lt;code&gt;maxTokens&lt;/code&gt;, &lt;code&gt;reasoning&lt;/code&gt; and &lt;code&gt;cost&lt;/code&gt;. These are facts about the endpoint, identical for everyone, and getting them wrong is what produces silent early compaction and fake $0 bills.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Never commit&lt;/strong&gt; the key. &lt;code&gt;"apiKey": "$OFOX_API_KEY"&lt;/code&gt; in the shared file, real value in each developer's shell profile or keychain.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pin the version you validated.&lt;/strong&gt; Pi ships roughly weekly, 0.82.1 through 0.84.2 in under a month. Record the version your config was tested against.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Give everyone the same base URL.&lt;/strong&gt; One endpoint means one model catalogue, one rate limit pool and one place to see spend, instead of a per-developer guess.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last point is the one teams skip, and it is the one that turns "which model are you on?" from a question into a lookup.&lt;/p&gt;

&lt;h3&gt;
  
  
  How Do You Point Every Harness at the Same Key?
&lt;/h3&gt;

&lt;p&gt;Every harness stores model access in its own dialect. Claude Code reads &lt;code&gt;ANTHROPIC_BASE_URL&lt;/code&gt; and &lt;code&gt;ANTHROPIC_AUTH_TOKEN&lt;/code&gt;. Codex CLI wants a &lt;code&gt;model_providers&lt;/code&gt; block in &lt;code&gt;config.toml&lt;/code&gt; and refuses anything that is not the Responses API. Cline has a settings pane. DeepSeek Harness wants a custom provider form or &lt;code&gt;DEEPSEEK_BASE_URL&lt;/code&gt;. Pi wants the JSON file above. Five tools, five places to rotate a key, five model lists drifting apart.&lt;/p&gt;

&lt;p&gt;They all speak HTTP against an OpenAI-compatible or Anthropic-compatible endpoint, so the fix is identical in each: one base URL, one key, and the model string as the only thing that changes. That is the entire reason the custom-provider form in every one of these tools has the same four fields.&lt;/p&gt;

&lt;p&gt;On ofox that endpoint is &lt;code&gt;https://api.ofox.ai/v1&lt;/code&gt; with &lt;code&gt;openai-completions&lt;/code&gt;, and one key reached 131 models on 2026-08-18, including Kimi K3 and MiniMax M3 alongside the DeepSeek, GLM and Claude entries used above. The equivalent setup for the other tools is in our Codex CLI custom provider guide, the OpenCode configuration walkthrough, and the Cursor, Claude Code and Cline setup.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Does Pi Compare to the Harness You Already Run?
&lt;/h2&gt;

&lt;p&gt;Same job, much smaller surface, and a config file that is honest about how little it assumes. Pi gives the model four tools and an extension API, where Claude Code gives it hooks, subagents, skills and MCP servers out of the box. Neither is better in the abstract. The question is whether you want assembly or assembly required.&lt;/p&gt;

&lt;p&gt;What the custom-provider path shows is where that minimalism has a price. No catalogue fetch, no pricing table, no protocol sniffing. Every one of the three problems in this post is Pi declining to guess something on your behalf, and every fix is you writing the fact down once.&lt;/p&gt;

&lt;p&gt;For where Pi sits against the rest of the field, including the OpenRouter usage data on which models people actually run inside each harness, see the nine-harness roundup. For the terminal agents specifically, Claude Code vs Codex CLI vs Cursor covers the trade in more depth.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/earendil-works/pi" rel="noopener noreferrer"&gt;Pi repository&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/earendil-works/pi/blob/main/packages/coding-agent/docs/models.md" rel="noopener noreferrer"&gt;Custom models configuration guide&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/earendil-works/pi/blob/main/packages/coding-agent/docs/providers.md" rel="noopener noreferrer"&gt;Provider list and &lt;code&gt;/login&lt;/code&gt; subscriptions&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/earendil-works/pi/blob/main/packages/coding-agent/docs/compaction.md" rel="noopener noreferrer"&gt;Compaction and branch summarization&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/earendil-works/pi/blob/main/packages/coding-agent/docs/quickstart.md" rel="noopener noreferrer"&gt;Pi quickstart&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://lucumr.pocoo.org/2026/7/4/better-models-worse-tools/" rel="noopener noreferrer"&gt;Armin Ronacher, "Better Models, Worse Tools" (2026-07-04)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://simonwillison.net/2026/Jul/4/better-models-worse-tools/" rel="noopener noreferrer"&gt;Simon Willison's note on the same post&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://pi.dev/" rel="noopener noreferrer"&gt;pi.dev&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.reddit.com/r/PiCodingAgent/" rel="noopener noreferrer"&gt;r/PiCodingAgent&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What is the Pi coding agent?
&lt;/h3&gt;

&lt;p&gt;A terminal coding agent from Armin Ronacher and Mario Zechner, MIT licensed, now developed under Earendil at github.com/earendil-works/pi. It gives the model four built-in tools (read, write, edit, bash) and ships almost nothing else, exposing an extension API instead of hooks, subagents and skills. As of 2026-08-18 the repository has 92,619 stars and the npm package pulls 1.37 million downloads a week.&lt;/p&gt;

&lt;h3&gt;
  
  
  Which npm package installs Pi?
&lt;/h3&gt;

&lt;p&gt;@earendil-works/pi-coding-agent, currently 0.84.2, engines node &amp;gt;=22.19.0. The older @mariozechner/pi package stops at 0.70.6 and gets a few hundred downloads a week; it is the pre-acquisition channel and installing it gets you a version from before the move to Earendil.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does Pi support third-party API endpoints?
&lt;/h3&gt;

&lt;p&gt;Yes, through ~/.pi/agent/models.json. A provider entry takes baseUrl, api, apiKey and a models array, where api is one of openai-completions, openai-responses, anthropic-messages or google-generative-ai. No code changes and no fork are required, and the file reloads every time the /model picker opens.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can Pi read an API key from an environment variable?
&lt;/h3&gt;

&lt;p&gt;Yes. apiKey accepts $VAR and ${VAR} interpolation, and it also accepts !command, which runs a shell command and takes stdout as the key. That second form is how you read from a system keychain instead of leaving a secret in a JSON file. Use $$ for a literal dollar sign.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does Pi need every model listed in models.json?
&lt;/h3&gt;

&lt;p&gt;No. Passing a model id that is not in the list prints Warning: Model not found for provider, then runs it as a custom model id. The catch is that an unlisted model inherits the defaults, 128,000 context and 16,384 max output, so a 1M-context model will compact far earlier than it needs to.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why does Pi show $0 cost for every request?
&lt;/h3&gt;

&lt;p&gt;Because a custom provider carries no pricing metadata. Token counts are recorded correctly in the session file, but every field under cost stays at zero until you add a cost block to the model entry with input, output, cacheRead and cacheWrite rates per million tokens.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can Pi point its built-in Anthropic provider at a gateway?
&lt;/h3&gt;

&lt;p&gt;Yes, if you give it the gateway's Anthropic Messages base path rather than its OpenAI one. Overriding baseUrl on the built-in anthropic provider keeps Pi's whole Claude catalogue with correct context windows and needs no models array. Stop the URL before the version segment, because Pi appends /v1/messages itself, and pointing it at the OpenAI path returns a 401 that is really a protocol mismatch.&lt;/p&gt;

&lt;h3&gt;
  
  
  What does unsupported message role: developer mean in Pi?
&lt;/h3&gt;

&lt;p&gt;It means the model entry has reasoning set to true, so Pi sends the system prompt as a developer role message, and something upstream refuses that role. Add compat with supportsDeveloperRole set to false to keep thinking, or set reasoning to false to drop it. On an OpenAI-shaped gateway this shows up only on Anthropic-family models.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://ofox.ai/blog/pi-coding-agent-custom-provider-setup-2026/" rel="noopener noreferrer"&gt;ofox.ai/blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>piagentharness</category>
      <category>codingagents</category>
      <category>apiguide</category>
    </item>
    <item>
      <title>DeepSeek Harness (dsh): Version, Updates, Stability (2026)</title>
      <dc:creator>Owen</dc:creator>
      <pubDate>Wed, 19 Aug 2026 01:10:28 +0000</pubDate>
      <link>https://dev.to/owen_fox/deepseek-harness-dsh-version-updates-stability-2026-5hf3</link>
      <guid>https://dev.to/owen_fox/deepseek-harness-dsh-version-updates-stability-2026-5hf3</guid>
      <description>&lt;h1&gt;
  
  
  DeepSeek Harness (dsh): Version, Updates, Stability (2026)
&lt;/h1&gt;

&lt;p&gt;DeepSeek Harness remains at version 0.1.0-rc.6, its launch build from August 13, 2026. The public repository has received no commits since that date, with no releases, tags, or active issue tracking. Pin your version before building anything with this tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  Current Status Summary
&lt;/h2&gt;

&lt;p&gt;The tool launched with six published versions between August 10-13, 2026, all available on npm under both &lt;code&gt;latest&lt;/code&gt; and &lt;code&gt;next&lt;/code&gt; dist-tags. The only branch in the repository is &lt;code&gt;master&lt;/code&gt;, and discussions have grown to 2,713+ threads while the codebase remains static since launch.&lt;/p&gt;

&lt;p&gt;Key facts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Current version:&lt;/strong&gt; 0.1.0-rc.6 (August 13, 2026 12:35 UTC)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Total versions published:&lt;/strong&gt; 6&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Last public commit:&lt;/strong&gt; August 13, 2026 11:38 UTC&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Releases/tags:&lt;/strong&gt; 0 / 0&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;License:&lt;/strong&gt; MIT&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Official status:&lt;/strong&gt; Developer preview with promised breaking changes&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Entry point:&lt;/strong&gt; &lt;code&gt;npx @deepseek-ai/dsh web&lt;/code&gt; on 127.0.0.1:3080&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Version History
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Version&lt;/th&gt;
&lt;th&gt;Published (UTC)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;0.0.1-rc.1&lt;/td&gt;
&lt;td&gt;2026-08-10 19:41&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;0.0.1-rc.2&lt;/td&gt;
&lt;td&gt;2026-08-11 15:24&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;0.0.1-rc.5&lt;/td&gt;
&lt;td&gt;2026-08-12 22:36&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;0.1.0-rc.2&lt;/td&gt;
&lt;td&gt;2026-08-13 09:48&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;0.1.0-rc.3&lt;/td&gt;
&lt;td&gt;2026-08-13 11:16&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;0.1.0-rc.6&lt;/td&gt;
&lt;td&gt;2026-08-13 12:35&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Every release predates or matches the public launch date.&lt;/p&gt;

&lt;h2&gt;
  
  
  Checking Status Independently
&lt;/h2&gt;

&lt;p&gt;Three commands verify current state without requiring a GitHub account:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm view @deepseek-ai/dsh dist-tags versions &lt;span class="nb"&gt;time&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; https://api.github.com/repos/deepseek-ai/deepseek-harness &lt;span class="se"&gt;\&lt;/span&gt;
  | jq &lt;span class="s1"&gt;'{pushed_at, has_issues, open_issues_count}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; https://api.github.com/repos/deepseek-ai/deepseek-harness/releases
curl &lt;span class="nt"&gt;-s&lt;/span&gt; https://api.github.com/repos/deepseek-ai/deepseek-harness/tags
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;As of August 17, these commands show no movement since the initial publication.&lt;/p&gt;

&lt;h2&gt;
  
  
  Production Readiness
&lt;/h2&gt;

&lt;p&gt;The project explicitly states it remains "in &lt;em&gt;developer preview&lt;/em&gt; and is iterating rapidly. &lt;strong&gt;THERE WILL BE COMPATIBILITY-BREAKING CHANGES.&lt;/strong&gt;" This is not merely marketing language — structural choices reinforce it.&lt;/p&gt;

&lt;p&gt;The repository disables issues entirely, routing all communication through Discussions and Discord. That works for signal gathering but provides no mechanism for tracking a regression you filed. The plugin architecture, built on Cordis, means breaking changes upstream can land in a surface you did not know you depended on.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Paradox of Transparency and Opacity
&lt;/h2&gt;

&lt;p&gt;The repository was created August 13 but arrived with complete history: 12,293 commits reaching back to June 10, 2026. The first commit reads "Initialize repo with README, AGENTS.md, and CLAUDE.md symlink." That is two months of visible development history.&lt;/p&gt;

&lt;p&gt;What stayed private is the surrounding process. Merge commits reference pull requests (#2519, #2520, #2521) in a separate, non-public &lt;code&gt;deepseek-harness&lt;/code&gt; organization. No branches exist beyond &lt;code&gt;master&lt;/code&gt;, and nothing has shipped since launch day.&lt;/p&gt;

&lt;p&gt;This creates an unusual situation: the code's development path is transparent while its future direction is opaque. "Iterating rapidly" is likely accurate but unobservable from outside. For consumers, that means you cannot distinguish "quiet because stable" from "quiet because actively developing."&lt;/p&gt;

&lt;h2&gt;
  
  
  Community Size and Patterns
&lt;/h2&gt;

&lt;p&gt;GitHub Discussions grew from roughly 620 threads three days earlier to 2,713 by August 17. Common themes:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Theme&lt;/th&gt;
&lt;th&gt;Requests&lt;/th&gt;
&lt;th&gt;Implication&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Native client instead of browser&lt;/td&gt;
&lt;td&gt;Real CLI, desktop app, editor extensions&lt;/td&gt;
&lt;td&gt;Web-UI-first design is the most contested choice&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Install failures&lt;/td&gt;
&lt;td&gt;Windows, Arch Linux, Termux reports&lt;/td&gt;
&lt;td&gt;Non-macOS platforms are unverified&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Third-party plugins&lt;/td&gt;
&lt;td&gt;Marketplaces, knowledge bases, UI skins&lt;/td&gt;
&lt;td&gt;Ecosystem outrunning core&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost visibility&lt;/td&gt;
&lt;td&gt;Per-turn tokens, peak-hour indicators&lt;/td&gt;
&lt;td&gt;Shipped build lacks this&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sandbox and permissions&lt;/td&gt;
&lt;td&gt;Multiple independent reports&lt;/td&gt;
&lt;td&gt;Preview disclaimer applies&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The ecosystem is expanding faster than the core, which carries a consequence: when plugins ship faster than the foundation they depend on, plugins are the first casualty when breaking changes arrive.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cost Transparency Gap
&lt;/h2&gt;

&lt;p&gt;The shipped build provides no per-turn token count or cost display. A highly upvoted Ideas thread filed August 14 requests exactly this. A community cost-tracker plugin appeared on launch day to fill the gap.&lt;/p&gt;

&lt;p&gt;The timing matters. DeepSeek shifted its entire V4 family to peak and off-peak pricing on August 16, 2026. Cache hits on &lt;code&gt;deepseek/deepseek-v4-pro&lt;/code&gt; increased 12.1x at peak; &lt;code&gt;deepseek/deepseek-v4-flash&lt;/code&gt; rose 5x. An agent harness that resends identical system prompts and file contents every turn is precisely the workload that pricing change targets.&lt;/p&gt;

&lt;p&gt;Until the harness surfaces cost data, you consult provider dashboards, proxy logs, or community plugins. None ship with dsh.&lt;/p&gt;

&lt;h2&gt;
  
  
  When dsh Is Safe to Use
&lt;/h2&gt;

&lt;p&gt;The practical distinction is what breaks when the next version appears:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Safe to explore:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Evaluating the plugin model on scratch repositories&lt;/li&gt;
&lt;li&gt;Comparing DeepSeek models against your own tasks&lt;/li&gt;
&lt;li&gt;Building or testing &lt;code&gt;dsh-plugin&lt;/code&gt; projects&lt;/li&gt;
&lt;li&gt;Local experiments you can redo in an afternoon&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Hold off on:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Anything with a delivery date attached&lt;/li&gt;
&lt;li&gt;Shared team configurations others will inherit&lt;/li&gt;
&lt;li&gt;CI or unattended scheduled jobs&lt;/li&gt;
&lt;li&gt;Work where an unexpected sandbox change would be costly&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The dividing line is whether an upgrade that breaks your setup costs you an afternoon or a deadline. DeepSeek has told you in capital letters which it is planning.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recommended Setup Approach
&lt;/h2&gt;

&lt;p&gt;Before integrating dsh into production workflows:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Pin the version:&lt;/strong&gt; Use &lt;code&gt;npx @deepseek-ai/dsh@0.1.0-rc.6&lt;/code&gt; exclusively, never bare &lt;code&gt;dsh&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Control the endpoint:&lt;/strong&gt; Configure &lt;code&gt;DEEPSEEK_BASE_URL&lt;/code&gt; and &lt;code&gt;DEEPSEEK_API_KEY&lt;/code&gt; pointing to a base URL you manage, not a hardcoded vendor default&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Maintain a backup harness:&lt;/strong&gt; Keep a second working harness for deadline-critical work&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Monitor npm, not GitHub:&lt;/strong&gt; The next version appears on npm first&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Review release notes:&lt;/strong&gt; Reread the README after any upgrade, since breaking changes are announced there and nowhere else&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Point 2 carries particular weight. dsh reads environment variables for its built-in route, and the custom-provider form accepts any OpenAI-compatible endpoint. That swappability only works if you configure it from the start. On ofox, that means &lt;code&gt;deepseek/deepseek-v4-flash&lt;/code&gt; and &lt;code&gt;deepseek/deepseek-v4-pro&lt;/code&gt; behind a single key, so on the day dsh breaks or pricing moves again, the model string is the only thing you edit.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What is the current version?&lt;/strong&gt;&lt;br&gt;
0.1.0-rc.6, published August 13, 2026. Both npm's &lt;code&gt;latest&lt;/code&gt; and &lt;code&gt;next&lt;/code&gt; tags point to it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Are there GitHub releases or tags?&lt;/strong&gt;&lt;br&gt;
No. Zero releases and zero tags exist. Version history lives entirely on npm.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why are issues disabled?&lt;/strong&gt;&lt;br&gt;
The maintainers route all communication through GitHub Discussions and Discord instead, eliminating formal issue tracking.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is dsh a fork?&lt;/strong&gt;&lt;br&gt;
Neither fork nor mirror. The repository was created August 13 with full commit history from June 10, but the surrounding development process remains private.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does it run on Windows or Linux?&lt;/strong&gt;&lt;br&gt;
It ships as a Node package, so it starts anywhere Node runs. However, community reports of install failures on Windows, Arch Linux, and Termux are frequent enough to treat non-macOS platforms as unverified.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I pin to a known-good version?&lt;/strong&gt;&lt;br&gt;
Yes: &lt;code&gt;npx @deepseek-ai/dsh@0.1.0-rc.6&lt;/code&gt;. Because no upstream tags exist, the npm version string is your only safeguard against promised breaking changes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is dsh free?&lt;/strong&gt;&lt;br&gt;
The harness itself is MIT licensed and free. Inference costs depend on your provider's rates. DeepSeek's pricing moved to a peak/off-peak structure on August 16, 2026.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/deepseek-ai/deepseek-harness" rel="noopener noreferrer"&gt;deepseek-ai/deepseek-harness on GitHub&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/deepseek-ai/deepseek-harness/blob/master/README.md" rel="noopener noreferrer"&gt;dsh README, developer preview notice&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/deepseek-ai/deepseek-harness/discussions" rel="noopener noreferrer"&gt;DeepSeek Harness Discussions&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://registry.npmjs.org/@deepseek-ai/dsh" rel="noopener noreferrer"&gt;@deepseek-ai/dsh on npm registry&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/cordiverse/cordis" rel="noopener noreferrer"&gt;Cordis plugin runtime&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://api-docs.deepseek.com/quick_start/pricing" rel="noopener noreferrer"&gt;DeepSeek API pricing&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;




&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://ofox.ai/blog/deepseek-harness-dsh-version-updates-stability-production-2026/" rel="noopener noreferrer"&gt;ofox.ai/blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>deepseek</category>
      <category>deepseekharness</category>
      <category>codingagents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>DeepSeek API Price Increase: Up to 12x, Peak Hours (2026)</title>
      <dc:creator>Owen</dc:creator>
      <pubDate>Tue, 18 Aug 2026 13:42:02 +0000</pubDate>
      <link>https://dev.to/owen_fox/deepseek-api-price-increase-up-to-12x-peak-hours-2026-1hgg</link>
      <guid>https://dev.to/owen_fox/deepseek-api-price-increase-up-to-12x-peak-hours-2026-1hgg</guid>
      <description>&lt;h1&gt;
  
  
  DeepSeek API Price Increase: Up to 12x, Peak Hours (2026)
&lt;/h1&gt;

&lt;p&gt;DeepSeek implemented significant API pricing changes effective August 16, 2026 at 16:00 UTC, introducing a peak/off-peak billing model replacing previous flat rates. The most dramatic increase affected V4 Pro cache hits, which jumped 12.1x at peak ($0.003625 to $0.044 per million tokens).&lt;/p&gt;

&lt;h2&gt;
  
  
  Key Changes
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Effective Date:&lt;/strong&gt; 2026-08-16, 16:00 UTC (2026-08-17, 00:00 Beijing time)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Peak Hours (7 hours daily):&lt;/strong&gt; 01:00–04:00 and 06:00–10:00 UTC&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Billing Model:&lt;/strong&gt; Off-peak rates are 50% of peak rates across all tiers&lt;/p&gt;

&lt;h3&gt;
  
  
  Pricing Comparison Table
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tier&lt;/th&gt;
&lt;th&gt;Old Flat&lt;/th&gt;
&lt;th&gt;New Off-peak&lt;/th&gt;
&lt;th&gt;New Peak&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;V4 Flash cache hit&lt;/td&gt;
&lt;td&gt;$0.0028&lt;/td&gt;
&lt;td&gt;$0.007&lt;/td&gt;
&lt;td&gt;$0.014&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;V4 Flash cache miss&lt;/td&gt;
&lt;td&gt;$0.14&lt;/td&gt;
&lt;td&gt;$0.22&lt;/td&gt;
&lt;td&gt;$0.44&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;V4 Flash output&lt;/td&gt;
&lt;td&gt;$0.28&lt;/td&gt;
&lt;td&gt;$0.66&lt;/td&gt;
&lt;td&gt;$1.32&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;V4 Pro cache hit&lt;/td&gt;
&lt;td&gt;$0.003625&lt;/td&gt;
&lt;td&gt;$0.022&lt;/td&gt;
&lt;td&gt;$0.044&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;V4 Pro cache miss&lt;/td&gt;
&lt;td&gt;$0.435&lt;/td&gt;
&lt;td&gt;$0.66&lt;/td&gt;
&lt;td&gt;$1.32&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;V4 Pro output&lt;/td&gt;
&lt;td&gt;$0.87&lt;/td&gt;
&lt;td&gt;$1.98&lt;/td&gt;
&lt;td&gt;$3.96&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Price Increase Multipliers
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tier&lt;/th&gt;
&lt;th&gt;Off-peak vs Old&lt;/th&gt;
&lt;th&gt;Peak vs Old&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;V4 Flash cache hit&lt;/td&gt;
&lt;td&gt;2.5x&lt;/td&gt;
&lt;td&gt;5.0x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;V4 Flash cache miss&lt;/td&gt;
&lt;td&gt;1.6x&lt;/td&gt;
&lt;td&gt;3.1x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;V4 Flash output&lt;/td&gt;
&lt;td&gt;2.4x&lt;/td&gt;
&lt;td&gt;4.7x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;V4 Pro cache hit&lt;/td&gt;
&lt;td&gt;6.1x&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;12.1x&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;V4 Pro cache miss&lt;/td&gt;
&lt;td&gt;1.5x&lt;/td&gt;
&lt;td&gt;3.0x&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;V4 Pro output&lt;/td&gt;
&lt;td&gt;2.3x&lt;/td&gt;
&lt;td&gt;4.6x&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Real-World Impact: Sample Monthly Bill
&lt;/h2&gt;

&lt;p&gt;For a typical agent workload (300M input tokens at 90% cache hit rate, 20M output):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scenario&lt;/th&gt;
&lt;th&gt;Cache Hits&lt;/th&gt;
&lt;th&gt;Cache Misses&lt;/th&gt;
&lt;th&gt;Output&lt;/th&gt;
&lt;th&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Old flat rate&lt;/td&gt;
&lt;td&gt;$0.76&lt;/td&gt;
&lt;td&gt;$4.20&lt;/td&gt;
&lt;td&gt;$5.60&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$10.56&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;New, all off-peak&lt;/td&gt;
&lt;td&gt;$1.89&lt;/td&gt;
&lt;td&gt;$6.60&lt;/td&gt;
&lt;td&gt;$13.20&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$21.69&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;New, all peak&lt;/td&gt;
&lt;td&gt;$3.78&lt;/td&gt;
&lt;td&gt;$13.20&lt;/td&gt;
&lt;td&gt;$26.40&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$43.38&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Peak Hours by Time Zone (Aug 2026)
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Location&lt;/th&gt;
&lt;th&gt;Peak Window 1&lt;/th&gt;
&lt;th&gt;Peak Window 2&lt;/th&gt;
&lt;th&gt;Workday Impact&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Beijing/Singapore (UTC+8)&lt;/td&gt;
&lt;td&gt;09:00–12:00&lt;/td&gt;
&lt;td&gt;14:00–18:00&lt;/td&gt;
&lt;td&gt;7 of 9 hours&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tokyo/Seoul (UTC+9)&lt;/td&gt;
&lt;td&gt;10:00–13:00&lt;/td&gt;
&lt;td&gt;15:00–19:00&lt;/td&gt;
&lt;td&gt;6 of 9 hours&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;India (UTC+5:30)&lt;/td&gt;
&lt;td&gt;06:30–09:30&lt;/td&gt;
&lt;td&gt;11:30–15:30&lt;/td&gt;
&lt;td&gt;4.5 of 9 hours&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Berlin/Paris (UTC+2, CEST)&lt;/td&gt;
&lt;td&gt;03:00–06:00&lt;/td&gt;
&lt;td&gt;08:00–12:00&lt;/td&gt;
&lt;td&gt;3 of 9 hours&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;London (UTC+1, BST)&lt;/td&gt;
&lt;td&gt;02:00–05:00&lt;/td&gt;
&lt;td&gt;07:00–11:00&lt;/td&gt;
&lt;td&gt;2 of 9 hours&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;New York (UTC-4, EDT)&lt;/td&gt;
&lt;td&gt;21:00–00:00 prev&lt;/td&gt;
&lt;td&gt;02:00–06:00&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0 of 9 hours&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;San Francisco (UTC-7, PDT)&lt;/td&gt;
&lt;td&gt;18:00–21:00 prev&lt;/td&gt;
&lt;td&gt;23:00–03:00&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0 of 9 hours&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Third-Party Hosts Comparison
&lt;/h2&gt;

&lt;h3&gt;
  
  
  V4 Flash (on OpenRouter, 2026-08-17)
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Host&lt;/th&gt;
&lt;th&gt;Input&lt;/th&gt;
&lt;th&gt;Output&lt;/th&gt;
&lt;th&gt;Cache Read&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;DeepInfra&lt;/td&gt;
&lt;td&gt;$0.08&lt;/td&gt;
&lt;td&gt;$0.18&lt;/td&gt;
&lt;td&gt;$0.016&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DigitalOcean&lt;/td&gt;
&lt;td&gt;$0.08&lt;/td&gt;
&lt;td&gt;$0.252&lt;/td&gt;
&lt;td&gt;$0.0252&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GMICloud&lt;/td&gt;
&lt;td&gt;$0.084&lt;/td&gt;
&lt;td&gt;$0.168&lt;/td&gt;
&lt;td&gt;$0.0168&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;DeepSeek&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.44&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$1.32&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.014&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Finding:&lt;/strong&gt; 26 of 28 third-party listings undercut DeepSeek's peak input rate.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Break-even analysis:&lt;/strong&gt; At 94% cache hit rate off-peak, DeepInfra becomes cheaper; at 99.4% peak. Previously, this threshold was 77%.&lt;/p&gt;

&lt;h3&gt;
  
  
  V4 Pro
&lt;/h3&gt;

&lt;p&gt;DeepSeek's off-peak rate ($0.66 input) undercuts all nine available third-party listings.&lt;/p&gt;

&lt;h2&gt;
  
  
  Historical Context
&lt;/h2&gt;

&lt;p&gt;DeepSeek has adjusted pricing three times in eighteen months:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Date&lt;/th&gt;
&lt;th&gt;Change&lt;/th&gt;
&lt;th&gt;Direction&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2025-02-26&lt;/td&gt;
&lt;td&gt;Off-peak discounts introduced (16:30–00:30 UTC)&lt;/td&gt;
&lt;td&gt;Down&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2025-09-05&lt;/td&gt;
&lt;td&gt;New pricing, off-peak discounts ended&lt;/td&gt;
&lt;td&gt;Up&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2026-08-16&lt;/td&gt;
&lt;td&gt;Peak/off-peak billing (off-peak = 50% of peak)&lt;/td&gt;
&lt;td&gt;Up&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This represents the first instance where time-of-day pricing increased rates rather than offered discounts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Cache Hits Increased Most
&lt;/h2&gt;

&lt;p&gt;The previous cache hit rate was an outlier: V4 Flash hits at $0.0028 were 50x cheaper than a miss. That disproportionate pricing supported the cost model for agent loops with 90%+ cache hit rates. The new structure reflects market realignment, where caching remains valuable but no longer dominates billing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Caching Still Viable?
&lt;/h2&gt;

&lt;p&gt;Yes. A cache hit remains approximately 31x cheaper than a miss at peak ($0.014 vs $0.44 for Flash). However, this advantage narrowed from the previous 50x ratio. Output tokens now represent a larger portion of total costs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key Considerations
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Timestamp Ambiguity:&lt;/strong&gt; DeepSeek has not clarified whether requests billing during peak windows are determined by start time or completion time. Requests crossing 04:00 or 10:00 UTC boundaries represent an open question.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Off-Peak as Partial Relief:&lt;/strong&gt; Off-peak rates represent a smaller increase, not a true discount. Even at off-peak, V4 Flash output costs $0.66 (2.4x the old flat $0.28).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No Grandfather Clause:&lt;/strong&gt; Existing balances top up at new rates; no legacy pricing applies.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Gateway Strategy:&lt;/strong&gt; Using an OpenAI-compatible API aggregator allows switching between providers via model string changes alone, avoiding lock-in to any single vendor's pricing schedule.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Does the increase differ between V4 Flash and V4 Pro?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Yes substantially. Cache hit increases ranged from 5x (Flash) to 12.1x (Pro) at peak.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How does billing handle requests crossing peak boundaries?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Unknown. DeepSeek's documentation omits this; measure actual invoices for verification.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Are third-party hosts now cheaper?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For V4 Flash: nearly uniformly yes (26 of 28 endpoints). For V4 Pro: no — DeepSeek undercuts all current alternatives even at off-peak rates.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is prompt caching still worthwhile?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Absolutely, though with diminished leverage. The economics shifted toward output token optimization.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Did previous balances retain old rates?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;No documented grandfather provision exists. All consumption applies current pricing.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published on &lt;a href="https://ofox.ai/blog/deepseek-api-price-increase-new-rates-peak-hours-cache-cost-2026/" rel="noopener noreferrer"&gt;ofox.ai/blog&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>deepseek</category>
      <category>apipricing</category>
      <category>pricing</category>
      <category>costoptimization</category>
    </item>
  </channel>
</rss>
