<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: MT_Notes</title>
    <description>The latest articles on DEV Community by MT_Notes (@mt_notes).</description>
    <link>https://dev.to/mt_notes</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4056352%2Faad51cd1-ecc7-4a90-9d4e-e49ff9b8e23c.png</url>
      <title>DEV Community: MT_Notes</title>
      <link>https://dev.to/mt_notes</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/mt_notes"/>
    <language>en</language>
    <item>
      <title>Space Bunny Tops OpenRouter: A Public Stress Test, and a Lesson for the Routing Layer</title>
      <dc:creator>MT_Notes</dc:creator>
      <pubDate>Mon, 28 Sep 2026 10:39:41 +0000</pubDate>
      <link>https://dev.to/mt_notes/space-bunny-tops-openrouter-a-public-stress-test-and-a-lesson-for-the-routing-layer-akk</link>
      <guid>https://dev.to/mt_notes/space-bunny-tops-openrouter-a-public-stress-test-and-a-lesson-for-the-routing-layer-akk</guid>
      <description>&lt;h2&gt;
  
  
  The Backdrop
&lt;/h2&gt;

&lt;p&gt;In late September, an anonymous model quietly appeared on OpenRouter and OpenCode: stealth/space-bunny-alpha. No vendor name, no model card, zero pricing — yet it shipped with a one-million-token context window, text, image and video input, tool calling, and adjustable reasoning effort. Within days it climbed to the top of daily usage on both platforms. The community traced its fingerprint through tokenizer behavior and error formats, and nearly everything points to MiniMax's unreleased M3.1 — which MiniMax has never confirmed. On September 27, MiniMax switched on M3.1-Flash-Preview inside its own coding agent, MiniMax Code: again no model card, no public API, no price. This complete anonymous pre-launch trail deserves a close look from anyone orchestrating multiple models.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fingerprinting: How the Community Pinned Down an Anonymous Model
&lt;/h2&gt;

&lt;p&gt;Within hours of the stealth launch, developers ran three kinds of fingerprint tests. First, tokenizer comparison: feeding the same text corpus to Space Bunny and MiniMax's known model family and checking whether token splits match. One study on OpenCode found a 24/24 token match against the MiniMax family; a broader measurement set reported 50/50. Second, adversarial probes: deliberately malformed inputs to compare error formats and truncation behavior. Third, breadcrumbs in official code: MiniMax's open-source MiniMax Code repository already contained test code referencing MiniMax-M3.1, a one-million-token context, and low/high/max effort tiers — closely matching Space Bunny's published traits. The evidence chain is strong, but the technically correct description remains: a MiniMax-family fingerprint, specific version unconfirmed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Vendors Love Going Anonymous
&lt;/h2&gt;

&lt;p&gt;The pattern is no accident. Launching anonymously and for free outsources stress testing to real traffic across the internet: no model card means no accountability for benchmark numbers; no price means no backlash when the free window ends; and if the model underperforms, pulling the anonymous route leaves no public record. For Chinese labs, OpenRouter has become a de facto pre-launch hotline. MiniMax's M3, open-sourced in June, is a natively multimodal MoE with 428 billion total parameters and 23 billion active, scoring 80.5% on SWE-bench Verified and 59.0% on the harder SWE-Bench Pro — VentureBeat reported it beat GPT-5.5 and Gemini 3.1 Pro on that benchmark at 5–10% of the cost. With that track record, using free anonymous routes to collect real-world load data for M3.1 makes perfect sense. Meanwhile the host platform itself just leveled up: OpenRouter was announced as acquired by Stripe, with media reports putting the figure around $7.5 billion, cementing the aggregation layer's position as the traffic gateway.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three Ledges Callers Must Clear
&lt;/h2&gt;

&lt;p&gt;For anyone wiring an anonymous endpoint into production, the risks concentrate in three places. First, mutable identity: today it points to M3.1; tomorrow the underlying checkpoint may silently change, invalidating every performance conclusion you measured. Second, data terms: OpenRouter's page explicitly warns that the anonymous provider may retain prompts and completions, while OpenCode's free route is separately labeled zero-retention and no-training — same model, two routes, completely different terms. Third, uncertain lifetime: the endpoint may vanish when the preview window ends, and a free price list can reset to nothing at any time. The community consensus is pragmatic: treat Space Bunny as an evaluation endpoint, benchmark it against your workhorse models on your own repositories, but never bind your critical path to it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where a Unified Gateway Fits: Evaluation and Fallback, Each in Its Place
&lt;/h2&gt;

&lt;p&gt;Endpoints like this — unidentified, disposable, free but with unclear terms — are exactly why multi-model calls need a unified interface. Take router.accels.tech as an example. Accels, a Singapore-based company, focuses on three things: stability (workhorse models run on official and reliable hosted endpoints with automatic failover), model coverage (mainstream and hot new models live in a single catalog, wired in as they launch), and unified billing (one key, one invoice — no registering and reconciling across every platform). A fitting use for this very topic: evaluate the anonymous model and your primary model side by side through the unified interface, changing almost nothing in your code — just the model string.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://router.accels.tech/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;your-accels-key&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;stealth/space-bunny-alpha&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Review the boundary-condition handling in this code.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once evaluation passes, put the anonymous endpoint into the fallback chain: your primary model first, the free anonymous model after. If the stealth route goes offline, requests automatically land on the primary model and the business never notices. Evaluation, comparison, and fallback all happen under the same base_url.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;Anonymous launches will only multiply: vendors want free real-world load data, aggregation platforms want exclusive first-look stories, and both sides get what they need. For developers, the point is not guessing whether it really is M3.1 — it is making sure your invocation layer can swap, compare, and fall back at any time. Consolidate evaluation and routing behind one unified gateway, and switching models is just editing a string. Check whether your usual models are already on router.accels.tech, and toss the next anonymous hit into your evaluation queue while you are at it. If this analysis helped, tell us in the comments whether you have ever wired in an anonymous endpoint.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sources&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Web Pulse — MiniMax slips a new coding model into its agent tool without a price tag: &lt;a href="https://wpnews.pro/news/minimax-slips-a-new-coding-model-into-its-agent-tool-without-a-price-tag" rel="noopener noreferrer"&gt;https://wpnews.pro/news/minimax-slips-a-new-coding-model-into-its-agent-tool-without-a-price-tag&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;BlockBeats — OpenRouter 上线新匿名模型 Space Bunny，多项指纹指向 MiniMax M3.1: &lt;a href="https://www.theblockbeats.info/flash/368770" rel="noopener noreferrer"&gt;https://www.theblockbeats.info/flash/368770&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;BuildFastWithAI — Space Bunny Review: 1M Context, Coding, Speed &amp;amp; Is It Worth Using?: &lt;a href="https://blog.buildfastwithai.com/space-bunny-review" rel="noopener noreferrer"&gt;https://blog.buildfastwithai.com/space-bunny-review&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;AGI Hunt — AI News Daily 2026-09-28: &lt;a href="https://agihunt.info/en/daily/2026-09-28" rel="noopener noreferrer"&gt;https://agihunt.info/en/daily/2026-09-28&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>apigateway</category>
      <category>accels</category>
    </item>
    <item>
      <title>Three New Models in 48 Hours: "Switching Models" Is Now a Four-Question Form</title>
      <dc:creator>MT_Notes</dc:creator>
      <pubDate>Thu, 24 Sep 2026 07:32:18 +0000</pubDate>
      <link>https://dev.to/mt_notes/three-new-models-in-48-hours-switching-models-is-now-a-four-question-form-1l9l</link>
      <guid>https://dev.to/mt_notes/three-new-models-in-48-hours-switching-models-is-now-a-four-question-form-1l9l</guid>
      <description>&lt;h2&gt;
  
  
  1. Background: the price tables are done, the incidents are not
&lt;/h2&gt;

&lt;p&gt;SpaceXAI shipped Grok 4.7 on September 21. Anthropic shipped Claude Opus 5.5 on September 22, and OpenAI followed roughly ninety minutes later with GPT-6 Sol and GPT-6 Luna. Three price tables in three days, and the press has covered the price war three times over. But the things that actually cause production incidents are not prices. They are four others: the request protocol, the long-context pricing cliff, the default effort level, and how many tokens a model actually burns.&lt;br&gt;
For anyone running multi-model routing and automatic fallback, this round of releases asks four fill-in-the-blank questions. None of them is about price, and any one of them can make a "50% cheaper" claim impossible to reconcile.&lt;/p&gt;
&lt;h2&gt;
  
  
  2. Technical Deep Dive
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;2.1 Protocol divergence: Opus 5.5's four 400s&lt;/strong&gt;&lt;br&gt;
Anthropic's own migration docs list four breaking changes, and the first three also apply to Fable 5.1:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F914sft3i7n1au3i688qk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F914sft3i7n1au3i688qk.png" alt=" " width="800" height="574"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;There is one more change that raises no error but alters the response shape: the short notes a model writes between tool calls now come back as progress-update thinking blocks, whose text is empty under the default display of omitted. A UI that streamed those notes as progress simply goes quiet between tool calls. Nothing errors; the user just stops seeing anything. To bring them back, set thinking.display to updates (the beta header thinking-display-updates-2026-08-18) or summarized.&lt;br&gt;
GPT-6's protocol constraint sits somewhere else. OpenAI's model pages point built-in tools and function calling at the Responses API; on Chat Completions, function calling requires reasoning effort to be none. That matters enormously to gateways: plenty of compatibility layers normalize every vendor onto /v1/chat/completions and fan out from there, so the assumption that only the model field needs to change breaks the moment tools are involved.&lt;br&gt;
&lt;strong&gt;2.2 Cliff divergence: three thresholds, three multipliers&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyye17u33sle52ygt5zmp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyye17u33sle52ygt5zmp.png" alt=" " width="800" height="458"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The two OpenAI models put the cliff at 272,000 tokens; Grok puts it at 200,000. That is a 72,000-token gap. The multipliers differ too: GPT-6 doubles input but only lifts output by half, while Grok doubles everything. Which means "which of these models has the cheaper input price" has no answer in long-context work unless you also say which band the request lands in. A 300K-input request bills at double on Sol, and at double on Grok 4.7 with output rising as well. This is a cliff, not a marginal rate: splitting 300K into two 150K requests is often cheaper than sending it once.&lt;br&gt;
One easily missed detail: Grok 4.7's cache read is $0.50 below 200K, but you have to ask for the hit. Set prompt_cache_key on the Responses API, or send the x-grok-conv-id header on Chat Completions. Without the hint, cache hits are unreliable, and cached input costs 75% less than fresh input.&lt;br&gt;
&lt;strong&gt;2.3 Default effort and token consumption&lt;/strong&gt;&lt;br&gt;
Opus 5.5's default effort dropped from high on Opus 5 to medium. GPT-6 Sol and Luna default to medium. Grok 4.7 defaults to high, yet several headline benchmarks are reported at xhigh. Three vendors, three different defaults, which means "send the same prompt to all three" is not a like-for-like comparison to begin with. Anthropic's migration doc is blunt about it: set effort explicitly first, then re-run your sweep.&lt;br&gt;
Token consumption deserves even more attention. Grok 4.7's per-token price is identical to Grok 4.6 — $$2 / $$6, with $$0.50 cache reads, all below 200K — but Artificial Analysis measured roughly 81,000 output tokens per Intelligence Index task in xHigh mode, against about 36,000 for Grok 4.6 in High mode. That is 125% more output. On the same index, a task works out to about $$3.74 on Grok 4.7 against about $1.99 on GPT-5.6 Sol. The unit price did not rise; the unit task did. Since reasoning tokens bill as output, the effort setting is currently the single biggest cost control you have.&lt;br&gt;
&lt;strong&gt;2.4 Measured: one slug, five endpoints, 1.75x spread in what you actually pay&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcgmllbhdqsodcs0tmn9y.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcgmllbhdqsodcs0tmn9y.png" alt=" " width="800" height="370"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;These figures come from OpenRouter's Claude Opus 5.5 model page, over a window that includes September 23. All five endpoints list exactly the same price, yet cache hit rates run from 80.6% to 91.7% and the effective input price runs from $$0.580 to $$1.016 — a 1.75x spread. The page's overall weighted average input price is $$0.8623, under a quarter of the $$4 list price, while the weighted average output price still sits at $$20.13 because output carries no cache discount. Several endpoints are excluded from the default routing pool, including a $$4.40 / $$22 tier priced 10% above list and a 2x Anthropic Fast tier at $$8 / $40.&lt;br&gt;
The same page also reports availability: 99.92% for the model, against 98.48% without routing. List price is a horizontal line; what you actually pay is a distribution. Cost decisions require holding both variables at once — which endpoint served the request, and whether the cache hit.&lt;br&gt;
&lt;strong&gt;2.5 In practice: turn the four questions into one table&lt;/strong&gt;&lt;br&gt;
In code, the four questions collapse into a capability table plus a small rewrite step:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;from openai import OpenAI

client = OpenAI(base_url="https://router.accels.tech/v1", api_key="your-accels-key")

# Protocol constraints, cliff thresholds and default effort live with the model ID
MODELS = {
    "claude-opus-5-5": {"effort": "medium", "forced_tool": False, "cliff": None},
    "gpt-6-sol":       {"effort": "medium", "forced_tool": True,  "cliff": 272_000},
    "grok-4.7":        {"effort": "high",   "forced_tool": True,  "cliff": 200_000},
}

def build(model, messages, tools=None, must_call_tool=False, effort=None):
    caps = MODELS[model]
    kwargs = {
        "model": model,
        "messages": messages,
        "output_config": {"effort": effort or caps["effort"]},
    }
    if tools:
        kwargs["tools"] = tools
        # Fall back to auto where forcing is unsupported, then validate the result
        kwargs["tool_choice"] = "required" if (must_call_tool and caps["forced_tool"]) else "auto"
    return kwargs
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three takeaways. Set effort explicitly rather than trusting a default, because the vendors disagree on what the default is. auto does not guarantee a tool call, so check and retry. And check the input token count against the cliff before sending — if you are over the line, split the request or compact the history rather than waiting for the vendor to warn you.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Where It Lands: Four Questions, One Entry Point
&lt;/h2&gt;

&lt;p&gt;Every question above points at the same thing: switching models is no longer a string edit. It means maintaining protocol compatibility, cliff thresholds, effort levels and per-endpoint effective pricing at the same time. That is exactly why a model gateway like router.accels.tech has become practical in this release cycle. Accels is a Singapore-based company, and it puts multiple vendors' models behind a single OpenAI-compatible entry point:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Stable: one base_url. New models don't mean maintaining a separate SDK and key rotation per vendor, and when an endpoint errors or rate-limits, traffic can move to a healthy one. Your regression script only touches the model field.&lt;/li&gt;
&lt;li&gt;Complete model coverage: Claude, GPT, Grok and other mainstream models sit behind one entry point, so the MODELS table above can cover every candidate in your fallback chain — cliff parameters like 272K and 200K declared alongside everything else.&lt;/li&gt;
&lt;li&gt;Unified billing: usage across vendors lands on one bill, so comparing what "Opus 5.5 at medium" and "Grok 4.7 at xhigh" actually cost on the same task set doesn't require reconciling several invoices and several caching conventions.
A usage pattern that fits today's topic: sample 200 real production requests, send them through the single entry point to all three models at low, medium and high effort, then measure the 400 rate, tool-call hit rate, cache hit rate and cost per task, and let the numbers set your migration pace:
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;import concurrent.futures
from openai import OpenAI

client = OpenAI(base_url="https://router.accels.tech/v1", api_key="your-accels-key")
MODELS = ["claude-opus-5-5", "gpt-6-sol", "grok-4.7"]
EFFORTS = ["low", "medium", "high"]

def run(model, effort, sample):
    resp = client.chat.completions.create(
        model=model,
        messages=sample["messages"],
        output_config={"effort": effort},
    )
    u = resp.usage
    return {
        "model": model, "effort": effort,
        "input": u.prompt_tokens, "output": u.completion_tokens,
        "cached": u.prompt_tokens_details.cached_tokens,
    }

jobs = [(m, e, s) for m in MODELS for e in EFFORTS for s in samples[:200]]
with concurrent.futures.ThreadPoolExecutor(max_workers=8) as pool:
    rows = list(pool.map(lambda a: run(*a), jobs))

# Aggregate by (model, effort): cache hit rate, cost per task, 400s, tool-call failures
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Check the console's model list for the exact model IDs and available fields.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Closing
The real signal in this round is not in the price tables. Opus 5.5 made thinking a default that cannot be switched off, GPT-6 tied function calling to the Responses API, and Grok 4.7 traded an unchanged unit price for more output tokens. Each vendor moved one step — on protocol, on cliffs, on effort, on token burn — and getting any one of those four questions wrong makes the "half the price" arithmetic unreconcilable. Rather than firefighting every upgrade, write the capability table, the cliff thresholds and the adaptation layer into your router now. If you are about to migrate, run the old and new models side by side through router.accels.tech with a single key before you move production traffic.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Sources&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Anthropic: What's new in Claude Opus 5.5 — &lt;a href="https://platform.claude.com/docs/en/models/opus-5-5/whats-new-opus-5-5" rel="noopener noreferrer"&gt;https://platform.claude.com/docs/en/models/opus-5-5/whats-new-opus-5-5&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Anthropic Release Notes (Opus 5.5 launch: $$4 / $$20, 1M context, 128k output, the four 400s) — &lt;a href="https://releasebot.io/updates/anthropic" rel="noopener noreferrer"&gt;https://releasebot.io/updates/anthropic&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;OpenRouter: Claude Opus 5.5 model page (endpoint pricing, cache hit rates, weighted effective price, availability) — &lt;a href="https://openrouter.ai/anthropic/claude-opus-5-5" rel="noopener noreferrer"&gt;https://openrouter.ai/anthropic/claude-opus-5-5&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;OpenAI: GPT-6 Sol model page (pricing and long-context tier) — &lt;a href="https://developers.openai.com/api/docs/models/gpt-6-sol" rel="noopener noreferrer"&gt;https://developers.openai.com/api/docs/models/gpt-6-sol&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;S5 Labs: GPT-6 Sol and Luna — Prices and API Differences (272K cliff, Chat Completions function-calling constraint) — &lt;a href="https://s5labs.io/resources/insights/gpt-6-sol-luna-pricing-agent-workflows" rel="noopener noreferrer"&gt;https://s5labs.io/resources/insights/gpt-6-sol-luna-pricing-agent-workflows&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;DataNorth: SpaceXAI releases Grok 4.7 ($$2 / $$6, 500K, Fast tier, output tokens per task) — &lt;a href="https://datanorth.ai/news/spacexai-releases-grok-4-7" rel="noopener noreferrer"&gt;https://datanorth.ai/news/spacexai-releases-grok-4-7&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Codersera: Grok 4.7 Complete Guide (200K cliff, $0.50 cache read, caching hints) — &lt;a href="https://codersera.com/blog/grok-4-7-complete-guide-2026" rel="noopener noreferrer"&gt;https://codersera.com/blog/grok-4-7-complete-guide-2026&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>apigateway</category>
    </item>
    <item>
      <title>Two Price Sheets in One Day, the Same $0.20 for Cache Reads: Why Multi-Model Routing Math Just Changed</title>
      <dc:creator>MT_Notes</dc:creator>
      <pubDate>Thu, 24 Sep 2026 07:27:18 +0000</pubDate>
      <link>https://dev.to/mt_notes/two-price-sheets-in-one-day-the-same-020-for-cache-reads-why-multi-model-routing-math-just-27ok</link>
      <guid>https://dev.to/mt_notes/two-price-sheets-in-one-day-the-same-020-for-cache-reads-why-multi-model-routing-math-just-27ok</guid>
      <description>&lt;h2&gt;
  
  
  1. Background
&lt;/h2&gt;

&lt;p&gt;On the afternoon of Tuesday, September 22, Anthropic and OpenAI shipped new models within about an hour of each other: Claude Opus 5.5 from Anthropic, and GPT-6 Sol plus GPT-6 Luna from OpenAI. Both labs put the price cut in the headline:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;OpenAI priced Sol at $2 input / $10 output per million tokens and Luna at $0.10 / $0.50, a 50% cut against GPT-5.6 promotional pricing. The flagship Astra is unchanged at $10 / $50. OpenAI confirmed to the press that these are permanent prices, not promotional ones.&lt;/li&gt;
&lt;li&gt;Anthropic priced Opus 5.5 at $4 / $20 (Opus 5 was $5 / $25), cut cache reads from $0.50 to $0.20, and says typical workloads cost 40% less than Opus 5 at default settings.
Most coverage concluded that "Opus 5.5 is still twice the price of Sol." For coding agents and long conversations, that is only half right. One number is identical on both sheets: cache reads cost $0.20 per million tokens. And both labs say plainly that cache reads are where most agentic and coding spend goes.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  2. The Technical Picture
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;2.1 Three price columns side by side, and where cache prices come from&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyittyaa8jy9dlki546k9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyittyaa8jy9dlki546k9.png" alt=" " width="798" height="223"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;(Per million tokens.)&lt;br&gt;
On input and output, Opus 5.5 is exactly 2x Sol. On cache reads, they are equal. On cache writes, both charge 125% of the input rate: $2.50 for Sol, $5.00 for Opus 5.5.&lt;br&gt;
The number worth remembering is the ratio between cache reads and the input price. The whole GPT-6 family sets it at 10% of input, a 90% discount. Opus 5.5 sets it at 5% of input, a 95% discount. So the $0.20 collision isn't coordination between the two labs. It's two independent policies -- "input costs twice as much" and "the read discount cuts twice as deep" -- landing on the same number by coincidence. That ratio outlives the absolute prices: when the next generation ships, you don't need to memorize a new table, just the multiplier.&lt;br&gt;
&lt;strong&gt;2.2 Caching is already the default state, not an optional discount&lt;/strong&gt;&lt;br&gt;
This shows up most clearly in real traffic. OpenRouter's model page publishes the first-day distribution for GPT-6 Luna: on OpenAI's own endpoint, roughly 85% of input tokens were served from cache, with a hit rate near 86%. The result is a weighted-average input price of $0.0318 -- less than a third of the $0.10 list price.&lt;br&gt;
Put another way, the price sheet describes a usage pattern that barely exists. A production agent reuses the same system prompt, tool definitions, and repository context on nearly every call, and that is precisely the part caching covers. Any cost estimate that bills every input token at the list price will be systematically too high.&lt;br&gt;
&lt;strong&gt;2.3 Re-running the math on an agent session&lt;/strong&gt;&lt;br&gt;
Take a coding-agent session of 50 steps. Each step reuses a 60K cached prefix, appends 2K of new content (billed as a cache write), and produces 1K of output. For the whole session: 3M cache reads, 100K cache writes, 50K output.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frl1v6xx3sxe9czk37hj1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frl1v6xx3sxe9czk37hj1.png" alt=" " width="799" height="222"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;On the same workload, the Opus 5.5 vs. Sol gap narrows from 2x on the sticker to about 1.56x, and the entire gap sits in the cache-write and output columns. The cache-read column is identical.&lt;br&gt;
Then add tokens per task. Anthropic says Opus 5.5 at its default medium effort matches GPT-6 Astra on Terminal-Bench 4.0 at about 40% of the cost per task, and beats Astra on FrontierCode at roughly one fifth of the cost. OpenAI says Sol beats Opus 5 on AutomationBench at about 9% of the cost per task. The two "cheaper" claims were measured on different benchmarks against different rivals, and nobody has run Sol against Opus 5.5 head-to-head. The takeaway: list prices are good for ruling models out, not for making routing decisions.&lt;br&gt;
&lt;strong&gt;2.4 Caching has two boundaries, and both are harder than price&lt;/strong&gt;&lt;br&gt;
The first is time. OpenAI is unusually explicit about the cost here: the cache discount applies only to shared prefixes reused inside a 30-minute window. Step away from a session for a meeting and come back, and that 60K prefix is no longer a cache hit -- it is billed at the input rate again. The longer the agent task and the more human checkpoints it has, the easier this boundary is to trip.&lt;br&gt;
The second is ownership. Caches are stored per vendor-and-model pair. Switch vendor or switch model mid-session and the prefix is written fresh: the same 60K prefix costs 60K x $5 / 1M = $0.30 on Anthropic's side, while the identical step on Sol would have been $0.012 as a cache read -- roughly 25x less. Switching once is fine. Switching per step hands back everything you saved.&lt;br&gt;
Together the two boundaries say one thing: a cache is a local asset, not a global one.&lt;br&gt;
OpenAI pushed that second boundary outward inside its own family. GPT-6 caches now survive changes to reasoning effort and to tool definitions, and developers get explicit cache breakpoints, a Prompt Caching Dashboard, a diagnostics tool for misses, and a prewarming interface. GitHub reports that over the past several months these changes cut the share of prompt tokens requiring fresh processing by more than half. In engineering terms: turning effort up and down inside one vendor is now close to free, while swapping models to change difficulty still costs you a cache rebuild.&lt;br&gt;
&lt;strong&gt;2.5 So routing needs two layers&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Pick the vendor per session. Decide once when the task starts. Try Opus 5.5 for long migrations and code audits, Luna for high-volume extraction and summarization, Sol for everyday coding -- then avoid switching for the rest of the session.&lt;/li&gt;
&lt;li&gt;Tune the tier per step. Inside a session, handle changes in difficulty with reasoning effort rather than a model swap, so the cache stays warm.&lt;/li&gt;
&lt;li&gt;Calibrate on cache hits. Watch the cached_tokens field in each response, not the price sheet.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  3. In Practice: Session-Level Vendor Choice as One Parameter
&lt;/h2&gt;

&lt;p&gt;This strategy only holds if switching vendors doesn't mean a new SDK, a new key, and another bill to reconcile. Accels is a Singapore-based company, and router.accels.tech exists to solve exactly those three problems:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Stable. One OpenAI-compatible endpoint serves models from multiple vendors, so you don't maintain separate integration and retry logic for each vendor's peaks and hiccups.&lt;/li&gt;
&lt;li&gt;Complete catalog. The GPT-6 family and the Claude family sit behind the same interface. When a new model ships, you change one model string -- use the names listed in the console.&lt;/li&gt;
&lt;li&gt;Unified billing. Usage across vendors lands in a single bill, so answering "what did this agent session actually cost?" doesn't mean stitching two dashboards together.
Here is a minimal example that picks a model per session, keeps the vendor fixed inside it, and logs cache hits on every step:
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;from openai import OpenAI

client = OpenAI(base_url="https://router.accels.tech/v1", api_key="your-accels-key")

# Session-level routing: choose once at the start, keep it for the whole session to preserve the cache
ROUTES = {
    "migration": "claude-opus-5-5",  # long migrations / code audits
    "coding":    "gpt-6-sol",        # everyday coding agents
    "extract":   "gpt-6-luna",       # high-volume extraction / summarization
}

def run_session(task_type: str, system_prompt: str, steps: list[str]):
    model = ROUTES[task_type]
    # Put the static parts first: system prompt, tool definitions, repo context
    messages = [{"role": "system", "content": system_prompt}]
    for step in steps:
        messages.append({"role": "user", "content": step})
        resp = client.chat.completions.create(model=model, messages=messages)
        messages.append({"role": "assistant", "content": resp.choices[0].message.content})
        details = getattr(resp.usage, "prompt_tokens_details", None)
        cached = getattr(details, "cached_tokens", 0) if details else 0
        print(f"{model} prompt={resp.usage.prompt_tokens} cached={cached}")
    return messages
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two practical tips. Put the parts that never change at the very front so the prefix stays long and stable. And log cache hits on every step, then use those numbers rather than list prices to judge which route actually saves money.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Closing
&lt;/h2&gt;

&lt;p&gt;The two price sheets of September 22 look like a price war. For engineers they raise a more specific problem: caching is now the biggest line item in agent costs, and caches are tied to a vendor, tied to a model, and expire in 30 minutes. So good multi-model orchestration is no longer "pick the cheapest model at every step." It is "pick the right vendor for each session, change effort rather than models inside it, and keep the cache warm." Anthropic has said Sonnet 5.5 and Haiku 5.5 follow in the coming weeks, so these sheets will change again. Make the model name a config parameter and route through router.accels.tech for one endpoint and one bill. Then when the next price cut lands, you change one line of config and rerun the comparison.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sources&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;OpenAI release notes: GPT-6 Sol and GPT-6 Luna (30-minute cache window, prewarming, cache breakpoints) - &lt;a href="https://releasebot.io/updates/openai" rel="noopener noreferrer"&gt;https://releasebot.io/updates/openai&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Unite.AI: OpenAI Introduces GPT-6 Sol and Luna With 50% Lower API Prices (caching dashboard and GitHub figures) - &lt;a href="https://www.unite.ai/openai-introduces-gpt-6-sol-and-luna-with-50-lower-api-prices" rel="noopener noreferrer"&gt;https://www.unite.ai/openai-introduces-gpt-6-sol-and-luna-with-50-lower-api-prices&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;VentureBeat: OpenAI releases GPT-6 Sol and Luna models, slashing API costs 50% or more - &lt;a href="https://venturebeat.com/technology/openai-releases-gpt-6-sol-and-luna-models-slashing-api-costs-50-or-more" rel="noopener noreferrer"&gt;https://venturebeat.com/technology/openai-releases-gpt-6-sol-and-luna-models-slashing-api-costs-50-or-more&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Anthropic: Introducing Claude Opus 5.5 (official pricing and the 40% cost claim) - &lt;a href="https://www.anthropic.com/claude-opus-5-5" rel="noopener noreferrer"&gt;https://www.anthropic.com/claude-opus-5-5&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;VentureBeat: Anthropic releases Claude Opus 5.5, beating Fable 5.1 on key agentic benchmarks at 60% cheaper API price - &lt;a href="https://venturebeat.com/technology/anthropic-releases-claude-opus-5-5-beating-fable-5-1-on-key-agentic-benchmarks-at-60-cheaper-api-price" rel="noopener noreferrer"&gt;https://venturebeat.com/technology/anthropic-releases-claude-opus-5-5-beating-fable-5-1-on-key-agentic-benchmarks-at-60-cheaper-api-price&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;OpenRouter model page: GPT-6 Sol / GPT-6 Luna pricing, provider endpoints, and cache hit distribution - &lt;a href="https://openrouter.ai/models/openai/gpt-6-luna" rel="noopener noreferrer"&gt;https://openrouter.ai/models/openai/gpt-6-luna&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>apigateway</category>
    </item>
    <item>
      <title>Model Is Not a Model: Google AX Turns "Which Provider" Into a Cluster-Level Config Object</title>
      <dc:creator>MT_Notes</dc:creator>
      <pubDate>Tue, 22 Sep 2026 09:16:57 +0000</pubDate>
      <link>https://dev.to/mt_notes/model-is-not-a-model-google-ax-turns-which-provider-into-a-cluster-level-config-object-lim</link>
      <guid>https://dev.to/mt_notes/model-is-not-a-model-google-ax-turns-which-provider-into-a-cluster-level-config-object-lim</guid>
      <description>&lt;h2&gt;
  
  
  1. The news
&lt;/h2&gt;

&lt;p&gt;On September 20, Google open-sourced AX (Agent Executor) under its own GitHub organization: a declarative orchestration runtime for agent workloads, Apache-2.0, with an API group of ax.io/v1alpha1. It hit number one on Hacker News that day. The project describes itself in one line — declare an agentic task, and AX sandboxes it, wires up its workspace, fences its network, and helps you run it at scale.&lt;br&gt;
Set it next to the other route that landed a week earlier and the contrast is the story. On September 10, OpenAI packaged the Codex harness into the public beta Agents API: sessions, context compaction, model calls and tool coordination are all run by OpenAI, while you supply the model, instructions, tools and environment. The cost is that your orchestration layer is bound to OpenAI's implementation, data residency is US-only during the beta, and Zero Data Retention is unsupported. AX goes the other direction — you run the harness, and Google gives you the runtime that carries it. The cost is equally concrete: a Kubernetes cluster, ko, a container registry your cluster can pull from, and a reachable Agent Substrate Control API (in-cluster default api.ate-system.svc.cluster.local:443).&lt;br&gt;
The line in the AX docs that explains the motive is worth keeping: agents are neither stateless microservices nor run-to-completion batch jobs — they accumulate state, need strict isolation, call out to model APIs and tool servers, and can burn money in a loop if nobody is watching.&lt;br&gt;
For anyone writing multi-model call code every day, though, the part of AX worth reading is not its complaint about Kubernetes. It is the fourth primitive: Model.&lt;/p&gt;
&lt;h2&gt;
  
  
  2. The technical part
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;2.1 Four primitives, one file&lt;/strong&gt;&lt;br&gt;
Everything in AX is expressed as ax.io/v1alpha1 manifests, applied with a single ax apply -f. Each primitive owns one concern:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa1y9mp1oxsjtthn35h3o.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa1y9mp1oxsjtthn35h3o.png" alt=" " width="800" height="347"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;None of these is novel on its own; the combination is the point. They are declarative and come up together in one multi-document YAML, so a task boots with its repos cloned, its tools wired and its network fenced, skipping a token-burning cold-start bootstrap.&lt;br&gt;
2.2 Model is not a model, it is a named configuration&lt;br&gt;
The official docs are blunt about this: a Model is not a model, it is a named model configuration — which provider to call, which model identifier to use, provider-specific generation parameters, and a reference to the Kubernetes secret holding the API key.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;apiVersion: ax.io/v1alpha1
kind: Model
metadata:
  name: default-model
  atespace: default
spec:
  provider: anthropic
  model: claude-opus-5
  secretKey:
    name: anthropic-api-secret
    key: ANTHROPIC_API_KEY
  parameters:
    maxTokens: 16000
    temperature: 0.9
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The credential itself lands in a Secret with one command:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;kubectl create secret generic anthropic-api-secret \
  --from-literal=ANTHROPIC_API_KEY="sk-ant-..."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Making it a resource buys cluster-level manageability: the configuration lives in one place rather than scattered across every agent's environment. Rotating a key, pinning a new model version, tightening a parameter — each is one ax apply instead of a hunt through task definitions. AX's own components read it too, for instance when planning a workspace from a goal.&lt;br&gt;
&lt;strong&gt;2.3 A sandbox contract you can introspect&lt;/strong&gt;&lt;br&gt;
AX is just as specific about the sandbox. Every task container starts with ax-task-runner as PID 1, which brings up a metadata and guest-management daemon on port 80, then launches spec.command as a child process and passes the daemon's address in through AX_METADATA_URL (something like &lt;a href="http://127.0.0.1:80" rel="noopener noreferrer"&gt;http://127.0.0.1:80&lt;/a&gt;). Code inside the sandbox can therefore look itself up without any SDK:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;curl -s "$AX_METADATA_URL/metadata/v1alpha1/ax/task"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two more endpoints on that port are worth remembering: /metadata/v1alpha1/ax/workspaces returns every bound Workspace as a multi-document stream in binding order, and /readyz answers 503 while the workspace is initializing, flipping to 200 only once clones, MCP config and skills are all in place — the most reliable signal that a sandbox can actually do work. Add spec.debug: true and the same port also serves Agent Substrate's guest services (process and filesystem), which is what ax ssh rides on. It is off by default, because it amounts to opening arbitrary command execution inside the sandbox.&lt;br&gt;
One more detail: if a Workspace binding carries a goal, the runner hands that goal to an Antigravity agent on first boot to finish environment setup. That agent needs GEMINI_API_KEY in the container and gets ten minutes by default, adjustable via AX_BOOTSTRAP_TIMEOUT. Describing an environment in plain English and letting an agent build it is a platform capability here, not just an orchestrator option.&lt;br&gt;
&lt;strong&gt;2.4 Gateway: egress gets modeled&lt;/strong&gt;&lt;br&gt;
Gateway owns the other half: the network boundary. It declares the listeners a task exposes and an egress allowlist measured on two axes, host and port — host: "*" with port: 443 allows any host on 443:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;apiVersion: ax.io/v1alpha1
kind: Gateway
metadata:
  name: default-gateway
spec:
  listeners:
    - name: http
      port: 8080
      protocol: HTTP
  egress:
    allowlist:
      hosts:
        - host: "*"      # allow everything on 443; tighten this in production
          port: 443
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The official example comment says it themselves: tighten this in production. Given that you are executing code a probabilistic model wrote, making egress an explicit allowlist enforced at the infrastructure layer beats relying on the agent's own code to behave.&lt;br&gt;
&lt;strong&gt;2.5 The gap worth staring at: no endpoint field in Model&lt;/strong&gt;&lt;br&gt;
At this point the real issue surfaces. As documented today, Model.spec covers provider, model, secretKey and parameters, and the providers the docs demonstrate are google and anthropic. There is no custom endpoint or base URL field.&lt;br&gt;
That reads less like an oversight in the docs and more like a design stance: "where models come from" is treated as a built-in platform capability. provider is a supported enum value and the credential arrives from a Secret; it is not a string where you can drop an arbitrary gateway address.&lt;br&gt;
Which sets the cost structure for an agent fleet spanning several vendors: N provider values, N Kubernetes Secrets, N hostnames in the Gateway allowlist, N wire-protocol dialects, and N invoices that do not reconcile. To A/B three vendors' models inside one cluster, what you change is not a field but the whole chain.&lt;br&gt;
&lt;strong&gt;2.6 It is still alpha&lt;/strong&gt;&lt;br&gt;
The README carries a warning in plain terms: core concepts, protocols and specifications are still being actively refined, and major breaking changes are likely before a stable release. The other half of the runtime is not an officially supported Google product either.&lt;br&gt;
AX relies on a separate project, Agent Substrate, for sandboxed execution, supporting both gVisor and microVM isolation. A third party that read through that repository notes it states plainly that it is not an officially supported Google product, and that it is still pre-v1 and pre-GA. So read AX as alpha infrastructure: worth reading now, worth waiting on for production dependencies.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Where this lands: collapsing N to 1
&lt;/h2&gt;

&lt;p&gt;Put the Model primitive, the host/port-granular egress allowlist and the missing endpoint field together, and the cost structure of multi-model orchestration gets rearranged.&lt;br&gt;
Behind a router such as router.accels.tech, that N collapses to 1:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;One hostname in the egress allowlist. A single router.accels.tech plus 443 entry in Gateway is enough, and the coarse granularity stops mattering — one hostname was all you needed.&lt;/li&gt;
&lt;li&gt;One Secret. Credentials collapse to a single key in a Kubernetes Secret, and switching models turns from "add a provider, add a Secret, rebuild the image" into editing a string. How complete the catalog is decides how many options that path can cover, instead of whatever one vendor's shelf happens to hold.&lt;/li&gt;
&lt;li&gt;One billing view. The hardest question for a fleet is what an eval run cost and to whom. A router consolidates usage across vendors into one metering view, turning a post-hoc spreadsheet merge into a query.
The sandbox-side wiring is a few lines — Model still describes which model the platform itself calls, while the runner points the actual traffic at the router (Accels is a Singapore-based company):
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;import os, yaml, requests
from openai import OpenAI

# The sandbox metadata server returns the current Task's full spec and status
spec = yaml.safe_load(
    requests.get(f"{os.environ['AX_METADATA_URL']}/metadata/v1alpha1/ax/task").text
)

client = OpenAI(
    api_key=os.environ["ACCELS_API_KEY"],
    base_url="https://router.accels.tech/v1",
)

resp = client.chat.completions.create(
    model="claude-opus-5",
    messages=[{"role": "user", "content": "Write a fix plan for this failed build"}],
)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Complete model coverage means the value in that model slot is not bounded by one vendor's catalog. Stability means this single egress path cannot become the fleet's single point of failure. Unified billing means that after an ax suspend you can say exactly what you saved. In a runtime built around the fact that agents burn money in a loop, all three are operational metrics rather than marketing words.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Closing
&lt;/h2&gt;

&lt;p&gt;AX is alpha today, depends on Agent Substrate, will break its API, and should not be anyone's production dependency this month. But it makes one thing explicit: once agents become workloads you schedule, "which model" and "where it may reach" stop being two constants in application code and become configuration objects the platform has to declare, audit and rotate.&lt;br&gt;
That shift puts concrete demands on the calling side — as few egress hosts as possible, as few credentials as possible, as consistent a metering view as possible. If you are choosing that layer for your own agent platform, it is worth counting your current chain against those three and seeing how large N actually is.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sources&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;google/ax — Google's open agentic orchestration runtime&lt;/li&gt;
&lt;li&gt;AX — agentexecutor.io&lt;/li&gt;
&lt;li&gt;AX core concepts: Model / Gateway / Workspace / Task — docs/concepts.md&lt;/li&gt;
&lt;li&gt;AX manifests and full field reference — docs/manifests.md&lt;/li&gt;
&lt;li&gt;Inside the AX sandbox: metadata server and guest services — docs/sandbox.md&lt;/li&gt;
&lt;li&gt;agent-substrate/substrate — the sandboxed execution runtime behind AX&lt;/li&gt;
&lt;li&gt;Introducing the Agents API — OpenAI, 2026-09-10&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>apigateway</category>
      <category>google</category>
    </item>
    <item>
      <title>Why the Cheaper-per-Token Model Costs More per Point of Capability: Step 5 Preview, MiniMax M3, and the New Denominator</title>
      <dc:creator>MT_Notes</dc:creator>
      <pubDate>Mon, 21 Sep 2026 05:58:39 +0000</pubDate>
      <link>https://dev.to/mt_notes/why-the-cheaper-per-token-model-costs-more-per-point-of-capability-step-5-preview-minimax-m3-and-lc1</link>
      <guid>https://dev.to/mt_notes/why-the-cheaper-per-token-model-costs-more-per-point-of-capability-step-5-preview-minimax-m3-and-lc1</guid>
      <description>&lt;h2&gt;
  
  
  1. Background
&lt;/h2&gt;

&lt;p&gt;On September 20, StepFun released its flagship foundation model Step 5 Preview: a sparse MoE architecture with 600B total parameters and only 27B activated per token, a 1M-token context window, and native text plus vision input. The third-party API opened the same day, with full weights promised for October 15. Artificial Analysis scored it 44 on its Intelligence Index, placing it in the global top three among open models. StepFun also pushed a more quotable number: a per-task cost roughly one-eighth that of Claude Opus 5.&lt;br&gt;
In almost the same window, another open-weight model, MiniMax M3, was pushed to prominent positions in several gateway and price-comparison catalogs: 428B total parameters, around 23B active, MiniMax Sparse Attention (MSA), a 1M-token context, and multimodal data mixed in from the very first pretraining step. It is far cheaper than Step 5 Preview, at $0.30 versus $1.00 per million input tokens and $1.20 versus $2.70 per million output tokens.&lt;br&gt;
So the question becomes concrete: is the model that is two to three times cheaper per token actually cheaper to run? Compare the two price sheets over a common denominator and the answer is no.&lt;/p&gt;
&lt;h2&gt;
  
  
  2. The technical picture
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;2.1 Change the denominator from tokens to tasks&lt;/strong&gt;&lt;br&gt;
Artificial Analysis runs the same evaluation set through both models and converts the tokens actually produced into money. Step 5 Preview lands at about $0.71 per index task; MiniMax M3 lands at about $0.51. Measured on raw throughput, M3 is roughly 28% cheaper.&lt;br&gt;
Divide by each model's score and the ranking flips: 44 against 29. That works out to about 1.6 cents of capability per point for Step 5 Preview and about 1.8 cents for M3. The model that costs 3.3x more per token is the cheaper one per unit of capability.&lt;br&gt;
That denominator deserves to be called out on its own, because a per-token price hides two things entirely: how much a model talks (verbosity) and how many tokens it spends thinking (reasoning overhead). A chatty cheap model can easily outspend a concise expensive one.&lt;br&gt;
One version trap is worth flagging. Artificial Analysis recalibrated its Intelligence Index on September 7, 2026, compressing scores across the board. Comparing a score from a screenshot taken before that date against a current score will produce the wrong conclusion.&lt;br&gt;
&lt;strong&gt;2.2 Two counterintuitive clauses in M3's price sheet&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F41bl1xpfkkyl0k9z0ir5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F41bl1xpfkkyl0k9z0ir5.png" alt=" " width="799" height="180"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The first row is an ordinary tier. The second row is the point. The vendor's rule reads: the whole request is priced according to the tier that its input token count falls into. This is not a surcharge on the excess; it is the entire request repriced. For a coding agent that keeps accumulating history into its context, the moment that context crosses 512K, the full input, the full output, and even the cache-hit rate all double. That is not a linear increase. It is a cliff.&lt;br&gt;
The second clause is a cache write price of $0.00. Most vendors charge more for writing to cache than for ordinary input, because writing is extra work and only a hit earns a discount. M3 turns the write into a free action, which means you no longer have to bet on your hit rate. You do not need to calculate how many hits it takes to amortize the write cost, because the first write already breaks even. The engineering conclusion is blunt: if a prefix can be pinned into cache, pin it, and stop deliberating.&lt;br&gt;
&lt;strong&gt;2.3 One model id, 49 hosting providers, a 3.2x price spread&lt;/strong&gt;&lt;br&gt;
M3 is an open-weight model, so the same weights can be hosted by many parties. Comparison catalogs show M3 quoted across 49 hosts, with input prices running from $0.225 to $1.20, a 3.2x spread. For contrast, GPT-5.5's spread over the same period is 33.3x.&lt;br&gt;
More expensive does not automatically mean worse. One host quotes $0.40 / $2.00, which is 33% and 67% above the official rates, yet its measured behavior on the relay path shows a 1.03s median time to first token and a 92.6% cache hit rate. Keep the sources straight: benchmark scores belong to the weights, while measured time to first token and hit rate belong to the host. They are not interchangeable.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;from openai import OpenAI

client = OpenAI(
    api_key="your-accels-key",
    base_url="https://router.accels.tech/v1",
)

resp = client.chat.completions.create(
    model="minimax-m3",
    messages=[
        {"role": "system", "content": STABLE_PREFIX},   # pin the prefix, change it rarely
        {"role": "user", "content": question},
    ],
)

u = resp.usage
print(u.prompt_tokens, u.completion_tokens)
print(getattr(u, "prompt_tokens_details", None))   # how much hit the cache
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What matters here is not the reply but the usage object: the tokens that hit the cache show up separately in the detail fields. Log that number weekly and you will forecast your bill far better than you can by staring at a rate card.&lt;br&gt;
It is also worth separating the two senses of "available." M3's weights have been downloadable since June 12, with an additional MXFP8 quantized release. Step 5 Preview currently exists only as a hosted endpoint; its Hugging Face repository is an empty shell and the weights are due October 15. One is a file you can hold. The other is a date in a letter of intent.&lt;/p&gt;
&lt;h2&gt;
  
  
  3. Where this lands: why the call layer should converge on one interface
&lt;/h2&gt;

&lt;p&gt;Once unit price, caching semantics, tier rules and hosting variance are all moving at the same time, maintaining a price sheet by hand stops being realistic. That is why we converge our calls onto router.accels.tech. Accels is a Singapore-based company, and Singapore is becoming the concentration point for AI infrastructure in the region: per the FT, OpenAI is in talks to lease roughly 100,000 square feet at Shaw Tower, and Anthropic opens its own Singapore office in October, following Tokyo, Bengaluru, Seoul and Sydney.&lt;br&gt;
Three things land concretely.&lt;br&gt;
Complete model coverage. For models like the ones above that are open-weight but served by many hosts, one key reaches them by model id, with no need to wire up each host separately. Model selection stops being "integrate one more vendor" and becomes "change one string."&lt;br&gt;
Unified billing. If cache-hit rates, tier boundaries and host differences are scattered across several invoices, you can never work out how much a given change actually saved. Unified billing turns cost per task from an estimate into a queryable number, which is exactly the denominator this article opened with.&lt;br&gt;
Stability. A cache hit rate depends on the previous request landing on the same host. One mid-run failure restarts the entire cache chain, and input that should have billed at the hit rate is rebilled at full price. On this path, link stability converts directly into money.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# Long-context workhorse plus a light router, one key, one invoice
for model in ["minimax-m3", "qwen3.8-flash-next", "gpt-5.6-luna"]:
    r = client.chat.completions.create(model=model, messages=msgs)
    print(model, r.usage)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The boundary deserves stating: a gateway will not decide how to slice your prefix or when to compact your context. That remains your architectural call. What it removes is duplicated integration work and inconsistent accounting, not design responsibility.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Closing
&lt;/h2&gt;

&lt;p&gt;In 2026, model selection is hard to answer with "which one is cheaper." What actually determines the bill is three numbers: cost per task, cache hit rate and host stability. Step 5 Preview pushes the Pareto frontier outward by a notch, running unattended 24-hour loops that take an H100 kernel from scratch to 508 TFLOPS, against Claude Opus 5's 493 TFLOPS in the same experiment. MiniMax M3 uses open weights plus free cache writes to turn long context from something you can afford into something you use routinely. Neither price sheet states its own denominator clearly.&lt;br&gt;
Rather than keep doing the arithmetic by hand, stand up the unified interface layer first and optimize the prefix structure afterwards. Spend the effort on where the prefix should be cut, not on rewriting auth and retry logic for the fourth time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sources&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Step 5 Preview: 600B total / 27B active, AA Intelligence Index 44 (National Business Daily)&lt;/li&gt;
&lt;li&gt;Small parameters, high performance: StepFun's new model enters the global open-source top three (Shanghai Observer)&lt;/li&gt;
&lt;li&gt;Step 5 Preview Goes Live: 600B Params (AI DAMN)&lt;/li&gt;
&lt;li&gt;Step 5 Preview vs MiniMax M3: cost per task and per index point (OrcaRouter)&lt;/li&gt;
&lt;li&gt;MiniMax M3 model page: rates, cache pricing, long-context tiers (LLM Abacus)&lt;/li&gt;
&lt;li&gt;MiniMax M3 multi-host quotes and spread (LLM Pricing)&lt;/li&gt;
&lt;li&gt;MiniMax M3 measured endpoint: time to first token and cache hit rate (Requesty)&lt;/li&gt;
&lt;li&gt;SWE-bench Pro leaderboard with live prices (AnotherWrapper)&lt;/li&gt;
&lt;li&gt;OpenAI and Anthropic push up Singapore office rents (Financial Times via AI Weekly)&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>apigateway</category>
      <category>minimax</category>
    </item>
    <item>
      <title>When Your Provider Swaps the Model Underneath You: DeepSeek V4.1 Flash, Alias Routing, and an 890-Byte Cache</title>
      <dc:creator>MT_Notes</dc:creator>
      <pubDate>Fri, 11 Sep 2026 07:10:36 +0000</pubDate>
      <link>https://dev.to/mt_notes/when-your-provider-swaps-the-model-underneath-you-deepseek-v41-flash-alias-routing-and-an-3e00</link>
      <guid>https://dev.to/mt_notes/when-your-provider-swaps-the-model-underneath-you-deepseek-v41-flash-alias-routing-and-an-3e00</guid>
      <description>&lt;h2&gt;
  
  
  &lt;strong&gt;1. What Happened&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;On September 10, DeepSeek released V4.1 Flash: a 552B-parameter MoE model built on an entirely new Causal-Encoder-Decoder (CED) architecture, with native multimodal vision, a 1M-token context, and MIT-licensed open weights on Hugging Face. For most developers the launch was just another headline. The part worth pausing on is the quiet retirement notice buried in the announcement:&lt;br&gt;
From 12:00 Beijing time on September 14, until V4.1 Pro launches, every request to deepseek-v4-pro will be routed to V4.1 Flash and billed at V4.1 Flash prices.&lt;br&gt;
Note what is actually happening here: nothing in your code changes — the model: "deepseek-v4-pro" string stays, the endpoint stays, the SDK stays — but the model behind it is replaced, and so is the bill. This time it moves in your favor. The legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp are also being routed to the new model on a temporary basis. The provider is using its own API as an alias-routing layer, folding an expensive tier into a cheap one. This differs from the third-party price calendars we covered last week: this time the model itself is being swapped at the routing layer, and most client applications have no way to notice.&lt;/p&gt;
&lt;h2&gt;
  
  
  &lt;strong&gt;2. The CED Architecture: Why This Model Is Cheap for Structural Reasons&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;V4.1 Flash's cost advantage is not a subsidy. It comes from three architectural decisions.&lt;br&gt;
First, an asymmetric encoder-decoder split. The 40 layers divide into a 20-layer causal encoder and a 20-layer decoder. Long inputs are summarized once by the encoder; the decoder's global KV states are projected directly from the encoder's final-layer hidden states, dropping prefill complexity from roughly O(NL) to about O(NL/2) for sequences much longer than the window. This explains a seemingly odd configuration: each input token activates only about 8B parameters, while decoding activates about 16B. The predecessor, V4 Flash, was a 284B model activating roughly 13B per token. The model doubled in size, yet the input stage computes with fewer active parameters.&lt;br&gt;
Second, cache compression. KV-cache HBM requirements drop to one quarter of the previous generation, SSD requirements to one eighth, and versus DeepSeek's first-generation model the cache has shrunk 437 times — about 890 bytes per token. The levers are FP4 storage and cross-layer sharing of attention indices. For agentic workloads this is the decisive number: whether long conversation histories can stay resident in cache and be reused repeatedly defines the entire cost curve.&lt;br&gt;
Third, pricing the cache gap into the rate card. Off-peak prices: 0.02 CNY per million tokens on cache hit, 1.0 on miss, 4.0 for output; peak hours double all three. The gap between a hit and a miss is 50x — cache hit rate is no longer just a performance metric, it is a first-class input to the pricing function.&lt;br&gt;
On capability, DeepSeek's own reported numbers: 74.2 on DeepSWE v1.1 (versus Claude Opus 5.0's 74.0) and 90.6 on Terminal-Bench 2.1 (versus 89.1). Read the footnotes, though: the same model scored anywhere from 65.6 to 74.2 on DeepSWE across eight different harnesses, a spread of nearly nine points, and on the harder Terminal-Bench 4.0 it reached only 31.2 versus Opus 5.0's 51.8. Treat it as an excellent cheap-tier agentic workhorse, not as a free frontier model, and the positioning is accurate.&lt;/p&gt;
&lt;h2&gt;
  
  
  &lt;strong&gt;3. Three Moves for Your Routing Layer&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Provider-side alias routing is already happening; client routing layers should catch up:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Verify the model field in responses. Alias routing means the requested name and the served model may diverge; logging, monitoring, and cost attribution should key off the actual model in the response.&lt;/li&gt;
&lt;li&gt;Put cache hit rate on the cost dashboard. With a 50x spread between hit and miss, raising an agent workflow's hit rate from 40% to 80% nearly halves input costs.&lt;/li&gt;
&lt;li&gt;Plan for forced provider-side migrations. After September 14 the billing tier behind deepseek-v4-pro changes. Your routing layer should alert on price-snapshot drift instead of discovering it on the month-end invoice.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2&gt;
  
  
  &lt;strong&gt;4. Putting It to Work: Harvesting the Architecture Dividend Behind a Unified Interface&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;V4.1 Flash is live on the DeepSeek API under the model name deepseek-flash, and the weights are open. If your application already sits on top of multiple models, adopting it requires no business-code changes — point your base_url at a unified gateway:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;your-accels-key&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://router.accels.tech&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="c1"&gt;# single entry point; upstream alias routing is handled for you
&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;deepseek-flash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Summarize this 1M-token repo changelog and flag three high-risk commits&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The value of a unified gateway shows up precisely at moments like this, when a provider reshuffles its tiers: stability — upstream alias routing and model retirements are absorbed by the gateway, so your calls do not break on an API migration; complete model coverage — new and old names like deepseek-flash and deepseek-v4-pro live in one catalog, and switching is just a string change; unified billing — hits, misses, peak and off-peak schedules get folded into one consistent billing surface, so comparing costs across models no longer means reconciling rate cards by hand. Run the same agent workload as an A/B between deepseek-flash and glm-5.3 by changing a single field, and reuse the rest of the pipeline untouched.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Closing&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;V4.1 Flash turns "cheap" from an operations problem into an architecture problem: an 890-byte-per-token cache, asymmetric activation, and a 50x cache-price spread all point to the same lesson — the cost levers in the agent era live in the cache and the routing layer, not in buying a more expensive model. And a provider running alias routing itself is a reminder that model names in an API are degrading from contracts into suggestions. Is your routing layer ready for that?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sources&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;DeepSeek official announcement: DeepSeek V4.1 Flash&lt;/li&gt;
&lt;li&gt;Tencent News / WenAI: DeepSeek-V4.1-Flash analysis (437x smaller cache, V4 Pro retirement)&lt;/li&gt;
&lt;li&gt;DataNorth AI: DeepSeek releases DeepSeek-V4.1-Flash&lt;/li&gt;
&lt;li&gt;Baidu Baike: DeepSeek V4.1 Flash (peak/off-peak pricing)&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>accels</category>
      <category>deepseek</category>
      <category>apigateway</category>
    </item>
    <item>
      <title>The August 31 Cliff: One of Those Three Deadlines No Longer Exists</title>
      <dc:creator>MT_Notes</dc:creator>
      <pubDate>Tue, 25 Aug 2026 08:40:17 +0000</pubDate>
      <link>https://dev.to/mt_notes/the-august-31-cliff-one-of-those-three-deadlines-no-longer-exists-5f7f</link>
      <guid>https://dev.to/mt_notes/the-august-31-cliff-one-of-those-three-deadlines-no-longer-exists-5f7f</guid>
      <description>&lt;h2&gt;
  
  
  The lead: a price increase that keeps getting reposted, months after it was cancelled
&lt;/h2&gt;

&lt;p&gt;Six days to August 31. The date keeps coming up in developer channels because three things land on the same Monday: Moonshot's full platform sunset of kimi-k2.5 and the moonshot-v1 series, the retirement of GPT-5.4 and GPT-5.4 mini from ChatGPT-signed-in Codex, and the expiry of Claude Sonnet 5's introductory pricing, taking it to $3/$15.&lt;br&gt;
The first two are real. The third was struck by Anthropic itself back on August 10. The pricing docs now say it plainly: the $2/$10 input/output rate "is now the standard price," and the increase to $3/$15 previously scheduled for September 1 "will not occur." The Sonnet 5 launch post carries a matching edit note from the same day.&lt;br&gt;
Yet as of August 24, a fair number of industry dailies, migration checklists and "model retirement calendars" still list that price increase, unchanged, under August 31. That is the more interesting story. Once your architecture depends on four or five upstreams at once, model ID lifecycle becomes a real engineering surface — and the information on that surface is drifting systematically in secondhand sources.&lt;/p&gt;
&lt;h2&gt;
  
  
  1. Three deadlines, two of which actually break requests
&lt;/h2&gt;

&lt;p&gt;Facts first, because "will error" and "you should migrate" are entirely different claims:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1dr919x94apfwetywlfl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1dr919x94apfwetywlfl.png" alt=" " width="799" height="355"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That second row is the one people misread. OpenAI's help center scopes it narrowly: the retirement applies to Codex sessions signed in with a ChatGPT account, and what needs updating is workspace defaults, saved model settings, managed configurations, custom agents and scheduled tasks. Codex authenticated with your own API key, and direct OpenAI API calls, are explicitly out of scope. So the same gpt-5.4 string may keep working fine in your CI script on September 1 while failing outright in a colleague's Codex scheduled task.&lt;br&gt;
The Moonshot row is a plain wall. New accounts have been unable to select either family since K3 shipped on July 16; August 31 is the final close for existing users. Worth noting: K2.5's open weights on Hugging Face are unaffected — what is being retired is the hosted API, not the model. Teams that self-host can ignore the date entirely.&lt;/p&gt;
&lt;h2&gt;
  
  
  2. The quieter axis: the tokenizer, not the rate card
&lt;/h2&gt;

&lt;p&gt;The Sonnet 5 increase was cancelled. Something else was not, and it appears on no pricing table.&lt;br&gt;
Anthropic's docs carry a note that Claude 4.7 and later models use a newer tokenizer that "produces approximately 30% more tokens for the same text." The Sonnet 5 launch post narrows it: roughly 1.0x to 1.35x depending on content type. Simon Willison filled in that range by measurement — the same Universal Declaration of Human Rights went from 2,356 tokens to 3,341 in English (1.42x), 1.33x in Spanish, 1.28x for a 4,279-line Python file, and essentially unchanged in Simplified Chinese (1.01x).&lt;br&gt;
Which means your cost model has two axes and most people watch only one. Per-token rates held steady, but the token count for identical text moved, so the invoice moves. Run it the other way: if your workload is mostly Chinese, that multiplier is close to 1, and the "+30%" you copied off someone's blog simply does not apply to you. The same multiplier also quietly shrinks the context window — a 1M-token window now holds less actual text.&lt;br&gt;
The conclusion is simple and unglamorous: do not cite token counts someone else measured. Send your own real payload to both the old and new model, compare the returned input_tokens, and that delta is your budget correction factor.&lt;/p&gt;
&lt;h2&gt;
  
  
  3. The real cost of migration is model choice, not code
&lt;/h2&gt;

&lt;p&gt;Kimi's API stays OpenAI-compatible, so at the code level migration is a one-string change. That is exactly the trap. The path of least effort is a global find-and-replace of kimi-k2.5 with kimi-k3 — and K3 is the $3/$15 flagship, roughly three times the tier k2.5 lived in. Route your entire legacy volume there and you are paying triple for calls that never needed frontier capability, for zero quality gain.&lt;br&gt;
The right move is to split by workload: routine coding and chat go to kimi-k2.7-code, which is not on the retirement list at all; only calls that genuinely need frontier reasoning, native vision or the full ~1M context go to kimi-k3 — with prompt caching turned on there, since K3 cache-hit input runs $0.30 per million tokens against $3.00 uncached, one tenth.&lt;br&gt;
Put differently, a "one string" migration actually requires you to answer three questions: where are all my call sites, what is the load shape at each one, and what is the marginal cost of each candidate ID. Most teams discover, six days out, that they cannot answer any of the three.&lt;/p&gt;
&lt;h2&gt;
  
  
  4. Landing it: collapse model IDs into one layer instead of scattering them through code
&lt;/h2&gt;

&lt;p&gt;Every pain point above points at the same structural issue. Model IDs, upstream base URLs, auth schemes and billing units — if those four things live directly in your application code, then every upstream retirement is a repo-wide grep plus a regression pass.&lt;br&gt;
This is what a model routing layer is for. Take wrouter.ai: once multiple upstreams sit behind a single OpenAI-compatible endpoint, you get three concrete things.&lt;br&gt;
Stability. Your application knows one base_url. When an upstream changes endpoints, revises auth, or sunsets an ID, that is absorbed inside the routing layer rather than propagating into business code.&lt;br&gt;
A complete model catalogue. Anthropic, OpenAI, Moonshot, Google and DeepSeek sit side by side on the same endpoint, so switching is one model field. That matters most during a deadline migration: you can fire the old and new IDs at real payloads, compare token counts and output quality, then decide — instead of changing blind and shipping.&lt;br&gt;
Unified billing. No reconciling four rate cards in four different units, and no stitching a migration's cost delta together across four consoles.&lt;br&gt;
Concretely, against that "don't cite someone else's token counts" conclusion:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;from openai import OpenAI

client = OpenAI(
    api_key="YOUR_WROUTER_KEY",
    base_url="https://wrouter.ai/v1",
)

payload = open("your_real_prompt.txt").read()

# One real payload, sent to the retiring ID and each migration candidate
for model in ["kimi-k2.5", "kimi-k2.7-code", "kimi-k3"]:
    r = client.chat.completions.create(
        model=model,
        messages=[{"role": "user", "content": payload}],
    )
    print(model, r.usage.prompt_tokens, r.usage.completion_tokens)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run it and you get three rows of numbers. Those three rows are your budget correction factor for this migration — considerably more applicable to you than any blog's "+30%". The same pattern crosses vendors: put claude-sonnet-4-6 and claude-sonnet-5 in the same loop and you have measured the tokenizer multiplier for your own content shape.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;The real lesson of August 31 is not "remember three dates." It is that a model ID is an external dependency that expires, and that secondhand information about it is not trustworthy. A price increase the vendor cancelled two weeks earlier is still circulating as a to-do item, which tells you the information half-life in this space is shorter than most teams' migration cycles.&lt;br&gt;
Three things you can do. Put each upstream's first-party deprecation page into your monitoring. Log which model ID every request actually used, so the next retirement is a diff you schedule rather than a fire you fight. And lift model IDs out of application code into a layer you control. The first two are discipline; the third is architecture.&lt;br&gt;
If you are migrating for next Monday's two real deadlines, start by putting the old and new IDs behind the same endpoint and running that snippet against your own payload. Once you have the numbers, the decision makes itself.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sources&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Claude Platform pricing docs ($$2/$$10 now standard; September 1 increase will not occur; 4.7+ tokenizer ~+30% tokens): &lt;a href="https://platform.claude.com/docs/en/about-claude/pricing" rel="noopener noreferrer"&gt;https://platform.claude.com/docs/en/about-claude/pricing&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Introducing Claude Sonnet 5 (August 10 edit note; 1.0–1.35x tokenizer footnote): &lt;a href="https://www.anthropic.com/news/claude-sonnet-5" rel="noopener noreferrer"&gt;https://www.anthropic.com/news/claude-sonnet-5&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Simon Willison, "What's new in Claude Sonnet 5" (measured tokenizer comparison table): &lt;a href="https://simonwillison.net/2026/Jun/30/claude-sonnet-5/" rel="noopener noreferrer"&gt;https://simonwillison.net/2026/Jun/30/claude-sonnet-5/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Kimi API model list (kimi-k2.5 and moonshot-v1 series, full platform sunset August 31): &lt;a href="https://platform.kimi.ai/docs/models" rel="noopener noreferrer"&gt;https://platform.kimi.ai/docs/models&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;OpenAI Help Center: scope of the GPT-5.4 / GPT-5.4 mini Codex retirement and migration targets: &lt;a href="https://help.openai.com/en/articles/11369540" rel="noopener noreferrer"&gt;https://help.openai.com/en/articles/11369540&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
    </item>
    <item>
      <title>After the Harness Went Open Source: The Agent Skeleton Is No Longer the Moat — the Model Is the Swappable Part</title>
      <dc:creator>MT_Notes</dc:creator>
      <pubDate>Fri, 21 Aug 2026 08:47:00 +0000</pubDate>
      <link>https://dev.to/mt_notes/after-the-harness-went-open-source-the-agent-skeleton-is-no-longer-the-moat-the-model-is-the-3l6p</link>
      <guid>https://dev.to/mt_notes/after-the-harness-went-open-source-the-agent-skeleton-is-no-longer-the-moat-the-model-is-the-3l6p</guid>
      <description>&lt;h2&gt;
  
  
  The setup: two skeletons dropped in one week
&lt;/h2&gt;

&lt;p&gt;On August 19, OpenAI Developers published a post with an unusually blunt title: Codex as a platform: build on the open agent harness. The next day, Greg Brockman amplified it on X and resurfaced a May case study: a tax-preparation system built on Codex by Thrive Holdings and the accounting network Crete Professionals Alliance (rebranded as Current in June) processed 7,000 returns and cut accountants' preparation time by roughly a third.&lt;br&gt;
The weight of this news isn't in a model. It's in the word "harness."&lt;br&gt;
A harness is the layer wrapped around the model: the agent loop, context gathering, tool execution, sandboxing, approval flows, multi-turn state management. For two years, labs treated it as a core asset and kept it hidden. Now OpenAI has laid the entire Codex harness on GitHub (openai/codex) under Apache-2.0 — readable, modifiable, commercially embeddable, no copyleft obligations.&lt;br&gt;
The timing is almost too neat. Six days earlier, DeepSeek open-sourced its own agent runtime, DeepSeek Harness v0.1, under MIT, built on the Cordis plugin system, which collected 23,000 GitHub stars within hours.&lt;br&gt;
Two frontier labs gave away their skeletons in a single week. For developers, that's a structural shift: the hardest part of an agent system is becoming public infrastructure, and the model has been demoted to a component you can swap at will.&lt;/p&gt;


&lt;h2&gt;
  
  
  1. What is a harness actually worth? A benchmark answered
&lt;/h2&gt;

&lt;p&gt;If your reaction is "it's just a loop with some tool calls," another OpenAI post is worth reading: How enabling two settings tripled our ARC-AGI-3 scores.&lt;br&gt;
They didn't change models. They didn't retrain anything. They flipped two switches at the harness layer: preserved reasoning and context compaction. GPT-5.6 Sol's ARC-AGI-3 score went from 13.3% to 38.3%, close to a 3x jump. Output token count dropped to roughly one-sixth of the original.&lt;br&gt;
That number should make anyone building agents pause. Same model, same benchmark — and purely because the outer orchestration changed, both capability and cost improved by close to an order of magnitude.&lt;br&gt;
Put differently: we've been trained to attribute "results aren't good enough" to "the model isn't good enough," then reach for a more expensive model. The open harness tells you that a meaningful slice of that spend was recoverable through architecture all along.&lt;/p&gt;
&lt;h2&gt;
  
  
  2. Three integration surfaces: exec, SDK, app-server
&lt;/h2&gt;

&lt;p&gt;What Codex opened up isn't a vague "framework" but three clearly separated surfaces. Picking the wrong one costs you a lot of pointless engineering, so it's worth getting straight:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhdld0v247lanaou7depv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhdld0v247lanaou7depv.png" alt=" " width="799" height="378"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The app-server is the headline. It is both a protocol and a long-lived process, composed internally of a stdio reader, a message processor, a thread manager, and core threads, with the thread manager spinning up one core session per thread. The protocol is bidirectional — the server can initiate requests (for instance, when it needs a human approval) and pause the turn until the client responds.&lt;br&gt;
OpenAI is candid about an early wrong turn: they first tried exposing Codex as an MCP server, but found MCP's semantics hard to stretch across the rich interactions an IDE needs — diff updates, workspace exploration, streamed reasoning — which is why they built a JSON-RPC protocol instead. That's a useful lesson for anyone building an agent platform: MCP is a good fit for treating an agent as a callable tool, and a poor fit for making an agent the spine of a product.&lt;br&gt;
To make the pattern concrete, OpenAI shipped a sample app called Relay: a fictional logistics operations dashboard where the agent pulls live data through the app's own MCP tools, the user clicks suggested actions like "Compare recovery options" rather than facing a blank prompt box, and any write action (rebooking a shipment) routes through human approval before it executes. GitHub and JetBrains embed Codex in their own workflows; Cisco uses it inside App Builder — all downstream of the same app-server protocol.&lt;/p&gt;
&lt;h2&gt;
  
  
  3. Open skeletons push all the cost pressure down to the model layer
&lt;/h2&gt;

&lt;p&gt;There's a boundary worth stating plainly: the harness is free; inference is not.&lt;br&gt;
OpenAI's own documentation is explicit — you can read and modify the code freely, but to run anything you still authenticate to a model. The open-source repo handles agent threads, tool execution, configuration, and approvals; model access requires a ChatGPT account or API setup.&lt;br&gt;
So the landscape now looks like this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Skeleton layer: commoditized, open source, zero cost, freely replaceable (Codex harness / DeepSeek Harness / Claude Agent SDK)&lt;/li&gt;
&lt;li&gt;Model layer: differentiated, mostly closed, billed per token, violently volatile in price
And what happened at the model layer this month? DeepSeek's V4 family moved to peak/off-peak pricing at 16:00 UTC on August 16, with peak-hour output rising from $$0.87 to $$3.96 per million tokens and cache-hit input rising more steeply still. GLM-5.3 landed on the API on August 18 at $$1.40/$$4.40. Grok 4.6 arrived on Amazon Bedrock on August 19 at $$2/$$6, with a 500K context window and four reasoning-effort settings.
Stack those two facts and the conclusion is clean: once the skeleton stops being a moat, your engineering center of gravity shifts from "how do I write an agent loop" to "how do I keep the backend swappable in a violently shifting model market."
Which happens to be exactly the thing harness architectures are structurally good at and operationally worst at — because every new model vendor means another API key, another invoice, another base URL, and another bet on somebody else's uptime.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  4. In practice: collapse the harness's model exit into one endpoint
&lt;/h2&gt;

&lt;p&gt;The Codex harness, DeepSeek Harness, and most agent frameworks share a useful engineering property: model access is injected through configuration, not hardcoded. That gives you a clean point of convergence.&lt;br&gt;
The move is straightforward — point every harness instance at a single base URL and let the routing layer handle multi-model orchestration. Using wrouter.ai as the example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;from openai import OpenAI

client = OpenAI(
    api_key="wr-***",
    base_url="https://wrouter.ai/v1",
)

# One client, work assigned by task difficulty
# Cheap tier: the harness's high-frequency small steps (read a file, run grep, format a diff)
cheap = client.chat.completions.create(
    model="deepseek-v4-flash",
    messages=[{"role": "user", "content": "Summarize the intent of this diff"}],
)

# Flagship tier: the one step that genuinely needs long-horizon reasoning
strong = client.chat.completions.create(
    model="claude-opus-5",
    messages=[{"role": "user", "content": "Refactor this module and propose a migration plan"}],
)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you're going the Codex app-server route, the idea is identical — point the harness's model provider config at the routing endpoint, and the harness's internal thread, approval, and sandbox logic stays untouched:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# ~/.codex/config.toml
[model_providers.wrouter]
name = "wrouter"
base_url = "https://wrouter.ai/v1"
env_key = "WROUTER_API_KEY"

[profiles.daily]
model_provider = "wrouter"
model = "gpt-5.6-sol"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The payoff maps onto three things wrouter.ai is built for:&lt;br&gt;
Stability. Agent failure modes are not chat failure modes. In chat, a single 429 means the user hits retry. In an agent, one timeout can send a forty-minute multi-step task back to the start. A single bad step contaminates the whole trajectory. A routing layer that presents one consistent surface while upstreams wobble is a practical way to take the edge off that long-task fragility.&lt;br&gt;
Complete model coverage. This week alone produced three endpoints worth testing (GLM-5.3, Grok 4.6 on Bedrock, DeepSeek V4 Pro 0813). Registering with each vendor, clearing verification, and wiring up environment variables is enough friction to kill the evaluation before it starts. A complete catalog means A/B testing is a model-string change, not a new vendor account.&lt;br&gt;
Unified billing. This matters more in the harness era than it did before. A single agent task can fire dozens of model calls spanning cheap and flagship tiers. When that spend is scattered across four or five vendor invoices, you simply cannot compute "what does one invocation of this feature cost." One bill lets you see every tier of the harness's consumption in one table, then decide which step to downgrade and which one earns the flagship.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;On the surface, open-sourcing the Codex harness looks like OpenAI handing out another tool. In practice it's a public bet on where competition moves next: the next round of differentiation happens at the orchestration layer, not only at the model layer. Anthropic is betting the same way with the Claude Agent SDK and MCP.&lt;br&gt;
For developers, though, the implication runs the other direction. When the orchestration layer is free and universally available, whatever differentiation you build into the skeleton gets flattened fast. What actually determines your product's cost and reliability becomes the swappable model interface underneath it — how steadily it connects, how completely it covers the field, how clearly it accounts for itself.&lt;br&gt;
If you're wiring the Codex harness or DeepSeek Harness into your own product, clean up the model exit before you write a line of agent loop. Point base_url at wrouter.ai, run the whole catalog through one key and one invoice, then go back and tune your harness — that's the time this open-source release actually saves you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sources&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;OpenAI Developers, Codex as a platform: build on the open agent harness (2026-08-19) &lt;a href="https://developers.openai.com/blog/codex-as-a-platform" rel="noopener noreferrer"&gt;https://developers.openai.com/blog/codex-as-a-platform&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;OpenAI, Unlocking the Codex harness: how we built the App Server &lt;a href="https://openai.com/index/unlocking-the-codex-harness/" rel="noopener noreferrer"&gt;https://openai.com/index/unlocking-the-codex-harness/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;OpenAI, How enabling two settings tripled our ARC-AGI-3 scores &lt;a href="https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/" rel="noopener noreferrer"&gt;https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;RuntimeWire, OpenAI pitches Codex for tax prep after a 7,000-return pilot (2026-08-20) &lt;a href="https://runtimewire.com/article/openai-codex-tax-prep-7000-return-pilot" rel="noopener noreferrer"&gt;https://runtimewire.com/article/openai-codex-tax-prep-7000-return-pilot&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;explainx.ai, Codex as a Platform: OpenAI Opens Up Its Agent Harness to Builders (2026-08-20) &lt;a href="https://explainx.ai/blog/codex-as-a-platform-open-agent-harness-august-2026" rel="noopener noreferrer"&gt;https://explainx.ai/blog/codex-as-a-platform-open-agent-harness-august-2026&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;DataNorth, DeepSeek releases V4-Pro-0813 and open sources Harness v0.1 &lt;a href="https://datanorth.ai/news/deepseek-releases-v4-pro-0813-and-harness-v0-1" rel="noopener noreferrer"&gt;https://datanorth.ai/news/deepseek-releases-v4-pro-0813-and-harness-v0-1&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;x.ai, Grok 4.6 on Amazon Bedrock (2026-08-19) &lt;a href="https://x.ai/news/grok-4-6-amazon-bedrock" rel="noopener noreferrer"&gt;https://x.ai/news/grok-4-6-amazon-bedrock&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;VentureBeat, GLM-5.3 hits the API at $$1.4/$$4.4 per million tokens (2026-08-19) &lt;a href="https://venturebeat.com/technology/glm-5-3-hits-the-api-at-1-4-4-4-per-million-tokens" rel="noopener noreferrer"&gt;https://venturebeat.com/technology/glm-5-3-hits-the-api-at-1-4-4-4-per-million-tokens&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>harness</category>
      <category>apigateway</category>
    </item>
    <item>
      <title>GPT-5.6 Sol Ultrafast: When Model Inference Becomes a Configurable "Speed Tier"</title>
      <dc:creator>MT_Notes</dc:creator>
      <pubDate>Thu, 20 Aug 2026 08:40:01 +0000</pubDate>
      <link>https://dev.to/mt_notes/gpt-56-sol-ultrafast-when-model-inference-becomes-a-configurable-speed-tier-349g</link>
      <guid>https://dev.to/mt_notes/gpt-56-sol-ultrafast-when-model-inference-becomes-a-configurable-speed-tier-349g</guid>
      <description>&lt;p&gt;OpenAI dropped a one-two punch today. On one hand, the Ultrafast tier for GPT-5.6 Sol entered limited preview, powered by Cerebras, delivering up to 14x the standard speed with 750 output tokens per second. On the other, the same model quietly appeared with a 50% limited-time discount on OpenRouter and Vercel, while AI coding platform Devin pushed an even steeper 70% promotional rate. Industry analyst SemiAnalysis put it bluntly: these platforms represent a tiny fraction of OpenAI's total usage, yet they are the primary data sources third parties use to estimate market share, suggesting the discounts may be a calculated exercise in "data narrative."&lt;br&gt;
For developers, both developments point to the same shift: model selection is evolving from "which model" to "which tier, through which gateway."&lt;/p&gt;
&lt;h2&gt;
  
  
  1. Ultrafast Is a New Speed Class, Not a New Model
&lt;/h2&gt;

&lt;p&gt;OpenAI explicitly defines Ultrafast as a "new speed class rather than a separate model." This means developers are still calling the same GPT-5.6 Sol weights, but the inference pipeline has been re-architected. Cerebras' Wafer-Scale Engine slashes inter-chip communication latency, compressing what was previously a multi-GPU batch into a near single-chip response rhythm, ultimately achieving throughput of up to 750 output tokens per second.&lt;br&gt;
To put that in perspective: standard GPT-5.6 Sol outputs at roughly 50–60 tokens/s. Ultrafast pushes that to 750 tokens/s, a 14x jump. For a 3,000-token technical summary, the standard tier makes you wait 50–60 seconds; Ultrafast finishes in about 4 seconds. In real-time interactive scenarios, that is a qualitative leap, not just a quantitative one.&lt;br&gt;
Early customers in the preview include Jane Street, Podium, Basis, and Rogo. John Crepezzi, AI Assistants lead at Jane Street, said the speed increase "enables different ways of using the models, and makes it practical for developers to work in a more focused and productive way alongside them." Internally, OpenAI is testing Ultrafast for incident response, real-time log reading, trace analysis, conversation synthesis, and fix validation during active outages, as well as for research workflows that previously required overnight batch jobs.&lt;/p&gt;
&lt;h2&gt;
  
  
  2. The Three-Dimensional Trade-Off: Capability, Cost, and Latency
&lt;/h2&gt;

&lt;p&gt;Traditionally, developers traded off capability against cost: stronger models commanded higher per-token prices. Ultrafast formally introduces "latency" as a third axis, creating a capability × cost × latency decision space.&lt;br&gt;
Which scenarios are latency-sensitive enough to pay the premium? OpenAI's examples include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Real-time market signal analysis (price windows may last only seconds)&lt;/li&gt;
&lt;li&gt;Complex multi-turn live customer support (users will not wait 30 seconds)&lt;/li&gt;
&lt;li&gt;Inventory validation and exception handling during e-commerce checkout&lt;/li&gt;
&lt;li&gt;Instant diagnostic assistance for engineers during system outages
The common thread: the cost of waiting exceeds the cost of compute. When "slow" causes business loss, paying a premium for "fast" is rational.
But Ultrafast remains in narrow preview, and OpenAI has not announced pricing. Based on industry norms, ultra-low-latency tiers typically cost 2–5x the standard rate. That means developers need finer-grained routing: standard tier for simple queries, Ultrafast for complex and time-sensitive tasks, rather than a one-size-fits-all approach.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  3. Discount Tactics and the "Market Share Narrative"
&lt;/h2&gt;

&lt;p&gt;In interesting contrast to Ultrafast, GPT-5.6 Sol is seeing aggressive discounts on third-party platforms. Devin offers 70% off API costs, while OpenRouter and Vercel provide 50% limited-time discounts. SemiAnalysis notes that OpenRouter and Vercel represent a small share of OpenAI's total API volume, yet they are the primary data sources used by third-party observers (such as Artificial Analysis and LangChain's model usage reports) to estimate market share. By discounting at these "data windows," OpenAI can artificially inflate adoption statistics in the metrics that analysts watch, without touching official API pricing.&lt;br&gt;
The practical impact on developers: the same model, same weights, can vary in price by several multiples depending on the gateway. Call directly through OpenAI's official API and you pay list price; route through OpenRouter or Devin and you might get half price or even a third. This fragmentation makes "where you call from" as important as "what you call."&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi86ruhoiws2mzrx1vv9d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi86ruhoiws2mzrx1vv9d.png" alt=" " width="800" height="275"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;
  
  
  4. A Unified Gateway: Let Speed and Price Stop Being Configuration Nightmares
&lt;/h2&gt;

&lt;p&gt;Faced with "same model, multiple speed tiers, multiple price gateways," what developers really need is not memorizing which platform is discounting today, but an automatic adapter with a unified interface. This is where a model routing hub comes in.&lt;br&gt;
Take wrouter.ai as an example. It provides a unified endpoint compatible with the OpenAI format:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;from openai import OpenAI

client = OpenAI(
    base_url="https://wrouter.ai/v1",
    api_key="your_wrouter_key"
)

# Standard tier: daily Q&amp;amp;A, document summarization
response = client.chat.completions.create(
    model="gpt-5.6-sol",
    messages=[{"role": "user", "content": "Summarize the core arguments of this paper"}]
)

# Ultrafast tier: real-time interaction, incident diagnosis
response_fast = client.chat.completions.create(
    model="gpt-5.6-sol-ultrafast",
    messages=[{"role": "user", "content": "Analyze the anomaly in this log"}]
)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With a single base_url switch, developers can migrate seamlessly between standard and ultrafast tiers without changing model invocation logic in their business code. Going further, simple latency detection combined with cost thresholds enables automatic routing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;def smart_route(prompt, max_latency_ms=2000):
    """
    If standard tier latency exceeds the threshold,
    automatically fall back to a cheaper alternative;
    if the task is marked urgent, go straight to ultrafast.
    """
    if prompt.get("urgent"):
        return "gpt-5.6-sol-ultrafast"
    # Real implementation can combine historical latency sampling with budget constraints
    return "gpt-5.6-sol"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The core value of wrouter.ai rests on three pillars:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Model completeness: Major frontier models (OpenAI, Anthropic, Google, Zhipu, DeepSeek, etc.) are unified under one key.&lt;/li&gt;
&lt;li&gt;Stability fallback: When one platform hits rate limits or a promotion ends, traffic automatically switches to prevent business disruption.&lt;/li&gt;
&lt;li&gt;Unified billing: Regardless of whether the underlying call goes through the official API, OpenRouter, or another channel, invoices are delivered in a single format, ending finance reconciliation headaches.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  5. Closing Thoughts
&lt;/h2&gt;

&lt;p&gt;The launch of GPT-5.6 Sol Ultrafast signals that large-model inference has officially entered the "speed tiering" era. The good news for developers is that choices are multiplying; the bad news is that decisions are getting more complex. Standard tier, ultrafast tier, third-party discount gateways, different platforms' hidden terms, these variables layered together make manual management nearly impossible.&lt;br&gt;
The answer remains the same old advice: build an abstraction layer above the model layer. Let routing algorithms decide "which path to take," let a unified interface shield you from "how complex the path is," and focus on your business itself. When latency becomes part of the product experience, whoever can make optimal model decisions at the millisecond level will gain the edge in the next wave of interactive AI.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>When GLM-5.3 Landed at Dawn: Why a Routing Layer Became the Default Choice</title>
      <dc:creator>MT_Notes</dc:creator>
      <pubDate>Wed, 19 Aug 2026 07:59:37 +0000</pubDate>
      <link>https://dev.to/mt_notes/when-glm-53-landed-at-dawn-why-a-routing-layer-became-the-default-choice-i9f</link>
      <guid>https://dev.to/mt_notes/when-glm-53-landed-at-dawn-why-a-routing-layer-became-the-default-choice-i9f</guid>
      <description>&lt;h2&gt;
  
  
  Lead: China's frontier model just joined the first tier overnight
&lt;/h2&gt;

&lt;p&gt;In the early hours of August 19, 2026, Zhipu officially opened the API for its new base model GLM-5.3. On the Artificial Analysis Intelligence Index, GLM-5.3 scored 60, putting it on the same shelf as the closed-source flagships Claude Fable 5 and GPT-5.6 Sol, and tying Moonshot's Kimi K3 for the open-source crown.&lt;br&gt;
The chart that really matters is the two-axis one: intelligence on the y-axis, average cost per completed task on the x-axis. At the same intelligence level, GLM-5.3 sits at the lowest single-task cost in the frontier group, pushing the Pareto frontier of "intelligence versus cost" noticeably outward. The lab's positioning is direct: frontier capability at the lowest per-task price.&lt;br&gt;
On the official timeline, the model weights will be released as open source next Friday. From now until the weekend, developers have two parallel windows: call the closed-source API today, and migrate to local or private cloud deployment once the weights drop.&lt;br&gt;
The past week has been the densest stretch of frontier releases in 2026. In just four days, SpaceXAI shipped Grok 4.6, Google shipped Gemini 3.7 Flash, DeepSeek took V4 Pro to GA, and Zhipu opened GLM-5.3. When "high intelligence" and "low unit price" are both pushed to the extreme, the advantage window of any single model keeps shrinking. What developers want is no longer "one more new model" but "one place to hold all of them."&lt;br&gt;
That is exactly why routing-layer gateways like wrouter.ai have been mentioned more and more often over the past six months.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. What exactly is GLM-5.3 strong at? Three key numbers
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1.1 Intelligence Index 60: across the frontier line&lt;/strong&gt;&lt;br&gt;
The Artificial Analysis Intelligence Index aggregates knowledge, reasoning, coding, and agentic evaluations to measure how a model performs on real, complex tasks. A 60 is not a record-smasher — Claude Opus 5 still leads at 63, and Claude Fable 5 and GPT-5.6 Sol sit in the same 60 band as GLM-5.3 — but it marks a clean inflection point: an open-source model has stably and reproducibly entered the frontier band.&lt;br&gt;
For a developer, "60" means something concrete: when you put GLM-5.3 in production and run real workloads, it will not be rejected by users for "falling short of the frontier tier." It means the model can be written into architecture docs with a straight face, plugged into ROI tables, and shown in quarterly reviews.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1.2 The Pareto frontier moves: same intelligence, lower cost&lt;/strong&gt;&lt;br&gt;
If "60" is where GLM-5.3 sits on the capability curve, its position on the two-axis cost-versus-intelligence chart is the more interesting one. At the same intelligence tier, GLM-5.3 has the lowest per-task cost among the frontier group. Artificial Analysis pegs GPT-5.6 Luna at around $0.7/task and GLM-5.2 close behind; GLM-5.3 is clearly aiming to push another notch lower.&lt;br&gt;
For a product that processes tens of millions of tokens a day, "per-task cost" multiplied by "task count" is the real bill. Bringing "frontier capability" down to a price that long-tail developers can actually afford is the real value of this generation of open-source models.&lt;br&gt;
It is worth noting that AA Index and price are not linearly related. A model's price advantage only means something when it can reliably complete tasks at the same intelligence tier — and the fact that GLM-5.3 can wear both labels, frontier intelligence and lowest per-task cost, is precisely because its post-training efficiency has been pushed to the limit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1.3 Coding + defensive security + long-horizon tasks: the use cases are pre-chosen&lt;/strong&gt;&lt;br&gt;
GLM-5.3 shares the same base model as the previous GLM-5.2, with gains coming from post-training. The three capabilities the lab highlights are: complex coding, defensive cybersecurity, and long-horizon tasks. That map almost mirrors what Grok 4.6 (long-horizon agent stability) and Gemini 3.7 Flash (coding price-performance) are selling at the same time.&lt;br&gt;
The industry consensus is now obvious: the second half of 2026 is no longer about who can hit the highest MMLU score, but about whose agent can stably run through a 200-step complex task. "Long-horizon task capability" is moving from a nice-to-have to a must-have, and every model that claims to be frontier has to prove itself on this axis.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Day-one integration: ZCode, GLM Coding Plan, and enterprise users
&lt;/h2&gt;

&lt;p&gt;GLM-5.3 is not following the "look great on a paper, slowly trickle into products" rhythm. From the moment the API went live, it landed in two specific products:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;ZCode: Zhipu's developer-facing coding platform; GLM-5.3 is already the default model.&lt;/li&gt;
&lt;li&gt;GLM Coding Plan: the enterprise coding subscription, priced the same as GLM-5.2. That means enterprise users get a near-zero-cost upgrade.
For an enterprise IT decision-maker, the "same price, new model" policy matters more than the "new model" headline: no budget re-approval, no competitive benchmarking, just swap GLM-5.2 for GLM-5.3 in production.
For individual developers, the more practical path is: call the API to validate prompts now, then switch to local inference or private cloud once the weights are open-sourced next Friday. Both paths are reachable through the same routing layer.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  3. Why a gateway is worth more than ever
&lt;/h2&gt;

&lt;p&gt;When the model ecosystem has three parallel tracks — open-source week, closed-source flagships, and long-horizon agents — the real developer pain has shifted from "which model is the strongest" to:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;How to switch models without rewriting code? Every vendor's API protocol, parameter names, and call details differ slightly.&lt;/li&gt;
&lt;li&gt;How to route intelligently across models? Coding on GLM-5.3, writing on Claude Fable 5, long-horizon agents on Grok 4.6.&lt;/li&gt;
&lt;li&gt;How to unify scattered billing and monitoring? Different vendors use different billing units, cycles, and rate-limit policies.&lt;/li&gt;
&lt;li&gt;How to keep up with weekly releases? This week alone brought GLM-5.3, Grok 4.6, and Gemini 3.7 Flash.
wrouter.ai is designed for exactly these four questions. It exposes a single OpenAI-compatible endpoint: change the base URL to &lt;a href="https://wrouter.ai/v1" rel="noopener noreferrer"&gt;https://wrouter.ai/v1&lt;/a&gt; and you can seamlessly switch between GLM-5.3, Claude Fable 5, GPT-5.6 Sol, Gemini 3.7 Flash, Grok 4.6, DeepSeek V4 Pro, and more, with API keys, billing, and rate limits unified in a single dashboard.
A few common code patterns (the official OpenAI SDK is enough, no extra dependencies required):
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;from openai import OpenAI

client = OpenAI(
    base_url="https://wrouter.ai/v1",
    api_key="YOUR_WROUTER_KEY"
)

# Scenario 1: complex coding task on GLM-5.3
resp = client.chat.completions.create(
    model="glm-5.3",
    messages=[{"role": "user", "content": "Refactor the 5 O(n^2) blocks in this Python script to O(n)."}]
)

# Scenario 2: long-horizon agent task on Grok 4.6
resp = client.chat.completions.create(
    model="grok-4.6",
    messages=[{"role": "user", "content": "Act as my research assistant and follow this GitHub issue until it is resolved."}]
)

# Scenario 3: complex analysis / long-form writing on Claude Fable 5
resp = client.chat.completions.create(
    model="claude-fable-5",
    messages=[{"role": "user", "content": "Based on this research report, write a 3,000-word market analysis."}]
)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A few engineering notes worth calling out:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Stability: wrouter.ai maintains primary-backup failover and load balancing across multiple upstream vendors. When a single upstream fails, traffic is automatically rerouted, so production does not stop because "that one model's API went down."&lt;/li&gt;
&lt;li&gt;Model completeness: coverage spans OpenAI, Anthropic, Google, xAI, Zhipu, DeepSeek, Alibaba, ByteDance, and other major vendors across domestic and overseas markets. New models are usually onboarded within 24-48 hours of release.&lt;/li&gt;
&lt;li&gt;Unified billing: priced by tokens and model tier, with all vendor bills merged into one wrouter.ai dashboard — developers only reconcile one invoice.
Here is a side-by-side view of a few representative frontier models currently available through wrouter.ai:&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdxmasm243mxs6dbmlufs.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdxmasm243mxs6dbmlufs.png" alt=" " width="800" height="510"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Closing: open-source week is not just new models, it is "affordable intelligence"
&lt;/h2&gt;

&lt;p&gt;GLM-5.3 going live and its weights about to be open-sourced is one of the most symbolic events in the open-source ecosystem of August 2026. Its meaning goes beyond "yet another Chinese model squeezing into the first tier": it is the moment "the ticket to frontier capability" is taken off the procurement desk and handed to anyone who writes code.&lt;br&gt;
For most developers, the "dizzying variety" of the model ecosystem is exactly the reason a routing layer exists. With GLM-5.3, Grok 4.6, and Gemini 3.7 Flash all shipping back to back, and more likely open-sourced next Friday, handing "integration" and "switching" to a stable, complete, and billing-unified middle tier is the more realistic engineering choice.&lt;br&gt;
The bar is being pushed down, the tools are being consolidated, and what is left for developers is the freedom to focus on the product itself.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>apigateway</category>
      <category>glm</category>
    </item>
    <item>
      <title>GPT-5.6 Multi-Agent v2 Goes Live: When Agents Start Picking Models for You, How Should Your API Routing Layer Adapt?</title>
      <dc:creator>MT_Notes</dc:creator>
      <pubDate>Wed, 19 Aug 2026 07:54:53 +0000</pubDate>
      <link>https://dev.to/mt_notes/gpt-56-multi-agent-v2-goes-live-when-agents-start-picking-models-for-you-how-should-your-api-acc</link>
      <guid>https://dev.to/mt_notes/gpt-56-multi-agent-v2-goes-live-when-agents-start-picking-models-for-you-how-should-your-api-acc</guid>
      <description>&lt;h2&gt;
  
  
  Introduction: From "Choosing Generals" to "Assigning Tasks"
&lt;/h2&gt;

&lt;p&gt;For the past two years, developers have felt like shoppers in a supermarket that keeps expanding: GPT-4, Claude, Gemini, DeepSeek, Qwen... Each model comes with its own API shape, billing dimension, and capability curve. Once a product is built, the recurring nightmare is rarely that the model is too weak; it is that "the model we tuned last month has already been overtaken, and the code has to change again."&lt;br&gt;
Around August 16, OpenAI rolled out GPT-5.6 Multi-Agent v2 to all Codex users. Its most understated yet paradigm-shifting change is this: the main agent can now automatically delegate subtasks to different models, and each sub-agent can set its own reasoning intensity. OpenAI President Greg Brockman summed it up plainly: this is a step "toward saying goodbye to manually picking models."&lt;br&gt;
Behind that sentence lies a broader migration in the AI application layer: model selection is shifting from human experience to system scheduling. The developer's job is no longer to maintain a hard-coded model mapping table, but to build a routing layer where models can come and go freely.&lt;/p&gt;
&lt;h2&gt;
  
  
  1. What Exactly Did Multi-Agent v2 Change?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1.1 Architecture: The Main Agent as "Foreman," Sub-Agents by Strength&lt;/strong&gt;&lt;br&gt;
The core design of GPT-5.6 Multi-Agent v2 can be captured in one sentence: break tasks down and automatically match them to model tiers by difficulty and cost.&lt;br&gt;
In the current model lineup available to ChatGPT and Codex, the roles are roughly:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdzq8197sofgu95at9gbl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdzq8197sofgu95at9gbl.png" alt=" " width="800" height="440"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The main agent no longer requires developers to explicitly specify model before each call. Instead, it automatically selects Sol, Terra, or Luna based on the subtask's complexity, context length, latency requirements, and tool dependencies. Each sub-agent can also independently configure reasoning_effort, enabling differentiated inference within the same model family.&lt;br&gt;
Three weeks earlier, Luna had been rejected by the system for multi-agent delegation because it lacked inter-agent communication support, prompting posts like "Give us back Luna" on GitHub and the OpenAI community. The v2 update fixed this, meaning the lightweight model is now truly part of the automatic scheduling pool, not just a fallback for the main agent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1.2 Cost Control: 20% of Steps Eat 80% of the Compute Budget&lt;/strong&gt;&lt;br&gt;
One key figure from OpenAI is that only about 20% of steps in complex tasks need the strongest model; the rest can go to cheaper tiers. It sounds like another Pareto distribution, but it has serious engineering implications.&lt;br&gt;
A few publicly verified examples:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Hypha AI used Luna for document extraction, retaining about 98% of GPT-5.5's accuracy at 1/18 the cost.&lt;/li&gt;
&lt;li&gt;Browser Use ran Luna on 106 of the hardest browser tasks, completing 78% for about $14, while the strongest model cost about $235 to reach 80%.&lt;/li&gt;
&lt;li&gt;PlayerZero cut inference costs by 64% and response time by 90% on a multi-agent engineering code-retrieval task, while improving F1 by 5 points.&lt;/li&gt;
&lt;li&gt;On ARC-AGI-3, Sol jumped from 13.3% to 38.3% after enabling "reasoning persistence across turns + long-context compression," while output tokens dropped by about 6x.
These numbers point to the same conclusion: cost optimization is not about switching to a cheaper model, but about using the right model in the right place.
&lt;strong&gt;1.3 Performance Foundation: Long Conversations No Longer Freeze&lt;/strong&gt;
Multi-agent parallelism only works if the platform can handle long contexts and high concurrency. OpenAI published internal benchmarks: in a 741-turn, 231 MB Codex session, app load speed dropped from 27.62 seconds to 1.66 seconds, heap memory growth fell by 87.8%, network requests dropped from 894 to 16, and conversation entry loading dropped from 15,529 to 64.
The logic is lazy loading: opening a conversation no longer renders the entire history, only the state slices currently needed. For enterprise agent deployments, this is a threshold-level improvement. Long tasks must not crash, and multi-agent concurrency must hold up, before the platform can carry real workflows.
&lt;strong&gt;1.4 Direct Takeaways for Developers&lt;/strong&gt;
Multi-Agent v2 does not mean writing less code; it means writing code in the right places:&lt;/li&gt;
&lt;li&gt;Stop hard-coding model = "gpt-5.6-sol" in business logic; leave model selection to the routing layer.&lt;/li&gt;
&lt;li&gt;Use reasoning_effort instead of temperature for sampling control, which is the recommended approach for the GPT-5.6 family.&lt;/li&gt;
&lt;li&gt;Design agents around tasks that are "parallelizable and summarizable," rather than stuffing all context into a single call.&lt;/li&gt;
&lt;li&gt;Watch concurrency and nesting limits: Codex defaults to agents.max_threads=6 and agents.max_depth=1. Pushing these too aggressively makes token usage and latency grow exponentially.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  2. Hands-On: Building a Model-Agnostic Agent Router with a Unified Interface
&lt;/h2&gt;

&lt;p&gt;Below is a minimal Python example showing how to combine "task grading + unified base_url." The idea mirrors OpenAI Multi-Agent v2: let a small model do initial triage, then decide whether to call a stronger model.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;from openai import OpenAI
import os

client = OpenAI(
    api_key=os.getenv("WROUTER_API_KEY"),
    base_url="https://wrouter.ai/v1",
)

def route_task(prompt: str) -&amp;gt; dict:
    # Step 1: lightweight model grades the task
    router_resp = client.chat.completions.create(
        model="gpt-5.6-luna",
        messages=[{
            "role": "system",
            "content": "You are a task router. Reply with one word: easy, medium, or hard."
        }, {"role": "user", "content": prompt}],
        max_tokens=5,
    )
    level = router_resp.choices[0].message.content.strip().lower()

    # Step 2: pick execution model by level
    model_map = {
        "easy": "gpt-5.6-luna",
        "medium": "gpt-5.6-terra",
        "hard": "gpt-5.6-sol",
    }
    model = model_map.get(level, "gpt-5.6-terra")

    exec_resp = client.chat.completions.create(
        model=model,
        messages=[{"role": "user", "content": prompt}],
        reasoning_effort="medium",
    )
    return {
        "level": level,
        "model": model,
        "content": exec_resp.choices[0].message.content,
    }

if __name__ == "__main__":
    result = route_task("Refactor this FastAPI project to support async database connection pools.")
    print(f"Routed level: {result['level']}, actual model: {result['model']}")
    print(result["content"][:500])
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The key here is not the grading itself, but that all models go through the same base_url. When the main agent needs to dispatch subtasks to different models, a single entry point avoids writing separate authentication, retry, and billing logic for every provider.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Production Context: Why a Routing Layer Matters in a Multi-Model World
&lt;/h2&gt;

&lt;p&gt;Multi-Agent v2 sends a clear signal: future application model calls will look more and more like microservices, multiple models running in parallel, scaling by load, and priced by capability. But it also means the number of model endpoints, billing dimensions, and failure modes developers must manage will multiply.&lt;br&gt;
This is where a unified API gateway becomes valuable. Take wrouter.ai as an example; it maps directly onto the needs of the Multi-Agent v2 era:&lt;br&gt;
Stability. When the main agent dispatches subtasks in parallel to multiple models, any upstream provider's rate limit or transient failure can slow the entire workflow. A unified gateway can use load balancing and automatic retries to minimize the impact of single-point jitter on the agent system.&lt;br&gt;
Model completeness. A multi-agent system will not be tied to OpenAI alone. Sol for coding, Claude for long text, DeepSeek for low-cost reasoning, Qwen for Chinese scenarios, if each has its own SDK, the agent's orchestration logic gets polluted by vendor differences. A unified interface lets developers treat different models as the same pool of "compute resources."&lt;br&gt;
Unified billing. When 20% of steps use a flagship model and 80% use a lightweight model, the bill comes from multiple providers, currencies, and billing granularities. Unified billing makes cost attribution traceable and makes per-task or per-agent budgets feasible.&lt;br&gt;
In other words, OpenAI handles automatic model selection at the application layer, while developers still need a model-agnostic access plane at the infrastructure layer. The latter does not decide which model to use, but it determines whether you are free to use any model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion: Let Agents Be Smart, and Keep Yourself Unlocked
&lt;/h2&gt;

&lt;p&gt;The launch of GPT-5.6 Multi-Agent v2 is not another leaderboard refresh. It removes "model selection" from the developer's shoulders. For ordinary developers, this means you can focus more on task decomposition and business logic, and less on which provider just cut prices or released a new benchmark.&lt;br&gt;
But see the other side of the coin: when agent systems start automatically switching between models, the coupling points between your code and those models become more hidden. If routing, authentication, and billing are still scattered across each vendor's SDK, the operational complexity will quickly eat the flexibility that "automatic model selection" promises.&lt;br&gt;
So the next step is clear: let agents be smart, and let a unified interface keep you free.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>apigateway</category>
    </item>
    <item>
      <title>The Model Became a Plugin. The Bill Didn't.</title>
      <dc:creator>MT_Notes</dc:creator>
      <pubDate>Tue, 18 Aug 2026 07:59:07 +0000</pubDate>
      <link>https://dev.to/mt_notes/the-model-became-a-plugin-the-bill-didnt-3oap</link>
      <guid>https://dev.to/mt_notes/the-model-became-a-plugin-the-bill-didnt-3oap</guid>
      <description>&lt;h2&gt;
  
  
  1. What actually happened this week
&lt;/h2&gt;

&lt;p&gt;On August 13, with no launch event, DeepSeek shipped two things at once: the GA build of V4-Pro (0813), and an open-source agent runtime called DeepSeek Harness (CLI name dsh). One is the model. The other is the shell the model runs inside. Four days later, the attention split in a way few predicted. The model got buried under the pricing story, while the shell — MIT licensed, written in Node.js, still carrying a "developer preview" warning — climbed toward six figures in GitHub stars, with a community plugin ecosystem forming over the weekend.&lt;br&gt;
Harness has exactly one design idea: everything is a plugin. Model adapters, the tool registry, the session log, sandboxes, the filesystem, and the agent loop itself are all plugins, and all of them are replaceable. The documentation's phrase "no privileged core to patch" is the whole point: extending Harness does not mean editing Harness. It means mounting another plugin beside the existing ones. Underneath sits the Cordis meta-framework, derived from a paper on a programming paradigm for spatiotemporal composability.&lt;br&gt;
Two other things landed the same week, pointing the same direction:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Around August 15, Codex's Multi Agents v2 gained cross-model delegation: a capable orchestrator such as GPT-5.6 Sol can hand narrowly scoped grunt work to the faster, cheaper Luna. Per the developer who surfaced the change, Luna workers are "pure sub agents" that cannot message each other or spawn further agents, and the default behavior still clones the parent's model and settings. Getting the new path currently requires prompting for it.&lt;/li&gt;
&lt;li&gt;Agent Plugins 1.0 reached general availability in GitHub Copilot, refined with Vercel, AWS, Anysphere, GitHub, Microsoft and OpenAI, with Google joining as a core maintainer on launch day. The spec is deliberately small: a plugin.json manifest, an optional skills/ directory, an optional mcp.json.
Read together, all three say the same thing: the model is being demoted to a configuration value, and the runtime is being promoted to an architectural decision.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  2. What's swappable and what isn't
&lt;/h2&gt;

&lt;p&gt;For two years the default mental model has been that the weights matter and everything around them is glue. Harness's four presets invert that assumption.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fajtm8cgps8u1gvdv4nmn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fajtm8cgps8u1gvdv4nmn.png" alt=" " width="800" height="346"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Note minimal. When DeepSeek published V4-Pro-0813's numbers on public coding-agent benchmarks — 87.9 on Terminal Bench 2.1, 62.7 on DeepSWE, 74.1 on Toolathlon-Verified — the model was running inside Harness in minimal mode for some of them. Which means a "model score" has contained a harness contribution all along. Swap the shell and the number moves. That isn't contamination; it's the honest shape of the thing. Agent capability was always a joint product of model and runtime.&lt;br&gt;
code mode deserves a closer look. Conventional function calling is a loop: the model emits a tool call, the runtime executes it, the result is appended to context, the model emits the next one. Five operations mean five full round trips, each one re-processing the accumulated context. Code mode compiles the tool surface into a TypeScript SDK, hands it to the model, and lets the model write a single program. What you save is not only latency but four prefills you no longer pay for.&lt;br&gt;
On the swappable side, the Harness provider catalog covers Anthropic, OpenAI, AWS Bedrock, Azure, Google's enterprise agent platform and DeepSeek's own endpoint — plus an explicit slot for custom OpenAI-compatible gateways. It even ships two subagent providers that hand work directly to Claude Code and Codex, both off by default, both resolving the vendor binary from your PATH so you supply the install and the login. There is an MCP client, Agent Client Protocol support, and it reads AGENTS.md and CLAUDE.md.&lt;br&gt;
A Chinese lab shipped an agent framework that runs a competitor's model with zero friction. That looks like generosity. It is closer to a bet: models will commoditize; harnesses won't. Changing models is a config line. Changing runtimes means rewriting session storage, rerunning regressions, retraining a team. VentureBeat put it bluntly — for enterprise developers, Harness may be the more consequential half of the August 13 announcement.&lt;br&gt;
There is an irony here. In the same week Harness made models trivially swappable, DeepSeek's billing became harder to swap. From 16:00 UTC on August 16 (midnight Beijing time on August 17), the V4 family moved to peak/off-peak pricing. V4-Pro output went from a flat $0.87 per million tokens to $1.98 off-peak and $3.96 at peak; cache-hit input went from $0.003625 to $0.022 off-peak and $0.044 at peak — roughly 5x even at the discounted rate, about 12x at peak. And long sessions, repository analysis and subagent fan-out — exactly what Harness is built for — are the most cache-intensive workloads there are.&lt;br&gt;
So developers get an odd pairing: the runtime layer has never been this portable, and the cost layer has never demanded this much arithmetic. The model is a plugin. The bill isn't.&lt;/p&gt;
&lt;h2&gt;
  
  
  3. Finishing the sentence "everything is a plugin"
&lt;/h2&gt;

&lt;p&gt;For "everything is a plugin" to hold in practice, one piece is still missing: something coherent behind the plugin. Otherwise you list five providers in config and inherit five keys, five invoices, five quota alarms and five rate-limit dialects. The configuration unified; the operations didn't.&lt;br&gt;
That is the slot a model gateway fills. wrouter.ai exposes a single OpenAI-compatible entry point, which is precisely what Harness's "custom OpenAI-compatible gateway" plugin expects:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;from openai import OpenAI

client = OpenAI(api_key="wr-***", base_url="https://wrouter.ai/v1")

# Orchestrator: planning and decomposition
plan = client.chat.completions.create(
    model="deepseek-v4-pro",
    messages=[{"role": "user", "content": "Analyze the auth module in this repo and list refactor steps"}],
)

# Workers: well-scoped grunt work on a cheaper model, same key, same invoice
for step in parse_steps(plan.choices[0].message.content):
    client.chat.completions.create(
        model="deepseek-v4-flash",
        messages=[{"role": "user", "content": step}],
    )
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Wiring it into a Harness model plugin is the same move: point the provider at &lt;a href="https://wrouter.ai/v1" rel="noopener noreferrer"&gt;https://wrouter.ai/v1&lt;/a&gt;, then assign model names by role inside the preset — a strong model for the orchestrator, a fast one for subagents, a third-party model for cross-checking during evaluation. Three properties do the work here. Interface stability means a preview-stage project that openly promises breaking changes at least won't break on the model side. Model coverage means "changing models is one line" holds across vendors, not just within one vendor's product line. Unified billing means peak/off-peak rates, cache hit ratios and the orchestrator-versus-worker split are legible in one place, instead of being reconstructed by subtraction across five invoices at month end.&lt;br&gt;
One practical note: since Harness records everything the model sees in an append-only session log, and billing is now time-of-day dependent, write the model name, token counts and request timestamp into that same log. When you need to answer "how much did last week's runaway agent session cost, and which step was expensive," that log will be more useful than any invoice.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Closing
&lt;/h2&gt;

&lt;p&gt;Harness went viral this week not because it invented a feature, but because it wrote down as architecture something everyone was already doing quietly: the model is a replaceable part, and the runtime is the asset. Codex's cross-model delegation and the Agent Plugins 1.0 packaging format are the same sentence in different words.&lt;br&gt;
The takeaway for developers is concrete. Assume you will replace every model you currently use within six months, then design your call layer for that assumption. Make model names configuration. Make providers plugins. Collapse billing into one place. That way, the next time a vendor silently repoints an endpoint at 2 a.m. or announces a 12x increase on cached input, the change on your side is a string.&lt;br&gt;
If you're building that layer now, wrouter.ai can serve as the unified entry point — closing out half the cost and stability problems of multi-model orchestration so you can go back to tuning the part that actually differentiates you: the agent loop.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sources&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;DeepSeek Harness repository (MIT, developer preview): &lt;a href="https://github.com/deepseek-ai/deepseek-harness" rel="noopener noreferrer"&gt;https://github.com/deepseek-ai/deepseek-harness&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;The New Stack, "DeepSeek open sources an agent harness where everything is a plugin": &lt;a href="https://thenewstack.io/deepseek-harness-open-source-plugins/" rel="noopener noreferrer"&gt;https://thenewstack.io/deepseek-harness-open-source-plugins/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;The Register, "DeepSeek's innovative harness treats everything as a plug-in": &lt;a href="https://www.theregister.com/ai-and-ml/2026/08/14/deepseeks-innovative-harness-treats-everything-as-a-plug-in/5288095" rel="noopener noreferrer"&gt;https://www.theregister.com/ai-and-ml/2026/08/14/deepseeks-innovative-harness-treats-everything-as-a-plug-in/5288095&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;VentureBeat, "DeepSeek Harness launches as open-source rival to Claude Code": &lt;a href="https://venturebeat.com/technology/deepseek-harness-launches-as-open-source-rival-to-claude-code-alongside-v4-pro-on-api-with-higher-prices" rel="noopener noreferrer"&gt;https://venturebeat.com/technology/deepseek-harness-launches-as-open-source-rival-to-claude-code-alongside-v4-pro-on-api-with-higher-prices&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;HPCwire, "DeepSeek Open-Sources the Missing Layer Between AI Models and Agents": &lt;a href="https://www.hpcwire.com/2026/08/14/deepseek-open-sources-the-missing-layer-between-ai-models-and-agents/" rel="noopener noreferrer"&gt;https://www.hpcwire.com/2026/08/14/deepseek-open-sources-the-missing-layer-between-ai-models-and-agents/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;RuntimeWire, "OpenAI lets GPT-5.6 Sol delegate grunt work to cheaper Luna agents": &lt;a href="https://runtimewire.com/article/openai-lets-gpt-5-6-sol-delegate-grunt-work-to-cheaper-luna-agents" rel="noopener noreferrer"&gt;https://runtimewire.com/article/openai-lets-gpt-5-6-sol-delegate-grunt-work-to-cheaper-luna-agents&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;DeepSeek API changelog (peak/off-peak pricing): &lt;a href="https://api-docs.deepseek.com/zh-cn/updates" rel="noopener noreferrer"&gt;https://api-docs.deepseek.com/zh-cn/updates&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>apigateway</category>
      <category>harness</category>
    </item>
  </channel>
</rss>
