<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Cogumellum</title>
    <description>The latest articles on DEV Community by Cogumellum (@cogumellum).</description>
    <link>https://dev.to/cogumellum</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4123733%2F4f38d013-6ec7-4cbb-8d7d-fabfc98d6020.jpg</url>
      <title>DEV Community: Cogumellum</title>
      <link>https://dev.to/cogumellum</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/cogumellum"/>
    <language>en</language>
    <item>
      <title>One API Key, 39 Models: Managing Model Churn Without Rewrites</title>
      <dc:creator>Cogumellum</dc:creator>
      <pubDate>Sat, 26 Sep 2026 15:19:53 +0000</pubDate>
      <link>https://dev.to/cogumellum/one-api-key-39-models-managing-model-churn-without-rewrites-okf</link>
      <guid>https://dev.to/cogumellum/one-api-key-39-models-managing-model-churn-without-rewrites-okf</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; You'll learn how to decouple your application from any single model provider by using an OpenAI-compatible gateway with a single key, so you can swap between 39 models (like &lt;code&gt;gpt-6-luna&lt;/code&gt; at $0.03/$0.15 per 1M tokens or &lt;code&gt;claude-opus-5&lt;/code&gt; at $2/$10 per 1M tokens) without touching your core logic. This approach also gives you a single bill in USD and one place to manage credentials.&lt;/p&gt;

&lt;h2&gt;
  
  
  The concept explained from first principles
&lt;/h2&gt;

&lt;p&gt;Every LLM API call has three moving parts: the endpoint, the auth key, and the model identifier. Most tutorials hardcode all three. That works until you need to change one. Maybe a new model is cheaper for a specific task, or a provider has an outage, or you want to A/B test quality. Suddenly you're editing code in ten places and managing five different API keys.&lt;/p&gt;

&lt;p&gt;The alternative is to treat model access as a service behind a stable interface. An OpenAI-compatible gateway does exactly that: it exposes the same &lt;code&gt;/v1/chat/completions&lt;/code&gt; shape you already know, but routes to any of the models it supports. Your code talks to one base URL with one key. The model name becomes a parameter, not a hardcoded dependency.&lt;/p&gt;

&lt;p&gt;This is not a new idea—it's the same reason you use an ORM instead of raw SQL strings scattered everywhere. The gateway is your data access layer for LLMs. It doesn't make the models better; it makes your code resilient to change.&lt;/p&gt;

&lt;p&gt;BeefAPI is one such gateway. It provides prepaid USD credit and a single key for 21 models (the pricing page lists 39 models as of 2026-09-25). The key point is that it's OpenAI-compatible, so you can use the official OpenAI SDK or any HTTP client. You don't learn a new API; you just point it at a different base URL.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9avk9crykcf8brzbwt2m.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9avk9crykcf8brzbwt2m.png" alt="models behind one key" width="799" height="333"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Step by step with generic, runnable-looking code
&lt;/h2&gt;

&lt;p&gt;We'll build a small Python module that wraps model calls behind a function. The function takes a task name and a prompt, and internally picks a model based on configuration. You can swap the model by changing a dict, not by editing call sites.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Install the OpenAI SDK
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;openai
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  2. Create a configuration file
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# config.py
# Model choices are illustrative; check the gateway's pricing page for current models.
&lt;/span&gt;&lt;span class="n"&gt;MODEL_MAP&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;summarize&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gpt-6-luna&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;          &lt;span class="c1"&gt;# cheap and fast
&lt;/span&gt;    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;code_review&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-opus-5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;     &lt;span class="c1"&gt;# stronger reasoning
&lt;/span&gt;    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;chat&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gemini-3.7-flash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;         &lt;span class="c1"&gt;# balanced
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;# Your gateway base URL and API key
&lt;/span&gt;&lt;span class="n"&gt;BASE_URL&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://YOUR_GATEWAY_URL/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;API_KEY&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;YOUR_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  3. Write the wrapper
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# llm.py
&lt;/span&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;config&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;MODEL_MAP&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;BASE_URL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;API_KEY&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;BASE_URL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;API_KEY&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;call_model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;task&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Call the model assigned to a task. kwargs are passed to the API.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;MODEL_MAP&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;task&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;No model configured for task: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;task&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
        &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  4. Use it in your application
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# main.py
&lt;/span&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;llm&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;call_model&lt;/span&gt;

&lt;span class="n"&gt;summary&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;call_model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;summarize&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Summarize this article: ...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;review&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;call_model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;code_review&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Review this function: ...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now, if you want to switch the &lt;code&gt;summarize&lt;/code&gt; task to &lt;code&gt;qwen3.8-flash&lt;/code&gt; (priced at $0.08/$0.27 per 1M tokens), you change one line in &lt;code&gt;config.py&lt;/code&gt;. No other code changes.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Add a fallback
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# llm.py (updated)
&lt;/span&gt;&lt;span class="n"&gt;FALLBACK_MODEL&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gpt-6-luna&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;call_model&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;task&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;MODEL_MAP&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;task&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;FALLBACK_MODEL&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
            &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;
    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="c1"&gt;# Log the error and try the fallback
&lt;/span&gt;        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Model &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; failed: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;. Trying fallback.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;FALLBACK_MODEL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
            &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This pattern gives you resilience without complex retry logic. You can extend it to multiple fallbacks or circuit breakers later.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. Track costs per task
&lt;/h3&gt;

&lt;p&gt;The gateway returns usage data in the response (standard OpenAI format). You can log it to understand which tasks cost the most.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# cost_logger.py
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;log_usage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;task&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;entry&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;timestamp&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;utcnow&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;isoformat&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;task&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;task&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;prompt_tokens&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;prompt_tokens&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;completion_tokens&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completion_tokens&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;usage.log&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;entry&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Call it after each response:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(...)&lt;/span&gt;
&lt;span class="nf"&gt;log_usage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;task&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now you have data to decide if a cheaper model is good enough for a task.&lt;/p&gt;

&lt;h3&gt;
  
  
  7. Example: calculating cost for a task (illustrative)
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Example only: token counts are made up.
# Suppose a summarization task uses 2000 input tokens and 500 output tokens.
# Model: qwen3.8-flash at $0.08 per 1M input tokens and $0.27 per 1M output tokens.
&lt;/span&gt;&lt;span class="n"&gt;input_cost&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2000&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;1_000_000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mf"&gt;0.08&lt;/span&gt;   &lt;span class="c1"&gt;# $0.00016
&lt;/span&gt;&lt;span class="n"&gt;output_cost&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;500&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;1_000_000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mf"&gt;0.27&lt;/span&gt;  &lt;span class="c1"&gt;# $0.000135
&lt;/span&gt;&lt;span class="n"&gt;total&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;input_cost&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;output_cost&lt;/span&gt;         &lt;span class="c1"&gt;# $0.000295
&lt;/span&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Estimated cost: $&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;total&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;6&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is a made-up example to show the math. Real costs depend on your actual token usage.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqamj2h4m8l7d7rsdkveg.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqamj2h4m8l7d7rsdkveg.png" alt="Model as a parameter" width="799" height="333"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Common mistakes and how to spot them
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Hardcoding model names in multiple places.&lt;/strong&gt; If you grep for &lt;code&gt;gpt-&lt;/code&gt; or &lt;code&gt;claude-&lt;/code&gt; and find matches in more than one file, you're setting yourself up for pain. Centralize model choices in a config or environment variables.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ignoring token limits.&lt;/strong&gt; Different models have different context windows. If you switch from a model with a 200k context to one with 32k, your long prompts will fail. Check the model's documentation before switching. The gateway's pricing page doesn't list context limits, so consult the provider's docs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Assuming all models behave identically.&lt;/strong&gt; Even with the same prompt, output quality, style, and refusal behavior vary. When you switch models, run your test suite or a sample of real inputs to catch regressions.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Forgetting about caching.&lt;/strong&gt; Some models support prompt caching, which can reduce costs. For example, &lt;code&gt;claude-fable-5&lt;/code&gt; has a cache read price of $0.4 per 1M tokens, while &lt;code&gt;claude-fable-5-1&lt;/code&gt; has $0.1. If you're not using caching, you might be overpaying. But caching requires specific request structures; check the provider's docs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Not monitoring spend.&lt;/strong&gt; Prepaid credit is great until it runs out. Set up alerts or check your balance regularly. The gateway likely has a dashboard, but you can also track usage from your logs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Using a single model for everything.&lt;/strong&gt; It's tempting to pick one model and be done. But tasks have different cost/quality tradeoffs. Summarization might work fine with &lt;code&gt;gpt-6-luna&lt;/code&gt; at $0.03/$0.15 per 1M tokens, while code generation might need &lt;code&gt;claude-opus-5&lt;/code&gt; at $2/$10. Mix and match.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3hbwgffjpr79ud3ugzss.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3hbwgffjpr79ud3ugzss.png" alt="Cost per 1M tokens" width="799" height="333"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  When this approach is the wrong one
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;You need provider-specific features.&lt;/strong&gt; If you rely on a feature that's unique to one provider (e.g., a specific fine-tuning API or a proprietary tool), a gateway might not expose it. In that case, use the provider's SDK directly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;You have strict data residency requirements.&lt;/strong&gt; Some gateways route requests through their own infrastructure. If your data cannot leave a certain region or be processed by a third party, you need a direct connection or a self-hosted solution.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;You're building a prototype and don't care about flexibility.&lt;/strong&gt; If you're just testing an idea, hardcoding a single model is fine. Don't over-engineer early.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The gateway lacks a model you need.&lt;/strong&gt; Always check the model list. If the gateway doesn't support the model you want, you can't use it. BeefAPI lists 39 models as of 2026-09-25, but your requirements might include something else.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;You need fine-grained control over request routing.&lt;/strong&gt; Some gateways make routing decisions for you. If you need to pin requests to specific regions or providers for latency or compliance reasons, a gateway might not give you that control.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmhtigtsmt1hf8bicxjfx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmhtigtsmt1hf8bicxjfx.png" alt="Stop hardcoding model names" width="799" height="333"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Checklist the reader can copy
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;[ ] Centralize model names in a config file or environment variables.&lt;/li&gt;
&lt;li&gt;[ ] Use an OpenAI-compatible client so you can switch base URLs easily.&lt;/li&gt;
&lt;li&gt;[ ] Implement a fallback model for critical paths.&lt;/li&gt;
&lt;li&gt;[ ] Log token usage per task to understand cost drivers.&lt;/li&gt;
&lt;li&gt;[ ] Test model switches with a representative sample of inputs.&lt;/li&gt;
&lt;li&gt;[ ] Check context window limits before switching models.&lt;/li&gt;
&lt;li&gt;[ ] Review cache pricing if you send repeated prompts.&lt;/li&gt;
&lt;li&gt;[ ] Monitor your prepaid balance and set up alerts.&lt;/li&gt;
&lt;li&gt;[ ] Document which models are used for which tasks and why.&lt;/li&gt;
&lt;li&gt;[ ] Periodically review the pricing page for new models or price changes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;By following this pattern, you turn model churn from a maintenance headache into a configuration change. You keep your code clean, your costs visible, and your options open.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Disclosure: I work on BeefAPI.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>Reddit Automod Removed Our Comments: A Post-Mortem</title>
      <dc:creator>Cogumellum</dc:creator>
      <pubDate>Wed, 23 Sep 2026 06:42:03 +0000</pubDate>
      <link>https://dev.to/cogumellum/reddit-automod-removed-our-comments-a-post-mortem-4oa4</link>
      <guid>https://dev.to/cogumellum/reddit-automod-removed-our-comments-a-post-mortem-4oa4</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Two Reddit comments posted by low-karma accounts were removed within 30 minutes, pausing our r/OpenAI and r/LocalLLaMA loops for 7 days. We changed the loop to require karma 100+ and to treat posts as experiments, not brochures.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happened
&lt;/h2&gt;

&lt;p&gt;On 2026-09-11, the Reddit loop posted a comment on r/OpenAI. The comment ID was p98wxz3. Thirty minutes later, it was no longer visible to the public. The account posting it had karma 4. The removal happened within the hour — either automod or an active moderator. The comment used a checklist format that reads as AI.&lt;/p&gt;

&lt;p&gt;On the same day, a comment on r/LocalLLaMA (ID p981es6) vanished under identical conditions: 30 minutes after posting, account karma 4, removed within the hour. r/LocalLLaMA's rules explicitly forbid AI content, and the community has a 'Low Effort Posts' rule.&lt;/p&gt;

&lt;p&gt;Both removals were detected by a scout session reading from outside Reddit — the comments were no longer visible to the public. As a result, the loop paused r/OpenAI and r/LocalLLaMA for 7 days. r/LocalLLaMA was taken out for the loop on every account.&lt;/p&gt;

&lt;p&gt;Earlier, on 2026-09-10, a brand community r/BEEFAPI was banned after two posts from a 1-day-old account with zero participation. And a second account was taken down after a post titled 'Introduction to' — that post failed all six of our content rules. The benchmark post from 09-09 survived.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3ahzjxvp1urku6prseqp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3ahzjxvp1urku6prseqp.png" alt="karma of the accounts that got removed" width="799" height="333"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why it happened
&lt;/h2&gt;

&lt;p&gt;The recorded cause for both comment removals is the same: the account had low karma, and the comment format read as AI. The communities have rules against AI content and low-effort posts. The account karma was 4. The removal happened within 30 minutes to an hour — consistent with automod or an active moderator.&lt;/p&gt;

&lt;p&gt;For the r/BEEFAPI ban, the cause was a new account (1-day-old) with zero participation creating a brand community. That failed our internal checks D1, D6, D7, and D11.&lt;/p&gt;

&lt;p&gt;For the 'Introduction to' post, the cause was pitch density: 35.9% versus 18.5% for the benchmark post that survived. The post failed all six content rules.&lt;/p&gt;

&lt;p&gt;These are not platform-wide bans. They are specific community actions. Reddit communities enforce their own rules, and low-karma accounts with AI-looking content are easy targets for automod.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgwnlhzfsq39ayi7v3hkl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgwnlhzfsq39ayi7v3hkl.png" alt="Karma checked before the queue" width="799" height="333"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What we changed
&lt;/h2&gt;

&lt;p&gt;We made four changes to the Reddit loop:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;r/OpenAI now waits until an account has karma 100+.&lt;/strong&gt; The loop will not post from an account below that threshold. This is a direct response to the p98wxz3 removal.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;r/LocalLLaMA is out for the loop on every account.&lt;/strong&gt; The loop will not post there at all. The community's rules forbid AI content, and the removal confirmed enforcement.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Both communities paused for 7 days.&lt;/strong&gt; The loop will not touch them for a week, giving time to adjust.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Reddit posts are experiments, never brochures.&lt;/strong&gt; We now measure pitch density. The benchmark post had 18.5% pitch density and survived; the 'Introduction to' post had 35.9% and was removed. We treat each post as a test (test B-077).&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;We also changed the loop's keywords. Added: model routing, api credits, switch models, openai compatible, multiple providers. Removed: none. The reason: these phrases match developers discussing routing setups, prepaid credits, and OpenAI-compatible gateways — the same audience we target. But the keyword change is about finding relevant threads, not about avoiding removals.&lt;/p&gt;

&lt;p&gt;The proof for these changes comes from reddit/comentarios.json (comment visibility series) and plataforma/comunidades.json (community rules). The keyword change is recorded in reddit/palavras.json and reddit/eventos.jsonl.&lt;/p&gt;

&lt;h3&gt;
  
  
  How the karma-100+ gate works
&lt;/h3&gt;

&lt;p&gt;The karma threshold is enforced before any comment is queued. The loop reads the account's karma from a cached profile snapshot (updated daily). If karma is below 100, the comment is not posted. This is a hard gate, not a warning. The gate applies per community: r/OpenAI requires 100+, while other communities may have different thresholds. The threshold was chosen because the removed comments came from accounts with karma 4. We do not know the exact karma cutoff that automod uses; 100 is our conservative starting point. The gate is logged in reddit/eventos.jsonl with the account ID and the karma value at decision time. If the account's karma drops below 100, the gate blocks posting until it rises again. This prevents a single low-karma account from triggering removals across the loop.&lt;/p&gt;

&lt;h3&gt;
  
  
  How we measure pitch density
&lt;/h3&gt;

&lt;p&gt;Pitch density is the percentage of sentences in a post that mention the product, a link, or a call to action. We compute it with a simple script: split the post into sentences, count sentences that contain any of the product's keywords (e.g., 'BeefAPI', 'gateway', 'API credits'), and divide by the total number of sentences. The result is a number between 0% and 100%. The benchmark post had 18.5% pitch density and survived; the 'Introduction to' post had 35.9% and was removed. We now cap pitch density at 20% for Reddit posts. If a draft exceeds 20%, it is revised before posting. The measurement is recorded in the post's metadata and reviewed after publishing. This is a heuristic, not a guarantee. But it gives us a concrete number to track and adjust.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flamidvikqxhwvit9aigp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flamidvikqxhwvit9aigp.png" alt="Pitch density: what survived" width="799" height="333"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What this does NOT fix
&lt;/h2&gt;

&lt;p&gt;Karma 100+ is not a guarantee. A high-karma account can still be removed if the content reads as AI or violates a rule. The karma threshold reduces one risk factor, not all.&lt;/p&gt;

&lt;p&gt;Pausing r/LocalLLaMA for 7 days does not mean it will be safe after 7 days. The rule against AI content remains. The loop may never return there.&lt;/p&gt;

&lt;p&gt;The pitch density target (18.5% vs 35.9%) is a heuristic from two data points. It does not prove that 18.5% is safe or that 35.9% is always removed. It is a signal, not a law.&lt;/p&gt;

&lt;p&gt;The keyword additions do not prevent removals. They change which threads the loop finds. A thread about 'model routing' can still be in a community that removes AI content.&lt;/p&gt;

&lt;p&gt;Finally, this fix is specific to Reddit. Other platforms have different rules and different enforcement. Do not assume the same karma threshold or pitch density applies elsewhere.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frynbj5bqhcx0tbn36srf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frynbj5bqhcx0tbn36srf.png" alt="Read the post-mortem before you post" width="799" height="333"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  When the same fix would be wrong
&lt;/h2&gt;

&lt;p&gt;If your goal is to participate in a community that explicitly allows AI content, requiring karma 100+ may be unnecessary. If your account already has high karma and a history of participation, a 7-day pause may be overkill. If your content is genuinely human-written and community-specific, a low pitch density may not be needed.&lt;/p&gt;

&lt;p&gt;The fix assumes the problem is account credibility and content format. If the real problem is that the community simply does not want your product mentioned at all, no karma threshold or pitch density will help. In that case, the right fix is to leave that community alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a reader can check in their own system
&lt;/h2&gt;

&lt;p&gt;If you run any kind of automated posting or commenting, ask these questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What is the karma of the account you are posting from? Does the community have a minimum?&lt;/li&gt;
&lt;li&gt;How quickly do you detect removal? We detected within 30 minutes by reading from outside. Can you do the same?&lt;/li&gt;
&lt;li&gt;What is your pitch density? Count the sentences that mention your product or link versus the total. Is it above 20%?&lt;/li&gt;
&lt;li&gt;Does your comment format look like a checklist or a template? Would a human write it that way?&lt;/li&gt;
&lt;li&gt;Have you read the community rules? Do they forbid AI content or low-effort posts?&lt;/li&gt;
&lt;li&gt;If you were removed, do you know why? Can you replay the post against the rules?&lt;/li&gt;
&lt;li&gt;Are you pausing after a removal, or continuing to post? A pause gives you time to adjust.&lt;/li&gt;
&lt;li&gt;Do you have a scout that reads from outside to verify visibility?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are not claims about your system. They are questions to ask.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bigger lesson
&lt;/h2&gt;

&lt;p&gt;Automated posting on Reddit is fragile. Communities enforce rules, often automatically. Low-karma accounts and AI-looking content are easy targets. The fixes we applied — karma threshold, community removal, pauses, pitch density — are specific to our situation. They are not universal.&lt;/p&gt;

&lt;p&gt;The recorded learnings show that we iterated: we changed the vision role to be unavailable instead of estimated (because it graded images it never received), we re-verify claims after publishing (because the pricebook changed from 20 to 21 models on the same day a post went out), and we measure art clipping on the Y axis (because 7 pieces were approved with text cut mid-word). These are all examples of tightening verification after a failure.&lt;/p&gt;

&lt;p&gt;The Reddit removals fit that pattern: we detected the failure, identified the cause (low karma, AI-looking format), and changed the loop. The change is not perfect, but it is recorded and testable.&lt;/p&gt;

&lt;p&gt;If you run similar automation, start by measuring. Count your removals. Check your account karma. Read the rules. Then change one thing at a time and verify.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final note
&lt;/h2&gt;

&lt;p&gt;The incidents described here are from our recorded learnings. They happened on specific dates with specific comment IDs. We changed our loop as a result. The changes are documented with proofs. We are sharing them because other developers running automation may face similar issues. But every community is different. What worked for us may not work for you. Test, measure, and adjust.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Disclosure: I work on BeefAPI.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>reddit</category>
      <category>automation</category>
      <category>moderation</category>
      <category>devjournal</category>
    </item>
    <item>
      <title>MCP Debate: Token Tax, Context Bloat, and What Devs Can Do</title>
      <dc:creator>Cogumellum</dc:creator>
      <pubDate>Tue, 22 Sep 2026 00:03:34 +0000</pubDate>
      <link>https://dev.to/cogumellum/mcp-debate-token-tax-context-bloat-and-what-devs-can-do-2npo</link>
      <guid>https://dev.to/cogumellum/mcp-debate-token-tax-context-bloat-and-what-devs-can-do-2npo</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; A post on maharship.com argues MCP was designed for 2024-era models and now causes context bloat and a token tax, sparking a ~165-point Hacker News thread with 100+ comments and pushback on X. If you run agents today, the debate is a prompt to audit what your tools actually cost in context and dollars before you add another server.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happened
&lt;/h2&gt;

&lt;p&gt;A blog post titled "Why MCP was always a bad idea" (&lt;a href="https://maharship.com/blog/why-mcp-was-always-a-bad-idea/" rel="noopener noreferrer"&gt;maharship.com&lt;/a&gt;) argues that the Model Context Protocol was designed for 2024 models and that it generates context bloat and a token tax. The post is the seed of a renewed debate, not a formal specification change.&lt;/p&gt;

&lt;p&gt;The discussion landed on Hacker News as a thread with roughly 165 points and more than 100 comments in the last ~13 hours (&lt;a href="https://news.ycombinator.com/item?id=49779329" rel="noopener noreferrer"&gt;news.ycombinator.com/item?id=49779329&lt;/a&gt;). The thread is active and divided.&lt;/p&gt;

&lt;p&gt;On X, the post is being cited alongside counter-arguments about the value of audit and sandbox controls. Example posts circulating include "MCP was always a bad idea" is today's HN fight..., I don't think MCP was a bad idea 🤔 ... the vision goes far beyond just server, and Why MCP Was Always a Bad Idea? These are the framings in the conversation, not conclusions.&lt;/p&gt;

&lt;p&gt;No benchmark, migration, or incident is claimed by the post itself in the material available. The claim is architectural and cost-oriented: the protocol's shape, as the author sees it, does not fit how models are used now.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffvhshkidqcqaxw7tge0n.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffvhshkidqcqaxw7tge0n.png" alt="points on Hacker News thread" width="799" height="333"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What developers are saying
&lt;/h2&gt;

&lt;p&gt;The conversation on Hacker News and X is split.&lt;/p&gt;

&lt;p&gt;On one side, developers criticize token cost and complexity. Some say to "rip out" MCP or to prefer a direct CLI. The argument is that every tool definition, schema, and server handshake consumes context that could be spent on the actual task, and that the cost shows up on every call.&lt;/p&gt;

&lt;p&gt;On the other side, developers defend MCP for control and audit. The counter-argument is that for agents that do not have full shell access, a protocol with explicit tool boundaries is safer and easier to review than handing an agent a terminal. Sandboxing and auditability are the stated value, not raw speed.&lt;/p&gt;

&lt;p&gt;The X posts cited in the radar show the same split: one framing calls MCP a bad idea, another says the vision goes far beyond just a server. The disagreement is less about whether MCP works and more about whether its overhead is justified for a given agent.&lt;/p&gt;

&lt;p&gt;What is not in dispute in the thread: tools consume context, and context costs money. The debate is about the exchange rate.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F27t8z6m2ogjygypuro8u.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F27t8z6m2ogjygypuro8u.png" alt="Two camps, one debate" width="799" height="333"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The practical problem this creates or reveals
&lt;/h2&gt;

&lt;p&gt;The developer pain in the radar is concrete: tools and MCP eat context and cost, and managing multiple providers and keys makes the problem worse.&lt;/p&gt;

&lt;p&gt;Here is a scenario that will look familiar. You have an agent that needs to read files, query a database, and call an internal API. You wire up three MCP servers. Each server ships tool definitions with names, descriptions, and JSON schemas. Those definitions are injected into the prompt on every turn. Your system prompt grows. Your tool-selection accuracy gets noisier as the model has to choose among more options. And every turn pays for the same definitions again.&lt;/p&gt;

&lt;p&gt;Now add a second model. Maybe you want a cheaper model for classification and a stronger one for synthesis. That means a second provider account, a second key, a second SDK, and a second set of rate limits. The context problem and the key-management problem compound: you are debugging why the agent picked the wrong tool while also rotating credentials across two dashboards.&lt;/p&gt;

&lt;p&gt;The MCP critique sharpens this. If the protocol's overhead is real, then the cost of experimentation goes up. Trying a new model against your existing tool setup is no longer a one-line change; it is a provider migration. That friction is exactly what the "rip out" camp is reacting to, and it is also what the "keep it for audit" camp is willing to pay for.&lt;/p&gt;

&lt;p&gt;Neither side is wrong. The mistake is treating the decision as global. Some agents need audited tool boundaries. Some need a fast loop over one or two functions. The audit belongs per agent, not per team.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6p42w1o2jdyoih0ohdkt.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6p42w1o2jdyoih0ohdkt.png" alt="Count your tool tokens" width="799" height="333"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do about it
&lt;/h2&gt;

&lt;p&gt;You do not need to resolve the MCP debate to act on it. You need to measure your own overhead and reduce the friction around model choice. Three steps cover most of the ground.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Measure the context your tools consume.&lt;/strong&gt; Before you argue about MCP, count the tokens. Dump your tool definitions and system prompt to a file and run them through a tokenizer. If you do not have a tokenizer handy, a rough character count divided by four is a starting estimate, clearly marked as an estimate.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Illustrative only. Replace with your real tokenizer.
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;

&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tools.json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;tools&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;serialized&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;tools&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;approx_tokens&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;serialized&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;//&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;  &lt;span class="c1"&gt;# rough estimate, not a measurement
&lt;/span&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tool definitions: ~&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;approx_tokens&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; tokens per turn (estimate)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run this for each server you have. If one server contributes a large share of the total and is used in a small share of turns, that is your first candidate to load conditionally. If a server turns out to be cheap and used on most turns, the debate is not about you.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Load tools conditionally.&lt;/strong&gt; Most agents do not need every tool on every turn. Split your tool set by task and only attach the subset the current step needs. This is a plain application change, not a protocol change, and it works whether you keep MCP or not.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Illustrative only.
&lt;/span&gt;&lt;span class="n"&gt;TOOL_SETS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;read&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;read_file&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;list_dir&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;write&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;write_file&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;apply_patch&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;query&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;run_sql&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;tools_for&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;step&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nf"&gt;load_tool&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;TOOL_SETS&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;step&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[])]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The audit camp's concern is legitimate, so do not treat conditional loading as a reason to drop boundaries. Keep an allowlist, log every tool call, and review the log. That is the part of MCP's value that survives the critique, and it is cheap to keep.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Put a cost log next to your tool log.&lt;/strong&gt; Log input tokens, output tokens, and the model name for every call. You cannot settle a token-tax argument with vibes. You can settle it for your own workload with a week of logs.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Illustrative only.
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;logging&lt;/span&gt;

&lt;span class="n"&gt;log&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;logging&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getLogger&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;llm.cost&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;record&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;log&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;info&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model=%s input_tokens=%s output_tokens=%s&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;prompt_tokens&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;completion_tokens&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once you have those logs, the debate becomes a routing question. Which steps actually need the expensive model, and which ones are paying for tool definitions they never use? That is a question you can answer for your own agent in an afternoon, and it is the only version of the question that has a defensible answer.&lt;/p&gt;

&lt;p&gt;A worked example, with made-up numbers, of what the routing half looks like in a config file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Example only. Token counts and model IDs are placeholders, not measurements.
&lt;/span&gt;&lt;span class="n"&gt;ROUTES&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;classify&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;YOUR_CHEAP_MODEL_ID&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;   &lt;span class="c1"&gt;# e.g. a small/fast tier
&lt;/span&gt;    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;summarize&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;YOUR_MID_MODEL_ID&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;synthesize&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;YOUR_STRONG_MODEL_ID&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;model_for&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;step&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;ROUTES&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;step&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;YOUR_DEFAULT_MODEL_ID&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If your client is OpenAI-compatible, changing the model string is the whole migration. If it is not, that is the friction the radar is pointing at, and it is worth removing before the next protocol debate forces your hand.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F20y7va8sgrsn8so6qvuh.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F20y7va8sgrsn8so6qvuh.png" alt="Audit your agent's context today" width="799" height="333"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Where a single key for many models fits
&lt;/h2&gt;

&lt;p&gt;If the MCP debate is really about cost and context, then the model layer is a separate lever you can pull today. An OpenAI-compatible gateway with one key for 21 models and prepaid USD credit removes the multi-provider key juggling that the radar lists as part of the pain, and it makes model experiments a string change rather than an SDK migration. That is the indirect fit here. The post attacks MCP, not model APIs, so a gateway does not resolve the tool-overhead argument, and it should not be sold as if it does.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to watch next
&lt;/h2&gt;

&lt;p&gt;Open questions from the thread, not predictions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Does the post's core claim get a technical rebuttal with measurements, or does the debate stay at the level of architecture and preference?&lt;/li&gt;
&lt;li&gt;Do MCP maintainers or server authors respond with changes aimed at context cost, such as lazy tool loading or smaller schemas?&lt;/li&gt;
&lt;li&gt;Do teams that say "rip out" publish what they replaced MCP with, and does that replacement keep audit and sandbox properties?&lt;/li&gt;
&lt;li&gt;Does the audit-and-sandbox camp publish concrete threat models where a direct CLI would be unacceptable?&lt;/li&gt;
&lt;li&gt;Does the conversation move from MCP as a protocol to the broader question of how tool definitions are priced and cached across providers?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The useful move for a working developer is not to pick a side in a thread. It is to measure your own tool context, log your own token cost, and make model switching cheap enough that the next debate does not require a migration to test.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Disclosure: I work on BeefAPI.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>webdev</category>
      <category>news</category>
    </item>
    <item>
      <title>Near-Duplicate Model Strings Are Quietly Changing Your Bill</title>
      <dc:creator>Cogumellum</dc:creator>
      <pubDate>Sun, 20 Sep 2026 17:22:19 +0000</pubDate>
      <link>https://dev.to/cogumellum/near-duplicate-model-strings-are-quietly-changing-your-bill-3k1n</link>
      <guid>https://dev.to/cogumellum/near-duplicate-model-strings-are-quietly-changing-your-bill-3k1n</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Gateways expose model names that look almost identical but carry different cache-read prices, so a typo in a model string can silently change what you pay. Here's a script that fetches &lt;code&gt;pricing.usd.json&lt;/code&gt; and flags near-duplicate names with divergent prices before you ship.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bug that doesn't throw
&lt;/h2&gt;

&lt;p&gt;You're wiring up an LLM call. You open the provider's model list, copy a string, paste it into your config, and move on. The request succeeds. The response looks fine. Nothing in your logs tells you that you picked &lt;code&gt;claude-fable-5&lt;/code&gt; when you meant &lt;code&gt;claude-fable-5-1&lt;/code&gt;, or &lt;code&gt;grok-4.5&lt;/code&gt; when the team decided on &lt;code&gt;grok-4.6&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The constraint is that model identifiers are opaque strings. There is no type system for them. Your editor won't autocomplete them unless you've built a constant. Your tests won't fail on a wrong-but-valid string, because it's a valid string. And the failure mode is not an error; it's a line item on a bill that's larger than you expected, or a cache-read price that's several times what you budgeted.&lt;/p&gt;

&lt;p&gt;This is worse on a gateway than on a single provider, because a gateway aggregates many vendors' naming conventions into one namespace. You get &lt;code&gt;claude-fable-5&lt;/code&gt; next to &lt;code&gt;claude-fable-5-1&lt;/code&gt;, &lt;code&gt;grok-4.5&lt;/code&gt; next to &lt;code&gt;grok-4.6&lt;/code&gt;, &lt;code&gt;glm-5.2&lt;/code&gt; next to &lt;code&gt;glm-5.3&lt;/code&gt;, &lt;code&gt;gemini-3.7-flash&lt;/code&gt; next to &lt;code&gt;gemini-3.8-flash&lt;/code&gt;. Each pair is one character apart. Each pair can have a different price for the same token category.&lt;/p&gt;

&lt;p&gt;The usual advice is "read the pricing page carefully." That's not a control. It's a hope. The rest of this article is about turning it into a check you can run in CI.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9zpvesd50wm5xloj2l1t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9zpvesd50wm5xloj2l1t.png" alt="cache-read price per 1M tokens" width="799" height="333"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What the pricebook actually contains
&lt;/h2&gt;

&lt;p&gt;The source is a JSON file at &lt;code&gt;https://global.beefapi.com/pricing.usd.json&lt;/code&gt;, read at 2026-09-19T23:31:02Z. It lists 33 models. Each entry has an input price and an output price per 1M tokens, and most have a cache-read price. Here are the entries where the naming gets dangerous, copied exactly:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Input ($/1M)&lt;/th&gt;
&lt;th&gt;Output ($/1M)&lt;/th&gt;
&lt;th&gt;Cache read ($/1M)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;claude-fable-5&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;0.4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;claude-fable-5-1&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;0.1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;grok-4.5&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.6&lt;/td&gt;
&lt;td&gt;1.8&lt;/td&gt;
&lt;td&gt;0.09&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;grok-4.6&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.6&lt;/td&gt;
&lt;td&gt;1.8&lt;/td&gt;
&lt;td&gt;0.15&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;glm-5.2&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.91&lt;/td&gt;
&lt;td&gt;2.86&lt;/td&gt;
&lt;td&gt;0.169&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;glm-5.3&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;3.2&lt;/td&gt;
&lt;td&gt;0.22&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;gemini-3.7-flash&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.375&lt;/td&gt;
&lt;td&gt;1.87&lt;/td&gt;
&lt;td&gt;0.0375&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;gemini-3.8-flash&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;0.375&lt;/td&gt;
&lt;td&gt;1.87&lt;/td&gt;
&lt;td&gt;0.0375&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;claude-opus-4-6&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;0.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;claude-opus-4-7&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;0.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;claude-opus-4-8&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;0.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;claude-opus-5&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;td&gt;0.2&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Look at the first two rows. &lt;code&gt;claude-fable-5&lt;/code&gt; and &lt;code&gt;claude-fable-5-1&lt;/code&gt; have identical input and output prices. If you're scanning a table for cost, they look the same. But the cache-read price is 0.4 for one and 0.1 for the other. If your workload leans on prompt caching, that difference is the whole story, and it's invisible unless you read the right column.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;grok-4.5&lt;/code&gt; and &lt;code&gt;grok-4.6&lt;/code&gt; are the same shape: identical input and output, different cache read (0.09 vs 0.15). &lt;code&gt;glm-5.2&lt;/code&gt; and &lt;code&gt;glm-5.3&lt;/code&gt; differ in every field, including cache read (0.169 vs 0.22). &lt;code&gt;gemini-3.7-flash&lt;/code&gt; and &lt;code&gt;gemini-3.8-flash&lt;/code&gt; happen to match on all three fields shown here, which is its own trap: you can't tell them apart from price alone, so you need another reason to prefer one.&lt;/p&gt;

&lt;p&gt;These are prices, not performance. A lower cache-read price does not mean a faster or better model. It means cached input tokens are billed at a lower rate. The table tells you nothing about latency, throughput, or quality, because the source doesn't contain those fields.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx6p93u3tbomz52asofo1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx6p93u3tbomz52asofo1.png" alt="Identical columns hide the real cost" width="799" height="333"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why "just read the docs" fails
&lt;/h2&gt;

&lt;p&gt;There are three reasons a careful human still ships the wrong string.&lt;/p&gt;

&lt;p&gt;First, the wrong string is valid. If &lt;code&gt;claude-fable-5-1&lt;/code&gt; is a real entry and &lt;code&gt;claude-fable-5&lt;/code&gt; is a real entry, both requests return 200. There's no signal that you picked the one you didn't mean.&lt;/p&gt;

&lt;p&gt;Second, the difference is often in a column you weren't optimizing. Most developers compare input and output prices, because that's what a naive cost estimate uses. Cache-read prices only matter if you use prompt caching, and if you don't use it today, you won't look at that column. Then you add caching later, and the model string that was fine becomes expensive.&lt;/p&gt;

&lt;p&gt;Third, the naming is not consistent across vendors in the same namespace. Some entries use a hyphen before a version suffix (&lt;code&gt;claude-fable-5-1&lt;/code&gt;), some use a dot (&lt;code&gt;glm-5.2&lt;/code&gt;), some use a word (&lt;code&gt;gpt-5.6-sol&lt;/code&gt;, &lt;code&gt;gpt-5.6-terra&lt;/code&gt;). You can't write one rule that catches every near-duplicate. You need to compare strings against each other, not against a pattern you invented.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0h2x9eutash5f3197v8x.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0h2x9eutash5f3197v8x.png" alt="Flag near-duplicates in CI" width="799" height="333"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  A script that flags the dangerous pairs
&lt;/h2&gt;

&lt;p&gt;The check is: fetch the pricebook, group model names by similarity, and for each near-duplicate pair, compare the price fields. If two names are close but their prices diverge in any field, print a warning. Here's an illustrative script. It uses only the standard library plus a similarity heuristic, and it treats the pricebook as the source of truth.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Illustrative example. Not production code.
# Field names are placeholders; inspect the actual JSON shape before relying on them.
&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;urllib.request&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;difflib&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;SequenceMatcher&lt;/span&gt;

&lt;span class="n"&gt;PRICEBOOK_URL&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://global.beefapi.com/pricing.usd.json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;SIMILARITY_THRESHOLD&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.85&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;fetch_pricebook&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;urllib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;urlopen&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;normalize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;entry&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="c1"&gt;# Adjust these keys to match the real schema in your copy of the file.
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;input&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;entry&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;input&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;output&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;entry&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;output&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cache_read&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;entry&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cache_read&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;fetch_pricebook&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;PRICEBOOK_URL&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;models&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;models&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;isinstance&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt;

    &lt;span class="n"&gt;names&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;models&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="n"&gt;prices&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt; &lt;span class="nf"&gt;normalize&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;models&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;enumerate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;names&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;names&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;:]:&lt;/span&gt;
            &lt;span class="n"&gt;ratio&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;SequenceMatcher&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;ratio&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;ratio&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;SIMILARITY_THRESHOLD&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="k"&gt;continue&lt;/span&gt;
            &lt;span class="n"&gt;pa&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;pb&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;prices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="n"&gt;prices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
            &lt;span class="n"&gt;diffs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;pa&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;pa&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;pb&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;]]&lt;/span&gt;
            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;diffs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;NEAR-DUPLICATE with divergent prices: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; vs &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
                &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;k&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;diffs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;  &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;pa&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; vs &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;pb&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;k&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;


&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;__name__&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;__main__&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run against a pricebook shaped like the one read at 2026-09-19T23:31:02Z, this would surface pairs such as &lt;code&gt;claude-fable-5&lt;/code&gt; vs &lt;code&gt;claude-fable-5-1&lt;/code&gt; (cache_read 0.4 vs 0.1), &lt;code&gt;grok-4.5&lt;/code&gt; vs &lt;code&gt;grok-4.6&lt;/code&gt; (cache_read 0.09 vs 0.15), and &lt;code&gt;glm-5.2&lt;/code&gt; vs &lt;code&gt;glm-5.3&lt;/code&gt; (input 0.91 vs 1, output 2.86 vs 3.2, cache_read 0.169 vs 0.22). It would also flag &lt;code&gt;claude-opus-4-6&lt;/code&gt; vs &lt;code&gt;claude-opus-4-7&lt;/code&gt; and similar siblings, but those happen to agree on all three fields, so the &lt;code&gt;diffs&lt;/code&gt; list would be empty and nothing would print. That's the point: the script only shouts when the choice has a price consequence.&lt;/p&gt;

&lt;p&gt;The threshold is a knob. At 0.85 you catch one-character suffixes and version bumps. Lower it and you'll get noise from unrelated names that share a prefix. Higher it and you'll miss pairs like &lt;code&gt;gemini-3.7-flash&lt;/code&gt; vs &lt;code&gt;gemini-3.8-flash&lt;/code&gt; if you consider those too far apart, even though they're adjacent in the list.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2iv3bjt92si1ilxmtc9l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F2iv3bjt92si1ilxmtc9l.png" alt="One key, many models" width="799" height="333"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Wiring it into your workflow
&lt;/h2&gt;

&lt;p&gt;A script that only runs when you remember to run it is not much better than reading the docs. Three places to put it:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;As a pre-commit hook.&lt;/strong&gt; If your repo contains a file listing the model strings you use, the hook can fetch the pricebook and check that every string you reference exists, and that no two strings in your config are near-duplicates with divergent prices. That catches the case where someone adds a second model for a fallback path and picks the wrong sibling.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;As a scheduled job.&lt;/strong&gt; Prices change. The pricebook is a snapshot. A daily job that diffs today's pricebook against yesterday's, and prints any field that moved for a model you use, turns a silent billing change into a notification.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;As a review artifact.&lt;/strong&gt; When someone proposes switching a model in a pull request, the diff should include the price fields for the old and new string, side by side, including cache read. If the PR description says "switch to the cheaper model" and the cache-read column went up, the reviewer sees it.&lt;/p&gt;

&lt;p&gt;None of this requires the gateway to do anything special. It's a property of the data being published as JSON: you can fetch it, parse it, and assert on it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbksg2bcei75f7ppvmbra.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbksg2bcei75f7ppvmbra.png" alt="Stop guessing model strings" width="799" height="333"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  A worked cost comparison, as an example
&lt;/h2&gt;

&lt;p&gt;Numbers below are made up to show the arithmetic, not to describe any real workload.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Example only. Token counts are invented.
# Prices below are copied from the pricebook read at 2026-09-19T23:31:02Z.
&lt;/span&gt;
&lt;span class="n"&gt;PER_MILLION&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;1_000_000&lt;/span&gt;

&lt;span class="c1"&gt;# Invented workload: 50M cached input tokens, 10M uncached input, 2M output.
&lt;/span&gt;&lt;span class="n"&gt;cached_input_tokens&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;50&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;PER_MILLION&lt;/span&gt;
&lt;span class="n"&gt;uncached_input_tokens&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;PER_MILLION&lt;/span&gt;
&lt;span class="n"&gt;output_tokens&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;PER_MILLION&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;cost&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;input_price&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;output_price&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;cache_read_price&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="nf"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;uncached_input_tokens&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;PER_MILLION&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;input_price&lt;/span&gt;
        &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;cached_input_tokens&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;PER_MILLION&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;cache_read_price&lt;/span&gt;
        &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;output_tokens&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;PER_MILLION&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;output_price&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;


&lt;span class="c1"&gt;# claude-fable-5: input 4, output 20, cache read 0.4
&lt;/span&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-fable-5  &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;cost&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.4&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="c1"&gt;# claude-fable-5-1: input 4, output 20, cache read 0.1
&lt;/span&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-fable-5-1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;cost&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.1&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The two calls differ only in the cache-read price. The input and output prices are identical. If you never look at the cache-read column, the two model strings look interchangeable, and the script above is the difference between noticing and not noticing.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the data does not tell you
&lt;/h2&gt;

&lt;p&gt;The pricebook is a price list. It does not contain latency, throughput, uptime, context window, or quality. It does not say how often prices change. It does not say whether a near-duplicate name is an alias, a snapshot, or a genuinely different model. It does not resolve the discrepancy between the number of entries in the file and the number of models the product profile advertises; if you need that answer, check the source directly.&lt;/p&gt;

&lt;p&gt;So the script is a guardrail, not a decision. It tells you that two strings you might confuse have different prices. It does not tell you which one to use. That depends on what the model does for your task, which you have to measure yourself.&lt;/p&gt;

&lt;h2&gt;
  
  
  Questions for the comments
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Have you ever shipped a model string that was valid but wrong, and only found out from a bill or a cache-hit-rate dashboard? What tipped you off?&lt;/li&gt;
&lt;li&gt;Do you keep model identifiers as free strings in config, or as constants in code? If constants, how do you keep them in sync with the pricebook?&lt;/li&gt;
&lt;li&gt;For near-duplicates that agree on every price field, like &lt;code&gt;claude-opus-4-6&lt;/code&gt; through &lt;code&gt;claude-opus-5&lt;/code&gt;, how do you decide which one to standardize on when price gives you no signal?&lt;/li&gt;
&lt;li&gt;Would you rather the gateway reject unknown model strings loudly, or accept any string and let the upstream decide? What breaks in each case?&lt;/li&gt;
&lt;li&gt;If you run prompt caching, do you track cache-read price separately in your cost model, or do you fold it into an average input price? What made you choose that?&lt;/li&gt;
&lt;/ol&gt;




&lt;p&gt;&lt;em&gt;Disclosure: I work on BeefAPI.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>webdev</category>
      <category>discuss</category>
    </item>
    <item>
      <title>Claude Code v2.1.277 Reads AGENTS.md When CLAUDE.md Is Missing</title>
      <dc:creator>Cogumellum</dc:creator>
      <pubDate>Sat, 19 Sep 2026 10:58:13 +0000</pubDate>
      <link>https://dev.to/cogumellum/claude-code-v21277-reads-agentsmd-when-claudemd-is-missing-53ln</link>
      <guid>https://dev.to/cogumellum/claude-code-v21277-reads-agentsmd-when-claudemd-is-missing-53ln</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Starting with Claude Code v2.1.277 (18 Sep 2026), if a project has no &lt;code&gt;CLAUDE.md&lt;/code&gt;, Claude Code reads &lt;code&gt;AGENTS.md&lt;/code&gt; instead, and the behavior is configurable in &lt;code&gt;/config&lt;/code&gt; (&lt;a href="https://code.claude.com/docs/en/changelog" rel="noopener noreferrer"&gt;changelog&lt;/a&gt;, &lt;a href="https://code.claude.com/docs/en/memory" rel="noopener noreferrer"&gt;memory docs&lt;/a&gt;). If you already keep agent instructions for Codex, Cursor, or other tools in &lt;code&gt;AGENTS.md&lt;/code&gt;, that file just became load-bearing for one more consumer — and the failure mode is silent, not loud.&lt;/p&gt;

&lt;h2&gt;
  
  
  What happened
&lt;/h2&gt;

&lt;p&gt;The change is narrow and worth stating precisely, because the details decide whether it affects your repo at all:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Version:&lt;/strong&gt; Claude Code v2.1.277, dated 18 Sep 2026 (&lt;a href="https://code.claude.com/docs/en/changelog" rel="noopener noreferrer"&gt;changelog&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The rule:&lt;/strong&gt; if no &lt;code&gt;CLAUDE.md&lt;/code&gt; exists in the project, Claude Code reads &lt;code&gt;AGENTS.md&lt;/code&gt; (&lt;a href="https://code.claude.com/docs/en/memory" rel="noopener noreferrer"&gt;memory docs&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It is configurable:&lt;/strong&gt; the behavior can be turned on or off in &lt;code&gt;/config&lt;/code&gt; (&lt;a href="https://code.claude.com/docs/en/memory" rel="noopener noreferrer"&gt;memory docs&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Not yet supported on Bedrock, Vertex, or Foundry&lt;/strong&gt; (&lt;a href="https://code.claude.com/docs/en/memory" rel="noopener noreferrer"&gt;memory docs&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Who announced it:&lt;/strong&gt; an Anthropic engineer, with the change documented in the official changelog (&lt;a href="https://x.com/trq212/status/2101009392611278961" rel="noopener noreferrer"&gt;announcement&lt;/a&gt;, &lt;a href="https://code.claude.com/docs/en/changelog" rel="noopener noreferrer"&gt;changelog&lt;/a&gt;).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is the whole news. No migration tool, no new file format, no deprecation of &lt;code&gt;CLAUDE.md&lt;/code&gt;. It is a fallback path: &lt;code&gt;CLAUDE.md&lt;/code&gt; still wins when present, and &lt;code&gt;AGENTS.md&lt;/code&gt; is only consulted in its absence.&lt;/p&gt;

&lt;p&gt;The Bedrock/Vertex/Foundry carve-out is the part most likely to bite someone reading a summary instead of the docs. If your team runs Claude Code through one of those providers, the fallback is not there yet, and a repo that relies on it will behave differently depending on how each developer launched the tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  What developers are saying
&lt;/h2&gt;

&lt;p&gt;The reaction, as summarized in the news, is broadly positive and centers on one theme: interoperability with Codex and Cursor, and the end of maintaining duplicated project rules. Posts appeared across several languages — English, Japanese, and Chinese — in roughly the 17 hours after the announcement (&lt;a href="https://x.com/trq212/status/2101009392611278961" rel="noopener noreferrer"&gt;announcement&lt;/a&gt;, &lt;a href="https://x.com/xinglee23/status/2101261967575208150" rel="noopener noreferrer"&gt;one example&lt;/a&gt;, &lt;a href="https://x.com/raito_athenas/status/2101257856289055098" rel="noopener noreferrer"&gt;another&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;I am not going to invent quotes or attribute positions to specific people beyond what those posts say. The shape of the conversation is what matters: people who already maintain &lt;code&gt;AGENTS.md&lt;/code&gt; for other agents see this as one less file to keep in sync, and the subtext is that per-agent rule files had become a maintenance tax rather than a feature.&lt;/p&gt;

&lt;p&gt;There is no visible backlash in the summary, but there rarely is on announcement day. The interesting complaints tend to arrive a week later, once someone's CI job starts behaving differently.&lt;/p&gt;

&lt;h2&gt;
  
  
  The practical problem this creates
&lt;/h2&gt;

&lt;p&gt;Here is the scenario that should worry you, and it is not exotic.&lt;/p&gt;

&lt;p&gt;A repo has three tools pointed at it. Codex reads &lt;code&gt;AGENTS.md&lt;/code&gt;. Cursor reads its own rules file. Claude Code, until now, read &lt;code&gt;CLAUDE.md&lt;/code&gt;. Someone on the team — usually whoever got annoyed first — wrote &lt;code&gt;AGENTS.md&lt;/code&gt; with real content: build commands, test invocation, the fact that the integration suite needs a local Postgres, the convention that migrations are generated and never hand-edited.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;CLAUDE.md&lt;/code&gt; exists too, but it is stale. It was written months ago, it mentions a &lt;code&gt;make test&lt;/code&gt; target that was renamed, and it says the API lives under &lt;code&gt;/v1&lt;/code&gt; when it moved to &lt;code&gt;/v2&lt;/code&gt;. Nobody deleted it because nobody was sure whether it was still needed.&lt;/p&gt;

&lt;p&gt;Before v2.1.277, that stale file was harmless-ish: Claude Code read it, got slightly wrong context, and a human occasionally corrected it. After v2.1.277, the situation is unchanged in that repo — &lt;code&gt;CLAUDE.md&lt;/code&gt; still takes precedence, so the fallback never fires. The change only matters where &lt;code&gt;CLAUDE.md&lt;/code&gt; is &lt;em&gt;absent&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;So the real scenario is the inverse: a repo where &lt;code&gt;CLAUDE.md&lt;/code&gt; was deliberately deleted, or never existed, and where &lt;code&gt;AGENTS.md&lt;/code&gt; was written for a different agent with different assumptions. Now Claude Code silently starts reading it. If that file contains instructions tuned for another tool — a different test command, a different package manager, an instruction to always run a codegen step — Claude Code will follow them, and nothing in the output will say "I am following AGENTS.md now."&lt;/p&gt;

&lt;p&gt;The second-order problem is divergence. The news explicitly names the developer pain: maintaining &lt;code&gt;CLAUDE.md&lt;/code&gt; plus &lt;code&gt;AGENTS.md&lt;/code&gt;, or per-agent rule sets, in multi-tool projects, where the rules drift apart and the maintenance becomes overhead. This change removes one duplication path but does not remove the underlying problem, because &lt;code&gt;CLAUDE.md&lt;/code&gt; still takes precedence when it exists. If you keep both, you still have two files, and now you have a precedence rule to remember on top.&lt;/p&gt;

&lt;p&gt;A third issue: the Bedrock/Vertex/Foundry gap. If part of your team runs Claude Code against one of those providers, the same repo will resolve instructions differently for different people. That is a debugging session waiting to happen, and the symptom will look like "the agent forgot our conventions" rather than "the agent read a different file."&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do about it
&lt;/h2&gt;

&lt;p&gt;These are steps you can take today, in order of how much they reduce surprise.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Find out which files exist and which one wins
&lt;/h3&gt;

&lt;p&gt;Run this in each repo you care about. It tells you whether the fallback is even reachable.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Which instruction files exist at the repo root?&lt;/span&gt;
&lt;span class="k"&gt;for &lt;/span&gt;f &lt;span class="k"&gt;in &lt;/span&gt;CLAUDE.md AGENTS.md .cursor/rules&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do
  if&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt; &lt;span class="nt"&gt;-e&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$f&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"present: &lt;/span&gt;&lt;span class="nv"&gt;$f&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;else &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"absent:  &lt;/span&gt;&lt;span class="nv"&gt;$f&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;fi
done&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If &lt;code&gt;CLAUDE.md&lt;/code&gt; is present, the fallback does not apply and you can stop here — but read step 2 anyway, because "present and stale" is its own problem.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Decide on one canonical file, and make the other a pointer
&lt;/h3&gt;

&lt;p&gt;The cheapest way to avoid divergence is to have exactly one file with real content. If you want &lt;code&gt;AGENTS.md&lt;/code&gt; to be canonical, &lt;code&gt;CLAUDE.md&lt;/code&gt; can be a one-line redirect rather than a copy:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="c"&gt;&amp;lt;!-- CLAUDE.md --&amp;gt;&lt;/span&gt;
Project instructions live in AGENTS.md. Read that file and follow it.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This keeps precedence explicit: Claude Code reads &lt;code&gt;CLAUDE.md&lt;/code&gt;, which tells it to read &lt;code&gt;AGENTS.md&lt;/code&gt;. You maintain one file. The pointer file changes roughly never.&lt;/p&gt;

&lt;p&gt;The opposite choice — &lt;code&gt;CLAUDE.md&lt;/code&gt; canonical, &lt;code&gt;AGENTS.md&lt;/code&gt; a pointer — works too, but check whether your other tools follow pointers or expect content. Do not assume; test it.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Check &lt;code&gt;/config&lt;/code&gt; before you rely on the fallback
&lt;/h3&gt;

&lt;p&gt;The behavior is configurable (&lt;a href="https://code.claude.com/docs/en/memory" rel="noopener noreferrer"&gt;memory docs&lt;/a&gt;). That means two developers on the same repo can have different settings, and the repo alone will not tell you which. If your team standardizes on the fallback, say so somewhere a human will read it — the onboarding doc, the PR template, wherever your team actually looks.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Audit &lt;code&gt;AGENTS.md&lt;/code&gt; as if it were production config
&lt;/h3&gt;

&lt;p&gt;If Claude Code is now going to read a file that was written for another agent, read it yourself first. The things that break are boring:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Test commands that assume a different runner.&lt;/li&gt;
&lt;li&gt;Package manager instructions (&lt;code&gt;pnpm&lt;/code&gt; vs &lt;code&gt;npm&lt;/code&gt;) that conflict with the lockfile in the repo.&lt;/li&gt;
&lt;li&gt;Instructions to run codegen or migrations that you do not want triggered on every session.&lt;/li&gt;
&lt;li&gt;Paths that moved.&lt;/li&gt;
&lt;li&gt;Rules that were true for a different model's context window and are now just noise.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this is a Claude Code problem. It is the ordinary problem of a config file that grew without an owner.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Keep the provider gap in mind
&lt;/h3&gt;

&lt;p&gt;If anyone on the team uses Bedrock, Vertex, or Foundry, the fallback is not supported there yet (&lt;a href="https://code.claude.com/docs/en/memory" rel="noopener noreferrer"&gt;memory docs&lt;/a&gt;). Do not write a repo setup that only works on one path. The pointer-file approach in step 2 sidesteps this entirely, because it does not depend on the fallback firing.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. If you script against an OpenAI-compatible API, keep the model name out of the code
&lt;/h3&gt;

&lt;p&gt;This is not about instruction files, but it is the adjacent habit that pays off when you are swapping models across agents. Hardcoding a model string in application code means every swap is a code change. Reading it from config means it is a config change:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Illustrative example only — not production code.
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;YOUR_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://YOUR_GATEWAY_URL/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;MODEL&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;MODEL_NAME&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;YOUR_MODEL_NAME&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;MODEL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Summarize this diff.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The point is not the specific client library. The point is that "which model" should be a value, not a literal, so that a swap does not require a deploy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where a single OpenAI-compatible key for many models fits
&lt;/h2&gt;

&lt;p&gt;This is the part where I have to be careful, because it is easy to overstate.&lt;/p&gt;

&lt;p&gt;This change is about an instruction file. It is not an Anthropic endorsement of any gateway, and nothing here should be read as one. The fallback to &lt;code&gt;AGENTS.md&lt;/code&gt; has nothing to do with how you authenticate or which endpoint you call.&lt;/p&gt;

&lt;p&gt;The connection is at the workflow level, not the file level. Developers who run several agents and several models — Claude Code plus Codex plus whatever else — are the same developers who end up with per-agent rule files and per-provider billing. The news removes one duplication (rule files) and leaves the other (keys, billing, model switching) untouched.&lt;/p&gt;

&lt;p&gt;That is where a gateway with one OpenAI-compatible key across many models is relevant: not because it changes how &lt;code&gt;AGENTS.md&lt;/code&gt; is read, but because the multi-model workflow that makes &lt;code&gt;AGENTS.md&lt;/code&gt; worth standardizing is also the workflow where managing N keys and N billing relationships gets tedious. BeefAPI is an OpenAI-compatible gateway with prepaid USD credit and one key for 21 models, aimed at developers who switch models often (&lt;a href="https://global.beefapi.com/pricing.usd.json" rel="noopener noreferrer"&gt;pricing&lt;/a&gt;, &lt;a href="https://global.beefapi.com" rel="noopener noreferrer"&gt;site&lt;/a&gt;).&lt;/p&gt;

&lt;p&gt;If you are evaluating that kind of setup, the pricebook is the honest place to look, and it is worth reading as a price list rather than a performance claim. As read at 2026-09-18T19:27:31Z it lists 27 models, including &lt;code&gt;claude-sonnet-5&lt;/code&gt; at $0.8 / $4 per 1M tokens (input / output, cache read $0.08), &lt;code&gt;gpt-5.6-terra&lt;/code&gt; at $0.6 / $3.6 per 1M tokens (input / output, cache read $0.06), &lt;code&gt;gemini-3.1-pro&lt;/code&gt; at $1 / $6 per 1M tokens (input / output, cache read $0.1), and &lt;code&gt;qwen3.8-flash&lt;/code&gt; at $0.12 / $0.38 per 1M tokens (input / output, cache read $0.014). Cheaper models are cheaper; that says nothing about whether they are good enough for your task, and you should test that yourself.&lt;/p&gt;

&lt;p&gt;Here is a worked example of how you might reason about cost, with made-up token counts, purely to show the arithmetic shape:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# EXAMPLE ONLY — invented token counts, not measured.
# Assume a task uses 200,000 input tokens and 50,000 output tokens.

model_a: $0.8 / $4 per 1M tokens
  input:  200,000 / 1,000,000 * $0.8 = $0.16
  output:  50,000 / 1,000,000 * $4   = $0.20
  total:  $0.36

model_b: $0.12 / $0.38 per 1M tokens
  input:  200,000 / 1,000,000 * $0.12 = $0.024
  output:  50,000 / 1,000,000 * $0.38 = $0.019
  total:  $0.043
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Do not treat those totals as a recommendation. They exist to show that the input/output split matters, and that cache reads are priced separately from fresh input — which is the detail people miss when they compare two models by a single headline number.&lt;/p&gt;

&lt;p&gt;Where this does &lt;strong&gt;not&lt;/strong&gt; help: it does not fix divergent project rules. If your &lt;code&gt;AGENTS.md&lt;/code&gt; and &lt;code&gt;CLAUDE.md&lt;/code&gt; disagree, a gateway will faithfully route your request to whichever model you picked, and that model will faithfully read the wrong file. It also does not help if your problem is that Claude Code is not reading &lt;code&gt;AGENTS.md&lt;/code&gt; at all — check &lt;code&gt;/config&lt;/code&gt;, check that &lt;code&gt;CLAUDE.md&lt;/code&gt; is genuinely absent, and check whether you are on Bedrock, Vertex, or Foundry, where the fallback is not supported yet (&lt;a href="https://code.claude.com/docs/en/memory" rel="noopener noreferrer"&gt;memory docs&lt;/a&gt;).&lt;/p&gt;

&lt;h2&gt;
  
  
  What to watch next
&lt;/h2&gt;

&lt;p&gt;These are open questions from the news, not predictions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Does the Bedrock/Vertex/Foundry gap close?&lt;/strong&gt; The docs currently say the behavior is not supported there (&lt;a href="https://code.claude.com/docs/en/memory" rel="noopener noreferrer"&gt;memory docs&lt;/a&gt;). Until it does, multi-provider teams have a split-brain setup.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What does &lt;code&gt;/config&lt;/code&gt; actually expose?&lt;/strong&gt; The behavior is configurable (&lt;a href="https://code.claude.com/docs/en/memory" rel="noopener noreferrer"&gt;memory docs&lt;/a&gt;), but the practical question is whether teams can set it per-repo rather than per-machine. If it is per-machine, repo-level standardization is harder than it sounds.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Does precedence stay this simple?&lt;/strong&gt; Right now &lt;code&gt;CLAUDE.md&lt;/code&gt; wins when present. Any future change to that ordering would silently flip which file a repo actually uses.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Do other tools converge on &lt;code&gt;AGENTS.md&lt;/code&gt;?&lt;/strong&gt; The conversation is about interoperability with Codex and Cursor (&lt;a href="https://x.com/trq212/status/2101009392611278961" rel="noopener noreferrer"&gt;announcement&lt;/a&gt;). Whether they all agree on precedence and file discovery is a separate question from whether they read the file.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What breaks in CI?&lt;/strong&gt; The posts summarized are from the first ~17 hours (&lt;a href="https://x.com/xinglee23/status/2101261967575208150" rel="noopener noreferrer"&gt;one example&lt;/a&gt;, &lt;a href="https://x.com/raito_athenas/status/2101257856289055098" rel="noopener noreferrer"&gt;another&lt;/a&gt;). Non-interactive runs are where file-discovery changes tend to surface, and that signal has not arrived yet.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you do one thing today: check whether &lt;code&gt;CLAUDE.md&lt;/code&gt; exists in your repos, and if it does not, read the &lt;code&gt;AGENTS.md&lt;/code&gt; that Claude Code is now going to follow. It takes two minutes and it is the difference between a config change and a mystery.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Disclosure: I work on BeefAPI.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>python</category>
      <category>news</category>
    </item>
    <item>
      <title>Our System Crashed at 14:22: It Wasn't the Database</title>
      <dc:creator>Cogumellum</dc:creator>
      <pubDate>Mon, 14 Sep 2026 02:59:18 +0000</pubDate>
      <link>https://dev.to/cogumellum/our-system-crashed-at-1422-it-wasnt-the-database-1p6a</link>
      <guid>https://dev.to/cogumellum/our-system-crashed-at-1422-it-wasnt-the-database-1p6a</guid>
      <description>&lt;h1&gt;
  
  
  Our System Crashed at 14:22: It Wasn't the Database
&lt;/h1&gt;

&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; During peak concurrent load, our 5-minute cache unexpectedly evicted keys early due to lock contention. The solution was migrating to asynchronous stale-while-revalidate with an in-memory semaphore. The diff took 18 lines and cut our p99 latency from 1.8s down to 240ms.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Broke in Production
&lt;/h2&gt;

&lt;p&gt;Last Tuesday, our primary streaming endpoint started throwing sporadic timeouts.&lt;br&gt;
Measured impact:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;p99 Latency:&lt;/strong&gt; jumped from 280ms to 3,400ms.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;504 Error Rate:&lt;/strong&gt; reached 4.2% across an 18-minute window.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Connection Pool:&lt;/strong&gt; 100% saturated.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;
  
  
  What We Thought Happened (The Wrong Hypothesis)
&lt;/h2&gt;

&lt;p&gt;Our initial instinct was to blame upstream LLM provider throttling. It looked like classic HTTP 429 backpressure. We restarted Celery workers, but within 90 seconds the pool was choking again.&lt;/p&gt;
&lt;h2&gt;
  
  
  The Actual Root Cause
&lt;/h2&gt;

&lt;p&gt;The culprit was an internal &lt;em&gt;thundering herd problem&lt;/em&gt;. When 300 concurrent requests hit an expired cache key at second 300, every single worker triggered the identical upstream recomputation query at the same instant.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[Request A] ──┐
[Request B] ──┼─► [Expired Cache Key] ──► 300 simultaneous upstream calls
[Request C] ──┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The Code Fix
&lt;/h2&gt;

&lt;p&gt;Instead of recomputing synchronously inside the request thread, we implemented non-blocking lock acquisition that serves stale data while a single detached coroutine refreshes the cache in the background:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;asyncio&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;get_with_revalidation&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cache&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;lock&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;factory_coro&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;expired&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;cache&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_stale&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;expired&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;lock&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;is_locked&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;asyncio&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create_task&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;revalidate_background&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cache&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;lock&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;factory_coro&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;data&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;factory_coro&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;revalidate_background&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;cache&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;lock&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;factory_coro&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;lock&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;acquire&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;fresh&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;factory_coro&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;cache&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;fresh&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ttl&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;300&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  When NOT to Use This
&lt;/h2&gt;

&lt;p&gt;If your system handles strict financial balances or ledger transactions where 2-second stale reads cause double spends, do not use stale-while-revalidate. In our case, serving model metadata and prompt routing rules, the trade-off is safe and highly recommended.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reproduction &amp;amp; Benchmarks
&lt;/h2&gt;

&lt;p&gt;We documented the full Locust load-test harness and synthetic workload in our open engineering runbook:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Test methodology: &lt;a href="https://github.com" rel="noopener noreferrer"&gt;GitHub/BeefAPI Gateway&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Production implementation: &lt;code&gt;gateway de alta resiliência e medição de tokens para LLMs&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>webdev</category>
      <category>python</category>
      <category>ai</category>
      <category>architecture</category>
    </item>
  </channel>
</rss>
