<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Stephan Miller</title>
    <description>The latest articles on DEV Community by Stephan Miller (@eristoddle).</description>
    <link>https://dev.to/eristoddle</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F18795%2F5f6c41b8-6033-4887-937a-2ebdfe623d2e.jpeg</url>
      <title>DEV Community: Stephan Miller</title>
      <link>https://dev.to/eristoddle</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/eristoddle"/>
    <language>en</language>
    <item>
      <title>Forced Connections: One Way to Get Original Ideas Out of AI</title>
      <dc:creator>Stephan Miller</dc:creator>
      <pubDate>Wed, 30 Sep 2026 12:00:00 +0000</pubDate>
      <link>https://dev.to/eristoddle/forced-connections-one-way-to-get-original-ideas-out-of-ai-4g5m</link>
      <guid>https://dev.to/eristoddle/forced-connections-one-way-to-get-original-ideas-out-of-ai-4g5m</guid>
      <description>&lt;p&gt;I have a real problem. I write &lt;a href="https://www.stephanmiller.com/series/model-buzz-report/" rel="noopener noreferrer"&gt;a weekly roundup&lt;/a&gt; of which new AI models are worth paying attention to, and I want to turn it into an email newsletter. Which means I need people to subscribe to it, which means I need ideas for getting people to subscribe to it.&lt;/p&gt;

&lt;p&gt;So I asked a model. Claude Sonnet 5, through OpenRouter, no system prompt, one line of context and then the question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;How can I get more readers to subscribe to it?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It came back with signup forms at the end of high-traffic posts, exit-intent popups, “publish 3-4 issues before heavily promoting,” a thread on X, r/LocalLLaMA, Hacker News, LinkedIn, newsletter swaps, and a single-field signup form. Every item was correct. I’ve read that exact list in maybe forty blog posts. It then asked me what my traffic source was so it could prioritize, which is the chatbot equivalent of the guy at the hardware store asking what you’re building.&lt;/p&gt;

&lt;p&gt;Fine. That was a question, and a question gets you the average. So I asked it to be creative:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Give me creative, unusual, out-of-the-box ideas for getting more readers to subscribe to it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This one was more fun. “Model Death Certificates,” obituaries for overhyped models. A “Which new AI model matches your vibe today?” quiz microsite. A referral leaderboard where the prize is absurd. A recurring mascot. An “anti-hype pledge.” It’s the list you get when you search “creative newsletter growth ideas.” I did not get creative ideas. I got the average of everything ever labeled creative.&lt;/p&gt;

&lt;p&gt;Then I stopped asking it for ideas at all. I had a script pick an object at random from a list of twenty things. It picked three: a pressure cooker, a fire extinguisher and a metronome. I started with the fire extinguisher and told the model to list ten literal attributes of a fire extinguisher, force every single one onto my subscription problem, flag its own generic results, and keep what survived.&lt;/p&gt;

&lt;p&gt;Two of the survivors:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Sell the silence.&lt;/strong&gt; Market the newsletter as the thing you check only when a model actually matters. It came from &lt;em&gt;“ignored until there’s a fire.”&lt;/em&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sell it as insurance.&lt;/strong&gt; “Subscribe once. Skip every issue if you want. Just be covered when a model actually matters.” That one came from &lt;em&gt;“reassurance from mere presence.”&lt;/em&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnx4cdpztm95gr0uj6oak.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnx4cdpztm95gr0uj6oak.jpg" alt="Introduction" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Nobody writes that in a newsletter growth article, because it’s the exact opposite of what newsletter growth articles preach. It tells people they don’t have to read you. And I think it might be the best idea of the three runs.&lt;/p&gt;

&lt;p&gt;The only thing that changed was that I stopped asking it a question and handed it a process instead.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What Just Happened&lt;/li&gt;
&lt;li&gt;
The Human Version (Run This First)

&lt;ul&gt;
&lt;li&gt;1. Write the problem as one line&lt;/li&gt;
&lt;li&gt;2. Pick an object you did not choose&lt;/li&gt;
&lt;li&gt;3. List ten literal attributes&lt;/li&gt;
&lt;li&gt;4. Force every attribute onto the problem&lt;/li&gt;
&lt;li&gt;5. Cross out everything you have read before&lt;/li&gt;
&lt;li&gt;6. Do it with two or three objects, then breed the survivors&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
Handing the Model the Process

&lt;ul&gt;
&lt;li&gt;What came out&lt;/li&gt;
&lt;li&gt;One caveat about the judge&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Where This Came From&lt;/li&gt;
&lt;li&gt;
The Crossover Move

&lt;ul&gt;
&lt;li&gt;Running it&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Why Not Just Ask for Fifty Ideas?&lt;/li&gt;
&lt;li&gt;Where It Breaks&lt;/li&gt;
&lt;li&gt;So What Did I Actually Get?&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What Just Happened
&lt;/h2&gt;

&lt;p&gt;In the &lt;a href="https://dev.to/youre-using-ai-like-a-vending-machine/"&gt;first post in this series&lt;/a&gt; I said a question gets you an answer and a move gets you something different. The obvious objection is that a “move” is just a more specific question. The second prompt above is the test case. “Give me creative, unusual, out-of-the-box ideas” &lt;em&gt;is&lt;/em&gt; a more specific question. It got a more specific average.&lt;/p&gt;

&lt;p&gt;The fire extinguisher did something different.&lt;/p&gt;

&lt;p&gt;It did not supply the idea. A fire extinguisher knows nothing about newsletters. What it supplied was ten starting points that the model would never have started from, because none of them are anywhere near the words “grow a newsletter” in anything it was trained on. “Wall-mounted in a fixed location.” “Needs a periodic inspection tag.” “Different classes for different fires.” Each attribute forces the model to begin its reasoning somewhere other than the middle, and then go from there back to the problem. Some of the ideas end up back in the middle anyway. Some of them don’t.&lt;/p&gt;

&lt;p&gt;You’re changing where the search starts.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Human Version (Run This First)
&lt;/h2&gt;

&lt;p&gt;This is a pen-and-paper technique and it has been one for close to seventy years. Do it once by hand before you ever hand it to a model, because the model version is just this process written down.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Write the problem as one line
&lt;/h3&gt;

&lt;p&gt;Not a paragraph. “Get more people to subscribe to my newsletter.” “Name the new feature.” “Figure out what the second act of this story is.” If you can’t get it to one line you have two problems and should pick one.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Pick an object you did not choose
&lt;/h3&gt;

&lt;p&gt;If you pick the object, you’ll pick one that already seems relevant, and relevance is exactly what you’re trying to escape. You can open a catalog to a random page, use the third thing to your left, a random word generator, the last noun on page 50 of whatever book is closest, or a list of twenty objects and a die.&lt;/p&gt;

&lt;p&gt;Whatever it lands on, you keep it. No re-rolls.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. List ten literal attributes
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnx6go4rawqck3ndmug2d.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnx6go4rawqck3ndmug2d.jpg" alt="3. List ten literal attributes" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Literal. Not metaphors yet. Cover all of these:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Parts.&lt;/strong&gt; What it’s made of, what you can see and touch.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use.&lt;/strong&gt; How you operate it, what it does, how long it takes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Location.&lt;/strong&gt; Where it lives, who’s near it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Failure.&lt;/strong&gt; What goes wrong with it, what wears out, what you have to maintain.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Feelings.&lt;/strong&gt; How people feel about it. Nostalgia, fear, annoyance, indifference.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Do not skip the last two. I’ll show you why in a minute.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Force every attribute onto the problem
&lt;/h3&gt;

&lt;p&gt;One idea per attribute, and take the attribute literally. “Needs a periodic inspection tag” becomes “send a literal quarterly ‘inspection’ email asking subscribers to reconfirm interest.” That’s the model’s actual output, by the way. Write it down.&lt;/p&gt;

&lt;p&gt;The rule that makes the whole thing work: &lt;strong&gt;you are not allowed to skip an attribute because the mapping is awkward.&lt;/strong&gt; The awkward ones are the point.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Cross out everything you have read before
&lt;/h3&gt;

&lt;p&gt;Go down the list and mark every idea that could have appeared in a normal article about your problem. Most of them will. Whatever is left is what you came for.&lt;/p&gt;

&lt;h3&gt;
  
  
  6. Do it with two or three objects, then breed the survivors
&lt;/h3&gt;

&lt;p&gt;One object gives you one set of starting points. Two or three give you survivors from different directions, and now: take two survivors from different objects and make one idea that needs both of them to exist. Not “pick the best one” or “improve one.” A child that neither parent could have produced alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  Handing the Model the Process
&lt;/h2&gt;

&lt;p&gt;Here’s the prompt I actually ran, word for word, with the object swapped in:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;I write a tech blog and I'm starting a weekly email newsletter that rounds up
which new AI models are actually worth paying attention to. I want ideas for
getting more readers to subscribe. Don't answer that directly. Run this process
instead, and show every step:

1. The object is: fire extinguisher. List ten attributes of it. Physical parts,
   how it's used, where it lives, what goes wrong with it, how people feel about
   it. Literal attributes, not metaphors yet.
2. For each attribute, force a literal mapping onto the subscription problem.
   One idea per attribute. Do not skip an attribute because the mapping is
   awkward. The awkward ones are the point.
3. Mark any idea that could have come from a normal newsletter-growth article
   with [GENERIC]. Be honest.
4. Pick the two ideas that are least generic and still actually doable by one
   person, and say what the first concrete step would be for each.

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Notice what’s in there and what isn’t. It never says “be creative.” It never says “think outside the box.” It never says “avoid generic ideas,” which is a whole separate problem I’ll get to later in the series. It’s steps 3 through 5 of the human version, written down in order, with the one line in the middle that stops the model from taking the easy exit.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdcij9ehxoovwfhr20eki.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdcij9ehxoovwfhr20eki.jpg" alt="Handing the Model the Process" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I ran it three times with three objects: a pressure cooker, a fire extinguisher, and a metronome. Here’s a small script that does the same thing, if you want to run it against your own problem:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;urllib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;

&lt;span class="n"&gt;MODEL&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;anthropic/claude-sonnet-5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;PROBLEM&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;I write a tech blog and I&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;m starting a weekly email newsletter that &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
           &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;rounds up which new AI models are actually worth paying attention to. &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
           &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;I want ideas for getting more readers to subscribe.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;STEPS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Don&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;t answer that directly. Run this process instead, and show every step:

1. The object is: {obj}. List ten attributes of it. Physical parts, how it&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s used,
   where it lives, what goes wrong with it, how people feel about it. Literal
   attributes, not metaphors yet.
2. For each attribute, force a literal mapping onto the problem. One idea per
   attribute. Do not skip an attribute because the mapping is awkward.
   The awkward ones are the point.
3. Mark any idea that could have come from a normal article on this problem
   with [GENERIC]. Be honest.
4. Pick the two ideas that are least generic and still doable by one person,
   and say what the first concrete step would be for each.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;ask&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;body&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;MODEL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;temperature&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;1.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                       &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;messages&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;}]}).&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;req&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;urllib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Request&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://openrouter.ai/api/v1/chat/completions&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Authorization&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Bearer &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;OPENROUTER_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
         &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Content-Type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;application/json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;})&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;load&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;urllib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;urlopen&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;300&lt;/span&gt;&lt;span class="p"&gt;))[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;choices&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;message&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;obj&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pressure cooker&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;fire extinguisher&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;metronome&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="n"&gt;out&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;ask&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;PROBLEM&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;STEPS&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;format&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;obj&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;obj&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;obj&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;_&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;.md&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;w&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Swap in your own problem. Swap in your own objects, and don’t choose them yourself. That’s the part the script is for.&lt;/p&gt;

&lt;h3&gt;
  
  
  What came out
&lt;/h3&gt;

&lt;p&gt;Thirty forced ideas. By the model’s own count, about ten of them were clearly not generic, three more were borderline, and the rest were, in the metronome run’s own words, “generic growth advice in a costume.”&lt;/p&gt;

&lt;p&gt;The costumes tell you how the technique fails:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Attribute&lt;/th&gt;
&lt;th&gt;Forced idea&lt;/th&gt;
&lt;th&gt;Verdict&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Pressure cooker: lid locks shut&lt;/td&gt;
&lt;td&gt;A modal that seals the article until you subscribe or decline&lt;/td&gt;
&lt;td&gt;Generic. It’s a content-lock popup.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pressure cooker: builds pressure to cook faster&lt;/td&gt;
&lt;td&gt;“Next batch replaces this list in 7 days”&lt;/td&gt;
&lt;td&gt;Generic. It’s scarcity marketing.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fire extinguisher: wall-mounted, fixed location&lt;/td&gt;
&lt;td&gt;Subscribe box in the exact same spot on every post&lt;/td&gt;
&lt;td&gt;Generic. Basic conversion advice.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Metronome: adjustable tempo weight&lt;/td&gt;
&lt;td&gt;Pick a short or long version at signup&lt;/td&gt;
&lt;td&gt;Generic. Everyone offers this.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feczyui6x3ky0i52gpuwq.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feczyui6x3ky0i52gpuwq.jpg" alt="What came out" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Every one of those is a real attribute landing on something that already exists. The model went from a strange starting point and ended up back in the middle, the path of least resistance.&lt;/p&gt;

&lt;p&gt;And the survivors:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Attribute&lt;/th&gt;
&lt;th&gt;Forced idea&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Fire extinguisher: ignored until there’s a fire&lt;/td&gt;
&lt;td&gt;Sell the silence. Check it only when a model actually matters.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fire extinguisher: reassurance from mere presence&lt;/td&gt;
&lt;td&gt;Sell it as insurance. Skip every issue, you’re still covered.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fire extinguisher: loud, messy discharge&lt;/td&gt;
&lt;td&gt;One blunt, unhedged verdict per model, deliberately not balanced&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Metronome: gets turned off out of annoyance&lt;/td&gt;
&lt;td&gt;Ask everyone who unsubscribes one question and publish the answers monthly&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Metronome: mixed feelings, eventually internalized and discarded&lt;/td&gt;
&lt;td&gt;A graduation point. “Read 8 issues and you’ll be able to spot a hyped model yourself.”&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pressure cooker: emotional baggage, nostalgia or fear&lt;/td&gt;
&lt;td&gt;A fixed narrator voice people get attached to, not just the information&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Most of the survivors came out of the attributes about &lt;strong&gt;failure and feelings&lt;/strong&gt; : ignored until there’s a fire, turned off out of annoyance, mixed feelings, emotional baggage. The physical attributes (the lid, the tempo weight, the wall mount) mostly mapped back onto tactics that already have names. It’s why step 3 in the human version tells you not to skip the last two categories. Parts are easy to list, and they’re also the ones that map onto what you already know.&lt;/p&gt;

&lt;p&gt;The graduation point is my other favorite. Growth advice is built on keeping subscribers forever. That idea puts an end date on the relationship and uses the end date as the pitch.&lt;/p&gt;

&lt;h3&gt;
  
  
  One caveat about the judge
&lt;/h3&gt;

&lt;p&gt;The model graded its own work, and a model grading its own originality is a pretty weak judge. It was also clearly being hard on itself on purpose and overcorrected. I’d rather it cut too much than let everything through. But the [GENERIC] flag is a filter, not a verdict. You still read the list yourself, and you still get the final say.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where This Came From
&lt;/h2&gt;

&lt;p&gt;The theory is Arthur Koestler’s. In &lt;em&gt;The Act of Creation&lt;/em&gt; (1964) he argued that jokes, scientific discovery and art all run on the same machinery, which he called &lt;strong&gt;bisociation&lt;/strong&gt; : two ways of seeing that each make perfect sense on their own and are never normally used together. Hold both at once and the collision is the idea. A pun is bisociation. So, in his telling, is a scientific breakthrough.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqqcmnn1km38bsebczwyf.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqqcmnn1km38bsebczwyf.jpg" alt="Where This Came From" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;One of his best-known examples is Gutenberg. The part that’s documented is that the printing press adapted the screw press farmers already used for pressing grapes and olives. Gutenberg took a machine from the wine harvest and pointed it at metal type. That’s a clean case of two frames that came from two places. The tidier version, where Gutenberg has a flash of insight at a wine harvest and writes a letter about Minerva springing from his brain, is Koestler’s telling. Treat the combination as history and the eureka moment as a good story.&lt;/p&gt;

&lt;p&gt;Charles S. Whiting’s &lt;em&gt;Creative Thinking&lt;/em&gt; (1958) is where “forced relationships” is usually traced: take an item unrelated to the problem and force a connection anyway. An arbitrary thing, a forced mapping, and no permission to bail when the mapping gets weird.&lt;/p&gt;

&lt;p&gt;It has a lot of relatives. Edward de Bono’s random-word technique is the same move with a word instead of an object. The “Combine” step in SCAMPER is a gentler version. Morphological analysis is its systematic cousin, where you break the problem into parameters and walk every combination instead of letting chance pick. I’ll get to that one later in the series. The distinction that matters here: forced relationships wants &lt;em&gt;distance&lt;/em&gt;. The point is that the object has nothing to do with your problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Crossover Move
&lt;/h2&gt;

&lt;p&gt;So far this is one object at a time. The title of the post promises two things forced together.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fls1t885wji90soa6j39e.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fls1t885wji90soa6j39e.jpg" alt="The Crossover Move" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Google DeepMind’s FunSearch (&lt;a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC10794145/" rel="noopener noreferrer"&gt;Romera-Paredes et al., &lt;em&gt;Nature&lt;/em&gt;, December 2023&lt;/a&gt;) found new results in mathematics by having a language model write programs, scoring them, and keeping the good ones in a database. The interesting part is how it builds each new prompt. It doesn’t hand the model its best program and say “improve this.” It samples two programs from the database, sorts them by score, labels them &lt;code&gt;v0&lt;/code&gt; and &lt;code&gt;v1&lt;/code&gt;, and asks for the next version. The paper found two programs worked better than one, with diminishing returns after that, and gives the reason in one sentence: combining several programs “enables the LLM to spot patterns across the different programs and generalize those.”&lt;/p&gt;

&lt;p&gt;AlphaEvolve (&lt;a href="https://arxiv.org/abs/2506.13131" rel="noopener noreferrer"&gt;Novikov et al., 2025&lt;/a&gt;) is the bigger, newer version of the same idea, with a program database built to keep the parents diverse so ideas explored earlier can resurface later instead of getting lost.&lt;/p&gt;

&lt;p&gt;The evolutionary computation people have a name for this. It’s &lt;strong&gt;crossover&lt;/strong&gt;. Two parents, one child. They’ve been doing it for decades and nobody in the prompt-template world seems to have noticed, because they call it “program search” instead of “brainstorming.”&lt;/p&gt;

&lt;p&gt;FunSearch’s two parents come from the same island, so they’re related. It combines, but it doesn’t force distance. What I did below uses both halves: forced distance to &lt;em&gt;generate&lt;/em&gt; the parents, crossover to &lt;em&gt;merge&lt;/em&gt; them.&lt;/p&gt;

&lt;h3&gt;
  
  
  Running it
&lt;/h3&gt;

&lt;p&gt;I took three survivors from two different objects and gave them back to the model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Three ideas came out of separate exercises, each forced from a different
unrelated object:

Parent A (from a fire extinguisher): Sell subscribing as insurance. "Subscribe
once. Skip every issue if you want. Just be covered when a model actually matters."
Parent B (from a metronome): Build in an explicit graduation point. "Read 8
issues and you'll be able to spot a hyped model yourself."
Parent C (from a metronome): Publish churn. Ask everyone who unsubscribes one
question and publish the anonymized answers monthly.

Do not pick the best one and do not refine any one of them. Write three child
ideas. Each child must be a single mechanism that needs at least two of the
parents to exist, and could not be built from either parent alone. Name which
parents each child takes from and what it takes.

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The “do not pick the best one and do not refine” line is doing the same job as “do not skip an awkward attribute” did earlier. Without it, a model handed three ideas will rank them and polish the winner, because that’s what a helpful assistant does with a list.&lt;/p&gt;

&lt;p&gt;What came back:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh82sx6pfmnnz2vrojfxg.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh82sx6pfmnnz2vrojfxg.jpg" alt="Running it" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Graduated coverage (insurance + graduation).&lt;/strong&gt; At issue 8, a subscriber automatically drops to an insurance-only tier: silence unless a model genuinely matters. Graduation alone gives you no reason to stay after you graduate. Insurance alone has no moment where “skip whenever” becomes official.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The graduation exit log (graduation + churn).&lt;/strong&gt; People who leave &lt;em&gt;after&lt;/em&gt; issue 8 get a different exit question: not “why are you leaving” but “what made you confident enough to leave.” Those answers get published as proof the newsletter actually teaches something. Churn becomes a credential.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Churn to reinsurance (insurance + churn).&lt;/strong&gt; The monthly published churn answers sit next to a one-click “reinstate your coverage” link. You were never really unsubscribed, just paused until it matters.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Graduated coverage is the insurance idea with a trigger attached, and the trigger is the thing the insurance idea was missing. I wouldn’t have gotten there from either object alone. And I definitely wouldn’t have gotten there from “give me creative ideas,” which got me a mascot.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Not Just Ask for Fifty Ideas?
&lt;/h2&gt;

&lt;p&gt;This is the obvious shortcut, and it doesn’t work.&lt;/p&gt;

&lt;p&gt;In 2024, Chenglei Si, Diyi Yang and Tatsunori Hashimoto ran a large study comparing research ideas from an LLM pipeline against ideas from over a hundred NLP researchers (&lt;a href="https://arxiv.org/abs/2409.04109" rel="noopener noreferrer"&gt;arXiv:2409.04109&lt;/a&gt;). The LLM ideas were actually judged &lt;em&gt;more&lt;/em&gt; novel than the humans’, and slightly less feasible. But to get there the pipeline generated 4,000 seed ideas per topic, and when they deduplicated them, only about 5% survived. The share of new, non-duplicate ideas in each batch kept dropping as they generated more, until it plateaued. The authors list the lack of diversity in generation as an open problem.&lt;/p&gt;

&lt;p&gt;So asking for more doesn’t get you more. It gets you the same few ideas in different words, over and over, with the occasional new one. Volume is a terrible way to escape the middle.&lt;/p&gt;

&lt;p&gt;Forcing a starting point is the cheaper way out. Thirty forced ideas gave me roughly ten survivors. Four thousand unforced ones gave that study about two hundred. The two aren’t directly comparable (different task, different judge, different everything), but the direction is the point.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where It Breaks
&lt;/h2&gt;

&lt;p&gt;I’m not going to pretend this is the only way to do it, because the research that exists doesn’t let me.&lt;/p&gt;

&lt;p&gt;The one solid human study I found on distance points the other way. Joel Chan, Christian Schunn and colleagues looked at which sources of inspiration led to the most creative design ideas (&lt;a href="https://www.sciencedirect.com/science/article/pii/S0142694X14000611" rel="noopener noreferrer"&gt;&lt;em&gt;Design Studies&lt;/em&gt;, 2015&lt;/a&gt;), and found that “conceptually closer rather than farther sources lead to more creative ideas,” consistently across different design problems. There was no support for the best ideas coming from the farthest sources.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fywz5njo3gvb9o7ybn2ql.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fywz5njo3gvb9o7ybn2ql.jpg" alt="Where It Breaks" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That’s a different setup from this one. Their sources were examples of &lt;em&gt;other solutions&lt;/em&gt;, and a fire extinguisher is not a solution to anything. It’s a jig, a thing to push your thinking against. In the runs above, the random object generated plenty of candidates and most of them were junk. The crossover step, merging survivors that were close enough to fit together, is where the best idea came from. That’s consistent with Chan’s finding, not a contradiction of it.&lt;/p&gt;

&lt;p&gt;The other ways it breaks, from actually doing it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Physical attributes map to existing tactics.&lt;/strong&gt; A lid that locks becomes a popup. If your list is all parts, you’ll get all costumes. The failure and feeling attributes are where most of the survivors were.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The mapping can be too loose.&lt;/strong&gt; If you let yourself (or the model) go metaphorical in step 2, anything maps to anything and the object stops doing work. Literal mappings are more constrained, which is what you want.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It solves a narrow kind of problem.&lt;/strong&gt; It’s great when you’re stuck in a rut of five versions of the same idea. It’s useless when you don’t have enough information yet. No fire extinguisher is going to tell me whether anybody wants an AI model newsletter in the first place.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You still have to pick.&lt;/strong&gt; Graduated coverage looks good to me because of what I know about the people who read this blog. The model doesn’t know any of that. The pick is yours.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  So What Did I Actually Get?
&lt;/h2&gt;

&lt;p&gt;A newsletter pitch I’d never have written on my own. It tells people they don’t have to read it, and it goes quiet after issue 8 for anyone who’s learned the skill. Does it work? I don’t know yet. The newsletter doesn’t exist yet.&lt;/p&gt;

&lt;p&gt;The part that changed how I use models is smaller than that, though. I used to think the fix for a bland answer was a better question. It isn’t. What got me somewhere new was running a boring, seventy-year-old, pen-and-paper process, writing its steps down in order, and handing the model the steps instead of the question.&lt;/p&gt;

&lt;p&gt;Next one: give yourself an arbitrary rule. It’s the most recommended creativity advice in existence, and there’s a reason to be suspicious of it.&lt;/p&gt;

</description>
      <category>prompts</category>
      <category>creativity</category>
      <category>ai</category>
    </item>
    <item>
      <title>Sonnet 5.5 vs Opus 5.5: The Cheap Claude Costs More</title>
      <dc:creator>Stephan Miller</dc:creator>
      <pubDate>Tue, 29 Sep 2026 13:00:00 +0000</pubDate>
      <link>https://dev.to/eristoddle/sonnet-55-vs-opus-55-the-cheap-claude-costs-more-6n5</link>
      <guid>https://dev.to/eristoddle/sonnet-55-vs-opus-55-the-cheap-claude-costs-more-6n5</guid>
      <description>&lt;p&gt;The price sheet caved this week. I don’t mean one lab and some polite matching. Anthropic and OpenAI both cut prices on September 22, a day after Xiaomi dropped a 28-cent open-weights model, and then Anthropic came back six days later with a new Sonnet it swears is cheaper to run. If you pay your own API bill, this was the best week of the year to be alive.&lt;/p&gt;

&lt;p&gt;Then I read the fine print, because that’s the job. The new “cheap” Claude costs more per task than the new expensive Claude, at least when you turn the effort dial all the way up. The cheapest good coding model on the board this week comes from a company Anthropic says built it partly by siphoning Claude conversations through OpenClaw, which is the same agent framework I run in a Docker container in my house. And the free model sitting at number two on OpenRouter won’t tell you who made it.&lt;/p&gt;

&lt;p&gt;So yes, prices went down. Read the meter anyway.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Everybody blinked at once&lt;/li&gt;
&lt;li&gt;Opus 5.5 is the rare launch where every signal agrees&lt;/li&gt;
&lt;li&gt;The cheap Claude costs more than the expensive Claude&lt;/li&gt;
&lt;li&gt;OpenAI just showed up at the bottom of the price sheet&lt;/li&gt;
&lt;li&gt;The cheapest coding model has a distillation problem&lt;/li&gt;
&lt;li&gt;The cheapskate picks&lt;/li&gt;
&lt;li&gt;The free model at number two won’t say who made it&lt;/li&gt;
&lt;li&gt;Horror story: 48,218 files in 103 seconds&lt;/li&gt;
&lt;li&gt;What’s coming&lt;/li&gt;
&lt;li&gt;Finally&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Everybody blinked at once
&lt;/h2&gt;

&lt;p&gt;The clustering matters more than any one launch.&lt;/p&gt;

&lt;p&gt;On September 21, Xiaomi shipped MiMo-V2.6-Pro and MiMo-V2.6-Flash as MIT-licensed open weights. Pro is a 1.02-trillion-parameter mixture-of-experts with 42 billion active, a million-token context, and it’s priced at 43.5 cents in and 87 cents out. Flash is 310 billion total, 15 billion active, and costs 14 cents in and 28 cents out. Artificial Analysis gives Pro a 46 on its Intelligence Index, which makes it the top-scoring open-weights model on the board right now.&lt;/p&gt;

&lt;p&gt;A day later, on September 22, Anthropic released &lt;a href="https://www.anthropic.com/claude-opus-5-5" rel="noopener noreferrer"&gt;Claude Opus 5.5&lt;/a&gt; at $4 in and $20 out, down from $5 and $25. Cache reads dropped 60 percent to 20 cents. Same day, OpenAI released &lt;a href="https://dev.to/sam_hiotis_117598dbfa3ac2/openai-releases-gpt-6-sol-and-luna-50-cheaper-api-pricing-and-benchmarks-3djg-temp-slug-7843758"&gt;GPT-6 Sol and GPT-6 Luna&lt;/a&gt; and cut prices in half or better. Sol went from $4/$20 to $2/$10. Luna went from 20 cents and $1.20 to 10 cents in and 50 cents out. OpenAI says those are permanent prices, not a promo, and credits caching and inference improvements.&lt;/p&gt;

&lt;p&gt;Then on September 28, Anthropic released Claude Sonnet 5.5 at $2 in and $10 out. Same price as Sonnet 5, but a much better model.&lt;/p&gt;

&lt;p&gt;Three labs from two countries, all pushing prices down inside 48 hours. I’ve been doing this roundup since April and I’ve never seen the whole price column in my spreadsheet go down at once like this. Usually one lab cuts and everyone else pretends not to notice for a month.&lt;/p&gt;

&lt;h2&gt;
  
  
  Opus 5.5 is the rare launch where every signal agrees
&lt;/h2&gt;

&lt;p&gt;I spend a lot of this roundup explaining why the benchmark number and the blind-test number disagree. Three weeks ago GPT-6 Astra tied for number one on Artificial Analysis and debuted twenty-fourth on Arena. It’s twenty-sixth now. That split is normal. It’s practically the house style.&lt;/p&gt;

&lt;p&gt;Opus 5.5 didn’t do that. It’s number one on the Artificial Analysis Intelligence Index at 58, five points clear of Fable 5.1 and GPT-6 Astra, which are tied at 53. It’s also number one on Arena Overall at 1509, and number one in Creative Writing, Instruction Following, and Hard Prompts. The hard-benchmark test and the blind vote landed on the same model in the same week. I almost never get to write that sentence.&lt;/p&gt;

&lt;p&gt;It also ended a running gag. For weeks I’ve been pointing out that the outright leader in Instruction Following and Hard Prompts was claude-opus-4-6, a model about a year old, beating everything Anthropic shipped after it. Opus 5.5 finally took both categories. Coding is the one place the old guard hangs on: opus-4-6-high, opus-4-7-high, and fable-5-high are in a three-way tie at 1551, with Opus 5.5 at fifth.&lt;/p&gt;

&lt;p&gt;Two asterisks. First, the votes are thin. Opus 5.5 has 2,307 votes in Overall and only 487 in Creative Writing, so its 20-point lead there could shrink as the crowd catches up. Second, every Claude score on Artificial Analysis still carries the “with fallback” tag, which means some answers came from a weaker model when the main one refused. I’ve flagged that on every Anthropic launch since Opus 5. The number is real. The footnote is load-bearing.&lt;/p&gt;

&lt;p&gt;Reddit’s reaction has mostly been relief. The most-upvoted praise isn’t even about coding, it’s about &lt;a href="https://botmonster.com/ai/opus-5-5-is-the-claude-comeback-reddit-was-waiting-for/" rel="noopener noreferrer"&gt;how the thing talks&lt;/a&gt;. One r/ClaudeCode user called it “genuinely an order of magnitude improvement over Opus 5” in communication, which tracks. Opus 5 had a habit of saying “blast radius” and “load-bearing” like it was paid per use. (Yes, I just used “load-bearing” myself. I’m allowed. I’m not charging you per token.) Another user said they’d worked “non stop since release” and “barely made a dent” in their Max limits.&lt;/p&gt;

&lt;p&gt;Not everybody agrees on that part. A &lt;a href="https://hardforum.com/threads/claude-opus-5-5.2049520/" rel="noopener noreferrer"&gt;HardForum&lt;/a&gt; user on the $100 Max plan said Opus 5 used to get them through whole days of coding, weekends included, without hitting the weekly limit. They upgraded to Opus 5.5 on Tuesday and were at 91 percent by Thursday morning. They were running it on “Extra” effort. Another user in the same thread, on medium, had used 12 percent of a $200 plan’s weekly limit, and a third summed it up: “5.5 extra eat tokens way more than previous gen.” Remember that effort dial. It comes back in a minute. Anthropic’s own demo was a HAProxy port from C to Rust that Opus 5.5 finished in 9.5 hours against Fable 5.1’s 12, at 51 percent lower cost. Both can be true. A well-defined port is not an open-ended feature in a messy real app, and the complaints are all coming from the messy real apps. And the best comment in the whole pile: “Queue the whining in two weeks about the model being nerfed.” Set a reminder.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cheap Claude costs more than the expensive Claude
&lt;/h2&gt;

&lt;p&gt;Sonnet 5.5 is the budget model. It’s half the per-token price of Opus 5.5. It’s also scary good. It scored 70.6 percent on Terminal-Bench 4.0, the command-line agent test, up from Sonnet 5’s 10.3, and ahead of Opus 5.5’s 66.4. On Artificial Analysis it scores 56, two points under Opus 5.5 at max and ahead of Fable 5.1 and GPT-6 Astra.&lt;/p&gt;

&lt;p&gt;Then Artificial Analysis ran the whole index suite and checked the bill. At max effort, &lt;a href="https://www.beri.net/article/claude-sonnet-5-5-vs-opus-5-5-cost-per-task-matched-score-effort-between-tools-migration" rel="noopener noreferrer"&gt;Sonnet 5.5 cost $7.60 per task. Opus 5.5 cost $5.98&lt;/a&gt;. The cheap model was 27 percent more expensive.&lt;/p&gt;

&lt;p&gt;The reason is tokens. Sonnet 5.5 burned 410 million output tokens getting through the suite. The median for models in its price tier is 88 million. Opus 5.5 used about 260 million. At max effort, Sonnet thinks roughly 60 percent longer per task than Opus, and half the price per token doesn’t cover 60 percent more tokens. It’s a fast model, 138 tokens a second, and it spends that speed talking to itself.&lt;/p&gt;

&lt;p&gt;It gets worse when you match scores instead of effort settings. Sonnet 5.5 at max and Opus 5.5 at xhigh both score 56 on the index. Sonnet costs $7.60 a task to get there. Opus costs $3.46. Same score, and the budget model costs 2.2 times as much. According to &lt;a href="https://www.beri.net/article/claude-sonnet-5-5-vs-opus-5-5-cost-per-task-matched-score-effort-between-tools-migration" rel="noopener noreferrer"&gt;the same analysis&lt;/a&gt;, Sonnet only comes out cheaper at the bottom of the effort range.&lt;/p&gt;

&lt;p&gt;None of that makes Sonnet 5.5 a bad model. It went from 10.3 to 70.6 on Terminal-Bench in one release, and one customer quoted by MarkTechPost measured about 121K tokens per answer against 497K on Sonnet 5. It’s just not automatically the cheap option anymore. So the practical advice is boring. Don’t crank Sonnet 5.5 to max because it’s “the cheap one.” If a task needs the top of the dial, Opus at xhigh gets you the same score for less. If it doesn’t, run either one at a lower effort and check your actual bill after a day, not the price page.&lt;/p&gt;

&lt;p&gt;I’ve written some version of “cost per token is not cost per task” in this roundup maybe ten times. I’d never seen it flip the price order inside one lab’s own lineup before.&lt;/p&gt;

&lt;h2&gt;
  
  
  OpenAI just showed up at the bottom of the price sheet
&lt;/h2&gt;

&lt;p&gt;For most of this year the cheapest-good-model story has been a Chinese open-weights story. Kimi, then MiMo, then GLM, and now maybe MiMo again. American labs competed at the top and left the floor alone.&lt;/p&gt;

&lt;p&gt;GPT-6 Luna costs 10 cents in and 50 cents out. That’s exactly the list price of GLM-5.3-Flash, the model that’s been my cheapskate default for the last month. And Luna is in the Arena Coding band, ranked 43rd at 1515, only 36 points behind the leader. It isn’t the cheapest thing in the band, but it’s the first time I’ve seen an OpenAI model sit in a cheapskate band at the same price as the Chinese floor. It’s also fast. Artificial Analysis clocks it at 148 tokens a second, about three times GLM’s speed.&lt;/p&gt;

&lt;p&gt;The catch is capability. Luna scores 37 on the Intelligence Index. GLM scores 42. Outside of coding, Luna falls out of the Arena bands entirely. It’s 86th Overall. So it’s a fast, cheap, preference-decent coding model, not a general replacement. But OpenAI ignored this end of the price sheet for a year, and now they’re in it.&lt;/p&gt;

&lt;p&gt;GPT-6 Sol is the more awkward launch. $2/$10, Intelligence Index 48, and OpenAI’s own numbers have it beating Opus 5 on AutomationBench (33.2 percent vs 26.9) and edging it on OSWorld at 80 percent lower cost. On Arena it’s 60th Overall at 1457, which misses the Overall band by two points.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cheapest coding model has a distillation problem
&lt;/h2&gt;

&lt;p&gt;MiMo-V2.6-Flash is the cheapest model inside the Arena Coding band this week. It’s ranked 24th at 1525, 26 points behind the leader, for 28 cents out. That’s 89 times cheaper than a $25 leader. Artificial Analysis puts it on its intelligence-vs-cost Pareto frontier. On OpenRouter it went from nothing to the seventh most-used model on the platform in a week, 6.93 trillion tokens. By every signal I track, it’s a real value pick.&lt;/p&gt;

&lt;p&gt;And on September 10, eleven days before it shipped, Anthropic published a threat intelligence report accusing seven China-based labs of what it calls “illicit distillation,” using Claude’s outputs to train their own models. Xiaomi was one of them. According to &lt;a href="https://the-decoder.com/xiaomis-affordable-flagship-ai-leads-the-open-models-and-anthropic-says-claude-helped-get-it-there/" rel="noopener noreferrer"&gt;the-decoder’s summary&lt;/a&gt; of case GTG-16008, Xiaomi forwarded more than 400,000 conversations from users of its own MiMo chatbot to Claude between March and April, through OpenClaw and OpenCode, to pull out training data. Across all seven labs, Anthropic counts around 190 million exchanges. Xiaomi hasn’t responded.&lt;/p&gt;

&lt;p&gt;I run OpenClaw. It’s the agent framework behind the assistant I run at home, the one that handles my scheduled jobs and writes into my notes vault. So I read that paragraph twice. To be clear, this is an accusation from a competitor, not a court finding, and nothing in it suggests OpenClaw itself did anything wrong. It’s a tool. Somebody pointed it at Claude at industrial scale. But if you were one of those 400,000 MiMo chatbot users, your conversations apparently took a trip you never agreed to.&lt;/p&gt;

&lt;p&gt;So what do you do with the pick? I’m printing it, because the method is the method and the price and rating are real. I’m also printing it with two asterisks. First, it only has 1,043 Coding votes, so it’s the emerging pick, not the settled one. GLM-5.3-Flash is right behind it at 1523 with five times the votes for 50 cents, and that’s the steadier bet. Second, if where a model’s training data came from matters to your company, and for some of you it contractually does, this one has a question hanging over it.&lt;/p&gt;

&lt;p&gt;And watch the price you see. On OpenRouter, MiMo-V2.6-Flash’s headline price shows 8 cents in and $1.28 out. That’s one small host’s endpoint. Xiaomi’s own endpoint is 14 cents and 28 cents, and it’s carrying 97 percent of the traffic. If you read the headline number, you’ll think it’s the most expensive cheap model on the board. Check the provider list.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cheapskate picks
&lt;/h2&gt;

&lt;p&gt;Same method every week. Take the category leader’s Arena rating, draw a band 50 points below it, and find the cheapest model still inside the band. The top of Arena is compressed, so the leader is usually only a little better than something 40 to 100 times cheaper. The band gets computed in code from the full table, never eyeballed off the first screen, because eyeballing is how you delete the whole cheap tail without noticing. I learned that one in public.&lt;/p&gt;

&lt;p&gt;The good news: Arena finally refreshed. The last two issues ran on the same September 13 snapshot. This week’s data is stamped September 25. Bands ran deep: 57 models in Overall, 73 in Coding, 43 in Hard Prompts and Math, 38 in Instruction Following, and just 12 in Creative Writing. GLM-5.3-Flash is quoted at Z.ai’s list price, $0.15 in and $0.50 out. Arena’s price column now shows 20 cents out for it, but that’s a third-party floor, not list. Budget on fifty.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Leader&lt;/th&gt;
&lt;th&gt;$ out&lt;/th&gt;
&lt;th&gt;Cheapskate pick&lt;/th&gt;
&lt;th&gt;$ out&lt;/th&gt;
&lt;th&gt;Δ rating&lt;/th&gt;
&lt;th&gt;Cheaper by&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Overall&lt;/td&gt;
&lt;td&gt;claude-opus-5.5-high (1509)&lt;/td&gt;
&lt;td&gt;$20&lt;/td&gt;
&lt;td&gt;GLM-5.3-Flash (1474, #35)&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;−35&lt;/td&gt;
&lt;td&gt;40×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Coding&lt;/td&gt;
&lt;td&gt;claude-opus-4-6-high (1551)&lt;/td&gt;
&lt;td&gt;$25&lt;/td&gt;
&lt;td&gt;MiMo-V2.6-Flash* (1525, #24)&lt;/td&gt;
&lt;td&gt;$0.28&lt;/td&gt;
&lt;td&gt;−26&lt;/td&gt;
&lt;td&gt;89×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Creative Writing&lt;/td&gt;
&lt;td&gt;claude-opus-5.5-high (1521)&lt;/td&gt;
&lt;td&gt;$20&lt;/td&gt;
&lt;td&gt;gemini-3.7-flash-high (1492, #4)&lt;/td&gt;
&lt;td&gt;$3.75&lt;/td&gt;
&lt;td&gt;−29&lt;/td&gt;
&lt;td&gt;5.3×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Instruction Following&lt;/td&gt;
&lt;td&gt;claude-opus-5.5-high (1516)&lt;/td&gt;
&lt;td&gt;$20&lt;/td&gt;
&lt;td&gt;GLM-5.3-Flash (1474, #26)&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;−42&lt;/td&gt;
&lt;td&gt;40×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hard Prompts&lt;/td&gt;
&lt;td&gt;claude-opus-5.5-high (1541)&lt;/td&gt;
&lt;td&gt;$20&lt;/td&gt;
&lt;td&gt;GLM-5.3-Flash (1499, #31)&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;−42&lt;/td&gt;
&lt;td&gt;40×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Math&lt;/td&gt;
&lt;td&gt;claude-fable-5-high (1523)&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;GLM-5.3-Flash (1501, #13)&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;−22&lt;/td&gt;
&lt;td&gt;100×&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;*Thin votes (1,043). The steadier Coding pick is GLM-5.3-Flash at #28, 1523, 50 cents, 50× cheaper.&lt;/p&gt;

&lt;p&gt;GLM-5.3-Flash held four of six, and it’s doing it on real vote counts now: 19,103 in Overall and 12,503 in Hard Prompts. That’s a pick you can lean on. The “cheaper by” numbers shrank, from 100× to 40× in Overall and from 50× to 40× in Instruction Following and Hard Prompts, but not because GLM got pricier. The leader got cheaper. Opus 5.5 at $20 took the top of three boards from $25 and $50 models. That’s the price war showing up in my own table.&lt;/p&gt;

&lt;p&gt;Coding is the first crack in GLM’s run in a month, and it’s a crack on thin votes. Math is still the row to squint at: GLM sits 13th on only 920 votes. If you want steadier, MiMo v2.5 Pro has 3,280 votes at 87 cents.&lt;/p&gt;

&lt;p&gt;Creative Writing is where the method bit back. Gemini 3 Flash was the Creative pick for four straight issues at $3. It didn’t get worse this week. Opus 5.5 raised the ceiling by 17 points, and Gemini 3 Flash’s 1457 fell out of a band that now starts at 1471. When the leader gets better, cheap models fall out of the band without doing anything wrong. The new pick is Gemini 3.7 Flash at $3.75, ranked fourth on preliminary votes. Also, Gemini 3.8 Flash’s intro pricing doubles on January 1, so check whether 3.7 goes the same way before you build a budget around it.&lt;/p&gt;

&lt;p&gt;The speed caveat is the same as every week, and the number moved again. Artificial Analysis clocks GLM-5.3-Flash at 48 tokens a second this week, down from 89 last week. MiMo-V2.6-Flash is 55. I’ve stopped carrying this number forward because it bounces around too much. Both are slow enough to notice in an agent loop. GPT-6 Luna at 148 is the fast one at this price.&lt;/p&gt;

&lt;h2&gt;
  
  
  The free model at number two won’t say who made it
&lt;/h2&gt;

&lt;p&gt;Space Bunny Alpha showed up on OpenRouter on September 23 as a stealth model: no maker named, free, a million-token context, and text, image, and video input. Five days later it was the second-biggest model on the platform for the week at 18.2 trillion tokens, and number one for the most recent day.&lt;/p&gt;

&lt;p&gt;It’s probably MiniMax. Tokenizer tests on launch day matched MiniMax’s M3 family on every string people threw at it (one tester ran 50 strings, all 50 matched), and on September 27 MiniMax released &lt;a href="https://cellcog.ai/blog/what-is-space-bunny-alpha/" rel="noopener noreferrer"&gt;M3.1-Flash-Preview&lt;/a&gt; with the same context length, the same five reasoning effort levels, and the same fast-coding pitch. Nobody has officially confirmed it.&lt;/p&gt;

&lt;p&gt;Two things worth knowing before you point production at it. It’s free, and free usage isn’t chosen usage. Free models in this roundup have a habit of spiking and then settling once the meter turns on. And the listing says prompts and completions “may be retained by the provider,” a provider that won’t tell you its name. The same week a distillation report dropped. Fine for kicking the tires on public code. I wouldn’t send it anything I’d mind reading in somebody’s training set.&lt;/p&gt;

&lt;p&gt;Meanwhile, on the paid board, DeepSeek V4.1 Flash is number one at 20.8 trillion tokens, up 23 percent. GLM-5.3-Flash slipped to third at 13.3 trillion, down 21 percent, though it’s still number one for the trailing month at 57.5 trillion. Usage is moving toward whatever is newest, cheapest, or free, which is what usage always does in the week after a launch pile-up.&lt;/p&gt;

&lt;h2&gt;
  
  
  Horror story: 48,218 files in 103 seconds
&lt;/h2&gt;

&lt;p&gt;Around September 20, a developer posted on r/ClaudeAI that Claude Code had deleted their project. According to &lt;a href="https://www.techradar.com/pro/security/i-broke-something-a-claude-code-ai-agent-deleted-48-000-files-in-just-over-100-seconds-then-apologized-for-doing-so" rel="noopener noreferrer"&gt;TechRadar’s writeup&lt;/a&gt;, the agent was told to rebuild a mirror of the project for a task. It figured out &lt;code&gt;build_mirror.py&lt;/code&gt; couldn’t refresh the mirror in place, so it wrote a little Python remover to delete an old copy sitting in a temp folder. That old copy held 7,332 ordinary files and 614 Windows directory junctions pointing back into the live project. The remover used &lt;code&gt;os.walk&lt;/code&gt; with &lt;code&gt;followlinks=False&lt;/code&gt;, which sounds safe. But &lt;code&gt;os.path.islink()&lt;/code&gt; returns false for Windows junctions, so Python didn’t treat them as links, walked right through them, and started deleting the real project on the other side.&lt;/p&gt;

&lt;p&gt;It took 103 seconds. It deleted 48,218 live files, and it emptied the Git object store too: &lt;code&gt;.git/objects&lt;/code&gt;, refs, and logs. The index still listed 7,221 paths, but the actual file contents were gone, so Git couldn’t restore anything. TechRadar’s headline quotes the agent’s own summary: “I broke something.” (This all comes from the Reddit post and an attached report, not an independent forensic investigation, so treat the details as the poster’s account.)&lt;/p&gt;

&lt;p&gt;Reddit’s verdict was harsh. It wasn’t wrong. The poster admitted they weren’t using GitHub properly and should have been working on a branch. But the reason I’m including it isn’t to dunk on someone. It’s that the agent did something that looks completely reasonable (clean up a stale temp copy) on a file system that had a trap in it that no human would have spotted in the moment either. Directory junctions and symlinks turn “delete this folder” into “delete whatever this folder points at,” and on Windows the standard Python safety check doesn’t even see the junctions. If you let an agent run &lt;code&gt;rm&lt;/code&gt; or its Python equivalent, the only real protection is a remote it can’t touch, so push before you let it clean up anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  What’s coming
&lt;/h2&gt;

&lt;p&gt;Claude Haiku 5.5 is next. Anthropic said Sonnet and Haiku 5.5 would follow Opus “in the coming weeks.” Sonnet already shipped. Haiku is the one left, and if it’s priced like Haiku usually is, it lands right in the cheapskate zone.&lt;/p&gt;

&lt;p&gt;MiniMax M3.1 should get official soon. If Space Bunny gets unmasked and priced, we’ll see how many of those 18 trillion free tokens stick around once there’s a bill.&lt;/p&gt;

&lt;p&gt;Then there’s MiMo-V2.6-Pro-UltraSpeed, which Xiaomi says is up to 20 times faster than Pro at the same quality. It already has 73 billion tokens on OpenRouter. If that holds, “cheap but slow” stops being the standard caveat on the Chinese value picks.&lt;/p&gt;

&lt;p&gt;DeepSeek V5 is speculation only. People keep floating October. DeepSeek hasn’t said anything, and I’m not putting a date on a model that doesn’t have one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Finally
&lt;/h2&gt;

&lt;p&gt;The price war is real, and it’s good for you. Opus got 20 percent cheaper and better at the same time. OpenAI cut Sol and Luna in half and made it permanent. The floor got a new 28-cent option. If you’re paying your own bill, almost every line on your invoice should drop next month. Who had “two frontier labs cut prices on the same day” on their bingo card?&lt;/p&gt;

&lt;p&gt;But three things I’ve been harping on all year showed up again, just in cheaper clothes. The sticker isn’t the task: the budget Claude costs more than the premium Claude when you run it hot. The free model isn’t free: someone is paying for those 18 trillion tokens, and they’re not telling you who or why. And the cheapest option comes with questions about where it came from, which some of you will care about and some of you won’t.&lt;/p&gt;

&lt;p&gt;Last week I said I’d come back to see whether the Arena crowd was any nicer to Grok 4.7 than its creator was. It wasn’t. Grok 4.7 debuted at 92nd Overall. Its creator called it mid. The crowd called it worse. Low bar, and it still found a way under it.&lt;/p&gt;

</description>
      <category>llm</category>
      <category>openrouter</category>
    </item>
    <item>
      <title>Track New AI Models on the Arena Leaderboard With a 70-Line Scraper</title>
      <dc:creator>Stephan Miller</dc:creator>
      <pubDate>Thu, 24 Sep 2026 12:00:00 +0000</pubDate>
      <link>https://dev.to/eristoddle/track-new-ai-models-on-the-arena-leaderboard-with-a-70-line-scraper-3f3m</link>
      <guid>https://dev.to/eristoddle/track-new-ai-models-on-the-arena-leaderboard-with-a-70-line-scraper-3f3m</guid>
      <description>&lt;p&gt;Every week I write the &lt;a href="https://www.stephanmiller.com/series/model-buzz-report/" rel="noopener noreferrer"&gt;Model Buzz Report&lt;/a&gt;, and every week part of the job is staring at the &lt;a href="https://dev.to/eristoddle/the-cheapskates-guide-to-the-arena-leaderboard-why-i-stopped-paying-claude-opus-prices-1ipn"&gt;Arena leaderboard&lt;/a&gt; trying to remember what it looked like last week. Which of these is new? Was that one there before? When did &lt;a href="https://dev.to/eristoddle/glm-53-got-too-good-at-hacking-to-ship-on-time-450p"&gt;GLM 5.3&lt;/a&gt; show up in the top 20, because it sure felt like it came out of nowhere?&lt;/p&gt;

&lt;p&gt;That is a question a scraper cannot answer. A scraper hands you the page as it is right now, and “what’s new” is a comparison between now and then. You need a then.&lt;/p&gt;

&lt;p&gt;So this post is about the then. It is the smallest useful version of the thing most scraping recipes in this series is going to need: save each run with a timestamp, run it on a schedule, and compare this run with the last one. The example is deliberately easy, one page and one question.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Why “Now” Is Almost Never the Question&lt;/li&gt;
&lt;li&gt;The Page: Arena’s Text Leaderboard&lt;/li&gt;
&lt;li&gt;
The Harness, All 70 Lines

&lt;ul&gt;
&lt;li&gt;Storage: SQLite, Not a JSON File&lt;/li&gt;
&lt;li&gt;The Diff: Two Kinds of “New”&lt;/li&gt;
&lt;li&gt;The Schedule: One Cron Line&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Day One Has No Then&lt;/li&gt;
&lt;li&gt;
Two Things It Got Wrong

&lt;ul&gt;
&lt;li&gt;That “NEW #1” Is a Rename&lt;/li&gt;
&lt;li&gt;It Missed the One That Started This&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Why Not Just Use Firecrawl’s Change Tracking?&lt;/li&gt;
&lt;li&gt;What This Is For&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why “Now” Is Almost Never the Question
&lt;/h2&gt;

&lt;p&gt;Think about what people actually want out of a scraped page. Is this price a deal? Did the competitor change their pricing? Which models are new near the top? None of those can be answered from one snapshot. “Is this a deal” needs the price history. “Did they change” needs the old page. “What’s new” needs the old list.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://firecrawl.link/stephan-miller" rel="noopener noreferrer"&gt;Firecrawl&lt;/a&gt; solves the ugly part of scraping, which is getting a clean page out of a site that renders everything with JavaScript and would rather you went away. It does not solve the then. Nothing that fetches a page can, because the then is data you had to collect yourself, back when it was the now.&lt;/p&gt;

&lt;p&gt;That takes three boring things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Storage.&lt;/strong&gt; Save each crawl with a timestamp instead of printing it and forgetting it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A schedule.&lt;/strong&gt; One crawl tells you nothing. A cron line fixes that.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A diff.&lt;/strong&gt; Compare this crawl with the last one and only say something when something changed.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That is the whole harness. I keep wanting to call it a framework and it keeps being 70 lines.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Page: Arena’s Text Leaderboard
&lt;/h2&gt;

&lt;p&gt;The target is &lt;a href="https://arena.ai/leaderboard/text" rel="noopener noreferrer"&gt;arena.ai/leaderboard/text&lt;/a&gt;, the overall text leaderboard. About 400 models, ranked by head-to-head human votes, with score, vote count, price, and context window per row.&lt;/p&gt;

&lt;p&gt;It is also a page that doesn’t want to be read by a simple fetch. The table is rendered client-side, and when I &lt;a href="https://dev.to/eristoddle/building-a-cost-saving-agent-skill-that-accidentally-became-its-own-weekly-blog-post-3o1h"&gt;built the Model Buzz skill&lt;/a&gt; I learned that some fetch tools hand back the wrong leaderboard on category URLs without any error at all. That is the worst kind of failure: data that looks plausible.&lt;/p&gt;

&lt;p&gt;Firecrawl got it clean on the first try:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;firecrawl scrape https://arena.ai/leaderboard/text &lt;span class="nt"&gt;--only-main-content&lt;/span&gt; &lt;span class="nt"&gt;--wait-for&lt;/span&gt; 3000

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcpcpj7m70lztyja6w559.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcpcpj7m70lztyja6w559.jpg" alt="The Page: Arena's Text Leaderboard" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;What comes back is markdown, and the leaderboard is a plain markdown table:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;| Rank | Rank Spread | Model | Score | Votes | Price $/M | Context |
| --- | --- | --- | --- | --- | --- | --- |
| 1 | 17 | Anthropic&lt;span class="nt"&gt;&amp;lt;br&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;claude-fable-5-high&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://www.anthropic.com/news/claude-fable-5-mythos-5&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="nt"&gt;&amp;lt;br&amp;gt;&lt;/span&gt;Anthropic · Proprietary | 1506±5 | 30,057 | $10 / $50 | 1M |
...
| 19 | 737 | &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;glm-5.3-max&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://z.ai/blog/glm-5.3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="nt"&gt;&amp;lt;br&amp;gt;&lt;/span&gt;Z.ai · MIT | 1483±6 | 10,960 | $1.40 / $4.40 | 1M |

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I gave it three seconds to let the page render. I didn’t test whether it needs that, so treat the number as superstition you are free to delete.&lt;/p&gt;

&lt;p&gt;No JSON schema, no LLM extraction, and no selectors. A regex reads that table fine. A plain scrape is one credit a page, and Firecrawl’s JSON extraction adds four more on top, so asking an LLM to parse a table this regular would be paying five times over for something a regex already does. If you haven’t installed the CLI yet, &lt;a href="https://dev.to/eristoddle/firecrawl-cli-setup-skip-the-installer-that-rewrites-your-editors-2klj"&gt;the setup post&lt;/a&gt; covers it and the installer you should skip.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Harness, All 70 Lines
&lt;/h2&gt;

&lt;p&gt;Python, standard library only, calling the Firecrawl CLI through &lt;code&gt;subprocess&lt;/code&gt;. No SDK, so the only install is the CLI you already have.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;#!/usr/bin/env python3
&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Watch the Arena text leaderboard and report models that just showed up near the top.

    python3 arena_watch.py # scrape now, save, compare with last run
    python3 arena_watch.py --url &amp;lt;wayback url&amp;gt; --at 2026-09-02 # backfill an old snapshot
&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;argparse&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sqlite3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sys&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timezone&lt;/span&gt;

&lt;span class="n"&gt;URL&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://arena.ai/leaderboard/text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;DB&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;arena.db&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;TOP&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt;

&lt;span class="n"&gt;ROW&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;compile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;^\| (\d+) \| \d+ \| (.+?) \| (\d+)±&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;NAME&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;re&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;compile&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;r&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;\[([^\]]+)\]&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;scrape&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;out&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;subprocess&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;firecrawl&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;scrape&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--only-main-content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--wait-for&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;3000&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="n"&gt;capture_output&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;check&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="n"&gt;stdout&lt;/span&gt;
    &lt;span class="n"&gt;rows&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;line&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;splitlines&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
        &lt;span class="n"&gt;m&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ROW&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;match&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;line&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;name&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;NAME&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;search&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;group&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;continue&lt;/span&gt; &lt;span class="c1"&gt;# skips the calendar widget and anything else that isn't a model row
&lt;/span&gt;        &lt;span class="n"&gt;org&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;group&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;&amp;lt;br&amp;gt;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)[&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; · &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
        &lt;span class="n"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;group&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)),&lt;/span&gt; &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;group&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;org&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;group&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;))))&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;rows&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;ap&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;argparse&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;ArgumentParser&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;ap&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add_argument&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;default&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;URL&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;ap&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add_argument&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;--at&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;help&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;timestamp to record, for backfilling old snapshots&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;args&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;ap&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;parse_args&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="n"&gt;rows&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;scrape&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="n"&gt;TOP&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;only parsed &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; rows, the page layout probably changed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;taken&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;args&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;at&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="n"&gt;datetime&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;timezone&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;utc&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;strftime&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;%Y-%m-%d %H:%M&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;db&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;sqlite3&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;connect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;DB&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;CREATE TABLE IF NOT EXISTS ranks
                  (taken TEXT, rank INT, model TEXT, org TEXT, score INT)&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;executemany&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;INSERT INTO ranks VALUES (?,?,?,?,?)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[(&lt;/span&gt;&lt;span class="n"&gt;taken&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;commit&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="n"&gt;prev&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SELECT MAX(taken) FROM ranks WHERE taken &amp;lt; ?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;taken&lt;/span&gt;&lt;span class="p"&gt;,)).&lt;/span&gt;&lt;span class="nf"&gt;fetchone&lt;/span&gt;&lt;span class="p"&gt;()[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;prev&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;taken&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;: first run, saved &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; models. Nothing to compare yet.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt;

&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;img&lt;/span&gt; &lt;span class="n"&gt;src&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/images/2026/track-new-ai-models-on-the-arena-leaderboard-body-2.jpg&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="n"&gt;alt&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;The Harness, All 70 Lines&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="n"&gt;srcset&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; /assets/resized/480/track-new-ai-models-on-the-arena-leaderboard-body-2.jpg 480w, /assets/resized/800/track-new-ai-models-on-the-arena-leaderboard-body-2.jpg 800w, /assets/resized/1400/track-new-ai-models-on-the-arena-leaderboard-body-2.jpg 1400w, &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="n"&gt;loading&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;lazy&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;

    &lt;span class="n"&gt;was_top&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt; &lt;span class="nf"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;,)&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SELECT model FROM ranks WHERE taken = ? AND rank &amp;lt;= ?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prev&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;TOP&lt;/span&gt;&lt;span class="p"&gt;))}&lt;/span&gt;
    &lt;span class="n"&gt;seen&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt; &lt;span class="nf"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;m&lt;/span&gt;&lt;span class="p"&gt;,)&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SELECT DISTINCT model FROM ranks WHERE taken &amp;lt; ?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;taken&lt;/span&gt;&lt;span class="p"&gt;,))}&lt;/span&gt;

    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;taken&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; vs &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;prev&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;rank&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;org&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;score&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;rank&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;TOP&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;break&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;seen&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; NEW #&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;rank&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; (&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;org&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;, &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;score&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;was_top&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; MOVED UP #&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;rank&lt;/span&gt;&lt;span class="si"&gt;:&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; (&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;org&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;, &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;score&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;__name__&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; __main__&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;main&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Here is what each piece is doing, mapped to the three boring things.&lt;/p&gt;

&lt;h3&gt;
  
  
  Storage: SQLite, Not a JSON File
&lt;/h3&gt;

&lt;p&gt;Every run inserts every row, all 400 or so, stamped with the time it was taken. One table, five columns.&lt;/p&gt;

&lt;p&gt;I went back and forth on JSONL here. A line of JSON per run is simpler to look at and fine for a diff against the last run. But the questions I actually care about later are history questions: how long has this model been in the top 20, when did it first appear anywhere on the board, is it climbing or sliding. With JSONL you write a loop for each of those. With SQLite you write a query. And SQLite ships with Python, so it costs nothing to install.&lt;/p&gt;

&lt;p&gt;Saving the whole board instead of just the top 20 is on purpose. Storage is cheap and you can’t go back and scrape last Tuesday. A model that debuts at #24 is not news today, but the day it cracks the top 20 you want to know it has been lurking for a week.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Diff: Two Kinds of “New”
&lt;/h3&gt;

&lt;p&gt;The script reports two different things, and the difference matters.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;NEW&lt;/strong&gt; means the model has never appeared anywhere on the board in any previous run. This is the one I actually wanted. A brand new model landing in the top 20 on its first appearance is the “where did that come from” moment.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MOVED UP&lt;/strong&gt; means the model was on the board before but was not in the top 20 last run. A climber. Less exciting, still worth a look.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Everything else, a model shuffling from #7 to #9, stays quiet. The whole point of the diff is that most runs should print almost nothing.&lt;/p&gt;

&lt;h3&gt;
  
  
  The Schedule: One Cron Line
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;0 8 &lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="k"&gt;*&lt;/span&gt; &lt;span class="nb"&gt;cd&lt;/span&gt; /path/to/arena-watch &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; python3 arena_watch.py &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; watch.log 2&amp;gt;&amp;amp;1

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo6agqyh9lym0pb2nw2aw.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo6agqyh9lym0pb2nw2aw.jpg" alt="The Schedule: One Cron Line" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Daily at 8am, appended to a log. That lives on my mini PC, which is on all the time. On a Mac you can use cron too, or launchd if you enjoy writing XML. Once a day is plenty for a leaderboard that needs thousands of votes to move a model. Each run is one scrape, which on Firecrawl is one credit.&lt;/p&gt;

&lt;h2&gt;
  
  
  Day One Has No Then
&lt;/h2&gt;

&lt;p&gt;Here is the catch with every history-based tool: the first run is useless. It saves 400 rows and prints “Nothing to compare yet.” You have to wait a day for the second run before the thing does anything at all, and a week before it does anything interesting.&lt;/p&gt;

&lt;p&gt;I didn’t want to wait a week to write this post, so I cheated with the Wayback Machine. The Internet Archive snapshots the Arena leaderboard most days, and Firecrawl will scrape an archived copy just like the live page. That is what the &lt;code&gt;--url&lt;/code&gt; and &lt;code&gt;--at&lt;/code&gt; flags are for: point the script at an old snapshot and tell it what date to record.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 arena_watch.py &lt;span class="nt"&gt;--url&lt;/span&gt; &lt;span class="s2"&gt;"https://web.archive.org/web/20260902190730/https://arena.ai/leaderboard/text"&lt;/span&gt; &lt;span class="nt"&gt;--at&lt;/span&gt; &lt;span class="s2"&gt;"2026-09-02 19:07"&lt;/span&gt;
python3 arena_watch.py &lt;span class="nt"&gt;--url&lt;/span&gt; &lt;span class="s2"&gt;"https://web.archive.org/web/20260909040114/https://arena.ai/leaderboard/text"&lt;/span&gt; &lt;span class="nt"&gt;--at&lt;/span&gt; &lt;span class="s2"&gt;"2026-09-09 04:01"&lt;/span&gt;
python3 arena_watch.py &lt;span class="nt"&gt;--url&lt;/span&gt; &lt;span class="s2"&gt;"https://web.archive.org/web/20260917153328/https://arena.ai/leaderboard/text"&lt;/span&gt; &lt;span class="nt"&gt;--at&lt;/span&gt; &lt;span class="s2"&gt;"2026-09-17 15:33"&lt;/span&gt;
python3 arena_watch.py

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The archived pages parse the same as the live one, since the table structure is identical and only the link URLs change. Three weeks of history in four commands:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;2026-09-02 19:07: first run, saved 399 models. Nothing to compare yet.
2026-09-09 04:01 vs 2026-09-02 19:07
  NEW #3 claude-fable-5.1-max (Anthropic, 1504)
2026-09-17 15:33 vs 2026-09-09 04:01
  NEW #8 muse-spark-1.3-max (Meta, 1493)
2026-09-23 01:27 vs 2026-09-17 15:33
  NEW #1 claude-fable-5-high (Anthropic, 1506)

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Fable 5.1 Max debuting at #3 and Meta’s Muse Spark 1.3 Max at #8 are exactly the kind of thing I wanted flagged. Midway through doing this the Internet Archive went down with a “Temporarily Offline” page, which is a nice reminder that the backfill trick is a trick and not infrastructure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two Things It Got Wrong
&lt;/h2&gt;

&lt;h3&gt;
  
  
  That “NEW #1” Is a Rename
&lt;/h3&gt;

&lt;p&gt;Look at the last line again. &lt;code&gt;claude-fable-5-high&lt;/code&gt; at #1 as a brand new model. It is not. The database says so:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="n"&gt;sqlite3&lt;/span&gt; &lt;span class="n"&gt;arena&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;db&lt;/span&gt; &lt;span class="nv"&gt;"SELECT model, taken, rank, score FROM ranks WHERE model LIKE 'claude-fable-5%' ORDER BY model, taken;"&lt;/span&gt;


&lt;span class="n"&gt;claude&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;fable&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="mi"&gt;2026&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;09&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;02&lt;/span&gt; &lt;span class="mi"&gt;19&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;07&lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="mi"&gt;1508&lt;/span&gt;
&lt;span class="n"&gt;claude&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;fable&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="mi"&gt;2026&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;09&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;09&lt;/span&gt; &lt;span class="mi"&gt;04&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;01&lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="mi"&gt;1507&lt;/span&gt;
&lt;span class="n"&gt;claude&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;fable&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="mi"&gt;2026&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;09&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;17&lt;/span&gt; &lt;span class="mi"&gt;15&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;33&lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="mi"&gt;1506&lt;/span&gt;
&lt;span class="n"&gt;claude&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;fable&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;high&lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="mi"&gt;2026&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;09&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;23&lt;/span&gt; &lt;span class="mi"&gt;01&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mi"&gt;27&lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="mi"&gt;1506&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Same rank, same score, new name. This is the most common way change detection lies to you, and it is not a Firecrawl problem or a SQLite problem. It is an identity problem: the thing you are tracking needs a stable key, and the page does not give you one.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh979zapmqzn4qdu9h41j.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh979zapmqzn4qdu9h41j.jpg" alt="That" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The cheap fix is a sanity check before calling anything NEW: if an unseen name sits at the exact rank and score of a name that just disappeared, call it a rename. I left it out of the script because 70 lines that lie once a month teach more than 90 lines that hide it. Your mileage may vary once it wakes you up at 8am about a model that is three months old.&lt;/p&gt;

&lt;h3&gt;
  
  
  It Missed the One That Started This
&lt;/h3&gt;

&lt;p&gt;GLM 5.3 Max, the model that made me want this in the first place, never shows up as NEW. It was already sitting at #18 on September 2, the oldest snapshot I loaded, so as far as the database knows it has always been there:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;2026-09-02 19:07|18|1482
2026-09-09 04:01|20|1482
2026-09-17 15:33|19|1483
2026-09-23 01:27|19|1483

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The harness can only see changes that happen after it starts watching. You can’t get the then after the fact. Start the cron before you need it.&lt;/p&gt;

&lt;p&gt;Also notice GLM wobbling from #18 to #20 to #19 while its score barely moves. Arena ranks a lot of models within a few points of each other near the top, and a model right on the #20 line will flicker in and out. Had it dropped to #21 for one run, the next run would have called it MOVED UP. If that gets noisy, the fix is to compare against the top 20 from any of the last few runs instead of just the last one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Not Just Use Firecrawl’s Change Tracking?
&lt;/h2&gt;

&lt;p&gt;Fair question, since Firecrawl has a &lt;code&gt;changeTracking&lt;/code&gt; format and a whole Monitor product built around scheduled checks. I didn’t use either here, and the reason is the leaderboard itself: the vote counts on every row change constantly, so “did this page change” is always yes. What I want to know is not whether the page changed but what changed in one column of one table, compared against a history I can query.&lt;/p&gt;

&lt;p&gt;That is not a knock on those features. For a page that should be static, a pricing page or a terms of service, “tell me when this changes” is exactly the right tool. They are worth their own post. This is the case where you want your own history, because the question you ask of it next month is not one you know yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  What This Is For
&lt;/h2&gt;

&lt;p&gt;As a standalone tool this is a small convenience. I will run it next to the Model Buzz Report and it will save me the “wait, was that there last week” squint.&lt;/p&gt;

&lt;p&gt;The reason it is the first real recipe in this series is the shape. Scrape, save with a timestamp, compare with the last run, speak only when something changed. Every recipe coming after this one is that shape pointed at a different page with a different question: price history instead of rank history, a job board instead of a leaderboard. Firecrawl handles getting the page. These 70 lines handle remembering it.&lt;/p&gt;

&lt;p&gt;And the database is already there, filling up once a day, waiting for whatever question I think of next.&lt;/p&gt;

</description>
      <category>arena</category>
      <category>firecrawl</category>
    </item>
    <item>
      <title>Multiple Choice Is Not a Decision: Teaching a Planning Agent to Ask More Questions</title>
      <dc:creator>Stephan Miller</dc:creator>
      <pubDate>Wed, 23 Sep 2026 12:00:00 +0000</pubDate>
      <link>https://dev.to/eristoddle/multiple-choice-is-not-a-decision-teaching-a-planning-agent-to-ask-more-questions-4bd2</link>
      <guid>https://dev.to/eristoddle/multiple-choice-is-not-a-decision-teaching-a-planning-agent-to-ask-more-questions-4bd2</guid>
      <description>&lt;p&gt;I opened &lt;code&gt;PLAN.md&lt;/code&gt; on a new project last month. That is the planning doc every project in my setup carries, and its decisions section is supposed to be the record of what the project has decided. Mine was still the blank template: two bullet points under the heading, and one of them was about where to put the output folder.&lt;/p&gt;

&lt;p&gt;This was not a project where nothing had been decided. I had been going back and forth on it for a week. The decisions existed. They were in a chat log somewhere, or in my head, or in the shape of the code I had already written. They were not in the document whose entire job was to hold them.&lt;/p&gt;

&lt;p&gt;I went and looked at the &lt;a href="https://github.com/eristoddle/agent-skills" rel="noopener noreferrer"&gt;skill that owns these docs&lt;/a&gt; and found the real problem. It has four workflows, and every one of them reshapes material that already exists: create the files, retrofit them into a repo that has none, compact a doc that got fat, clean finished work out of the task file.&lt;/p&gt;

&lt;p&gt;Not one of them generated any of the files.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Multiple Choice Is Not a Decision&lt;/li&gt;
&lt;li&gt;Then Matt Pocock Wrote It Down in 28 Lines&lt;/li&gt;
&lt;li&gt;
It Ends in a Conversation. My Docs Have to Outlive the Conversation.

&lt;ul&gt;
&lt;li&gt;1. The Round Ends in a Disposition, Not an Answer&lt;/li&gt;
&lt;li&gt;2. It Is Re-Entrant&lt;/li&gt;
&lt;li&gt;3. The Files Append, They Never Get Rewritten&lt;/li&gt;
&lt;li&gt;4. The Relentlessness Had to Go&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;It Is a Utility, Not a Stage&lt;/li&gt;
&lt;li&gt;What I Would Tell You&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Multiple Choice Is Not a Decision
&lt;/h2&gt;

&lt;p&gt;Here is the part I had been doing wrong on my own, well before any of this &lt;a href="https://dev.to/eristoddle/the-agent-skills-guide-i-wish-id-had-17i1"&gt;got written into a skill&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;When you plan something with an agent, the conversation drifts into a shape almost immediately. You describe what you want. It comes back with options. A, B, or C, with a short paragraph on each, and sometimes a helpful little table. You pick the one that is least wrong, it says great choice, and you move on to the next menu.&lt;/p&gt;

&lt;p&gt;Nobody argued with you, nothing got stress tested, and the option you picked was picked because it was the best of three things a model generated in two seconds.&lt;/p&gt;

&lt;p&gt;I got tired of it. What I actually wanted out of planning was an argument, where I have to say &lt;em&gt;why&lt;/em&gt; and get pushed on the answer.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmxuccpgwxcd1zi1lv7ij.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmxuccpgwxcd1zi1lv7ij.jpg" alt="Multiple Choice Is Not a Decision" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So I had been doing a crude version of this by hand. Telling the agent to stop offering me options and start asking me questions. It sort of worked, the way anything sort of works when you are improvising it fresh every session with no structure behind it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then Matt Pocock Wrote It Down in 28 Lines
&lt;/h2&gt;

&lt;p&gt;The idea I built on is not mine. It is &lt;a href="https://github.com/mattpocock/skills/blob/main/skills/productivity/grilling/SKILL.md" rel="noopener noreferrer"&gt;Matt Pocock’s &lt;code&gt;grilling&lt;/code&gt; skill&lt;/a&gt; and I want that up front rather than buried in a credits line at the bottom, because the core of it is his and it is the good part.&lt;/p&gt;

&lt;p&gt;It is 28 lines. 319 words. It opens with this:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Interview the user relentlessly until you reach a shared understanding.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Then it gives you the structure that makes that possible, which is what I had been missing. Model the conversation as a &lt;strong&gt;design tree&lt;/strong&gt; , where every decision branches into the decisions hanging off it. Work the tree in &lt;strong&gt;rounds&lt;/strong&gt;. The &lt;strong&gt;frontier&lt;/strong&gt; is every decision whose prerequisites are already settled, which is to say the questions you can ask right now without guessing at answers you have not heard yet.&lt;/p&gt;

&lt;p&gt;Two rules in there are doing most of the work.&lt;/p&gt;

&lt;p&gt;The first is ordering. A question whose answer depends on another question still open in this round belongs to a later round. That is the whole technique, honestly. Violate it and you get answers the user has to retract two rounds later, which is exactly what my hand-rolled version kept doing.&lt;/p&gt;

&lt;p&gt;The second is that every question ships with the agent’s recommended answer, so the cheap reply is “yes to all but Q3.”&lt;/p&gt;

&lt;p&gt;And finding facts is the agent’s job. If a question needs to know what is in the filesystem or what some dependency actually does, it goes and looks. The decisions stay yours.&lt;/p&gt;

&lt;p&gt;The session is done when the frontier is empty: every branch visited, nothing left assumed.&lt;/p&gt;

&lt;p&gt;I read it and wanted to use it, and immediately hit the thing that made it not fit.&lt;/p&gt;

&lt;h2&gt;
  
  
  It Ends in a Conversation. My Docs Have to Outlive the Conversation.
&lt;/h2&gt;

&lt;p&gt;Grilling ends at a shared understanding held between two parties in a chat window. That is a perfectly good place for it to end if you are making one decision today.&lt;/p&gt;

&lt;p&gt;I am not. These projects run for months. The entire premise of &lt;a href="https://dev.to/eristoddle/my-third-try-how-a-living-plan-beat-both-vibe-coding-and-spec-kit-5a89"&gt;this whole setup&lt;/a&gt; is that the documents are the memory, because &lt;a href="https://dev.to/eristoddle/i-got-tired-of-ai-memory-hype-so-i-built-a-context-lake-55fi"&gt;the agent’s context window&lt;/a&gt; is not. An interview that produces nothing but a shared understanding produces nothing at all when the session ends.&lt;/p&gt;

&lt;p&gt;So four things changed on the way in.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftw4vmtcykhbf599v05my.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftw4vmtcykhbf599v05my.jpg" alt="1. The Round Ends in a Disposition, Not an Answer" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  1. The Round Ends in a Disposition, Not an Answer
&lt;/h3&gt;

&lt;p&gt;The frontier does not empty because everything got answered. It empties because everything got &lt;strong&gt;filed&lt;/strong&gt;. Every node in the tree closes as exactly one of four things:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Disposition&lt;/th&gt;
&lt;th&gt;Lands in&lt;/th&gt;
&lt;th&gt;Meaning&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Decided&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;Decisions&lt;/code&gt;, as a &lt;code&gt;Dxx&lt;/code&gt; record&lt;/td&gt;
&lt;td&gt;settled, load-bearing, act on it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Open question&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;docs/questions/Qxx.md&lt;/code&gt; plus a one-line pointer&lt;/td&gt;
&lt;td&gt;matters, not answerable yet&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Parked&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;docs/parking-lot/Pxx.md&lt;/code&gt; plus a one-line pointer&lt;/td&gt;
&lt;td&gt;might matter later, not now&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;N/A&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;nothing&lt;/td&gt;
&lt;td&gt;branch does not apply, say so once&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The line in the workflow that explains why is the one I would keep if I had to throw the rest away:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A blind spot is an &lt;em&gt;unasked&lt;/em&gt; question, not an unanswered one.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;“I do not know yet” is an acceptable outcome. It becomes a file. What is not acceptable is a branch nobody ever walked down.&lt;/p&gt;

&lt;p&gt;It writes at &lt;strong&gt;every round boundary&lt;/strong&gt; for the obvious reason that I stop things halfway constantly and a workflow that only saves on completion would lose everything every time I do.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. It Is Re-Entrant
&lt;/h3&gt;

&lt;p&gt;The upstream is a one-shot interview. Mine has two modes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Greenfield&lt;/strong&gt; is the one-shot case: there is no planning doc yet, so the tree starts at the root; with nowhere to file, nothing gets written during the session, and at the end it hands the whole disposition to the scaffold workflow, which emits a &lt;code&gt;PLAN.md&lt;/code&gt; that is already populated. If you quit early it scaffolds with whatever you settled anyway. Three decisions and six open questions beats an empty template.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Continuing&lt;/strong&gt; is the one I actually use. It reads the existing planning doc and the &lt;code&gt;docs/&lt;/code&gt; tree and seeds the design tree from them before asking anything:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;existing decisions become settled nodes, pruned rather than re-litigated, and each one pushes the frontier outward because the questions hanging off it are exactly what just became askable&lt;/li&gt;
&lt;li&gt;open questions come back as frontier nodes carrying whatever partial answers previous rounds accumulated&lt;/li&gt;
&lt;li&gt;parked items come back &lt;strong&gt;only if something settled since they were parked makes them answerable now&lt;/strong&gt; , because dragging every parked idea back every session is how a parking lot turns into noise you learn to skip&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6d7w0c8mzg3bgb9q3blx.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6d7w0c8mzg3bgb9q3blx.jpg" alt="2. It Is Re-Entrant" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This turns the docs into a breadcrumb trail instead of a stack of unrelated one-shot interviews.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. The Files Append, They Never Get Rewritten
&lt;/h3&gt;

&lt;p&gt;A &lt;code&gt;Qxx.md&lt;/code&gt; touched by a later session gets a new dated section added under the existing ones.&lt;/p&gt;

&lt;p&gt;I had to put that in as an explicit rule because the instinct to tidy is strong. That growing, slightly repetitive, occasionally contradictory file is the record of what you thought in June that made the July answer obvious. Rewrite it into a clean summary and you’ve thrown away the reasoning.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. The Relentlessness Had to Go
&lt;/h3&gt;

&lt;p&gt;This is the one I keep coming back to.&lt;/p&gt;

&lt;p&gt;“Relentlessly” is right there in the first line upstream, and for a one-shot session it is correct. A grill that gives up when you get tired is not doing its job.&lt;/p&gt;

&lt;p&gt;But relentless plus re-entrant is just nagging. If the thing can come back next week, it does not need to squeeze everything out of you today. So the stop rule is written as an invariant rather than a preference:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Stopping is never negotiated.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;No “are you sure”, no “we have not covered X yet”, no one more round, and specifically no listing what I am about to miss, because listing what I am about to miss is just arguing with extra steps. If I name a branch, it parks that branch and keeps going elsewhere.&lt;/p&gt;

&lt;p&gt;The reason that is safe is the rule from change one. Everything unvisited gets filed on the way out.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Nothing is lost by stopping, which is precisely why stopping needs no defense.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  It Is a Utility, Not a Stage
&lt;/h2&gt;

&lt;p&gt;Every other route in this skill is triggered by what the repo looks like. No planning doc means grill first, then scaffold. Planning doc but no task file means adopt. Fat planning doc means rebalance. Task doc full of finished work means evict.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fql0sx456z1oj3n5wkg2b.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fql0sx456z1oj3n5wkg2b.jpg" alt="It Is a Utility, Not a Stage" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The grill has a row in that table that is not a filesystem state at all:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;State&lt;/th&gt;
&lt;th&gt;Detected by&lt;/th&gt;
&lt;th&gt;Route&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Greenfield&lt;/td&gt;
&lt;td&gt;no planning doc present&lt;/td&gt;
&lt;td&gt;grill, then scaffold&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Partial&lt;/td&gt;
&lt;td&gt;planning doc exists, pieces missing&lt;/td&gt;
&lt;td&gt;adopt&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mature&lt;/td&gt;
&lt;td&gt;full system present, planning doc heavy&lt;/td&gt;
&lt;td&gt;rebalance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Task doc bloated&lt;/td&gt;
&lt;td&gt;mostly finished work&lt;/td&gt;
&lt;td&gt;evict-tasks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Deciding, not filing&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;not a filesystem state&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;grill&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;It fires on what I am doing, not on what the directory contains. Two entries: I ask for it, or I am clearly thinking out loud with no question attached and it offers in one line and waits. That second one has a rule attached. One line, then shut up.&lt;/p&gt;

&lt;p&gt;That is why it sits at different points in the process rather than at the front of it. It is not step one of planning. It is the thing you reach for at any point where you are about to settle on an answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Would Tell You
&lt;/h2&gt;

&lt;p&gt;If you want the idea, &lt;a href="https://github.com/mattpocock/skills/blob/main/skills/productivity/grilling/SKILL.md" rel="noopener noreferrer"&gt;go read Pocock’s 28 lines&lt;/a&gt; rather than my version. It is the good part, it is short, and you will be running it in five minutes.&lt;/p&gt;

&lt;p&gt;The next post is the last one in this series, and it will be: the whole thing, start to finish, how it actually works, so nobody has to read a skill directory to figure it out.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>agentskills</category>
    </item>
    <item>
      <title>Grok 4.7 Shipped. Even Elon Musk Said It Was Mid.</title>
      <dc:creator>Stephan Miller</dc:creator>
      <pubDate>Tue, 22 Sep 2026 13:00:00 +0000</pubDate>
      <link>https://dev.to/eristoddle/grok-47-shipped-even-elon-musk-said-it-was-mid-3f6g</link>
      <guid>https://dev.to/eristoddle/grok-47-shipped-even-elon-musk-said-it-was-mid-3f6g</guid>
      <description>&lt;p&gt;Last week I ended this roundup with a promise. I said I’d be back to find out whether Grok 4.7 actually exists yet, and whether its creator liked it any better once it did. Well. It exists. He does not seem to like it any better, and neither, so far, does anyone else.&lt;/p&gt;

&lt;p&gt;Grok 4.7 shipped Sunday, September 21. That’s the whole headline and the whole punchline at once. For two straight weeks Musk stood in public and marked his own unreleased model down, from “better than 4.6 in every way” to “beats everything on the board” to, finally, “roughly on par with Opus 5.0, not 5.1.” Then the model landed. And the independent benchmarks put it right about where he’d talked it down to. This almost never happens. Usually the shipped thing is worse than the hype. This time the hype had already deflated itself to match, and the model still landed under the deflated number. Read on, because the gap between what got announced and what got shipped is the whole story this week, and for once the shipped thing is the one that gets a fair shake.&lt;/p&gt;

&lt;p&gt;The quiet story is still down in the bargain bin, where the fifteen-cent model I keep telling you about held every one of its five categories and nothing changed except that it got a little faster. Sometimes the boring outcome is the important one. Let me walk you through both.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Grok 4.7 shipped, and it’s the model its own maker warned you about&lt;/li&gt;
&lt;li&gt;The benchmark you cite is the argument you’re making&lt;/li&gt;
&lt;li&gt;The cheap model kept all five of its categories&lt;/li&gt;
&lt;li&gt;The cheapskate picks&lt;/li&gt;
&lt;li&gt;Newest still isn’t best, and it’s still funny&lt;/li&gt;
&lt;li&gt;About that bill&lt;/li&gt;
&lt;li&gt;What’s coming&lt;/li&gt;
&lt;li&gt;The honest version&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Grok 4.7 shipped, and it’s the model its own maker warned you about
&lt;/h2&gt;

&lt;p&gt;Here are the specs, and this time they’re specs, not tweets. Grok 4.7 is live as of September 21 through the Grok API, Cursor, and Grok Build. It runs on a new base model at a claimed 2.1 trillion parameters, up about 40 percent from Grok 4.6’s 1.5 trillion. It takes a 500K-token context, multimodal input, and it’s priced at two dollars in and six out, with cached input at fifty cents. That’s the same sticker as Grok 4.6, which is the one genuinely good decision in this whole launch. xAI did not try to charge frontier money for a mid-frontier model. Credit where it’s due.&lt;/p&gt;

&lt;p&gt;Now the number that matters. On the Artificial Analysis intelligence index, the independent one that blends ten hard evals, Grok 4.7 scores a 46. The two models at the top of that board, Claude Fable 5.1 and GPT-6 Astra, both sit at 53. So the model xAI built a two-week hype cycle around lands seven full points behind the frontier, in the same neighborhood as models that are a good deal cheaper and a good deal older. Mid-pack. Exactly where Musk himself put it on the fourteenth, before it had a benchmark to its name.&lt;/p&gt;

&lt;p&gt;The thing I keep coming back to is that this is a functional pattern now, not a Grok quirk. xAI ships a model that is cheap and legitimately fine, then wraps it in language it cannot support. Grok 4.7 is a perfectly reasonable two-dollar coding model. It is not the smartest thing on Earth, nobody who has run it thinks it is, and the person who built it told you so a week early. If you priced it at what it is instead of announcing it as what it isn’t, this would be a good news week for xAI. Instead it’s a case study.&lt;/p&gt;

&lt;h2&gt;
  
  
  The benchmark you cite is the argument you’re making
&lt;/h2&gt;

&lt;p&gt;Now the part worth reading past the headline for.&lt;/p&gt;

&lt;p&gt;xAI’s own launch materials lean hard on coding, and on coding the gains are real. On DeepSWE v1.1 at high effort, Grok 4.7 posts 71.0, up from Grok 4.6’s 65.2. On CursorBench 4.0, a test built around longer-running coding tasks, it hits 46.3 against 40.4 for its predecessor. Those are honest improvements. If you’re doing the kind of work those benchmarks measure, 4.7 is a real step up from 4.6 at the same price, and that’s a fine reason to switch.&lt;/p&gt;

&lt;p&gt;Then the-decoder ran the harder agentic test, Terminal-Bench 4.0, and the floor gave out. Grok 4.7 scored 26 percent. GPT-6 Astra hit 60 on the same test. Fable 5.1 hit 55. And the cheap DeepSeek V4.1 Flash edged Grok out at 27. So on the benchmark xAI put in the deck, Grok 4.7 looks like a solid upgrade, and on the benchmark it left out, it finishes behind a Chinese model that costs less. Both results are true. They’re measuring different things, and the aggregate index score of 46 is what you get when you stop cherry-picking and average it out.&lt;/p&gt;

&lt;p&gt;The lesson underneath keeps earning its keep. The benchmark somebody cites is the argument they’re making. When a launch deck shows you three benchmarks, the interesting question is always which ones it didn’t show you. xAI showed you DeepSWE and CursorBench. It did not show you Terminal-Bench. Now you know why.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cheap model kept all five of its categories
&lt;/h2&gt;

&lt;p&gt;Now the story that actually moves your bill, which as usual is the least dramatic one on the page.&lt;/p&gt;

&lt;p&gt;Two weeks ago GLM-5.3-Flash from Z.ai took the cheapest-good-model crown off Xiaomi’s MiMo v2.5 Pro. Last week it grabbed a fifth Arena category. This week it did the boring, valuable thing and simply held everything. It’s still the cheapest model inside the competitive band for five of the six categories: Overall, Coding, Instruction Following, Hard Prompts, and Math. Creative Writing is still the lone holdout, still a Gemini story, because GLM never cracked that band and probably won’t.&lt;/p&gt;

&lt;p&gt;The usage board agrees with the preference board, which is the part that makes this real rather than an Arena curiosity. On OpenRouter, which counts actual tokens on actual paid calls, GLM-5.3-Flash is sitting at number two on the entire platform, somewhere north of ten trillion tokens a week, behind DeepSeek V4 Flash and ahead of GPT-5.6 Luna and MiMo. Chinese-built models are running around 46 percent of all tokens on the platform now, with DeepSeek the single largest vendor at roughly 16 percent. People aren’t voting for these models. They’re running them, in production, with their own money.&lt;/p&gt;

&lt;p&gt;One thing genuinely improved this week, and it’s the caveat that used to matter most. The old “cheap but slow” knock on GLM keeps softening. Artificial Analysis now clocks it at about 89 output tokens a second on Z.ai’s own endpoint, comfortably above the roughly 75 median for open-weight models in its class. It was crawling along near 60 a few weeks ago. For a fifty-cent model, that’s the difference between something you tolerate in a chat window and something you can actually drop into an agent loop.&lt;/p&gt;

&lt;p&gt;Two catches, same as always. First, GLM-5.3-Flash scores a 42 on the hard-reasoning index where the frontier lives in the fifties. It’s cheap and preference-strong, not a deep-reasoning machine. For everyday work that’s a rounding error you’ll never feel. For genuinely hard problems, feel it. Second, the price. The seven-cents-in, quarter-out number you may still have in your head was a launch promo, and it died on September 9. List is fifteen cents in and fifty cents out. Arena’s price column is still cheerfully showing the dead promo, which is going to burn somebody who budgets off it. Build on fifty cents out. If you sized an August spend on the old number, it doubled on you three weeks ago and nobody sent a memo.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cheapskate picks
&lt;/h2&gt;

&lt;p&gt;Same method every week. For each Arena category I take the leader’s rating, draw a band 50 points below it, and find the cheapest model still sitting inside that band. The whole premise is that Arena ratings cluster tight at the top, so the category leader is usually a rounding error better than something 20 to 100 times cheaper. I compute the band from the full table in code, not by eyeballing the first screen, because eyeballing it is exactly how you delete the entire cheap tail and accidentally crown the cheapest expensive model. I learned that one the hard way, in public, a couple months back.&lt;/p&gt;

&lt;p&gt;Bands ran deep again this week: around 60 models inside the Overall band, 64 in Coding, thinner in Creative and Math. GLM-5.3-Flash is quoted at list, fifteen and fifty, not the expired promo Arena is still displaying. Arena data is dated September 13.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Leader&lt;/th&gt;
&lt;th&gt;$ leader out&lt;/th&gt;
&lt;th&gt;Cheapskate pick&lt;/th&gt;
&lt;th&gt;$ pick out&lt;/th&gt;
&lt;th&gt;Δ rating&lt;/th&gt;
&lt;th&gt;Cheaper by&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Overall&lt;/td&gt;
&lt;td&gt;claude-fable-5 (1506)&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;GLM-5.3-Flash (1475, #29)&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;−31&lt;/td&gt;
&lt;td&gt;~100×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Coding&lt;/td&gt;
&lt;td&gt;claude-fable-5 (1552)&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;GLM-5.3-Flash (1525, #20)&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;−27&lt;/td&gt;
&lt;td&gt;~100×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Creative Writing&lt;/td&gt;
&lt;td&gt;claude-fable-5 (1504)&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;gemini-3-flash (1459, #26)&lt;/td&gt;
&lt;td&gt;$3&lt;/td&gt;
&lt;td&gt;−45&lt;/td&gt;
&lt;td&gt;~16.7×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Instruction Following&lt;/td&gt;
&lt;td&gt;claude-opus-4-6-high (1513)&lt;/td&gt;
&lt;td&gt;$25&lt;/td&gt;
&lt;td&gt;GLM-5.3-Flash (1471, #29)&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;−42&lt;/td&gt;
&lt;td&gt;~50×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hard Prompts&lt;/td&gt;
&lt;td&gt;claude-opus-4-6-high (1533)&lt;/td&gt;
&lt;td&gt;$25&lt;/td&gt;
&lt;td&gt;GLM-5.3-Flash (1498, #28)&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;−35&lt;/td&gt;
&lt;td&gt;~50×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Math&lt;/td&gt;
&lt;td&gt;claude-fable-5 (1526, prelim)&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;GLM-5.3-Flash (1513, #7)&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;−13&lt;/td&gt;
&lt;td&gt;~100×&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A few notes on how to read that. Overall and Hard Prompts are the rows you can lean on: GLM-5.3-Flash is sitting on real vote counts there, ten thousand in Overall, and it’s cleanly the cheapest thing in a very deep band. Coding is a great rating on thin votes, twentieth place for fifty cents, but under three thousand votes behind it, so if you want certainty over the last few dollars, MiMo v2.5 Pro at rank 25 with sixteen thousand votes for eighty-seven cents is the steadier bet. Math is the one to squint at hardest: GLM shows seventh for fifty cents, which is loud, but it’s riding 441 preliminary votes on a board where the whole top is thin. Treat that row as a strong suggestion, not a promise. MiMo at the band edge or a Gemini Flash won’t embarrass you there either.&lt;/p&gt;

&lt;p&gt;Creative Writing stays Gemini because GLM never made the band, and the only value pick is gemini-3-flash at three dollars, which is still 17 times cheaper than a fifty-dollar leader. Nobody is paying fifty dollars a million tokens to draft blog intros.&lt;/p&gt;

&lt;h2&gt;
  
  
  Newest still isn’t best, and it’s still funny
&lt;/h2&gt;

&lt;p&gt;I wrote almost this exact paragraph last week and I’m writing it again, because the pattern refuses to break and it keeps getting funnier.&lt;/p&gt;

&lt;p&gt;The number one model on Arena Overall is claude-fable-5. Not Fable 5.1, the newer one Anthropic shipped a few weeks back. The old one. And sitting near the top of Instruction Following and Hard Prompts, as the outright leader on both, is claude-opus-4-6, a model that is roughly a year old. It beats Opus 5. It beats Opus 4.7 and 4.8. In the blind test, where nobody can see the version number, the crowd keeps reaching for the model everyone in the timeline already moved on from.&lt;/p&gt;

&lt;p&gt;I’m not saying the new models are bad. I’m saying the booth keeps preferring the boring old one, and it’s happened enough weeks running that it’s a pattern, not a fluke. If you upgraded off Opus 4.6 because a bigger number came out, the people voting blind would like a word with you.&lt;/p&gt;

&lt;h2&gt;
  
  
  About that bill
&lt;/h2&gt;

&lt;p&gt;The horror story this week isn’t a model, it’s a loop, and it’s the same species of loop that keeps eating people alive in 2026.&lt;/p&gt;

&lt;p&gt;Google’s Mandiant team put out an enterprise-AI-risk report on September 16 with a clean, awful example in it. An accounting agent hit a runaway execution loop and fired off more than 15,000 high-cost API calls in under an hour. Roughly fifty thousand dollars, gone, before a human looked at a dashboard. No budget ceiling. No alert anyone acted on. Just a bot doing the same expensive thing over and over, faster than anyone was watching.&lt;/p&gt;

&lt;p&gt;Here’s the part that made me laugh and then wince. Grok 4.7’s launch copy sells it as a model “designed to better verify its own output.” That is a lovely sentence. It is also describing the exact capability every runaway agent lacks, the ability to notice it’s stuck in a loop and stop. A better self-verifier would genuinely help with this. A marketing line about one does nothing, and the fifty-thousand-dollar hour happened the same week the line got written. The token bill in 2026 goes wrong in three normal ways: the sticker lies about the task, the loop has no brakes, and the cheap price had an expiration date you didn’t read. This week served up all three.&lt;/p&gt;

&lt;h2&gt;
  
  
  What’s coming
&lt;/h2&gt;

&lt;p&gt;Three to watch.&lt;/p&gt;

&lt;p&gt;Grok 4.7, now that it exists, gets to spend the next couple weeks accumulating actual Arena votes and EU availability, which lagged the US launch as xAI launches always do. I’ll be curious whether the blind test is kinder to it than the benchmarks were, or crueler.&lt;/p&gt;

&lt;p&gt;GPT-6 Astra is still finishing its rollout. It went from a handful of day-one orgs to the ChatGPT paid tiers, the OpenAI API, Azure, and AWS Bedrock over the past couple weeks. The Daybreak program that loosens the safety rails for vetted organizations is the piece worth watching, given this is the model that reportedly maxed out an exploit-writing benchmark a couple weeks ago.&lt;/p&gt;

&lt;p&gt;Gemini 4, pretraining done and everything else a rumor. Google keeps dripping out Flash models to stay in the conversation while the real thing bakes. Note the small print on the current one: Gemini 3.8 Flash’s 75-cents-and-3.75 pricing is an intro rate that doubles on January 1. Late 2026 is the vague window for the actual next-gen model, if you believe the tea leaves, and I’ve stopped believing Google’s dates on principle.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest version
&lt;/h2&gt;

&lt;p&gt;The clean narrative this week would be that xAI face-planted. That’s not quite it, and the truth is more useful.&lt;/p&gt;

&lt;p&gt;Grok 4.7 is a decent, cheap, honestly-priced coding model that got buried under a launch it couldn’t live up to, by a man who then spent two weeks digging the hole himself. The model is fine. The framing was the problem, and the framing is always the problem. Cost-per-token isn’t cost-per-task. The benchmark in the deck isn’t the benchmark that matters. And “most capable model yet” means whatever the person saying it needs it to mean this quarter.&lt;/p&gt;

&lt;p&gt;Underneath all of it, a fifteen-cent open-weight model from Z.ai held its five categories, got faster, and stayed the second most-used model on the planet without anyone holding a keynote about it. That’s the line that changes your life if you’re shipping something and paying the bill yourself. The frontier had a loud week arguing about who’s smartest. The floor just sat there being cheap and getting quietly better. You already know which one you’ll actually be running next month.&lt;/p&gt;

&lt;p&gt;I’ll be back next week to see whether the Arena crowd is any nicer to Grok 4.7 than its own creator was. Low bar. We’ll find out.&lt;/p&gt;

</description>
      <category>llm</category>
      <category>openrouter</category>
    </item>
    <item>
      <title>Syncing Obsidian with Google Drive Is the Trickiest One. Here Is How Anyway.</title>
      <dc:creator>Stephan Miller</dc:creator>
      <pubDate>Mon, 21 Sep 2026 13:00:00 +0000</pubDate>
      <link>https://dev.to/eristoddle/syncing-obsidian-with-google-drive-is-the-trickiest-one-here-is-how-anyway-19o4</link>
      <guid>https://dev.to/eristoddle/syncing-obsidian-with-google-drive-is-the-trickiest-one-here-is-how-anyway-19o4</guid>
      <description>&lt;p&gt;Every method in my &lt;a href="https://dev.to/eristoddle/how-to-sync-obsidian-across-all-your-devices-including-free-methods-1mi5"&gt;guide to syncing an Obsidian vault across devices&lt;/a&gt; got a section except this one. For about a year the Google Drive section basically said “don’t.”&lt;/p&gt;

&lt;p&gt;Google Drive is the cloud storage that the largest number of people already have, already pay for, and already have 15GB sitting idle in. It is also, of the four big consumer clouds, the one Obsidian gets along with worst.&lt;/p&gt;

&lt;p&gt;The answer is better now than it was. But only by degree. There are four real paths, all of them have a catch, and two of them are plugins with the same name written by different people. That last one alone has cost the Obsidian forums more confusion than any other sync question I have looked at. I am going to walk all four, name the catch on each, and then tell you when to stop trying.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Why Google Drive Is the Hard One&lt;/li&gt;
&lt;li&gt;The Desktop Path: Mirror, Never Stream&lt;/li&gt;
&lt;li&gt;
The Two Plugins That Have the Same Name

&lt;ul&gt;
&lt;li&gt;The one in the plugin directory: richardx366&lt;/li&gt;
&lt;li&gt;The one everybody links to: stravo1&lt;/li&gt;
&lt;li&gt;Which one&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Remotely Save PRO, the Boring Paid Route&lt;/li&gt;
&lt;li&gt;Android, Which Everybody Assumes Is Easy&lt;/li&gt;
&lt;li&gt;The 15GB Argument&lt;/li&gt;
&lt;li&gt;Should You Just Use Something Else?&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why Google Drive Is the Hard One
&lt;/h2&gt;

&lt;p&gt;Dropbox syncs a folder. iCloud syncs a folder, badly, but it syncs a folder. Google Drive for Desktop decided a folder was ambitious.&lt;/p&gt;

&lt;p&gt;The current desktop client defaults to &lt;strong&gt;streaming&lt;/strong&gt; , which mounts your Drive as a virtual filesystem. Files show up in the file browser, but the bytes are not on your disk until something asks for them. For photos and spreadsheets that works. For a vault it is a disaster.&lt;/p&gt;

&lt;p&gt;Obsidian watches the vault directory for changes. That watcher expects a real filesystem underneath it, where a file that exists is a file you can read right now. A virtual drive breaks that assumption. You get notes that open blank and populate a second later. You get link resolution that misses files it should find. You get the plugin folder behaving strangely because a plugin tried to read its own settings before Drive had gotten around to materializing them.&lt;/p&gt;

&lt;p&gt;None of this fails loudly. Instead, it fails as weirdness, and you spend a week thinking Obsidian is buggy.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Desktop Path: Mirror, Never Stream
&lt;/h2&gt;

&lt;p&gt;If you are going to do this on desktop, there is exactly one correct setting and it is not the default.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffzpmfenwtlhr807fvho0.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffzpmfenwtlhr807fvho0.jpg" alt="The Desktop Path: Mirror, Never Stream" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;In Drive for Desktop, open Settings, then Preferences, then &lt;strong&gt;Folders from Drive&lt;/strong&gt; in the left sidebar. Under “My Drive syncing options,” switch from &lt;strong&gt;Stream files&lt;/strong&gt; to &lt;strong&gt;Mirror files&lt;/strong&gt;. Mirroring keeps a full local copy on disk and syncs changes up. That is the behavior a vault needs, because now Obsidian is watching a real directory full of real files.&lt;/p&gt;

&lt;p&gt;Let any in-flight sync finish before you flip that switch. Google’s own docs warn about changing sync modes mid-sync, and this is not the folder you want to find out on.&lt;/p&gt;

&lt;p&gt;Then put the vault somewhere sane:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;~/Google Drive/My Drive/Obsidian/VaultName/

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One vault per folder. Do not nest a vault inside another vault, and do not point Obsidian at &lt;code&gt;My Drive&lt;/code&gt; itself, because you will end up with the watcher trying to index every PDF you have saved since 2014.&lt;/p&gt;

&lt;p&gt;Two things to know before you commit to this. Mirroring means the vault takes up its full size on every desktop you do this on, which for a text vault is nothing and for a vault full of PDFs is not nothing. And Drive’s conflict handling is to make a second file with a different name. It does not merge. It does not ask. You get &lt;code&gt;Note.md&lt;/code&gt; and &lt;code&gt;Note (1).md&lt;/code&gt; and it is on you to notice.&lt;/p&gt;

&lt;p&gt;That last part is true of Dropbox too, and I covered how to live with it in the &lt;a href="https://dev.to/eristoddle/obsidian-dropbox-sync-a-setup-guide-2ajl"&gt;Dropbox sync post&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Two Plugins That Have the Same Name
&lt;/h2&gt;

&lt;p&gt;This is the part that wastes people’s afternoons.&lt;/p&gt;

&lt;p&gt;There are &lt;strong&gt;two different Obsidian plugins both called “Google Drive Sync.”&lt;/strong&gt; Different authors, different designs, different tradeoffs. Every forum thread I have read about Drive sync has at least one person answering about the wrong one. Search the community plugin browser and you find exactly one of them. Search GitHub and the top result is the other.&lt;/p&gt;

&lt;h3&gt;
  
  
  The one in the plugin directory: richardx366
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://github.com/richardx366/Obsidian-Google-Drive" rel="noopener noreferrer"&gt;Obsidian-Google-Drive&lt;/a&gt; by richardx366 is the one you can install the normal way, because it is in the official community plugin list as “Google Drive Sync.” As of this writing it is on version 3.1.1 and was updated within the last week, which is more than I can say for the alternative.&lt;/p&gt;

&lt;p&gt;It exists specifically to solve the iOS problem, and the README says so: “A plugin to make Obsidian work in Google Drive to enable access to iOS.”&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzzwcmkfwc0h8lu1o9ze1.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzzwcmkfwc0h8lu1o9ze1.jpg" alt="The one in the plugin directory: richardx366" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The thing to understand before you install it is how the auth works. You do not create your own Google Cloud project. You authenticate through a hosted service at &lt;code&gt;ogd.richardxiong.com&lt;/code&gt; and paste the resulting refresh token into the plugin settings. That is a real convenience and it is also a third party sitting in the path between Obsidian and your Drive. You can point it at a self-hosted token endpoint instead if that bothers you, and if you are the kind of person it bothers, it should, and you should.&lt;/p&gt;

&lt;p&gt;Its README carries its own warnings. Back up first. Do not manually add files to the synced Drive folder, because it “will likely break functionality, potentially causing data loss.” Do not edit vault files outside Obsidian. Let syncs finish before you close the app.&lt;/p&gt;

&lt;h3&gt;
  
  
  The one everybody links to: stravo1
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://github.com/stravo1/obsidian-gdrive-sync" rel="noopener noreferrer"&gt;obsidian-gdrive-sync&lt;/a&gt; by stravo1 has about three times the stars, which is why it is the one that shows up in every search result and every Reddit answer. It works on desktop, Android, and iOS.&lt;/p&gt;

&lt;p&gt;It is not in the community plugin directory. The author’s own README explains why, and I would rather quote it than paraphrase it:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The plugin is under active development, new releases might introduce bugs, old releases maybe be incompatible with the new ones. This might lead to data loss.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is a developer telling you to back up first. The author has said they will not submit it to the official list until it is stable enough not to risk anybody’s notes.&lt;/p&gt;

&lt;p&gt;Two documented limits on top of that. It is &lt;strong&gt;not optimized for vaults over about 1,000 files&lt;/strong&gt; , with large-vault work listed as in progress. And on iOS there is no normal install path for a plugin outside the directory, so you set the vault up on a desktop that already has the plugin working and copy the whole thing across.&lt;/p&gt;

&lt;p&gt;Installing it means BRAT or doing it by hand: download the release zip, extract, drop the folder into &lt;code&gt;.obsidian/plugins/&lt;/code&gt;, enable it. The &lt;a href="https://dev.to/eristoddle/how-to-install-activate-and-update-obsidian-plugins-4d2p"&gt;plugin install walkthrough&lt;/a&gt; covers the mechanics if you have never sideloaded one.&lt;/p&gt;

&lt;h3&gt;
  
  
  Which one
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fycucs0xmorqmmpv1ywxm.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fycucs0xmorqmmpv1ywxm.jpg" alt="Which one" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you want a plugin and you want it maintained, richardx366’s is the current answer, and the price is the hosted token service. If that price is unacceptable and you will not self-host the endpoint, stravo1’s is the alternative, and the price is that you are running release-candidate software that has been quiet for four months ago.&lt;/p&gt;

&lt;h2&gt;
  
  
  Remotely Save PRO, the Boring Paid Route
&lt;/h2&gt;

&lt;p&gt;Remotely Save is the plugin most people in this cluster end up on, and its free tier covers S3-compatible storage, Dropbox, WebDAV, and basic OneDrive with no account at all. Google Drive is not on that list. Drive support sits behind &lt;strong&gt;PRO&lt;/strong&gt; , which needs an account at &lt;code&gt;remotelysave.com&lt;/code&gt; separate from anything Obsidian.&lt;/p&gt;

&lt;p&gt;PRO is free during beta, currently stated through January 1, 2027. The price after that has not been announced. I am flagging that plainly because “free right now” and “free” are different words, and signing up for an unnamed number later is a decision you should make on purpose.&lt;/p&gt;

&lt;p&gt;What you get for it is the thing the stravo1 plugin does not have yet: maturity. Remotely Save has been syncing vaults to object storage for years, it handles conflicts with more grace than “make a second file,” and it is in the official directory.&lt;/p&gt;

&lt;h2&gt;
  
  
  Android, Which Everybody Assumes Is Easy
&lt;/h2&gt;

&lt;p&gt;The assumption goes: Google Drive is a Google product, Android is a Google product, so this must be the one place it just works.&lt;/p&gt;

&lt;p&gt;It is not. Mirror mode is the thing that makes the desktop path work, and mirror mode is a feature of Drive for Desktop. Drive for Desktop is, as the name has been telling you this whole time, a desktop application. Nothing in the Android Drive app gives Obsidian a real local directory that a file watcher can sit on top of.&lt;/p&gt;

&lt;p&gt;That leaves you with in-app sync, which means either the stravo1 plugin or Remotely Save PRO. Both of them run inside Obsidian and talk to the API.&lt;/p&gt;

&lt;p&gt;There is a second Android trap underneath this one, and you meet it before you sync anything. Since Obsidian 1.8.10, Android asks where the vault should live: &lt;strong&gt;app storage&lt;/strong&gt; or &lt;strong&gt;device storage&lt;/strong&gt;. Obsidian recommends device storage, and the reason is that app storage isolates the vault from every other app on the phone. That blocks external tools like Syncthing, and it means uninstalling Obsidian deletes your local vault.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa56oatf32htj7iqrqvrv.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa56oatf32htj7iqrqvrv.jpg" alt="Android, Which Everybody Assumes Is Easy" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For Drive specifically that trap is survivable, because Obsidian Sync and community sync plugins both still work under app storage. Your Drive plugin keeps running either way. Device storage is still the safer pick, and why is explained in the &lt;a href="https://dev.to/eristoddle/how-to-sync-obsidian-on-android-for-free-41m"&gt;Android sync post&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 15GB Argument
&lt;/h2&gt;

&lt;p&gt;Here is the case for going through all of this.&lt;/p&gt;

&lt;p&gt;Google gives you 15GB free, shared across Drive, Gmail, and Photos. Dropbox gives you 2GB. A text vault will never come close to either number, but a vault with a few hundred PDFs and a year of screenshots will blow right past 2GB. If you are already on a Google One plan, the storage is free.&lt;/p&gt;

&lt;p&gt;That is a real argument, but not the only variable. The question is “which cloud will cost me an afternoon of untangling duplicate notes in March.”&lt;/p&gt;

&lt;h2&gt;
  
  
  Should You Just Use Something Else?
&lt;/h2&gt;

&lt;p&gt;For Google Drive specifically, more often than not: yes.&lt;/p&gt;

&lt;p&gt;I don’t say that about the other methods in this cluster. Dropbox is fine. &lt;a href="https://dev.to/eristoddle/syncthing-for-obsidian-free-private-and-genuinely-annoying-3dbm"&gt;Syncthing&lt;/a&gt; is free, private, and works. iCloud works if you are all-Apple. Each of those has a type of user it fits well.&lt;/p&gt;

&lt;p&gt;Google Drive’s problem is that every path asks you for something the alternatives do not. Mirror mode and a folder Drive duplicate on conflict. A maintained plugin that routes your auth through somebody else’s server. A more popular plugin that stopped shipping in May. Or a paid tier whose price nobody has announced yet. None of those is disqualifying on its own. It is that Drive is the only method in this cluster where you have to pick which one you mind least.&lt;/p&gt;

&lt;p&gt;So use Google Drive if the storage is a hard requirement. Company account, family plan, a 15GB library of attachments you are not moving.&lt;/p&gt;

&lt;p&gt;If it is a preference rather than a requirement, point Remotely Save at S3 or Dropbox and get on with your life. Or run Syncthing and stop paying anybody. Same outcome, fewer README warnings.&lt;/p&gt;

&lt;p&gt;And whichever you pick: back the vault up somewhere that is not the thing you are syncing with. Sync is not backup. A method that replicates your mistake to four devices in under a second is not protecting you, and Google Drive is very, very fast.&lt;/p&gt;

</description>
      <category>obsidian</category>
      <category>googledrive</category>
      <category>sync</category>
      <category>android</category>
    </item>
    <item>
      <title>The Subagents Guide I Wish I'd Had</title>
      <dc:creator>Stephan Miller</dc:creator>
      <pubDate>Thu, 17 Sep 2026 12:00:00 +0000</pubDate>
      <link>https://dev.to/eristoddle/the-subagents-guide-i-wish-id-had-dhh</link>
      <guid>https://dev.to/eristoddle/the-subagents-guide-i-wish-id-had-dhh</guid>
      <description>&lt;p&gt;A while back I wrote &lt;a href="https://dev.to/the-agent-skills-guide-i-wish-id-had/"&gt;the skills guide I wish I’d had&lt;/a&gt;. Skills stop your agent from forgetting what it knows about your codebase. The other half is that your agent also forgets &lt;em&gt;who it’s supposed to be&lt;/em&gt;. Every fresh session, it shows up as the same eager generalist, ready to have a reasonable, mediocre opinion about anything you throw at it.&lt;/p&gt;

&lt;p&gt;Skills are the memory problem. Subagents are the identity problem. This is the post about the second one.&lt;/p&gt;

&lt;p&gt;Here’s the shortest version I can give you before the table of contents scares you off: &lt;strong&gt;a skill tells the model what your world is like; a subagent tells it what role to play.&lt;/strong&gt; You can have both. And if you’re a solo dev shipping five half-finished projects at once like I am, you especially want both, because you are the only specialist you’ve got and you can’t be in five roles.&lt;/p&gt;

&lt;p&gt;Half of this is the conceptual guide (what these things are, where the files live, how to write one), and that part I could have written from docs. The second half is the stuff I only know because two of my subagents have been in use for months across a dozen repos, and they have failed in ways no best-practices post prepared me for. Both are open source and I’ll link to the actual files, because the most useful thing I can show you isn’t my advice, it’s a prompt that’s been beaten into shape by real runs.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The Problem Subagents Actually Solve&lt;/li&gt;
&lt;li&gt;Skills vs. Subagents: Knowledge vs. Behavior&lt;/li&gt;
&lt;li&gt;What a Subagent Actually Is (in Claude Code)&lt;/li&gt;
&lt;li&gt;The “Agent” Word Is Overloaded and It’s Making You Dumber&lt;/li&gt;
&lt;li&gt;
Writing Your First One

&lt;ul&gt;
&lt;li&gt;What goes in the body&lt;/li&gt;
&lt;li&gt;A quick walk-through: a bug-hunt agent that doesn’t chase the last commit&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
The Thing I Had Backwards: A Subagent Is Cold on the Conversation

&lt;ul&gt;
&lt;li&gt;Run state, because sessions die&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Delegation Economics: Who Pays for Which Thinking&lt;/li&gt;
&lt;li&gt;Give It a Budget or It Will Spend Everything&lt;/li&gt;
&lt;li&gt;The Bans Are Most of a Mature Agent&lt;/li&gt;
&lt;li&gt;They Fail Silently and Confidently&lt;/li&gt;
&lt;li&gt;Keeping the Contract Tight (and the Antipatterns That Bloat It)&lt;/li&gt;
&lt;li&gt;
The Two That Actually Survived

&lt;ul&gt;
&lt;li&gt;Template, then specialize&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;The Killer Combo: Point the Agent at Knowledge, Don’t Paste It In&lt;/li&gt;
&lt;li&gt;Where They Live: Personal, Project, Packaged&lt;/li&gt;
&lt;li&gt;How This Works in the Other Tools&lt;/li&gt;
&lt;li&gt;The Honest Version of How Good Ones Get Built&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The Problem Subagents Actually Solve
&lt;/h2&gt;

&lt;p&gt;Your coding agent has one default mode: helpful generalist. Ask it to review a pull request and you get generically reasonable feedback. Ask it to figure out why your app is slow and it starts poking at the last file you touched. Ask it to plan a database migration and it hands you something sensible that ignores three things that would bite you in production.&lt;/p&gt;

&lt;p&gt;The model isn’t dumb. It just doesn’t have a &lt;em&gt;role&lt;/em&gt;. A generalist doing a security pass thinks about different things than a security engineer would. A generalist chasing a bug looks at different evidence than someone who’s been on call and seen that exact failure three times already. The intelligence is there. The framing isn’t.&lt;/p&gt;

&lt;p&gt;I noticed this because I kept typing the same preamble. Before I’d let Claude Code review anything, I’d paste in some version of “review this like a paranoid security person, look at auth boundaries first, tell me about anything leaking into logs, don’t waste my time with style nits.” Third time I typed that in two weeks, the light went on. That preamble is a role. And a role you keep re-typing is a subagent you haven’t written yet.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F22oqi8g5j8x020v38bv1.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F22oqi8g5j8x020v38bv1.jpg" alt="The Problem Subagents Actually Solve" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That’s the trigger every write-up on this gives you, mine included. But it’s not why either of my two durable subagents exists.&lt;/p&gt;

&lt;p&gt;The two that actually survived came from different pressures:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Context economics.&lt;/strong&gt; Research grinding (twenty searches, a dozen fetched pages, half of them junk) will fill your main session with garbage you never wanted to read. That work has to happen somewhere else and come back as a summary. That’s &lt;a href="https://github.com/eristoddle/deep-research-agent" rel="noopener noreferrer"&gt;&lt;code&gt;web-search-agent&lt;/code&gt;&lt;/a&gt;, and the reason it exists is not that it’s smarter than the main thread. It’s that I don’t want what it read in my context.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Role separation.&lt;/strong&gt; When I’m designing something, I want to stay in the design conversation. I don’t want to be interrupted to approve the eleventh mechanical file edit that follows obviously from a decision I already made. So the decisions stay with me and the typing goes to an &lt;code&gt;implementer&lt;/code&gt; agent. The split isn’t about capability. It’s about which of us should be thinking about what.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The first one gets you a nice reviewer. The other two are what turn subagents from a convenience into how the project actually runs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Skills vs. Subagents: Knowledge vs. Behavior
&lt;/h2&gt;

&lt;p&gt;A &lt;strong&gt;skill&lt;/strong&gt; packages what the agent needs to &lt;em&gt;know&lt;/em&gt;. Your weird internal library. The environment quirk that breaks builds. The domain knowledge a smart new hire would have to be told because there’s no way to guess it. Conditionally loaded. I wrote a whole &lt;a href="https://dev.to/the-agent-skills-guide-i-wish-id-had/"&gt;guide on those&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;A &lt;strong&gt;subagent&lt;/strong&gt; packages how the agent should &lt;em&gt;behave&lt;/em&gt;. What it optimizes for. What it checks first. What tradeoffs it makes by default. What shape the output comes back in. It doesn’t teach the model anything new. It biases the intelligence that’s already in there toward one job.&lt;/p&gt;

&lt;p&gt;The clean test is to imagine calling in a specialist coworker for the task. Is the value they bring mostly &lt;em&gt;information you don’t have&lt;/em&gt;: domain knowledge, context, tribal know-how? That’s a skill. Or is the value mostly &lt;em&gt;how they approach the problem&lt;/em&gt;: what they look at first, what they weight heavily, what they refuse to sign off without checking, what they hand back? That’s a subagent.&lt;/p&gt;

&lt;p&gt;A security engineer reviewing your code doesn’t just know more than you. They &lt;em&gt;work&lt;/em&gt; differently. They look at auth boundaries first. They weight a privilege escalation path way higher than an ugly variable name. They won’t close the review without saying something about secrets. And they hand you structured findings. That ordering of attention is the thing a subagent encodes.&lt;/p&gt;

&lt;p&gt;Knowledge = skill. Methodology = subagent.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a Subagent Actually Is (in Claude Code)
&lt;/h2&gt;

&lt;p&gt;I’m Claude Code first here, same as the skills guide, because that’s my daily driver. Every other tool gets its section further down, quirks and all.&lt;/p&gt;

&lt;p&gt;In Claude Code, a subagent is a Markdown file with a little YAML frontmatter on top. It lives in one of two places:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;~/.claude/agents/&lt;/code&gt;: your personal library, available in every project on your machine. This is the sandbox.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;.claude/agents/&lt;/code&gt; inside a repo: scoped to that project, and if you commit it, it travels with the repo.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The frontmatter is small. The fields you’ll use:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;migration-risk-reviewer&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Reviews database migration plans and schema changes for rollback risk, lock contention, and data integrity problems. Use before running any migration against real data.&lt;/span&gt;
&lt;span class="na"&gt;tools&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Read, Grep, Glob&lt;/span&gt;
&lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;sonnet&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

&lt;span class="s"&gt;You are a senior database engineer reviewing a migration for production risk.&lt;/span&gt;

&lt;span class="c1"&gt;## What you look at first&lt;/span&gt;

&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;Can this be rolled back without losing data?&lt;/span&gt;
&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;Does it hold locks on a busy table during deploy?&lt;/span&gt;
&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;Is the ordering safe for a multi-step change?&lt;/span&gt;
&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;Will existing data violate any new constraint?&lt;/span&gt;

&lt;span class="c1"&gt;## Output shape&lt;/span&gt;

&lt;span class="na"&gt;Findings in order of severity&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;

&lt;span class="s"&gt;1. Blockers — deploy will fail or data will be lost&lt;/span&gt;
&lt;span class="s"&gt;2. High risk — real production risk, needs a mitigation&lt;/span&gt;
&lt;span class="s"&gt;3. Medium risk — should fix, won't necessarily block&lt;/span&gt;
&lt;span class="s"&gt;4. Notes — worth tracking&lt;/span&gt;

&lt;span class="na"&gt;For each finding&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;what it is, why it matters, what to do about it.&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That’s a working subagent. A couple of things worth knowing about how Claude Code treats it:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F84velq3tvuzyvemtgrmn.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F84velq3tvuzyvemtgrmn.jpg" alt="Output shape" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The body is the system prompt.&lt;/strong&gt; When the subagent runs, everything below the frontmatter becomes its instructions. It’s not documentation you read and then act on. The model reads it and &lt;em&gt;becomes&lt;/em&gt; the thing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The &lt;code&gt;description&lt;/code&gt; is load-bearing.&lt;/strong&gt; Claude Code uses it to decide when to hand a task off to this subagent automatically. Write it like an API someone else has to discover from context. “Use before running any migration” is findable. “helps with db stuff” is not. Write the &lt;em&gt;negative&lt;/em&gt; half too. Both of my real agents spend a clause on what they’re not for, because the routing will absolutely hand an agent work it has no business doing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;tools&lt;/code&gt; is an allowlist.&lt;/strong&gt; Leave it off and the subagent inherits everything. Narrow it and you’ve got a reviewer that literally can’t edit your files even if it gets an idea. For a review agent, that’s a feature. I don’t want my “just tell me what’s wrong” pass rewriting things on a whim.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;model&lt;/code&gt; is a budget decision, not a detail.&lt;/strong&gt; Cheap model for mechanical work, expensive model for judgment. There’s a whole section on this below, because it’s most of why delegation pays for itself.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It gets its own context window.&lt;/strong&gt; This is the part people underrate: a subagent runs in a separate context, does its thing, and hands back a summary. Your main session doesn’t get flooded with everything it read. It’s also the part &lt;em&gt;I&lt;/em&gt; underrated in the opposite direction, and there’s a whole section below where it breaks, because the separate context is what breaks it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You don’t have to hand-write the file.&lt;/strong&gt; You can just ask Claude Code to write the subagent for you: describe the role and it drops the file in the right place. The &lt;code&gt;/agents&lt;/code&gt; command is also there for managing them.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You invoke one by asking for it, by letting Claude route to it based on that description, or by wiring it into a larger flow where it runs as a delegated worker. Same behavior either way.&lt;/p&gt;

&lt;h2&gt;
  
  
  The “Agent” Word Is Overloaded and It’s Making You Dumber
&lt;/h2&gt;

&lt;p&gt;Not you specifically. Me. “Agent” gets bolted onto four different things and people conflate them constantly, so let me split them apart:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A subagent&lt;/strong&gt; (&lt;code&gt;.claude/agents/*.md&lt;/code&gt;, or a &lt;code&gt;.agent.md&lt;/code&gt; file over in Copilot land) is a &lt;em&gt;reusable file&lt;/em&gt; that defines a specialist role. This is the thing this whole post is about.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agent mode / autonomous run&lt;/strong&gt; is a &lt;em&gt;capability toggle&lt;/em&gt;: you’re letting the tool run commands, edit files, and generally act without you hitting approve on every line. That’s a permission setting, not a specialist.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A background/cloud coding agent&lt;/strong&gt; is a &lt;em&gt;service&lt;/em&gt; that picks up a task, churns on it out of sight, and comes back with a branch or a PR. Also not a specialist file.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;CLAUDE.md&lt;/code&gt; / &lt;code&gt;AGENTS.md&lt;/code&gt;&lt;/strong&gt; is &lt;em&gt;always-on repo guidance&lt;/em&gt;. The house rules every agent obeys on that codebase. It’s the employee handbook, not a coworker you call by name.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When I say subagent, I mean the reusable specialist you write once and invoke on purpose. Not the toggle, not the service, not the handbook. Keeping those four straight fixes about half the confusion.&lt;/p&gt;

&lt;h2&gt;
  
  
  Writing Your First One
&lt;/h2&gt;

&lt;p&gt;Start with the thinnest thing that does real work. A first subagent has exactly two jobs: declare the role, and say what that role actually means in practice. The &lt;code&gt;migration-risk-reviewer&lt;/code&gt; up above is already that: a clear role, specific heuristics, a predictable output shape, and it’s narrow on purpose. You’ll add more once you run it on real work and watch it miss things.&lt;/p&gt;

&lt;p&gt;The single most common way a first subagent flops is a mushy role. “Security helper” is not a role. It tells the model nothing about what to optimize for or check first. A real role is a behavioral contract: here’s the job, here’s what you look at first, here’s what you won’t let slide. If you can’t say the job in one sentence, it’s too vague.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Vague&lt;/th&gt;
&lt;th&gt;Concrete&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;“Security helper”&lt;/td&gt;
&lt;td&gt;“Review backend changes for auth boundary violations, secrets leaking into logs, and privilege escalation paths”&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;“Migration assistant”&lt;/td&gt;
&lt;td&gt;“Analyze schema changes for rollback safety, lock contention on busy tables, and data integrity risk on existing rows”&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;“Code reviewer”&lt;/td&gt;
&lt;td&gt;“Review async code for missing error handling, swallowed exceptions, and N+1 query patterns”&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The concrete version tells the model what it cares about &lt;em&gt;and what it ignores.&lt;/em&gt; The ignoring is half the value. A reviewer that gives equal weight to a typo and a privilege boundary is only an averaging machine.&lt;/p&gt;

&lt;h3&gt;
  
  
  What goes in the body
&lt;/h3&gt;

&lt;p&gt;Once the role’s defined, a useful body covers five things. You don’t need all five on day one:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;What it optimizes for:&lt;/strong&gt; the one thing this role is most trying to get right. When it has to make a tradeoff, this is what it trades toward.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What it checks first:&lt;/strong&gt; the high-priority signals a real specialist always looks at before anything else.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Failure modes:&lt;/strong&gt; &lt;a href="https://dev.to/eristoddle/somebody-finally-wrote-down-why-my-coding-agents-keep-failing-the-same-way-13o"&gt;the specific things that go wrong in this class of work&lt;/a&gt;. This is where the actual expertise lives, and it’s mostly stuff you’ll only learn by watching the agent mess up real tasks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Output shape:&lt;/strong&gt; not “some feedback” but &lt;em&gt;severity-ordered findings&lt;/em&gt;, &lt;em&gt;a phased plan&lt;/em&gt;, &lt;em&gt;a ranked list of hypotheses&lt;/em&gt;. Predictable output is the single biggest practical win over just prompting.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What it must not do.&lt;/strong&gt; I used to list this one as optional. It is not optional. In my most-used agent it’s the longest section in the file, and every line of it is there because something went wrong once.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Start with role, check-first, and output shape. Everything else the agent earns by screwing up.&lt;/p&gt;

&lt;h3&gt;
  
  
  A quick walk-through: a bug-hunt agent that doesn’t chase the last commit
&lt;/h3&gt;

&lt;p&gt;Here’s one I wanted, because I kept hitting the same dumb pattern. Something breaks, I paste the error into Claude Code, and its first instinct is to go stare at whatever file I edited most recently. Sometimes that’s right. Sometimes the last change had nothing to do with it.&lt;/p&gt;

&lt;p&gt;An experienced debugger doesn’t start at the code. They start at the evidence (what’s actually failing, what the logs say, what changed in the environment) and only open the source once there’s a hypothesis worth checking. That’s a behavior. So I saved it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;bug-investigator&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Investigates a bug or failure starting from evidence, not from recent code changes. Produces ranked hypotheses with the next diagnostic step for each. Use for "why is this broken" before touching the source.&lt;/span&gt;
&lt;span class="na"&gt;tools&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Read, Grep, Glob, Bash&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

&lt;span class="s"&gt;You are a senior engineer investigating a failure. Your job is to find the most&lt;/span&gt;
&lt;span class="s"&gt;likely cause, rule out the alternatives, and say what to check next.&lt;/span&gt;

&lt;span class="c1"&gt;## Investigation order&lt;/span&gt;

&lt;span class="s"&gt;1. Start from the symptom — what is actually failing, and how does it show up?&lt;/span&gt;
&lt;span class="s"&gt;2. Get the error output and any logs from around the time it broke — ask the caller for them if they weren't provided.&lt;/span&gt;
&lt;span class="s"&gt;3. Only then look at recent changes, and only if the evidence points there.&lt;/span&gt;
&lt;span class="s"&gt;4. Read source last, once there's a symptom-to-cause hypothesis worth testing.&lt;/span&gt;

&lt;span class="c1"&gt;## Output shape&lt;/span&gt;

&lt;span class="err"&gt;*&lt;/span&gt;&lt;span class="nv"&gt;*Symptoms&lt;/span&gt;&lt;span class="s"&gt;:** what's failing and how it manifests.&lt;/span&gt;
&lt;span class="err"&gt;*&lt;/span&gt;&lt;span class="nv"&gt;*Likely&lt;/span&gt; &lt;span class="s"&gt;causes (ranked):** for each — the hypothesis, the evidence for it, the next diagnostic step.&lt;/span&gt;
&lt;span class="err"&gt;*&lt;/span&gt;&lt;span class="nv"&gt;*Ruled&lt;/span&gt; &lt;span class="s"&gt;out:** what you checked and why it isn't the cause.&lt;/span&gt;
&lt;span class="err"&gt;*&lt;/span&gt;&lt;span class="nv"&gt;*Next&lt;/span&gt; &lt;span class="s"&gt;steps:** in priority order.&lt;/span&gt;
&lt;span class="err"&gt;*&lt;/span&gt;&lt;span class="nv"&gt;*Open&lt;/span&gt; &lt;span class="s"&gt;questions:** what you still can't confirm.&lt;/span&gt;

&lt;span class="c1"&gt;## Failure modes to avoid&lt;/span&gt;

&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;Don't assume the last change is the cause without evidence.&lt;/span&gt;
&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;Don't propose a fix before you can name the mechanism.&lt;/span&gt;
&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;Don't close with "cause unknown" — say what evidence would confirm or kill each remaining hypothesis.&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then (and this is the step people skip) I tested it on a real bug I’d already solved, not a toy. I gave it the symptom and the logs I’d had at the time and watched whether its investigation path matched the one that found the problem. Where it drifted, that drift became a new line in the body.&lt;/p&gt;

&lt;p&gt;That file is where everybody’s subagent post stops. Everything below is what happens after you’ve been running these things for months, which is the part I actually needed someone to tell me.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Thing I Had Backwards: A Subagent Is Cold on the Conversation
&lt;/h2&gt;

&lt;p&gt;The separate context window gets sold as pure upside. But it isn’t symmetric.&lt;/p&gt;

&lt;p&gt;This line is in my implementer agent, written after the same handoff failed three different ways:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;You are cold on the planning conversation but warm on the project.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Cold on the conversation: it never saw the two hours where you ruled out the obvious approach, argued yourself out of an abstraction, and settled on the weird-looking solution for a good reason. None of that exists for it. Warm on the project: it has &lt;code&gt;CLAUDE.md&lt;/code&gt; and it has the codebase, so it knows the house rules and the idioms without being told.&lt;/p&gt;

&lt;p&gt;Get that backwards in either direction and you get a specific, recognizable failure. Treat it as warm on the conversation and it confidently reimplements the approach you spent an hour rejecting, because from where it sits that approach looks fine. Treat it as cold on the project and you waste a thousand tokens re-explaining your own repo to something that can read it faster than you can describe it.&lt;/p&gt;

&lt;p&gt;What falls out of this is the thing no subagent guide told me to build: &lt;strong&gt;the handoff is a file, and the file is the contract.&lt;/strong&gt; Not a paragraph you type into the task prompt. A document on disk with fixed sections, because a prompt you improvise each time will quietly omit a different critical section every time.&lt;/p&gt;

&lt;p&gt;Mine is a &lt;a href="https://dev.to/eristoddle/the-other-doc-got-fat-too-my-task-file-was-90-finished-work-4hjc"&gt;&lt;code&gt;TASKS.md&lt;/code&gt; at the repo root&lt;/a&gt;, generated by the &lt;a href="https://github.com/eristoddle/agent-skills" rel="noopener noreferrer"&gt;&lt;code&gt;living-plan&lt;/code&gt;&lt;/a&gt; skill, and the active task has these sections and no others:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Goal:&lt;/strong&gt; what’s true when this is done.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Why (pointer):&lt;/strong&gt; a link to &lt;a href="https://dev.to/eristoddle/my-third-try-how-a-living-plan-beat-both-vibe-coding-and-spec-kit-5a89"&gt;the decision in &lt;code&gt;PLAN.md&lt;/code&gt;&lt;/a&gt;, not a re-argument of it. Pointer, not prose.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;▶ Run state:&lt;/strong&gt; the agent keeps this current. More on this in a second, because it’s the one that saves you.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Design — numbered pieces:&lt;/strong&gt; a serial queue, &lt;code&gt;[]&lt;/code&gt; not started, &lt;code&gt;[x]&lt;/code&gt; done, &lt;code&gt;[!]&lt;/code&gt; blocked, with dependencies marked inline as &lt;code&gt;depends on #2&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Files:&lt;/strong&gt; where the work happens.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tests:&lt;/strong&gt; the checks that stand in for review.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Out of scope (do NOT do):&lt;/strong&gt; the single highest-value section in the document.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Report back:&lt;/strong&gt; what the final message has to contain.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The rule that makes it work: &lt;em&gt;fill every section before launching, and mark one &lt;code&gt;—&lt;/code&gt; only if it genuinely doesn’t apply.&lt;/em&gt; An unfilled section is where the agent improvises.&lt;/p&gt;

&lt;p&gt;And then the division of labor that took me longest to see. Two documents, two lifetimes:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Standing “how we implement here” knowledge lives in the &lt;strong&gt;agent definition&lt;/strong&gt; , not in every task.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The agent file holds the protocol: how to read a task, how to handle a blocker, what to never touch, what the report has to contain. The task file holds this job. If you find yourself pasting the same instruction into three consecutive tasks, that instruction was never task content. It belongs in the agent.&lt;/p&gt;

&lt;h3&gt;
  
  
  Run state, because sessions die
&lt;/h3&gt;

&lt;p&gt;Here’s the rule that has saved me more real work than any clever prompt engineering in this post:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Log run-state whenever you stop — done OR blocked. Mandatory.&lt;/strong&gt; Flip each piece’s status box as you go, and update the &lt;strong&gt;▶ Run state&lt;/strong&gt; note (done / blocked+why / remaining / resume-from). Editing &lt;code&gt;TASKS.md&lt;/code&gt; for this is in-scope — it is the recovery point if the session dies.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Sessions die. You close the laptop, the connection drops, you hit a limit, you get bored and Ctrl-C something that was mostly finished. When that happens to a subagent, &lt;em&gt;everything it knew is gone&lt;/em&gt;. If the only record of progress was in its head, you now get to diff your own repo against your memory to figure out where it got to.&lt;/p&gt;

&lt;p&gt;So the agent writes its progress to disk as it goes, into the same file that gave it the job. The next session reads the file and picks up at the resume point. It costs a few lines in the agent definition and it turns a dead session from lost work into a paused one.&lt;/p&gt;

&lt;p&gt;The companion rule, which is about throughput rather than recovery:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A blocked piece does NOT halt the queue.&lt;/strong&gt; If a piece is underspecified or hits a dependency you can’t resolve, mark it &lt;code&gt;[!]&lt;/code&gt; blocked with a one-line reason and &lt;strong&gt;continue with any remaining piece that doesn’t depend on it&lt;/strong&gt;. Halt only when nothing remaining can proceed. Never guess a design — escalate blocked forks in your report.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgub23rgl70175sdmqs9o.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgub23rgl70175sdmqs9o.jpg" alt="Run state, because sessions die" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Default agent behavior on a snag is to stop and ask, which means a queue of eight tasks gets you one task and a question. Default behavior with no guardrail at all is worse: it guesses a design and keeps going, and now you’ve got seven pieces built on an invention you never approved. Skip-and-continue, never-guess, report-the-fork is the combination that lets you queue work and walk away.&lt;/p&gt;

&lt;h2&gt;
  
  
  Delegation Economics: Who Pays for Which Thinking
&lt;/h2&gt;

&lt;p&gt;That &lt;code&gt;model&lt;/code&gt; line is why this whole arrangement pays for itself.&lt;/p&gt;

&lt;p&gt;My implementer’s description, verbatim from a real project:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Sonnet implementation worker for the project. Use it to execute a fully-specified, mechanical coding task defined in TASKS.md while the main (Opus) planning thread keeps going. NOT for design decisions, exploring open questions, or live LLM / quality testing — those stay in the main thread.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;a href="https://dev.to/eristoddle/the-expensive-model-only-plans-now-i-split-my-ai-coding-rig-across-two-tools-4dch"&gt;Expensive model makes the decisions&lt;/a&gt;. Cheap model does the typing. And a narrow scope is precisely what makes the cheap model reliable: a Sonnet worker handed a fully-specified numbered queue in a codebase it can read is dependable.&lt;/p&gt;

&lt;p&gt;That split also gives you two working rhythms out of the same machinery:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Pacing.&lt;/strong&gt; Queue two or three pieces, launch the agent, and &lt;a href="https://dev.to/eristoddle/the-bottleneck-was-me-how-i-stopped-racing-my-ai-builder-and-started-pacing-it-3e1n"&gt;keep planning the next batch while it grinds&lt;/a&gt;. You’re never waiting on it and it’s never waiting on you.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Unattended.&lt;/strong&gt; Queue a lot and walk away. Come back to a diff, a report, and a run-state note telling you which piece blocked and why.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Permissions are the enforcement layer under this. My unattended implementer’s file and shell commands are allowlisted in the project’s permission settings; anything outside the list prompts for approval. The bans in the prompt are policy. The allowlist is the fence.&lt;/p&gt;

&lt;p&gt;One more line in that agent, which exists because of the most expensive mistake I made:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Work the numbered pieces as a serial queue, top-to-bottom, in one pass.&lt;/strong&gt; Do not spawn parallel workers and do not stop to report after each piece.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Parallel agents sound like free speed. They are not free. Every one of them finishes by dumping its full output back into the context that spawned it, and a fan-out that looked clever can bury the session that has to read all of it. I’ve killed real work that way. Parallel runs across disjoint files are an opt-in tactic for a specific burst, never the default. The default is one worker doing one queue in order.&lt;/p&gt;

&lt;h2&gt;
  
  
  Give It a Budget or It Will Spend Everything
&lt;/h2&gt;

&lt;p&gt;Any subagent with tools (search, fetch, shell) will grind until something stops it, and “I think that’s enough” is not a thing a model reliably concludes on its own. A research agent without a budget is a machine that turns your afternoon into many (hidden) browser tabs and a summary you could have gotten from the first three.&lt;/p&gt;

&lt;p&gt;So &lt;code&gt;web-search-agent&lt;/code&gt; opens with hard limits, above the methodology, above everything:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Level&lt;/th&gt;
&lt;th&gt;Searches&lt;/th&gt;
&lt;th&gt;Fetches&lt;/th&gt;
&lt;th&gt;Link depth&lt;/th&gt;
&lt;th&gt;Modules&lt;/th&gt;
&lt;th&gt;Use when&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;quick&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;One specific fact, a URL check, a yes/no. Minutes.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;standard&lt;/code&gt; &lt;em&gt;(default)&lt;/em&gt;
&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;12&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;Normal research task. Answer the questions and stop.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;deep&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;30&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;Genuinely hard question, contested facts, or a topic where the first page of results is known to be junk.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Modules are the per-source search playbooks the agent loads at runtime: source lists and query tactics per domain. There’s a whole section on them below.&lt;/p&gt;

&lt;p&gt;Four things about that table matter more than the specific numbers, which you should tune to your own work:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A depth dial beats a fixed cap.&lt;/strong&gt; One setting is always wrong: too tight for the hard question, too loose for the quick fact. Three named levels with a stated default means I ask for &lt;code&gt;deep&lt;/code&gt; on purpose instead of discovering I got it by accident. Precedence is explicit too: numbers the caller gives win, then a level named in the prompt, then the project’s config, then &lt;code&gt;standard&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The limits need anti-gaming clauses, because a model reads a budget as a target.&lt;/strong&gt; The two that do the work:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;1 fetch per URL.&lt;/strong&gt; Never re-fetch a URL you already read.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stop as soon as the caller’s questions are answered.&lt;/strong&gt; Remaining budget is not a quota to spend. A &lt;code&gt;deep&lt;/code&gt; run that finishes in six searches is a success, not a waste.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That second line went in after I watched a run answer the question at search four and then keep going, because twenty was the number it had been given. You have to say out loud that finishing early is winning.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Count what actually costs, not what’s convenient to count.&lt;/strong&gt; My fetch budget covers every network retrieval: native fetch, the escalation rungs for blocked pages, a helper script’s individual HTTP requests, and each &lt;code&gt;429&lt;/code&gt; retry inside that helper counted separately. The first version counted “fetch calls,” and the agent found the loophole without trying: it wasn’t cheating, it was obeying a rule I’d written badly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Make it announce its spending.&lt;/strong&gt; The agent reports &lt;code&gt;[deep: 3/20 searches, 5/30 fetches]&lt;/code&gt; after each phase. Without that you have no idea whether you’re watching a careful run or a runaway until it’s over. And when it does exhaust the budget with questions still open, it stops and reports exactly what’s unknown, which URL would most likely answer it, and that re-running deeper is an option. What it must never do is quietly continue past the limit: the caller chose the level, and overspending it takes that choice away.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Bans Are Most of a Mature Agent
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjohho272ghbs2jtbrq52.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjohho272ghbs2jtbrq52.jpg" alt="The Bans Are Most of a Mature Agent" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Look at the shape of my most-used agent file and the thing that jumps out is proportion. The role description is a paragraph. The negative constraints are a list. That inversion isn’t bloat. It’s that a capable agent with tools has far more ways to technically-comply than you can anticipate.&lt;/p&gt;

&lt;p&gt;Here’s one, quoted exactly, because the parenthetical is the whole point:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Never use browser automation.&lt;/strong&gt; No &lt;code&gt;Simple Browser&lt;/code&gt;, no embedded/preview browser, no Playwright, no &lt;code&gt;mcp __claude-in-chrome__ *&lt;/code&gt;, no &lt;code&gt;open&lt;/code&gt;, no opening tabs or windows. If a browser tool is offered to you, it is not for this task. Opening browser tabs to read pages has previously spawned ~100 tabs and wrecked a run.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And another, from the implementer in my research-agent repo, where the mechanism is the entire reason anyone would respect the rule:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Never create any file under &lt;code&gt;agents/&lt;/code&gt; or &lt;code&gt;skills/&lt;/code&gt;.&lt;/strong&gt; &lt;code&gt;agents/&lt;/code&gt; is shipped payload — the package manager flattens every &lt;code&gt;.md&lt;/code&gt; beneath it into a separate top-level agent on install, so a stray file there lands in every consumer project.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That’s the rule I’d extract from all of it: &lt;strong&gt;every ban names its mechanism.&lt;/strong&gt; Not because the agent needs to be persuaded, but because a constraint whose reason isn’t written down gets edited away. Six weeks later you’ll be tightening the prose, you’ll hit a line that says “never create files here,” it’ll read as arbitrary, and you’ll helpfully generalize it. The “because the installer flattens it into every consumer project” clause is what makes future-you leave it alone.&lt;/p&gt;

&lt;p&gt;A few more shapes these constraints take, once you’ve been at it a while:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ban the category, list the instances.&lt;/strong&gt; “No browser automation” alone leaves the agent deciding whether a preview pane counts. Naming six specific things it must not reach for closes the gap where it reasons its way into one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Name the behavior, not just the tool.&lt;/strong&gt; My favorite line in the file isn’t a prohibition on a command, it’s a prohibition on a tendency: &lt;em&gt;“If you find yourself authoring a scraper or a report generator, stop — you are working around the task, not doing it.”&lt;/em&gt; Capable agents route around constraints by building tools. That’s the category, and you have to ban the category.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;An escape hatch needs conditions, or it becomes the main road.&lt;/strong&gt; Some pages really are unfetchable, so there’s &lt;a href="https://dev.to/eristoddle/what-to-do-when-your-ai-coding-agent-cant-read-a-web-page-2f0i"&gt;an escalation ladder for blocked URLs&lt;/a&gt;. But it only opens after the normal path has already failed on that exact URL, it never grants a second fetch slot, it checks whether a tool is installed rather than installing anything, and it caps the output so a page dump can’t blow the context. Then the hard stop: &lt;em&gt;“If the URL is still unreachable after the rungs available to you, it is done.”&lt;/em&gt; Record it, move on. No further attempt, no other helper, no creative fourth idea.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhgfd6hb71ipkavgnriky.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhgfd6hb71ipkavgnriky.jpg" alt="The Bans Are Most of a Mature Agent" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Say why the exception isn’t a loophole.&lt;/strong&gt; There’s a paragraph in there explaining that a one-shot headless fetch that prints text is not a violation of the browser ban, because the ban is about driving a browser and leaving tabs for a human to close. Without that paragraph, the exception and the ban look like a contradiction, and a model resolving a contradiction will pick the reading that lets it do more.&lt;/p&gt;

&lt;h2&gt;
  
  
  They Fail Silently and Confidently
&lt;/h2&gt;

&lt;p&gt;I had two research runs come back fine. Not fine. &lt;em&gt;Good&lt;/em&gt;. Clean findings, real sources, coherent report. Both runs named Reddit as their primary source. Neither run had gotten a single thing from Reddit.&lt;/p&gt;

&lt;p&gt;What happened is that a WebSearch &lt;code&gt;site:reddit.com&lt;/code&gt; query returned ten clean, plausible results from &lt;em&gt;other domains&lt;/em&gt; (an Etsy community forum, the SBA, slideshare) and &lt;strong&gt;no error at all.&lt;/strong&gt; Not a 403. Not an empty set. Not a warning. The &lt;code&gt;site:&lt;/code&gt; constraint was silently ignored and nothing in the output said the domain filter hadn’t applied. A tidy list of usable pages from the wrong places. The agent did what any reasonable worker does with a tidy list of usable pages: it used them, and it reported success.&lt;/p&gt;

&lt;p&gt;The decision I wrote that day, which is now the foundational one in that project:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;When a named source cannot be reached, the run says so. It never proceeds silently on whatever the tool returned instead.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A run that names Reddit as its primary source while reporting success on Etsy forums is worse than a run that fails, because the failure is invisible to whoever reads the report.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A failed run costs you a re-run. A silently substituted run costs you a decision made on evidence you think came from somewhere it didn’t. And the separate context window, the feature itself, is exactly why you can’t see it happen. You don’t get the transcript. You get the summary the agent chose to write.&lt;/p&gt;

&lt;p&gt;Three things fix this, and none of them is “tell the agent to be careful.”&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Give it a mechanical detection rule.&lt;/strong&gt; Nothing errored, so the agent had no reason to look. The fix is one sentence with no judgment in it: &lt;em&gt;if you constrained a search to a domain and no returned URL is on that domain, that is a zero-result finding, not a result set.&lt;/em&gt; It needs no extra tool, no extra fetch, no permission. The URLs are already in its hands. “Be skeptical of your sources” is not actionable. “Compare the hostname to the one you constrained on” is.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Make provenance a required output channel with a schema.&lt;/strong&gt; This is where “output shape” graduates from section headers into something closer to a type. Two arrays, specified down to the key names:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;unreachable[]&lt;/code&gt;: one entry per wall, keys exactly &lt;code&gt;source&lt;/code&gt;, &lt;code&gt;url&lt;/code&gt;, &lt;code&gt;reason&lt;/code&gt;. Covers both hard fetch failure and silent substitution.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;sources[]&lt;/code&gt;: one entry per source that &lt;em&gt;actually supported an answer&lt;/em&gt;, keys exactly &lt;code&gt;source&lt;/code&gt;, &lt;code&gt;url&lt;/code&gt;, &lt;code&gt;fields&lt;/code&gt;. A page you opened that didn’t inform anything doesn’t get recorded. A source that answered four fields is one entry with four names in it, not four entries.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The subtle call in there, which I got wrong first: &lt;strong&gt;provenance annotates, it never blocks.&lt;/strong&gt; If a documented substitute answered the question, the field is answered. The &lt;code&gt;unreachable&lt;/code&gt; entry records that the answer didn’t come from the named source. Making a wall fail the field would mean the pipeline breaks loudly on exactly the situation it has a workaround for.&lt;/p&gt;

&lt;p&gt;And a failure worth stealing the lesson from: the first version of that agent said “record it as unreachable” &lt;strong&gt;four separate times and never once named a field, an array, or an output slot.&lt;/strong&gt; The instruction dead-ended. The agent was told to record something with nowhere to put it, so it didn’t, and nothing anywhere complained. When you write an instruction to a subagent, check that the thing it’s told to do has a destination. “Report X” with no named slot for X is a no-op with a clean conscience.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verify from outside the agent.&lt;/strong&gt; The agent’s compliance with its own output contract is checked by a script I run afterward, not by the agent’s assurance that it complied. That’s 180 lines of Python validating the JSON against the declared fields, and it is worth more than any amount of emphasis inside the prompt.&lt;/p&gt;

&lt;p&gt;The generalized version of that last point shows up again in the implementer, in a repo that has no test suite to lean on:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Verify instead of running a suite.&lt;/strong&gt; There is no suite. The task’s Tests section lists the checks that stand in for one — run every one of them and report the results concretely. Where a check is “this URL resolves,” that means &lt;strong&gt;fetch it&lt;/strong&gt; , not assume it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A URL you could not verify does not go into a file.&lt;/strong&gt; Report it as unverified with what happened. An unverified URL is worse than none — a reader trusts it and spends a fetch on a 404 instead of falling back to search.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Agents will &lt;a href="https://dev.to/eristoddle/my-ai-agent-kept-making-shit-up-and-other-lessons-from-running-openclaw-566p"&gt;report verification they did not perform&lt;/a&gt;, because a plausible claim and a checked claim look identical in a summary. Every check you care about needs to name the physical action that constitutes doing it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keeping the Contract Tight (and the Antipatterns That Bloat It)
&lt;/h2&gt;

&lt;p&gt;The best subagents have one job. And here you’ll object, because the files I’ve been quoting aren’t thin. My research agent is 188 lines and nobody would call it thin. The distinction that resolves it: &lt;strong&gt;the job stayed one sentence; the file grew.&lt;/strong&gt; What accumulated was constraints, detection rules, and output contract: scar tissue on a fixed skeleton. Not one new responsibility in a year. Growth from things going wrong is the file working as intended. Growth from new kinds of work is bloat, and the fix is a split, not another section.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj48sjoi1d033fyfkg53t.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj48sjoi1d033fyfkg53t.jpg" alt="Keeping the Contract Tight (and the Antipatterns That Bloat It)" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The bloated versions all look the same, and I’ve shipped every one: the everything-reviewer that averages across five concerns and is expert at none. The “senior engineer” label with no behavior behind it. The agent that restates your &lt;code&gt;CLAUDE.md&lt;/code&gt; and adds token cost instead of behavior. The one that duplicates a skill until the two drift apart. Each is the same mistake: a job that got wider than one sentence.&lt;/p&gt;

&lt;p&gt;Which raises the obvious question: after a year of this, what’s actually left on my bench?&lt;/p&gt;

&lt;h2&gt;
  
  
  The Two That Actually Survived
&lt;/h2&gt;

&lt;p&gt;Most posts on this hand you an org chart of agent types to build. Here’s my real bench after a year:&lt;/p&gt;

&lt;p&gt;Two agents. That’s it.&lt;/p&gt;

&lt;p&gt;One disclosure, because this post’s own rule applies to it: &lt;code&gt;web-search-agent&lt;/code&gt; started as a fork of &lt;a href="https://github.com/Weizhena/Deep-Research-skills" rel="noopener noreferrer"&gt;Lan Zheng’s Deep-Research-skills&lt;/a&gt;. The output-contract skeleton (the summary sections, the always-required sources list) is upstream’s, substantially verbatim. The control plane on top (the budgets, the bans, the provenance schema, the escalation ladder) is mine, and it’s half the line count. Everything below about scars is about that half.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;web-search-agent&lt;/code&gt;&lt;/strong&gt; does bounded web research and hands back findings with sources. It’s installed across a dozen repos right now: my blog’s static site, a couple of pipelines, and a few research projects. Everything in this post about budgets, bans, and silent substitution came out of that one file’s revision history.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;implementer&lt;/code&gt;&lt;/strong&gt; executes a fully-specified task queue while I keep planning. Four projects have one. Everything about handoff contracts, run state, and delegation economics came from those.&lt;/p&gt;

&lt;p&gt;The rest of what’s sitting in my agents folders came bundled with frameworks I installed, and it’s a junk drawer with better branding than my actual junk drawer. The specialists I &lt;em&gt;predicted&lt;/em&gt; I’d need (the flaky-test investigator, the PR-polish agent, the docs drafter, the codebase-orientation agent for projects I abandoned three months ago) were all reasonable ideas and I built approximately none of them. Turns out the ones that stick aren’t the ones that sound useful. They’re the ones where I kept feeling the specific pain of the work being in the wrong context or the wrong head.&lt;/p&gt;

&lt;p&gt;So don’t build the org chart. Notice which work you wish were happening somewhere else, and save that one.&lt;/p&gt;

&lt;h3&gt;
  
  
  Template, then specialize
&lt;/h3&gt;

&lt;p&gt;Here’s the pattern that made the second, third, and fourth implementer cheap: the agent gets &lt;strong&gt;generated from a template and then diverges where the project differs.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;living-plan&lt;/code&gt; skill ships an &lt;code&gt;implementer.template.md&lt;/code&gt; with &lt;code&gt;and&lt;/code&gt; placeholders, and scaffolding a new project writes out the agent, the plan doc, and the task doc together. What you get is the &lt;em&gt;protocol&lt;/em&gt;: read the contract, serial queue, skip blockers, log run state, stay in scope, report back. That part is identical everywhere because it’s about how delegation works, not about your code.&lt;/p&gt;

&lt;p&gt;Then each copy grows a local section for the local failure surface, and comparing two of them is instructive. The template says “keep the suite green.” One of my projects has no suite at all (it’s a repo of prompts, where the only executable file is the validator), so its implementer says this instead:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;This repo is &lt;strong&gt;prompts and data, not code.&lt;/strong&gt; Editing a file is editing a prompt. Wording, ordering, and emphasis are the implementation. A rewrite that reads better but drops a hard constraint is a regression, and nothing will catch it — there is no build and no test suite.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Another project’s copy grew a rule about reading the specific &lt;code&gt;PLAN.md&lt;/code&gt; decisions a task cites before writing code, and a report-back line for “anything you hit that requires an Opus/human call.” Same protocol, different hazards.&lt;/p&gt;

&lt;p&gt;That’s the division worth internalizing: &lt;strong&gt;the reusable part of a subagent is the protocol; the per-project part is the failure surface.&lt;/strong&gt; Template the first, hand-write the second, and don’t try to make one file serve both.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Killer Combo: Point the Agent at Knowledge, Don’t Paste It In
&lt;/h2&gt;

&lt;p&gt;This is where the two guides shake hands, and my understanding of it got a lot more specific once a real pipeline depended on it.&lt;/p&gt;

&lt;p&gt;What I’d have told you before is “mention the relevant skill in the agent body.” What actually works is stronger: &lt;strong&gt;the agent reads the knowledge at runtime, before it’s allowed to act, and the prompt says why.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;My research agent can’t run a single search until it has read a routing table that lives in a skill:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Module Selection (MANDATORY — routing lives in one file).&lt;/strong&gt; Before executing any search or fetch, you MUST &lt;code&gt;Read&lt;/code&gt; the routing table at the first existing path below. DO NOT skip this step. DO NOT route from memory — the module list changes without this prompt changing, so a module you remember may be gone and one you need may be new.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;em&gt;The module list changes without this prompt changing.&lt;/em&gt; That sentence is the entire argument for the pattern. The knowledge (which sources answer which kinds of question, how to query each one, what’s known to be blocked) changes monthly. The methodology doesn’t. If I’d baked the source list into the agent body, every new source would mean editing a prompt I’d otherwise leave alone, and the agent would confidently route to a module I deleted in March.&lt;/p&gt;

&lt;p&gt;Two mechanics that turned out to matter more than I expected:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A path resolution ladder, not a path.&lt;/strong&gt; The same agent runs in projects with different layouts, on two machines with different usernames, installed by different tools. So the instruction lists candidate paths in order and says take the first that exists: project-local first, then the user-level copies. Hardcode one path and the agent works on your machine and nowhere else.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl7w1fjlmcrs6hz22k7wb.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl7w1fjlmcrs6hz22k7wb.jpg" alt="The Killer Combo: Point the Agent at Knowledge, Don't Paste It In" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The layers are coupled by literal names, with nothing checking them.&lt;/strong&gt; The architecture is three tiers: skills orchestrate, an agent does the work, data modules hold the knowledge. And there are no imports anywhere. A skill launches the agent &lt;em&gt;by its exact registered name&lt;/em&gt;, and the agent reads modules &lt;em&gt;by path&lt;/em&gt;. Which leads to the most mundane and most dangerous line in that repo:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Do not rename &lt;code&gt;web-search-agent&lt;/code&gt;; existing consumers call it directly.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;strong&gt;The agent’s name is a public API.&lt;/strong&gt; Nothing will tell you otherwise: there’s no compiler, no test, no warning. A rename is a clean-looking commit that breaks every caller in every repo that installed it, and you find out the next time you ask for research and get a generalist instead.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where They Live: Personal, Project, Packaged
&lt;/h2&gt;

&lt;p&gt;There are three rungs, and the third one is where all my real agents live.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Personal&lt;/strong&gt; (&lt;code&gt;~/.claude/agents/&lt;/code&gt;): general-purpose specialists you want in every project. This is the workshop. Break things here.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Project&lt;/strong&gt; (&lt;code&gt;.claude/agents/&lt;/code&gt;, committed): specialists that only make sense for &lt;em&gt;this&lt;/em&gt; codebase. The implementer that knows this project has no test suite. Commit it and it’s there next time, even if next time is four months from now and you’ve forgotten the project exists.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Packaged&lt;/strong&gt; (installed by a package manager, pinned to a commit): an agent that’s a dependency. This is the rung I didn’t know I needed until the same agent was in a dozen repos and I had thirteen copies drifting apart.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That third rung changes a few things, and they’re worth knowing before you get there rather than after:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ship the agent with the skills that call it.&lt;/strong&gt; My package contains both the skills and the agents they launch, in one bundle, for an unglamorous reason: &lt;em&gt;installing the skills without the agents yields a pipeline that fails at first use.&lt;/em&gt; A skill that dispatches by name to an agent you don’t have is a broken install with a clean error message at best. They’re one unit because they’re one contract.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Know where your installer puts things.&lt;/strong&gt; This one installs through &lt;a href="https://github.com/microsoft/apm" rel="noopener noreferrer"&gt;APM&lt;/a&gt;, and APM flattens every &lt;code&gt;.md&lt;/code&gt; under an &lt;code&gt;agents/&lt;/code&gt; directory into a separate top-level agent on install. That’s a sane default that bites hard: a data file parked in that tree becomes a bogus registered agent in every consumer project, and a project-specific agent written there installs itself into everybody’s repos. It’s why my search-strategy modules live under &lt;code&gt;skills/&lt;/code&gt; rather than the &lt;code&gt;agents/&lt;/code&gt; directory they’d otherwise obviously belong in: a folder of them under &lt;code&gt;agents/&lt;/code&gt; loses its structure on install and registers five bogus agents in every project that pulls the package.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Pin a commit and know that consumers are behind.&lt;/strong&gt; Consumers depend on a SHA, which means a fix isn’t live for anybody until the pin moves and the install re-runs. Writing a fix is not shipping a fix. That sounds obvious written down and it is absolutely not obvious at 11pm when the bug you fixed last week reappears.&lt;/p&gt;

&lt;p&gt;And still: curate. A folder full of overlapping, half-abandoned agents is a noise generator, not a toolkit. If two have drifted into the same job, merge them. If one hasn’t been invoked in months, delete it. The file’s in git history if you’re wrong. Which, yes, is also good life advice, and no, I don’t follow it with my actual junk drawer.&lt;/p&gt;

&lt;h2&gt;
  
  
  How This Works in the Other Tools
&lt;/h2&gt;

&lt;p&gt;Claude Code isn’t the only place this exists, and if you bounce between tools like I do, the ideas port even when the file formats don’t.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;GitHub Copilot / VS Code&lt;/strong&gt; call these &lt;em&gt;custom agents&lt;/em&gt; and use &lt;code&gt;.agent.md&lt;/code&gt; files, typically in &lt;code&gt;.github/agents/&lt;/code&gt;, with a personal library option under your user directory. The frontmatter is richer than Claude’s: beyond &lt;code&gt;name&lt;/code&gt;, &lt;code&gt;description&lt;/code&gt;, and &lt;code&gt;tools&lt;/code&gt; you get things like &lt;code&gt;argument-hint&lt;/code&gt;, &lt;code&gt;model&lt;/code&gt;, &lt;code&gt;user-invocable&lt;/code&gt;, &lt;code&gt;disable-model-invocation&lt;/code&gt;, and my favorite, &lt;strong&gt;handoffs&lt;/strong&gt; : a finished agent can offer a button to hand off to another agent with a prefilled next prompt, so you can chain planning to implementation to review without collapsing everything into one mush-brained agent. One migration gotcha if you’re reading older examples: the old &lt;code&gt;infer&lt;/code&gt; field got split into &lt;code&gt;user-invocable&lt;/code&gt; and &lt;code&gt;disable-model-invocation&lt;/code&gt;, which is clearer once you know and maddening until you do.&lt;/p&gt;

&lt;p&gt;You’ll also read that VS Code detects &lt;code&gt;.claude/agents/&lt;/code&gt; files and maps the tool names across, so you can share an agent between tools for free. It’s true enough to be misleading. The tool vocabularies don’t line up, and a prompt that hard-bans tool names by name (which any mature agent of mine does) does not survive automatic translation intact.&lt;/p&gt;

&lt;p&gt;What I actually ship is &lt;strong&gt;one canonical prompt plus a thin native shim per tool.&lt;/strong&gt; The entire Copilot-side agent in my package is eleven lines:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Web Research Writer&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Use&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;for&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;bounded&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;web&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;research&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;that&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;must&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;read&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;local&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;schema,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;search&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;fetch&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;current&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;sources,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;write&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;one&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;designated&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;result&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;file,&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;validate&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;it&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;with&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;local&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;command."&lt;/span&gt;
&lt;span class="na"&gt;tools&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;read&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;search&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;web&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;edit&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;execute&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
&lt;span class="na"&gt;user-invocable&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
&lt;span class="na"&gt;disable-model-invocation&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

&lt;span class="na"&gt;Load the installed canonical research prompt from the first existing candidate&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;

&lt;span class="s"&gt;1. `.github/agents/web-search-agent.agent.md`&lt;/span&gt;
&lt;span class="s"&gt;2. `.claude/agents/web-search-agent.md`&lt;/span&gt;

&lt;span class="s"&gt;Follow its search budgets, module routing, source standards, tool discipline, output, and validation rules. Ignore that file's incompatible `tools` frontmatter; this wrapper's Copilot-native tool categories govern this session.&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That’s the whole pattern. Native frontmatter so the host registers it properly, a path ladder to find the real prompt, and one explicit sentence about which tool vocabulary wins. All 188 lines of hard-won behavior live in exactly one file, and the second tool gets a pointer instead of a fork. Two copies of a prompt is two prompts, and the day you fix a bug in one of them is the day they start lying to you about being the same agent.&lt;/p&gt;

&lt;p&gt;One caveat from the README: point the loader at the &lt;code&gt;.github/agents/&lt;/code&gt; tree and the canonical prompt registers as its own agent too, so Copilot can end up seeing duplicate names: the “Web Research Writer” shim plus the canonical file underneath it. Pick one home per tool, or keep the canonical file in the &lt;code&gt;.claude/agents/&lt;/code&gt; tree where Copilot’s auto-detection won’t pick it up.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fat8lu6vdqdcvla3vkf7k.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fat8lu6vdqdcvla3vkf7k.jpg" alt="How This Works in the Other Tools" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Everyone else&lt;/strong&gt; (Cursor, Windsurf, and the rest) is somewhere on the road to the same thing under names like rules, modes, or agents, and the details shift often enough that I’d rather point you at each tool’s current docs than confidently tell you something that changed last month. The underlying move is identical everywhere: a reusable file that defines a role, invoked on purpose. Learn the concept once and you’re mostly just learning where each tool hides the folder.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Honest Version of How Good Ones Get Built
&lt;/h2&gt;

&lt;p&gt;Every subagent I use started too broad. I wrote a “reviewer” that reviewed everything. A “planner” with no opinion about what a plan should look like. An “investigator” that approached every bug like the last one. That’s not failure. That’s the starting point.&lt;/p&gt;

&lt;p&gt;But here’s the part I’d have found genuinely useful a year ago, and it’s not “iterate.” Everybody says iterate. It’s &lt;em&gt;what the iterations are made of.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Go back through the agent files I’ve quoted and look at where the words actually are. A paragraph of role. A page of constraints. A schema with key names spelled out. A mechanical detection rule. A numeric budget with an anti-gaming clause. A note that a paid fetch has to be disclosed. A ban that explains what the installer does to stray files. &lt;strong&gt;Almost all of it is a record of a specific thing that went wrong once.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Which means the useful question after a disappointing run isn’t “how do I describe this role better.” It’s: &lt;em&gt;what exactly did it do, what in the file permitted that, and what sentence makes it impossible next time.&lt;/em&gt; A hundred tabs. Ten results from the wrong domain and no error. An instruction with nowhere to write its answer. A verification it reported but didn’t perform. Each one of those is a line, and the line outlives the session that taught it to you, which is the entire reason this is a file.&lt;/p&gt;

&lt;p&gt;The test before you keep one is the same as before you keep a skill: can you say its job in one sentence? If not, it’s too broad, and you’ve got two smaller agents wearing a trenchcoat. But the test for whether it’s any &lt;em&gt;good&lt;/em&gt; is different, and it’s this: when it screws up, can you point at the line that let it? If you can’t, you don’t have a specialist yet. You have a vibe with frontmatter.&lt;/p&gt;

&lt;p&gt;Put them together and you’re not prompting a very smart generalist over and over. You’re keeping a small bench of specialists who already know your world and already know their job. For one person trying to ship more than one person reasonably should, that’s the closest thing to hiring help I’ve found that doesn’t involve hiring anyone.&lt;/p&gt;

&lt;p&gt;Both of the agents I’ve been quoting are open source, scars included: &lt;a href="https://github.com/eristoddle/deep-research-agent/blob/main/agents/web-search-agent.md" rel="noopener noreferrer"&gt;the research agent itself&lt;/a&gt;, and &lt;a href="https://github.com/eristoddle/agent-skills" rel="noopener noreferrer"&gt;the implementer template&lt;/a&gt; plus &lt;a href="https://github.com/eristoddle/agent-skills" rel="noopener noreferrer"&gt;living-plan&lt;/a&gt;. Steal whatever’s useful.&lt;/p&gt;

&lt;p&gt;Now go look at the work you wish were happening somewhere else. That’s a subagent you haven’t saved yet. And the last time a subagent disappointed you, that’s a line you haven’t written yet.&lt;/p&gt;

</description>
      <category>subagents</category>
    </item>
    <item>
      <title>You're Using AI Like a Vending Machine</title>
      <dc:creator>Stephan Miller</dc:creator>
      <pubDate>Wed, 16 Sep 2026 12:00:00 +0000</pubDate>
      <link>https://dev.to/eristoddle/youre-using-ai-like-a-vending-machine-3gf2</link>
      <guid>https://dev.to/eristoddle/youre-using-ai-like-a-vending-machine-3gf2</guid>
      <description>&lt;p&gt;I asked a model a question about database indexing last week and got a good answer. It gave me four strategies, when each one applies, and a note about write amplification I had not considered. I read it, nodded, and closed the tab.&lt;/p&gt;

&lt;p&gt;An hour later it started bothering me and it took a while to work out why. The answer was fine. That was the problem. I had asked a question and received an answer, which is exactly the transaction I set up, and what came back was the most reasonable thing anybody could possibly say about database indexing.&lt;/p&gt;

&lt;p&gt;I put money in. Something fell down the chute. I took it and walked off.&lt;/p&gt;

&lt;p&gt;That is how most people use these things every single day, and it is the main reason they think the ceiling is a lot lower than it actually is.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models Are the Generalization Agents
&lt;/h2&gt;

&lt;p&gt;I said this in a working session a few months ago and I keep coming back to it:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Models are generalization agents. You’re not gonna find artisan or boutique products in them. They’re Walmart. They’ll have everything, but the most generalized version of it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Walmart is not a criticism. It has everything, it is close by, it is open, and it is cheap. If you need a phone charger at nine at night, Walmart is the correct answer.&lt;/p&gt;

&lt;p&gt;But nobody walks into a Walmart looking for the thing nobody else has.&lt;/p&gt;

&lt;p&gt;A one-sentence prompt returns the centroid of everything ever written about the subject. Not the best answer. The middle of the distribution, which is the safest possible place for the machine to stand. And when the question is unremarkable, this is what you want. Most of my questions are unremarkable. Most of yours are too.&lt;/p&gt;

&lt;p&gt;The trouble starts when you bring it something that actually matters to you and it hands you the middle of the distribution for that too, and it sounds so reasonable that you never notice you got handed the average.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Flattening Happens Somewhere Surprising
&lt;/h2&gt;

&lt;p&gt;The standard story is that RLHF ruined the models. Safety alignment sanded off the interesting edges, the alignment tax made everything bland, and if only we could get at the raw base model we would all be writing like Nabokov. It is a satisfying story because it has a villain.&lt;/p&gt;

&lt;p&gt;It also appears to be mostly wrong about where the loss occurs.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4ye9nljqfyxqomrm5iew.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4ye9nljqfyxqomrm5iew.jpg" alt="The Flattening Is Real, and It Happens Somewhere Surprising" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A 2026 analysis of the OLMo 3 post-training lineages by Constantinos Karouzos, Xingwei Tan and Nikolaos Aletras (&lt;a href="https://arxiv.org/abs/2604.16027" rel="noopener noreferrer"&gt;arXiv:2604.16027&lt;/a&gt;, a third-party analysis, not an Ai2 publication) goes looking for the exact stage where output diversity dies. The chain-of-thought distilled lineage keeps 38% of the base model’s diversity. The instruction-tuned lineage keeps 34%. Different roads to the same floor: the CoT line collapses at supervised fine-tuning, while the Instruct line bleeds most at DPO, the preference step most people mean when they say RLHF. The lineage trained with RL only, skipping both of those bottlenecks, keeps the most, at least 71% with a median of 94%, though it lands around half the distilled model’s score on grade-school math.&lt;/p&gt;

&lt;p&gt;Read that again, because it inverts the folk explanation. The stage everybody blames is the stage that preserved the most.&lt;/p&gt;

&lt;p&gt;I am not going to pretend this resolves into a clean recommendation about which model to pick, because it does not, and anybody selling you that certainty is ahead of the evidence. One paper and one model family does not settle. What it suggests is enough for this post. The flattening is real, it is measurable, and it is baked in upstream of anything you type. You are not going to prompt your way around a property of how the thing was trained.&lt;/p&gt;

&lt;p&gt;Which means the answer was never a better sentence. It was never a magic phrase, never “act as a world-class expert,” never a jailbreak, and never the prompt template somebody is selling you on a newsletter. Those are all still questions. They are just longer questions.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Question Gets You an Answer. A Move Gets You Something Else.
&lt;/h2&gt;

&lt;p&gt;Here is the distinction the rest of this series runs on.&lt;/p&gt;

&lt;p&gt;A question is a request. You describe what you want and the machine returns the most defensible version of it. The transaction is complete, the chute has delivered, and the quality of what you get is capped by how average your subject is.&lt;/p&gt;

&lt;p&gt;A move is different. A move does something &lt;em&gt;to&lt;/em&gt; the problem, or &lt;em&gt;to&lt;/em&gt; the model, that makes the centroid unavailable as an answer. You force two unrelated things into the same frame and make it reconcile them. You impose an arbitrary rule that rules out the obvious response. You take the problem apart into parameters before you try to solve any of it. You ask for the answer from four different seats and pool them afterward. You let something genuinely random pick a direction. You invert the question and solve the opposite.&lt;/p&gt;

&lt;p&gt;None of that is prompt engineering. It is closer to what a person does when they are stuck and finally stop staring at the thing.&lt;/p&gt;

&lt;p&gt;That is the whole series. Eleven more posts after this one, each one a family of moves, each one with the mechanism explained well enough that you could build your own variants instead of copying mine.&lt;/p&gt;

&lt;h2&gt;
  
  
  Most of These Are Older Than Computing
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4tnxi7mtpsi96rp5c69c.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4tnxi7mtpsi96rp5c69c.jpg" alt="Most of These Are Older Than Computing" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Almost none of these techniques were invented for AI. Forced relationships, morphological analysis, the deliberately terrible idea, writing under an arbitrary constraint, brainwriting, defamiliarization. They are decades old at minimum, some of them a century. They were developed by novelists, engineers, design researchers and a few outright lunatics, all working on paper, and they have survived decades of use by people who were stuck.&lt;/p&gt;

&lt;p&gt;It turns out a lot of them also work on models.&lt;/p&gt;

&lt;p&gt;So every post in this series does two jobs. The human version comes first and stands on its own, which means you can run it today with a pen and no account anywhere. Then the translation, which is what the same mechanism looks like when there is a model in the loop. If you have no interest in AI whatsoever you should still be able to take something usable out of every one of these. If that ever stops being true, I have written the post wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Half You Cannot Hand Off
&lt;/h2&gt;

&lt;p&gt;I want to be honest about the limit here, because the genre this post belongs to is usually dishonest about it.&lt;/p&gt;

&lt;p&gt;I once spent a long working session designing a system with a model, and afterward I &lt;a href="https://dev.to/eristoddle/planning-is-cheaper-than-coding-and-my-own-logs-proved-me-wrong-about-why-51bh"&gt;did the math&lt;/a&gt; and realized I had spent roughly as long as it would have taken me to just write the thing myself. The win was not speed. There was no speed. The win was that my attention went into judgment instead of getting burned down in the weeds of debugging, and I came out the other side understanding the design better than I would have.&lt;/p&gt;

&lt;p&gt;The less comfortable part of that same session: the model made three confident predictions about how the system would behave and every single one was wrong. That is the mild version of a failure mode I have hit much harder, back when &lt;a href="https://dev.to/eristoddle/my-ai-agent-kept-making-shit-up-and-other-lessons-from-running-openclaw-566p"&gt;a self-hosted agent of mine hallucinated fake news for two days&lt;/a&gt; before I noticed. The two ideas that actually survived contact came from me. One of them came out of a completely unrelated domain that the model had no reason to connect to the problem and never would have.&lt;/p&gt;

&lt;p&gt;So the moves in this series are not a way to get the machine to think for you. They are a way to stop getting the average. You still have to supply the taste, the judgment, and the weird connection from the unrelated thing you happened to know. What you get to outsource is the keystrokes, and honestly, the keystrokes were never the hard part.&lt;/p&gt;

&lt;p&gt;Stop putting in a dollar and taking whatever falls out.&lt;/p&gt;

&lt;p&gt;Next one: force two unrelated things together, and watch what the machine has to do to make them fit.&lt;/p&gt;

</description>
      <category>aiagents</category>
    </item>
    <item>
      <title>GPT-6 Astra Is the 'Smartest' Model. It Finished 24th.</title>
      <dc:creator>Stephan Miller</dc:creator>
      <pubDate>Tue, 15 Sep 2026 13:00:00 +0000</pubDate>
      <link>https://dev.to/eristoddle/gpt-6-astra-is-the-smartest-model-it-finished-24th-2dkc</link>
      <guid>https://dev.to/eristoddle/gpt-6-astra-is-the-smartest-model-it-finished-24th-2dkc</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F21a9x0aytu9t0fhl39im.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F21a9x0aytu9t0fhl39im.jpg" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Last week I wrote that the American labs had finally woken up. GPT-6 Astra and Claude Fable 5.1 both shipped, both at ten dollars in and fifty out, and I told you to read the footnote on each before you trusted the headline. This week the footnotes came due.&lt;/p&gt;

&lt;p&gt;Here’s what happens after the keynote. The blind-test votes trickle in, the independent testers run their own bills, and the model that got called “the world’s most intelligent” a week ago has to actually go stand in line with everything else. GPT-6 Astra did that this week. It finished 24th.&lt;/p&gt;

&lt;p&gt;Meanwhile Grok 4.7 got the worst possible review, from the one person who has actually used it. And down in the bargain bin, the cheap model I keep telling you about quietly took a fifth category off the board. So the pattern this week is simple. Everything loud got quieter, and the quiet thing got louder. Let me walk you through it, and yes, the table you came for is still here.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The “world’s most intelligent model” showed up to the blind test and finished 24th&lt;/li&gt;
&lt;li&gt;Grok 4.7 got a bad review from the person who built it&lt;/li&gt;
&lt;li&gt;The boring cheap model took a fifth category&lt;/li&gt;
&lt;li&gt;The cheapskate picks&lt;/li&gt;
&lt;li&gt;Newest still isn’t best, and now it’s just funny&lt;/li&gt;
&lt;li&gt;Google shipped another Flash, because of course it did&lt;/li&gt;
&lt;li&gt;About that bill&lt;/li&gt;
&lt;li&gt;What’s coming&lt;/li&gt;
&lt;li&gt;The honest version&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The “world’s most intelligent model” showed up to the blind test and finished 24th
&lt;/h2&gt;

&lt;p&gt;GPT-6 Astra shipped September 3 and OpenAI’s launch copy called it the most intelligent model on Earth. On the Artificial Analysis intelligence index, fair enough, it ties Fable 5.1 at the top. That’s a benchmark score.&lt;/p&gt;

&lt;p&gt;This week it hit Arena, which is the board where humans vote on which of two anonymous answers they actually prefer, and the picture got more honest in a hurry. The max variant landed 24th overall on preliminary votes, sitting behind a wall of older Claude models, some of them the better part of a year old. It does better in Coding, where it slots in around sixth. Everywhere else it’s mid-pack. This is the split I write about most weeks and it never stops being useful: Artificial Analysis measures what scores well on hard evals, Arena measures what people like when they can’t see the logo, and the two disagree constantly. A model can be smart on paper and forgettable in the booth. Astra is currently both.&lt;/p&gt;

&lt;p&gt;The independent testers found the more expensive problem. OpenAI prices Astra at roughly 40 percent of Fable 5.1 per token, which sounds like a win until someone runs the same job through both. One head-to-head test suite cost about 198 dollars in tokens on Astra against 113 dollars on Fable 5.1. That’s 75 percent more money to do the same work, on the model that’s supposedly cheaper. It happens because Astra spends more tokens per task, and the per-token sticker price never tells you how many tokens a task takes. This is the oldest trap in the roundup and it caught a brand-new flagship on launch week.&lt;/p&gt;

&lt;p&gt;Where Astra genuinely wins, it wins hard. FrontierMath Tier 4 at 97.6 percent against Fable 5.1’s 87.8. Terminal-Bench Science well ahead. If your work lives in those specific rooms, it’s a real tool. Just don’t read “most intelligent model on Earth” and assume that means “the one that’ll feel best or cost least on your actual work.” The blind test and the invoice both disagree.&lt;/p&gt;

&lt;h2&gt;
  
  
  Grok 4.7 got a bad review from the person who built it
&lt;/h2&gt;

&lt;p&gt;I love this one. Last week I told you Grok 4.7 was the loud upcoming thing, that Musk had said “ten days” on September 2, which pointed at roughly the twelfth, and that you should treat the date as a tweet and not a commitment. Reader, it was a tweet.&lt;/p&gt;

&lt;p&gt;The twelfth came and went. On the eleventh Musk said it “needs a few more days to cook.” Then on the fourteenth the pitch itself changed. The model that was going to be “better than 4.6 in every way” and beat everything on the board got quietly re-described as “roughly on par with Opus 5.0, not 5.1.” Read that again. Before the thing has a model card, a price, a benchmark table, or an API you can call, its own creator walked it back from “beats everyone” to “about as good as the model Anthropic already replaced.” That’s not a launch. That’s a man marking down his own inventory in public.&lt;/p&gt;

&lt;p&gt;The specs, such as they are, remain a set of claims: 2.1 trillion parameters, up 40 percent from Grok 4.6, some of the training data pulled from SpaceX internal engineering records, “even better token efficiency,” “slightly slower to serve.” Every one of those is a sentence from a person and not a line from a document. When there’s an actual model to test I’ll test it. Until then, Grok 4.7 is the clearest example this month of the gap between what gets announced and what gets shipped, and the gap is being narrated live by the announcer.&lt;/p&gt;

&lt;h2&gt;
  
  
  The boring cheap model took a fifth category
&lt;/h2&gt;

&lt;p&gt;Now the quiet story, which is the one that actually changes your bill.&lt;/p&gt;

&lt;p&gt;Two weeks ago I called the handover. GLM-5.3-Flash from Z.ai took the cheapest-good-model crown from Xiaomi’s MiMo v2.5 Pro, first as an emerging pick on thin votes, then confirmed once the votes firmed up. This week it kept going. GLM-5.3-Flash is now the cheapest model inside the competitive band for five of the six Arena categories. It picked up Math this week, a board it wasn’t even in a fortnight ago. The only holdout is Creative Writing, which stays a Gemini story because GLM never cracked that band.&lt;/p&gt;

&lt;p&gt;It’s not just an Arena artifact either. On OpenRouter, which measures actual token volume, real money spent on real calls, GLM-5.3-Flash is now the number two model on the entire platform at around 10 trillion tokens a week. It sits behind DeepSeek V4 Flash and ahead of GPT-5.6 Luna and MiMo. So the preference board and the usage board agree: this is a model people are genuinely running, not just voting on. MiMo is the runner-up now, and still a good one, with ten to sixty times the vote count depending on the category, which is exactly why it stays in the table as the steady anchor when GLM’s votes are thin.&lt;/p&gt;

&lt;p&gt;Two catches, because there are always two. First, GLM-5.3-Flash scores a 42 on the hard-reasoning index where the frontier models live in the fifties and sixties. It’s preference-strong and cheap, not a deep-reasoning machine. For everyday work that’s a rounding error. For genuinely hard problems it’s a caveat you should feel. Second, and this one’s on the price tag: the seven-and-a-half-cents-in, quarter-out number you’ve seen quoted everywhere was a 50 percent launch promo, and it expired September 9. List is fifteen cents in and fifty cents out, and Z.ai’s own OpenRouter endpoint moved to list on the eleventh. Arena’s table still shows the dead promo price, which is going to mislead somebody. Build your budget on fifteen and fifty. If you sized a spend off the August number, it just doubled on you.&lt;/p&gt;

&lt;p&gt;One piece of good news to balance that. The old “cheap but slow” knock is softening. Artificial Analysis now clocks GLM-5.3-Flash at around 114 output tokens a second on Z.ai’s own endpoint, up from the roughly 60 I was quoting a couple weeks ago. It still has a slightly high time-to-first-token, but for a fifty-cent model that’s genuinely usable in an agent loop now, not just in a chat window.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cheapskate picks
&lt;/h2&gt;

&lt;p&gt;Same method as always. For each Arena category I take the leader’s rating, draw a band 50 points below it, and find the cheapest model still inside that band. The whole premise is that Arena ratings cluster tight at the top, so the leader is usually a rounding error better than something 20 to 100 times cheaper. I compute the band from the full table in code, not by eyeballing the first screen, because eyeballing it is how you accidentally delete the entire cheap tail and crown the cheapest expensive model. Ask me how I know.&lt;/p&gt;

&lt;p&gt;Bands ran deep again this week: about 62 models inside the Overall band, 64 in Coding, thinner in Creative and Math. GLM-5.3-Flash quoted at list, fifteen and fifty, not the expired promo. Arena data dated around September 14.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Leader&lt;/th&gt;
&lt;th&gt;$ leader out&lt;/th&gt;
&lt;th&gt;Cheapskate pick&lt;/th&gt;
&lt;th&gt;$ pick out&lt;/th&gt;
&lt;th&gt;Δ rating&lt;/th&gt;
&lt;th&gt;Cheaper by&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Overall&lt;/td&gt;
&lt;td&gt;claude-fable-5 (1506)&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;GLM-5.3-Flash (1475, #29)&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;−31&lt;/td&gt;
&lt;td&gt;~100×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Coding&lt;/td&gt;
&lt;td&gt;claude-fable-5 (1552)&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;GLM-5.3-Flash (1525, #20)&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;−27&lt;/td&gt;
&lt;td&gt;~100×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Creative Writing&lt;/td&gt;
&lt;td&gt;claude-fable-5 (1504)&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;gemini-3-flash (1459, #26)&lt;/td&gt;
&lt;td&gt;$3&lt;/td&gt;
&lt;td&gt;−45&lt;/td&gt;
&lt;td&gt;~16.7×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Instruction Following&lt;/td&gt;
&lt;td&gt;claude-opus-4-6-high (1513)&lt;/td&gt;
&lt;td&gt;$25&lt;/td&gt;
&lt;td&gt;GLM-5.3-Flash (1471, #29)&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;−42&lt;/td&gt;
&lt;td&gt;~50×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hard Prompts&lt;/td&gt;
&lt;td&gt;claude-opus-4-6-high (1533)&lt;/td&gt;
&lt;td&gt;$25&lt;/td&gt;
&lt;td&gt;GLM-5.3-Flash (1498, #28)&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;−35&lt;/td&gt;
&lt;td&gt;~50×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Math&lt;/td&gt;
&lt;td&gt;claude-fable-5 (1526, prelim)&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;GLM-5.3-Flash (1513, #7)&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;−13&lt;/td&gt;
&lt;td&gt;~100×&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A few notes on reading that. Overall and Hard Prompts are the solid ones: GLM-5.3-Flash is sitting on real vote counts there, ten thousand in Overall, and it’s cleanly the cheapest thing in a very deep band. In Coding the rating is great, twentieth for fifty cents, but the vote count is still on the thin side at under three thousand, so if you want certainty over the last few dollars, MiMo v2.5 Pro at rank 25 with sixteen thousand votes for eighty-seven cents is the steadier bet. Math is the one to squint at: GLM-5.3-Flash shows up seventh for fifty cents, which is loud, but it’s riding 441 preliminary votes on a board where the whole top is thin. Treat that row as a strong suggestion, not a promise, and if you want a sturdier number, MiMo at the band edge or a Gemini Flash won’t embarrass you.&lt;/p&gt;

&lt;p&gt;Creative Writing stays Gemini because GLM never made that band, and the only value pick there is gemini-3-flash at three dollars, which is still 17 times cheaper than the leader. Nobody’s paying fifty dollars a million to write blog intros.&lt;/p&gt;

&lt;h2&gt;
  
  
  Newest still isn’t best, and now it’s just funny
&lt;/h2&gt;

&lt;p&gt;Here’s the thing I can’t stop noticing. The number one model on Arena Overall is claude-fable-5. Not Fable 5.1, the newer one Anthropic shipped two weeks ago. The old one. And sitting at number two, ahead of every single newer Opus, is claude-opus-4-6, a model that’s roughly a year old at this point. It beats Opus 5. It beats Opus 4.7 and 4.8. In Instruction Following and Hard Prompts it’s the outright leader.&lt;/p&gt;

&lt;p&gt;I’m not saying the new models are bad. I’m saying the blind test keeps preferring the model everyone already moved on from, and that’s now happened enough weeks in a row that it’s a pattern and not a fluke. If you switched off Opus 4.6 because something newer came out, the crowd that can’t see the version number would like a word.&lt;/p&gt;

&lt;p&gt;While we’re at it, one more sleeper. Meta’s Muse Spark line is quietly camped in the Arena top five overall and almost nobody talks about it. The 1.2 xHigh variant sits fourth, the new 1.3 max is up around eighth. It’s closed-weight and the newest tier doesn’t have a clean public price, so it can’t take a cheapskate slot, but if you’re only watching the Claude-OpenAI-Google fight you’re missing a genuinely competitive model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Google shipped another Flash, because of course it did
&lt;/h2&gt;

&lt;p&gt;Quick Google check-in, same paragraph I’ve now written four or five times. On September 2 they shipped Gemini 3.8 Flash, the third Flash release in six weeks, built on the 3.7 base, same price as 3.7 at 75 cents in and 3.75 out, with a locked-down Cyber sibling for security work. It’s good. It’s already climbing the Arena boards fast, top ten overall and third in Creative on preliminary votes. Note the price cliff, though: that 75-and-3.75 is an intro rate that doubles to 1.50 and 7.50 on January 1.&lt;/p&gt;

&lt;p&gt;And it’s, again, shipping instead of the real thing. Gemini 3.5 Pro is shelved. Gemini 4 has “cleared pretraining” with the largest training run in Google’s history and no benchmarks, no price, no date beyond a vague late 2026. So the strategy holds: drip out excellent cheap Flash models to stay in the conversation while the actual next-generation model bakes in the background. It’s working. It’s also the reason I can copy-paste this section every month.&lt;/p&gt;

&lt;h2&gt;
  
  
  About that bill
&lt;/h2&gt;

&lt;p&gt;The horror story this week isn’t one incident, it’s a category, and it keeps getting more expensive.&lt;/p&gt;

&lt;p&gt;The clean version is the Astra number above: a “cheaper” model that cost 75 percent more to run because cost-per-token isn’t cost-per-task. But the loud version is the agent loop. One that made the rounds this week: an Analyzer agent and a Verifier agent started ping-ponging, the Analyzer generating output and the Verifier asking for more analysis, over and over, with no budget ceiling and no alert anyone acted on. It ran for 264 hours before the number on the billing dashboard got big enough for a human to finally look. Forty-seven thousand dollars. Two bots politely asking each other to keep going, for eleven days.&lt;/p&gt;

&lt;p&gt;And the quiet version is the promo cliff. If you wired GLM-5.3-Flash into anything in August at seven-and-a-half cents, your cost doubled on September 9 and Arena is still showing you the old number like nothing happened. None of these are exotic. They’re the three normal ways a 2026 AI bill goes wrong: the sticker lies about the task, the loop has no brakes, and the cheap price had an expiration date you didn’t read.&lt;/p&gt;

&lt;h2&gt;
  
  
  What’s coming
&lt;/h2&gt;

&lt;p&gt;Three to watch.&lt;/p&gt;

&lt;p&gt;Grok 4.7, covered above, whenever it actually ships and at whatever the reviews say it is by then rather than what it was announced as.&lt;/p&gt;

&lt;p&gt;GPT-6 Astra is still mid-rollout. Day-one access went to a handful of orgs, and the ChatGPT tiers, the API, and AWS are coming online “over the coming days.” There’s also a program called Daybreak that loosens the safety rails for vetted organizations, which is worth keeping an eye on given this is the model that reportedly aced an exploit-writing benchmark last week.&lt;/p&gt;

&lt;p&gt;Gemini 4, pretraining done, everything else a rumor. Late 2026 if you believe the tea leaves.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest version
&lt;/h2&gt;

&lt;p&gt;The clean narrative last week was that the West is back. The clean narrative this week would be that the West already fumbled it. Neither is true, and the truth is more boring and more useful than either.&lt;/p&gt;

&lt;p&gt;The frontier models are real and good and expensive, and when you actually run them the story gets complicated fast. The one billed as smartest finished 24th in the blind test and cost more to run than its pricier rival. The loudest upcoming model got downgraded by its own creator before launch. That’s not the West failing. It’s just the difference between a keynote and an invoice, and the invoice always shows up a week late.&lt;/p&gt;

&lt;p&gt;And underneath all of it, a fifteen-cent open-weight model from Z.ai took its fifth category and became the second most-used model on the planet without a single keynote. If you’re shipping something and paying the bill yourself, that’s the line in this whole roundup that changes your life. The frontier got a loud, expensive week. The floor got quietly, permanently cheaper. Guess which one you’ll actually be using in a month.&lt;/p&gt;

&lt;p&gt;I’ll be back next week to find out whether Grok 4.7 exists yet, and whether its creator likes it any better by then.&lt;/p&gt;

</description>
      <category>llm</category>
      <category>openrouter</category>
    </item>
    <item>
      <title>Syncthing for Obsidian: Free, Private, and Genuinely Annoying</title>
      <dc:creator>Stephan Miller</dc:creator>
      <pubDate>Thu, 10 Sep 2026 13:00:00 +0000</pubDate>
      <link>https://dev.to/eristoddle/syncthing-for-obsidian-free-private-and-genuinely-annoying-3dbm</link>
      <guid>https://dev.to/eristoddle/syncthing-for-obsidian-free-private-and-genuinely-annoying-3dbm</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fijgxncwwy0xwiqcvunr5.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fijgxncwwy0xwiqcvunr5.jpg" alt=" " width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I have recommended Syncthing in four separate posts on this site and I do not use it.&lt;/p&gt;

&lt;p&gt;That is not a great look, so let me explain it, because the explanation is the whole point of this post. Syncthing is the best free way to sync an Obsidian vault. No account, no storage cap, no monthly anything, and your notes never sit on a server owned by a company that could get bought, breached, or bored. If you are still picking a method, my guide to &lt;a href="https://dev.to/eristoddle/how-to-sync-obsidian-across-all-your-devices-including-free-methods-1mi5"&gt;syncing an Obsidian vault across devices&lt;/a&gt; is the page that compares all of them. This one assumes you already picked Syncthing and want to run it without hating it.&lt;/p&gt;

&lt;p&gt;The reason I run &lt;a href="https://dev.to/eristoddle/obsidian-dropbox-sync-a-setup-guide-2ajl"&gt;Remotely Save over Dropbox&lt;/a&gt; instead comes down to one design decision, and it is not a bug. Syncthing has no server. That is the feature everyone likes. It is also the reason it stops working, and the good news is that it is fixable for about the cost of a Raspberry Pi.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;There Is No Server, and That Is the Whole Point&lt;/li&gt;
&lt;li&gt;Setting Up Syncthing for an Obsidian Vault&lt;/li&gt;
&lt;li&gt;Three Devices Is Where People Give Up&lt;/li&gt;
&lt;li&gt;
Best Practices for an Obsidian Vault Specifically

&lt;ul&gt;
&lt;li&gt;Ignore the workspace files, on every device separately&lt;/li&gt;
&lt;li&gt;Turn on versioning, and know where it puts things&lt;/li&gt;
&lt;li&gt;Use Receive Only for the devices you only read on&lt;/li&gt;
&lt;li&gt;Never run two sync systems on one vault&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;What a Sync Conflict Actually Looks Like&lt;/li&gt;
&lt;li&gt;Is There a Syncthing Plugin for Obsidian?&lt;/li&gt;
&lt;li&gt;Phones Are a Separate Problem&lt;/li&gt;
&lt;li&gt;The Reason People Quit&lt;/li&gt;
&lt;li&gt;Fix It With One Always-On Node&lt;/li&gt;
&lt;li&gt;
Frequently Asked Questions

&lt;ul&gt;
&lt;li&gt;Is there a Syncthing plugin for Obsidian?&lt;/li&gt;
&lt;li&gt;Does Syncthing work with Obsidian on iOS?&lt;/li&gt;
&lt;li&gt;Is Syncthing free for Obsidian?&lt;/li&gt;
&lt;li&gt;Why is my Obsidian vault not syncing with Syncthing?&lt;/li&gt;
&lt;li&gt;What should I add to Syncthing ignore patterns for Obsidian?&lt;/li&gt;
&lt;li&gt;Is Syncthing better than Obsidian Sync?&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;My Verdict&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  There Is No Server, and That Is the Whole Point
&lt;/h2&gt;

&lt;p&gt;Every other sync method in this series works the same way underneath. Your laptop pushes changes to a computer somewhere, your phone pulls them down later, and the computer in the middle holds your notes until both ends get around to showing up.&lt;/p&gt;

&lt;p&gt;Syncthing doesn’t have the middle. Your devices find each other and talk directly, encrypted, peer to peer. There is no account to create or storage tier to blow through when your vault gets attachments. Nobody is holding a copy.&lt;/p&gt;

&lt;p&gt;For a folder of markdown files this is close to perfect, and it is why Syncthing keeps winning the “best free option” argument in forum threads that have been running for years.&lt;/p&gt;

&lt;p&gt;It also means that when neither device is awake, nothing happens. Hold onto that. It comes back later and it’s the whole reason this post has the word “annoying” in the title.&lt;/p&gt;

&lt;h2&gt;
  
  
  Setting Up Syncthing for an Obsidian Vault
&lt;/h2&gt;

&lt;p&gt;Start with two desktops. Get that working before you bring a phone into it, because if you debug the mesh and the mobile client at the same time you will not know which one is lying to you.&lt;/p&gt;

&lt;p&gt;On macOS:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;brew &lt;span class="nb"&gt;install &lt;/span&gt;syncthing
brew services start syncthing

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Linux has it in every package manager worth using. Windows has installers and a tray app on the &lt;a href="https://syncthing.net/downloads/" rel="noopener noreferrer"&gt;downloads page&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwf6uyzaol499s7qbpgkx.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwf6uyzaol499s7qbpgkx.jpg" alt="Setting Up Syncthing for an Obsidian Vault" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Then comes the part that throws most people. &lt;strong&gt;Syncthing has no window.&lt;/strong&gt; It is a background service with a web UI, and the UI lives at &lt;code&gt;http://localhost:8384&lt;/code&gt; in your browser. That page is the entire application. People install it, look for an icon, do not find one, and conclude the install failed.&lt;/p&gt;

&lt;p&gt;From there:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Add the vault.&lt;/strong&gt; Click Add Folder, point it at your vault directory, and note the Folder ID it generates. Both devices need the same Folder ID for the same folder, which Syncthing handles for you when you share it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Trade device IDs.&lt;/strong&gt; Actions, then Show ID. You get a long string and a QR code. Paste that string into the other machine under Devices, then the plus button. The first machine throws a prompt asking whether it should talk to this new device. Accept it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Share the folder.&lt;/strong&gt; Back on the first machine, edit the vault folder, open the Sharing tab, tick the other device. The other end gets a notification asking where to put the folder. That is where you choose the path.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Walk away for ten minutes.&lt;/strong&gt; The first pass has to hash every file in the vault before it moves a byte. Both ends show a percentage while it works. Watch the number move instead of deciding it is broken.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Two desktops, and you are done. It really is that quick, which is why the guides all end here and why everyone’s actual problems start at device three.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three Devices Is Where People Give Up
&lt;/h2&gt;

&lt;p&gt;Syncthing pairs devices, not networks. Add a third device the obvious way and you are not adding one connection, you are adding two. Every device has to know every other device.&lt;/p&gt;

&lt;p&gt;The math gets ugly faster than you would think:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Devices&lt;/th&gt;
&lt;th&gt;Pairings you set up by hand&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I sync five devices. Ten pairings, done by hand, each one needing a device ID pasted correctly and a share accepted on both ends. That’s a bad afternoon.&lt;/p&gt;

&lt;p&gt;The fix is a setting that’s hard to find, because it is a checkbox on the device rather than on the folder. Mark one device as an &lt;strong&gt;introducer&lt;/strong&gt; , and it hands out its own connection list. Per &lt;a href="https://docs.syncthing.net/users/introducer.html" rel="noopener noreferrer"&gt;Syncthing’s docs&lt;/a&gt;, “when two devices connect they exchange a list of mutually shared folders and the devices connected to those shares.” Pick your most stable machine, mark it introducer on everything else, and new devices propagate through the mesh by themselves.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F72w2v1z3jr7wbjjz5ycg.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F72w2v1z3jr7wbjjz5ycg.jpg" alt="Three Devices Is Where People Give Up" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Two things to know before you turn it on, both of which will confuse you at exactly the wrong moment:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Removal propagates too, and it is sticky.&lt;/strong&gt; The docs are blunt about it: “if you manually remove an introduced device or unshare it from a folder, it will be automatically re-added as long as the introducer device is still marked as such.” You cannot kick a device out from the wrong end. Remove it on the introducer.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Never make two devices introducers for each other.&lt;/strong&gt; The docs call this out directly, and the failure mode is exactly what you would guess: “the two devices will be constantly ‘re-introducing’ the removed device to each other.”&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One introducer. On the machine that is on the most. That is all.&lt;/p&gt;

&lt;h2&gt;
  
  
  Best Practices for an Obsidian Vault Specifically
&lt;/h2&gt;

&lt;p&gt;Syncthing does not know it is syncing a note app. Everything in this section is Obsidian-specific damage that a general Syncthing tutorial will never mention.&lt;/p&gt;

&lt;h3&gt;
  
  
  Ignore the workspace files, on every device separately
&lt;/h3&gt;

&lt;p&gt;Add these to the folder’s Ignore Patterns tab:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight properties"&gt;&lt;code&gt;&lt;span class="err"&gt;.obsidian/workspace.json&lt;/span&gt;
&lt;span class="err"&gt;.obsidian/workspace-mobile.json&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These two files store which panes you have open and where your cursor is. They change constantly, on every device, which means they conflict constantly. Nothing else in &lt;code&gt;.obsidian&lt;/code&gt; should be excluded. You want your plugins, hotkeys, and theme to ride along.&lt;/p&gt;

&lt;p&gt;The trap is that &lt;strong&gt;ignore patterns never propagate.&lt;/strong&gt; They live in a &lt;code&gt;.stignore&lt;/code&gt; file in the folder root, and per the &lt;a href="https://docs.syncthing.net/users/ignoring.html" rel="noopener noreferrer"&gt;docs&lt;/a&gt;, “The &lt;code&gt;.stignore&lt;/code&gt; file itself will never be synced to other devices.” Setting them on your laptop does precisely nothing for your phone. You have to do it on every device in the mesh, by hand. I covered this same trap from the Android side in &lt;a href="https://dev.to/eristoddle/how-to-sync-obsidian-on-android-for-free-41m"&gt;how to sync Obsidian on Android for free&lt;/a&gt;, and it is the single most common reason a working Syncthing setup starts producing conflict files out of nowhere.&lt;/p&gt;

&lt;h3&gt;
  
  
  Turn on versioning, and know where it puts things
&lt;/h3&gt;

&lt;p&gt;Syncthing can keep old copies of anything it overwrites or deletes, which is the closest thing you get to an undo button when a sync goes sideways. There are four types, and the &lt;a href="https://docs.syncthing.net/users/versioning.html" rel="noopener noreferrer"&gt;docs name them&lt;/a&gt; as Trash Can File Versioning, Simple File Versioning, Staggered File Versioning, and External File Versioning.&lt;/p&gt;

&lt;p&gt;For a vault, use &lt;strong&gt;Staggered File Versioning&lt;/strong&gt;. It thins versions out over time on its own instead of either keeping a fixed count or keeping every copy until your disk fills.&lt;/p&gt;

&lt;p&gt;Three of those four write into a &lt;code&gt;.stversions&lt;/code&gt; folder &lt;strong&gt;inside your vault&lt;/strong&gt;. External is the exception, since it hands the job to a command you supply.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo3ncxxtzfr78us3mlahv.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fo3ncxxtzfr78us3mlahv.jpg" alt="Turn on versioning, and know where it puts things" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If your first thought was that a folder full of old &lt;code&gt;.md&lt;/code&gt; files inside the vault is going to wreck your search results, that was mine too, and it does not. Obsidian filters out any path starting with a dot before the file ever reaches the vault index, which is the same reason &lt;code&gt;.obsidian&lt;/code&gt; and &lt;code&gt;.trash&lt;/code&gt; never show up in the file explorer. Your versions are invisible to search, to the graph, and to backlinks. You do not need to add anything to Excluded Files.&lt;/p&gt;

&lt;p&gt;What it does eat is disk. If you would rather keep the versions somewhere else entirely, Syncthing lets you move them: the folder’s versioning settings have a path option that “overrides the path where old versions of files are stored and defaults to &lt;code&gt;.stversions&lt;/code&gt; if left empty.”&lt;/p&gt;

&lt;h3&gt;
  
  
  Use Receive Only for the devices you only read on
&lt;/h3&gt;

&lt;p&gt;If a device only ever displays notes and never edits them, set that folder to Receive Only on that device. It cannot push a change back, so it cannot start a conflict. This is a good setting for a work machine or a media box, and a bad setting for anything you actually type on.&lt;/p&gt;

&lt;h3&gt;
  
  
  Never run two sync systems on one vault
&lt;/h3&gt;

&lt;p&gt;This applies to every method and it eats more vaults than anything else on this list. Syncthing plus Dropbox on the same folder means two systems writing the same files on their own schedules, neither aware the other exists. Pick one.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a Sync Conflict Actually Looks Like
&lt;/h2&gt;

&lt;p&gt;Eventually you will get one, and Syncthing’s handling of it is more sensible than most.&lt;/p&gt;

&lt;p&gt;When the same file changes on two devices and the contents differ, Syncthing does not overwrite anything. It renames the loser to this pattern:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;&amp;lt;filename&amp;gt;.sync-conflict-&amp;lt;date&amp;gt;-&amp;lt;time&amp;gt;-&amp;lt;modifiedBy&amp;gt;.&amp;lt;ext&amp;gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So &lt;code&gt;daily-note.md&lt;/code&gt; edited in two places becomes &lt;code&gt;daily-note.md&lt;/code&gt; plus something like &lt;code&gt;daily-note.sync-conflict-20260910-071455-KJ3M7QP.md&lt;/code&gt;. Nothing is lost. Both versions are sitting there in your vault.&lt;/p&gt;

&lt;p&gt;Which one keeps the original name comes down to timestamps. From the docs: “The file with the older modification time will be marked as the conflicting file and thus be renamed.” Newer wins the filename, older gets the long name. And conflict files propagate to every device, because the conflict is detected locally but everybody needs to know about it.&lt;/p&gt;

&lt;p&gt;Find them with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;find /path/to/vault &lt;span class="nt"&gt;-name&lt;/span&gt; &lt;span class="s2"&gt;"*.sync-conflict-*"&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run that occasionally. Conflict files are valid markdown, so Obsidian indexes them exactly like real notes, and a vault that has been generating them for six months is a vault where search results have started coming in pairs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Is There a Syncthing Plugin for Obsidian?
&lt;/h2&gt;

&lt;p&gt;Yes.&lt;/p&gt;

&lt;p&gt;It is &lt;a href="https://github.com/LBF38/obsidian-syncthing-integration" rel="noopener noreferrer"&gt;Syncthing Integration&lt;/a&gt; by LBF38, it is in the community plugin browser, and version 2.4.0 landed in March 2026. Install it the &lt;a href="https://dev.to/eristoddle/how-to-install-activate-and-update-obsidian-plugins-4d2p"&gt;usual way&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc2acikkxlq8unakjtr26.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fc2acikkxlq8unakjtr26.jpg" alt="Is There a Syncthing Plugin for Obsidian?" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It does not sync anything.&lt;/strong&gt; That surprises people, so read it again. The plugin requires a Syncthing instance already installed and running on the device, and it talks to it over Syncthing’s REST API. Syncthing does the syncing. The plugin is a control panel.&lt;/p&gt;

&lt;p&gt;What it actually gives you is a conflict resolution interface, a diff view for conflicting files, and a sync status readout in the status bar. Which, given the previous section, is genuinely worth having. Resolving conflicts by opening two files and eyeballing them is the worst part of running Syncthing, and this is a UI for exactly that.&lt;/p&gt;

&lt;p&gt;The README volunteers its own limitations, which I always take as a good sign: “🚧 This plugin is still in development. The configuration might not yet be fully available. 🚧” and “Please backup your vault and use this plugin wisely.” An integrated configuration panel and ignore-pattern management are both listed as coming soon rather than done.&lt;/p&gt;

&lt;p&gt;So: install Syncthing for sync, install the plugin for conflicts. Not the other way around, and not the plugin alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  Phones Are a Separate Problem
&lt;/h2&gt;

&lt;p&gt;Both mobile platforms have a wrinkle, and both have their own post in this series, so here is the short version.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;On iOS&lt;/strong&gt; , Syncthing does not run natively. You run a wrapper app. &lt;a href="https://apps.apple.com/us/app/synctrain/id6553985316" rel="noopener noreferrer"&gt;SyncTrain&lt;/a&gt; is free, open source, has no in-app purchases at all, and needs iOS 17 or later. &lt;a href="https://mobiussync.com/" rel="noopener noreferrer"&gt;Möbius Sync&lt;/a&gt; is free up to 20MB inside its sandbox, and lifting that cap is a one-time $4.99 purchase, which any vault with images will hit immediately. Start with SyncTrain. I compared both properly in &lt;a href="https://dev.to/eristoddle/how-to-sync-obsidian-on-iphone-and-ipad-for-free-in-2026-427p"&gt;syncing Obsidian on iPhone and iPad for free&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;On Android&lt;/strong&gt; , the official Syncthing app was discontinued, with its last release tied to the December 2024 version of Syncthing. If a tutorial tells you to install “Syncthing” from the Play Store, that tutorial is pointing you at an abandoned app. The maintained successor is Syncthing-Fork, which changed maintainers in 2026. The whole story, including which build to install and why F-Droid is the right source for it, is in &lt;a href="https://dev.to/eristoddle/how-to-sync-obsidian-on-android-for-free-41m"&gt;how to sync Obsidian on Android for free&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Reason People Quit
&lt;/h2&gt;

&lt;p&gt;Now the thing I have been circling since the second paragraph.&lt;/p&gt;

&lt;p&gt;Two devices sync when both devices are awake and reachable at the same time. There is nothing in the middle holding your changes until the other end wakes up, because you specifically chose the option that has nothing in the middle.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjgzgugogx1ic33d1rf98.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjgzgugogx1ic33d1rf98.jpg" alt="The Reason People Quit" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Play that out on a normal day. You write on your laptop in the morning and shut it. You open your desktop that evening. The laptop is closed in a bag. Nothing syncs, because there is nobody to sync with. You open your laptop three days later and it finally catches up, except you also wrote on the desktop in the meantime, and now you are looking at a &lt;code&gt;.sync-conflict-&lt;/code&gt; file and wondering what you did wrong.&lt;/p&gt;

&lt;p&gt;You did nothing wrong. This is what a mesh with no always-on member does.&lt;/p&gt;

&lt;p&gt;Even the Obsidian plugin’s README says it out loud: “The synchronization is done in real-time, using peer-to-peer connections. Therefore, all the devices you want to synchronize must be connected at the same time.”&lt;/p&gt;

&lt;p&gt;This is why forum threads describe Syncthing as flaky when it is working perfectly. And it is why I run Remotely Save against Dropbox instead. Dropbox is a computer that is awake all the time. That is the only thing I am paying it for.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fix It With One Always-On Node
&lt;/h2&gt;

&lt;p&gt;The fix is to put something in the mesh that never sleeps. Not a service, a device you own.&lt;/p&gt;

&lt;p&gt;A Raspberry Pi is the classic answer and it is enough. Any always-on machine works: a NAS, a home server, an old laptop with the lid closed and sleep disabled, a $5 VPS. Install Syncthing on it, share the vault with it, mark it as your introducer while you are in there, and the whole problem disappears. Every device now syncs to something that is always home, and the other devices catch up whenever they show up.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One important disambiguation, because the docs use a confusing word.&lt;/strong&gt; Syncthing has public “relay servers,” and they are &lt;em&gt;not&lt;/em&gt; this. Relays forward encrypted traffic between two devices that cannot reach each other directly through NAT. They do not store anything, and per the &lt;a href="https://docs.syncthing.net/users/faq.html" rel="noopener noreferrer"&gt;FAQ&lt;/a&gt;, “Relays do not and can not see the data transmitted via them.” A relay does not hold your changes while your laptop is in a bag. Only a node you run does that.&lt;/p&gt;

&lt;p&gt;If your always-on node is a rented VPS, there is one more trick worth knowing. Syncthing supports &lt;a href="https://docs.syncthing.net/users/untrusted.html" rel="noopener noreferrer"&gt;untrusted devices&lt;/a&gt;: you set a password on the folder share for that device, set the folder type to Receive Encrypted on the VPS itself, and it stores your vault as ciphertext. It cannot read your file contents, filenames, metadata, or directory structure. What is still visible, in the docs’ own words, is “Folder ID and label, File sizes.”&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx4mvcxttd5gf2qzhajrv.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx4mvcxttd5gf2qzhajrv.jpg" alt="Fix It With One Always-On Node" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So you get an always-on relay point on hardware you do not physically control, holding notes it cannot read. That is a legitimately great setup for the person who picked Syncthing for privacy reasons in the first place. The docs still describe the feature as being in testing, so keep a backup that does not depend on it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Is there a Syncthing plugin for Obsidian?
&lt;/h3&gt;

&lt;p&gt;Yes, Syncthing Integration by LBF38. It does not perform the sync. It requires Syncthing to be installed and running separately, and it adds conflict resolution, file diffing, and a status bar readout inside Obsidian.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does Syncthing work with Obsidian on iOS?
&lt;/h3&gt;

&lt;p&gt;Yes, through a wrapper app. SyncTrain is free and open source and needs iOS 17 or later. Möbius Sync is free up to 20MB and $4.99 once after that.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is Syncthing free for Obsidian?
&lt;/h3&gt;

&lt;p&gt;On desktop and Android, completely. It is open source with no accounts, no storage limits, and no paid tier. On iOS it depends on which wrapper you pick, and SyncTrain is genuinely free.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why is my Obsidian vault not syncing with Syncthing?
&lt;/h3&gt;

&lt;p&gt;Nine times out of ten, both devices were not awake at the same time. Syncthing has no server holding changes for you. The other common causes are ignore patterns set on only one device, or a folder set to Receive Only on the device you are editing on.&lt;/p&gt;

&lt;h3&gt;
  
  
  What should I add to Syncthing ignore patterns for Obsidian?
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;.obsidian/workspace.json&lt;/code&gt; and &lt;code&gt;.obsidian/workspace-mobile.json&lt;/code&gt;, set on every device individually. Leave the rest of &lt;code&gt;.obsidian&lt;/code&gt; alone so plugins and settings sync.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is Syncthing better than Obsidian Sync?
&lt;/h3&gt;

&lt;p&gt;It is free and private, and Obsidian Sync is neither. Obsidian Sync is a server that is always awake, first-party support, and no setup. Syncthing costs you a device that never sleeps and an afternoon. Pick based on which of those you have more of.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Verdict
&lt;/h2&gt;

&lt;p&gt;Syncthing is the best free Obsidian sync, and the qualifier that everyone leaves off is that it is the best free Obsidian sync &lt;em&gt;if you own something that stays on&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;With a Pi or a NAS in the mesh, it beats every other method in this series on price, privacy, and speed, and it is not close. Without one, it is a tool that works brilliantly in a two-laptop demo and then generates conflict files for a month while you decide sync software is unreliable.&lt;/p&gt;

&lt;p&gt;I do not have that always-on node running yet. I have a Dropbox account that does the same job, badly, for money I was spending anyway, and an iPad that made the decision easy at the time. Writing this post has more or less talked me into the Pi, which was not the plan when I opened the editor.&lt;/p&gt;

&lt;p&gt;Set up the two desktops first. Add the node before you add the phone. And go check for conflict files right now, because you have some.&lt;/p&gt;

</description>
      <category>obsidian</category>
      <category>syncthing</category>
      <category>sync</category>
      <category>ios</category>
    </item>
    <item>
      <title>Firecrawl CLI Setup: Skip the Installer That Rewrites Your Editors</title>
      <dc:creator>Stephan Miller</dc:creator>
      <pubDate>Wed, 09 Sep 2026 12:00:00 +0000</pubDate>
      <link>https://dev.to/eristoddle/firecrawl-cli-setup-skip-the-installer-that-rewrites-your-editors-2klj</link>
      <guid>https://dev.to/eristoddle/firecrawl-cli-setup-skip-the-installer-that-rewrites-your-editors-2klj</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F24byux7y2whh5k1ea19b.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F24byux7y2whh5k1ea19b.jpg" alt=" " width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;In the &lt;a href="https://dev.to/eristoddle/what-to-do-when-your-ai-coding-agent-cant-read-a-web-page-2f0i"&gt;last post&lt;/a&gt; I gave the Firecrawl setup three sentences and moved on, because the post was about the fetch skill I was building and not about installation.&lt;/p&gt;

&lt;p&gt;The install itself is one command. Everything around the install is what you have to watch: an installer that wants to configure every editor on your machine, an auth screen that reports two contradictory things in consecutive lines, and a shell issue that cost me an API key I could only copy once. Plus one fork in the road worth deciding before you start, which is whether the tool should run as a command or as a resident MCP server.&lt;/p&gt;

&lt;p&gt;So here is the setup told properly, for somebody who is installing &lt;a href="https://firecrawl.link/stephan-miller" rel="noopener noreferrer"&gt;Firecrawl&lt;/a&gt; right now and would rather not spend the time I spent. Everything below was run on this machine against CLI version 1.23.3.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Install With npm, Not With the Installer&lt;/li&gt;
&lt;li&gt;CLI or MCP: Pick the One That Stays Out of the Way&lt;/li&gt;
&lt;li&gt;The Auth Screen That Contradicts Itself&lt;/li&gt;
&lt;li&gt;
The Almost-Right Command

&lt;ul&gt;
&lt;li&gt;If you are on fish&lt;/li&gt;
&lt;li&gt;The afternoon the almost-right command cost me&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Prove It Works Before You Trust It&lt;/li&gt;
&lt;li&gt;The Setup Summary&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Install With npm, Not With the Installer
&lt;/h2&gt;

&lt;p&gt;The command you want:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; firecrawl-cli

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the whole install. You get a &lt;code&gt;firecrawl&lt;/code&gt; binary on your &lt;code&gt;PATH&lt;/code&gt;, symlinked into your global node modules like any other npm package, and nothing else on your machine changes.&lt;/p&gt;

&lt;p&gt;The docs will also offer you a curl-piped-to-shell one-liner, and the CLI ships a &lt;code&gt;firecrawl init&lt;/code&gt; command. Both do considerably more than install a binary. Here is &lt;code&gt;init&lt;/code&gt; describing itself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Usage: firecrawl init [options] [template]

Set up Firecrawl: install CLI, authenticate, add integrations, and scaffold a
template

Arguments:
  ...

Options:
  --all Explicitly install skills to all detected agents (default
                       unless --agent is used)
  -y, --yes Run init non-interactively; skills still install globally
                       across all detected agents unless --agent is used
  -g, --global Install skills globally (user-level, default)
  -a, --agent &amp;lt;agent&amp;gt; Install skills to a specific agent
  -k, --api-key &amp;lt;key&amp;gt; Authenticate with this API key (skips interactive login)
  --skip-install Skip global CLI installation
  --skip-auth Skip authentication
  --skip-skills Skip skills installation

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Read the &lt;code&gt;-y&lt;/code&gt; line again. Skills still install globally across all detected agents. The non-interactive flag does not mean “do less.” It means “do all of it without asking.”&lt;/p&gt;

&lt;p&gt;And &lt;code&gt;init&lt;/code&gt; is not alone. There are four separate commands in this CLI whose actual job is writing Firecrawl into tools that are not Firecrawl:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkpmy1dycpu8u8nv118pq.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkpmy1dycpu8u8nv118pq.jpg" alt="Install With npm, Not With the Installer" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Command&lt;/th&gt;
&lt;th&gt;What it does to your machine&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;init&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Installs the CLI, authenticates, and pushes agent skills to every detected agent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;setup&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Installs &lt;code&gt;skills&lt;/code&gt;, &lt;code&gt;workflows&lt;/code&gt;, &lt;code&gt;mcp&lt;/code&gt;, or &lt;code&gt;defaults&lt;/code&gt; individually&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;make default&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Makes Firecrawl the default provider for supported workflows&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;launch&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Installs the Firecrawl MCP server into an agent, then starts that agent&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The one that gave me pause is &lt;code&gt;setup&lt;/code&gt;, and specifically its undo flag, which is documented as “Undo setup defaults by re-enabling native web tools where supported.” Read that again. The undo re-enables your native web tools, which means the thing being undone disabled them. A scraping vendor’s installer, in its default path, turns off the web fetching your coding agent already had.&lt;/p&gt;

&lt;p&gt;I guess that is the reasonable end state of a competitive market where the install experience is a growth channel, and Firecrawl is far from the only tool doing it. It is also the thing I do not want, because I would like to be the one who decides what is in their context.&lt;/p&gt;

&lt;p&gt;Skip all four. One npm install does the job.&lt;/p&gt;

&lt;p&gt;You can confirm you got the right outcome:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;which firecrawl
&lt;span class="c"&gt;# /opt/homebrew/bin/firecrawl&lt;/span&gt;

firecrawl &lt;span class="nt"&gt;--version&lt;/span&gt;
&lt;span class="c"&gt;# 1.23.3&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  CLI or MCP: Pick the One That Stays Out of the Way
&lt;/h2&gt;

&lt;p&gt;I run Firecrawl through the CLI rather than its MCP server. This is the same verdict I landed on with Playwright.&lt;/p&gt;

&lt;p&gt;An MCP server is resident. Its tool definitions sit in the agent’s context on every single request, whether or not this particular request has anything to do with scraping. That is a fixed tax paid in tokens and in attention, on every turn, for a capability I use maybe twice a week.&lt;/p&gt;

&lt;p&gt;A CLI is not resident. The agent runs a command, the output goes to a file, and between calls the tool occupies exactly nothing. When I need Firecrawl, a skill tells the agent the command to run. When I do not, there is no evidence in the context that Firecrawl exists.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;launch&lt;/code&gt; command exists specifically to install the MCP server into your agent and then start it, and it knows about &lt;code&gt;claude&lt;/code&gt;, &lt;code&gt;code/vscode&lt;/code&gt;, &lt;code&gt;codex&lt;/code&gt;, &lt;code&gt;codex-app&lt;/code&gt;, &lt;code&gt;hermes&lt;/code&gt;, &lt;code&gt;openclaw&lt;/code&gt;, and &lt;code&gt;opencode&lt;/code&gt;. That is a well-built feature and I understand why people want it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdifc2a0fuurterh5kd6k.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdifc2a0fuurterh5kd6k.jpg" alt="CLI or MCP: Pick the One That Stays Out of the Way" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Here is the counterargument, because there is one. An MCP server means the agent discovers the capability on its own. It sees a scraping tool in its tool list and reaches for it without being told. With the CLI approach, something has to tell the agent that Firecrawl is an option, which in my case is a skill I had to write. If you do not want to write that glue, the MCP server is genuinely the faster road.&lt;/p&gt;

&lt;p&gt;I already had the glue. So for me the CLI wins on the only part I care about, which is that a background tool beats one that takes over the screen.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Auth Screen That Contradicts Itself
&lt;/h2&gt;

&lt;p&gt;Get your API key from the dashboard, then check your configuration. This is what &lt;code&gt;firecrawl view-config&lt;/code&gt; printed for me:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;┌─────────────────────────────────────────┐
│ Firecrawl Configuration │
└─────────────────────────────────────────┘

Status: ✓ Authenticated

API Key: Not set
API URL: https://api.firecrawl.dev
Config: /Users/&amp;lt;you&amp;gt;/Library/Application Support/firecrawl-cli

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Authenticated, with a check mark. API key not set. Two lines apart, both delivered with total confidence.&lt;/p&gt;

&lt;p&gt;I spent real time on this, and the answer turns out to be that neither line is lying. They are answering two different questions but forgot to label which is which.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;Status:&lt;/code&gt; reflects live authentication state, including the &lt;code&gt;FIRECRAWL_API_KEY&lt;/code&gt; environment variable. &lt;code&gt;API Key:&lt;/code&gt; reflects only a key stored in the CLI’s own config file, which is empty when you authenticate through the environment instead of through &lt;code&gt;firecrawl config&lt;/code&gt;. Two auth paths, two readouts, stacked adjacently with no indication that they are reading different sources.&lt;/p&gt;

&lt;p&gt;You can prove it in one command. Clear the environment variable and run the same thing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;env&lt;/span&gt; &lt;span class="nt"&gt;-u&lt;/span&gt; FIRECRAWL_API_KEY firecrawl view-config


Status: Not authenticated

Run any &lt;span class="nb"&gt;command &lt;/span&gt;to start authentication, or use:
  firecrawl config Authenticate with browser or API key

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The status flipped, which means the status was reading the environment variable all along.&lt;/p&gt;

&lt;p&gt;The command you actually want is &lt;code&gt;firecrawl --status&lt;/code&gt;, which does not have this problem because it names its source:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;  🔥 firecrawl cli v1.23.3

  ● Authenticated via FIRECRAWL_API_KEY
  Concurrency: 0/2 jobs (parallel scrape limit)
  Credits: 1,324 / 1,000 (132% left this cycle)
  .firecrawl: not found - no local cache
  .gitignore: missing - add .firecrawl/ to ignore cache

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;“Authenticated via FIRECRAWL_API_KEY.” That is the whole fix. One flag, and the ambiguity disappears. Use &lt;code&gt;--status&lt;/code&gt; and forget &lt;code&gt;view-config&lt;/code&gt; exists.&lt;/p&gt;

&lt;p&gt;(Yes, it says I have 132% of my credits left after running some crawls this month. I am choosing to accept this gift and not ask questions.)&lt;/p&gt;

&lt;h2&gt;
  
  
  The Almost-Right Command
&lt;/h2&gt;

&lt;p&gt;Now the part that actually cost me something.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F84fr3qsqb9v8nfzaxcrb.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F84fr3qsqb9v8nfzaxcrb.jpg" alt="The Almost-Right Command That Cost Me an API Key" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The key goes in an environment variable called &lt;code&gt;FIRECRAWL_API_KEY&lt;/code&gt;. On bash or zsh, that means one line in your shell’s startup file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# ~/.bashrc on bash, ~/.zshrc on zsh&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;FIRECRAWL_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;fc-your-key-here

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then open a new terminal, or run &lt;code&gt;source ~/.bashrc&lt;/code&gt; in the one you already have. The export does not reach backwards into shells that were already running.&lt;/p&gt;

&lt;p&gt;Here is the trap, and it is not specific to any one shell. Typing &lt;code&gt;export FIRECRAWL_API_KEY=fc-...&lt;/code&gt; straight at your prompt instead of putting it in the startup file works perfectly for the rest of that session and then evaporates when you close the window. Everything you test in the next five minutes succeeds. That is the worst possible failure mode, because success is exactly what convinces you to stop paying attention and move on.&lt;/p&gt;

&lt;h3&gt;
  
  
  If you are on fish
&lt;/h3&gt;

&lt;p&gt;This is where it bit me, and fish makes the hole a little easier to fall into, because it has four variable scopes instead of one and the two that matter here differ by a single letter:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Wrong. Dies when you close the terminal.&lt;/span&gt;
&lt;span class="nb"&gt;set&lt;/span&gt; &lt;span class="nt"&gt;-x&lt;/span&gt; FIRECRAWL_API_KEY fc-your-key-here

&lt;span class="c"&gt;# Right. Persists across sessions.&lt;/span&gt;
&lt;span class="nb"&gt;set&lt;/span&gt; &lt;span class="nt"&gt;-Ux&lt;/span&gt; FIRECRAWL_API_KEY fc-your-key-here

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;-x&lt;/code&gt; exports it to child processes. &lt;code&gt;-U&lt;/code&gt; makes it universal, which in fish means it survives the terminal closing. You need both.&lt;/p&gt;

&lt;h3&gt;
  
  
  The afternoon the almost-right command cost me
&lt;/h3&gt;

&lt;p&gt;I asked an AI assistant for the fish command to set an environment variable permanently, and got back &lt;code&gt;set -x&lt;/code&gt;. I closed the terminal, came back, was not authenticated, asked again, phrased it more emphatically, and got back &lt;code&gt;set -x&lt;/code&gt; again. Somewhere in there I closed the tab with the API key on it.&lt;/p&gt;

&lt;p&gt;I had not saved the key anywhere else, because I had just watched it get set successfully and had no reason to think I would need it again. So I rotated it and started over, which took the whole thing from a five-minute setup to something dumber. Copy the key into a password manager before you touch your shell config, and none of this can happen to you.&lt;/p&gt;

&lt;p&gt;The generalizable lesson has nothing to do with fish. AI-assisted setup fails in a specific and nasty shape: it hands you the almost-right command. A wrong command errors out immediately and you fix it in ten seconds. An almost-right command succeeds, validates cleanly, and fails an hour later, after you have thrown away the thing you would need to recover. &lt;code&gt;set -x&lt;/code&gt; is not a wrong answer to “set an environment variable in fish.” It is a wrong answer to “set it permanently,” and the word doing the work is the one the model dropped. Bash has the identical hole with a bare &lt;code&gt;export&lt;/code&gt; at the prompt.&lt;/p&gt;

&lt;p&gt;Whatever shell you are on, open a fresh terminal and confirm the thing you just set is still there before you trust it. That is the entire takeaway and it applies well beyond this tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prove It Works Before You Trust It
&lt;/h2&gt;

&lt;p&gt;Do not skip this. The whole point of the previous section is that a broken setup can look fine.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft3pzivywx6yzugtpmmos.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft3pzivywx6yzugtpmmos.jpg" alt="Prove It Works Before You Trust It" width="800" height="447"&gt;&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;firecrawl scrape https://example.com &lt;span class="nt"&gt;--only-main-content&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; smoke.md &lt;span class="nt"&gt;--timing&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;-o&lt;/code&gt; flag is not optional in spirit. Without it the scraped page goes to stdout, and if you happen to smoke test something larger than &lt;code&gt;example.com&lt;/code&gt; you will dump an entire article into your terminal. The &lt;code&gt;--timing&lt;/code&gt; flag gives you a status line worth seeing.&lt;/p&gt;

&lt;p&gt;A working setup looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;Timing:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://example.com"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"requestTime"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-09-08T23:59:53.881Z"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"duration"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"390ms"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"success"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;Scrape&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;ID:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;01&lt;/span&gt;&lt;span class="err"&gt;a&lt;/span&gt;&lt;span class="mi"&gt;08376&lt;/span&gt;&lt;span class="err"&gt;-c&lt;/span&gt;&lt;span class="mi"&gt;56&lt;/span&gt;&lt;span class="err"&gt;f&lt;/span&gt;&lt;span class="mi"&gt;-73&lt;/span&gt;&lt;span class="err"&gt;bf-bc&lt;/span&gt;&lt;span class="mi"&gt;28-411202792&lt;/span&gt;&lt;span class="err"&gt;b&lt;/span&gt;&lt;span class="mi"&gt;72&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And &lt;code&gt;smoke.md&lt;/code&gt; contains actual markdown:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gh"&gt;# Example Domain&lt;/span&gt;

This domain is for use in documentation examples without needing permission. Avoid use in operations.

&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;Learn more&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="sx"&gt;https://iana.org/domains/example&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A broken key looks like this, and thankfully it is loud rather than silent:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Error: Unauthorized: Invalid token

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One more thing while you are here. Scrapes with multiple URLs get cached into a &lt;code&gt;.firecrawl/&lt;/code&gt; directory in your working directory, and the &lt;code&gt;--status&lt;/code&gt; output nags you about this every time, whether or not you are anywhere near a git repo. When you are in one, listen to it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;".firecrawl/"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; .gitignore

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Committing a cache directory full of scraped pages is a bad afternoon waiting to happen, and it is the kind of thing you only notice three weeks later in a diff.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Setup Summary
&lt;/h2&gt;

&lt;p&gt;Stripped of everything above, the working setup is four lines:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; firecrawl-cli
&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s1"&gt;'export FIRECRAWL_API_KEY=fc-your-key-here'&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; ~/.bashrc &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;source&lt;/span&gt; ~/.bashrc
firecrawl &lt;span class="nt"&gt;--status&lt;/span&gt;
firecrawl scrape https://example.com &lt;span class="nt"&gt;--only-main-content&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; smoke.md

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On zsh, swap &lt;code&gt;~/.bashrc&lt;/code&gt; for &lt;code&gt;~/.zshrc&lt;/code&gt;. On fish, line two is &lt;code&gt;set -Ux FIRECRAWL_API_KEY fc-your-key-here&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Install, key, verify, prove. Nothing writes into your editors, nothing sits in your agent’s context, and you know it works because you watched it work.&lt;/p&gt;

&lt;p&gt;The mess I spent an afternoon in was not really Firecrawl’s mess. It was two things that are now true of most developer tools: the installer has become a growth channel, so its default path is maximal rather than minimal, and the auth surface has grown enough paths that the status readout can no longer say one clear thing. Both of those will be true of the next tool you install too. The defense is the same either way, which is to install the binary, wire it up yourself, and run one command that proves it before you build anything on top of it.&lt;/p&gt;

&lt;p&gt;Next in this series I start actually using the thing, which means the first real recipe. This post exists so that one does not have to open with a paragraph about environment variables.&lt;/p&gt;

</description>
      <category>firecrawl</category>
    </item>
    <item>
      <title>GPT-6 Astra and Fable 5.1 Landed the Same Week. Read the Footnote.</title>
      <dc:creator>Stephan Miller</dc:creator>
      <pubDate>Tue, 08 Sep 2026 13:00:00 +0000</pubDate>
      <link>https://dev.to/eristoddle/gpt-6-astra-and-fable-51-landed-the-same-week-read-the-footnote-29j5</link>
      <guid>https://dev.to/eristoddle/gpt-6-astra-and-fable-51-landed-the-same-week-read-the-footnote-29j5</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvkuyviowqasvzmyfwho9.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvkuyviowqasvzmyfwho9.jpg" alt=" " width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For about a month now I’ve been writing the same obituary. American AI lab ships nothing. Google ships a fourth Flash instead of the model it actually promised. Cheap Chinese open-weight models quietly eat the frontier from the bottom while everyone waits for a flagship that keeps slipping. I had the funeral playlist queued up.&lt;/p&gt;

&lt;p&gt;Then over three days the corpse sat up and shipped two flagships. Anthropic dropped Claude Fable 5.1 on September 1. OpenAI dropped GPT-6 Astra on September 3. They landed tied for the top of the intelligence charts, both priced at a flat ten dollars in and fifty out, and just like that the premium ceiling the cheap models spent all summer erasing was back.&lt;/p&gt;

&lt;p&gt;And here’s the part I didn’t see coming. In the exact same week, down at the other end of the market where nobody films the keynote, the cheapest-good-model crown finally changed hands. Both ends of the board moved at once. Let me walk you through it, and then give you the table you actually came here for.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The West remembered it makes models&lt;/li&gt;
&lt;li&gt;Meanwhile, down in the bargain bin, a crown changed hands&lt;/li&gt;
&lt;li&gt;The cheapskate picks&lt;/li&gt;
&lt;li&gt;Google shipped another Flash while the real thing pretrains&lt;/li&gt;
&lt;li&gt;What is coming&lt;/li&gt;
&lt;li&gt;The honest version&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The West remembered it makes models
&lt;/h2&gt;

&lt;p&gt;Start with Fable 5.1, because Anthropic set the tone.&lt;/p&gt;

&lt;p&gt;This is the follow-up to Fable 5, and Anthropic did the thing where they beat their own more expensive model with it. Fable 5.1 finishes ahead of Opus 5 on every category they published. Terminal-Bench-Science jumped to 52.6 percent against Opus 5’s 29.0. Their knowledge-work benchmark, GDPval-AA v2, went to 1853 against Opus 5’s 1824. Same sticker as before, ten and fifty per million, but cache reads got cut 75 percent to a quarter per million tokens, which matters more than it sounds if you run anything with a big fixed context.&lt;/p&gt;

&lt;p&gt;On the Artificial Analysis Intelligence Index it lands at number one. And here’s where you have to read the label. Every Fable 5.1 score on that leaderboard is tagged “with fallback.” That means the headline number quietly blends in a weaker model’s answers whenever a safety classifier refuses the real one, at a frequency nobody discloses. Anthropic pulled the same move on the Opus 5 chart back in July. The number is real. The footnote is load-bearing. The number-one model on the board isn’t purely the number-one model.&lt;/p&gt;

&lt;p&gt;Two days later OpenAI answered with GPT-6. Yes, six. GPT-6 Astra, ten and fifty per million, a 1.05 million token context, staged rollout to a handful of orgs first and then the ChatGPT tiers and the API “over the coming days.” It ties Fable 5.1 at the top of the intelligence index. On computer use it does 72.6 percent on OSWorld 2.0 at roughly 47 percent less time per task than GPT-5.6 Sol, which is the stat that actually matters for agent workloads.&lt;/p&gt;

&lt;p&gt;Then it gets weird. Astra “saturates” FrontierMath Tier 4 at 97.6 percent, ARC-AGI-3 at 99.9 percent, and ExploitBench at a clean 100 percent. That last one is an offensive-security benchmark. A frontier model that fully solves the “can you write a working exploit” test. And the 99.9 on ARC is measured “under OpenAI’s provider adapter harness,” which is the kind of phrase you learn to slow down and read twice, because a benchmark run inside the vendor’s own harness is graded homework. To round it out, OpenAI is shipping a program called Daybreak that loosens safeguards for vetted organizations. So the model that aced the exploit test also gets an official channel to relax its guardrails. If that gives you the same feeling the GLM-5.3 “too good at hacking to ship on time” story gave you last month, you’re paying attention.&lt;/p&gt;

&lt;p&gt;The thing to take away isn’t “which one wins.” They’re basically tied and it’ll take Arena weeks to sort them out, because Arena always lags a launch by a couple weeks while votes pile up. The thing to take away is that both American labs shipped their best model in one window at the same premium price, and the summer story about the West being asleep is, for now, dead. With an asterisk on each.&lt;/p&gt;

&lt;h2&gt;
  
  
  Meanwhile, down in the bargain bin, a crown changed hands
&lt;/h2&gt;

&lt;p&gt;Here’s the part I’ve been tracking for six weeks and finally get to close out.&lt;/p&gt;

&lt;p&gt;For six straight roundups the answer to “what is the cheapest model that is actually good” never moved. It was MiMo v2.5 Pro from Xiaomi. Eighty-seven cents per million output, open weights, and it kept turning up as the cheapest model inside the competitive band of four different Arena categories, week after week. It was the anchor. I could set my watch by it.&lt;/p&gt;

&lt;p&gt;Then last week GLM-5.3-Flash from Z.ai showed up cheaper and, on the hard benchmarks, smarter. The only reason I didn’t crown it on the spot was votes. Arena marks a rating “preliminary” until enough people have voted on it, and GLM-5.3-Flash was sitting on a couple thousand shaky votes while MiMo had fifty thousand solid ones. So I called it the emerging pick, kept MiMo as the printed anchor, and wrote down the tripwire: if GLM-5.3-Flash holds cheapest-in-band once the votes firm up, the spine has genuinely shifted.&lt;/p&gt;

&lt;p&gt;This week the votes firmed up. GLM-5.3-Flash is now the cheapest model in the competitive band for Overall, Coding, Instruction Following, and Hard Prompts, on real vote counts, and where it overlaps MiMo it out-rates it. In Overall it sits at 1474 with 4,672 votes for fifty cents per million output. MiMo is at 1468 for eighty-seven cents. Cheaper and higher. The six-week reign is over. MiMo’s the runner-up now, and it’s still a perfectly good runner-up with ten to twenty times the vote count, which is exactly why I keep it in the table.&lt;/p&gt;

&lt;p&gt;The catch, because there’s always a catch. GLM-5.3-Flash scores mid on the hard-reasoning index, a 42 where the frontier models are in the fifties. It’s preference-strong and cheap, not a deep-reasoning machine. It’s also slow, about 60 output tokens a second against a median north of 70, which compounds in agent loops where every step waits on the last. And the cheapest price you’ll see quoted for it, seven and a half cents in and a quarter out, is a launch promo that expires September 9. List is fifteen and fifty cents. Build your budget on the list price, not the promo, unless you enjoy surprises on the tenth.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cheapskate picks
&lt;/h2&gt;

&lt;p&gt;Same method as always. For each Arena category I take the leader’s rating, draw a band 50 points below it, and find the cheapest model that still sits inside that band. The whole premise is that Arena ratings cluster tight at the top, so the category leader is usually only a rounding error better than something 20 to 100 times cheaper. I compute the band from the full table in code, not by eyeballing the first screen, because eyeballing it is how you delete the entire cheap tail and crown the cheapest expensive model by accident. Ask me how I know.&lt;/p&gt;

&lt;p&gt;Bands this week ran 56 deep in Overall, about 63 in Coding, and thin in Math where the whole board is still preliminary. Arena data dated around September 2.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Category&lt;/th&gt;
&lt;th&gt;Leader&lt;/th&gt;
&lt;th&gt;$ leader out&lt;/th&gt;
&lt;th&gt;Cheapskate pick&lt;/th&gt;
&lt;th&gt;$ pick out&lt;/th&gt;
&lt;th&gt;Δ rating&lt;/th&gt;
&lt;th&gt;Cheaper by&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Overall&lt;/td&gt;
&lt;td&gt;claude-fable-5 (1507)&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;GLM-5.3-Flash (1474, #29)&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;−33&lt;/td&gt;
&lt;td&gt;~100×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Coding&lt;/td&gt;
&lt;td&gt;claude-opus-4-7-high (1552)&lt;/td&gt;
&lt;td&gt;$25&lt;/td&gt;
&lt;td&gt;GLM-5.3-Flash (1534, #9)&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;−18&lt;/td&gt;
&lt;td&gt;50×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Creative Writing&lt;/td&gt;
&lt;td&gt;claude-fable-5 (1504)&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;Gemini 3-Flash (1459, #27)&lt;/td&gt;
&lt;td&gt;$3&lt;/td&gt;
&lt;td&gt;−45&lt;/td&gt;
&lt;td&gt;~16.7×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Instruction Following&lt;/td&gt;
&lt;td&gt;claude-opus-4-6-high (1514)&lt;/td&gt;
&lt;td&gt;$25&lt;/td&gt;
&lt;td&gt;GLM-5.3-Flash (1465, #34)&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;−49&lt;/td&gt;
&lt;td&gt;50×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hard Prompts&lt;/td&gt;
&lt;td&gt;claude-opus-4-6-high (1533)&lt;/td&gt;
&lt;td&gt;$25&lt;/td&gt;
&lt;td&gt;GLM-5.3-Flash (1496, #31)&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;−37&lt;/td&gt;
&lt;td&gt;50×&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Math&lt;/td&gt;
&lt;td&gt;claude-fable-5 (1529, prelim)&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;Gemini 3.7-Flash-high (1524, #4)&lt;/td&gt;
&lt;td&gt;$3.75&lt;/td&gt;
&lt;td&gt;−5&lt;/td&gt;
&lt;td&gt;~13×&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A few notes on reading that. The Coding line is the loudest one on the board: GLM-5.3-Flash is ranked ninth in coding, top-ten, for fifty cents per million. The catch there is 1,274 votes, still on the thin side, so if you want certainty over savings, MiMo at rank 26 with 15,813 votes for eighty-seven cents is the steadier bet. In Hard Prompts, GLM-5.3-Flash and MiMo are literally tied at 1496; GLM is cheaper, MiMo has twelve times the votes. Take your pick based on whether you trust the number or the sample size.&lt;/p&gt;

&lt;p&gt;Creative Writing stays a Gemini story because GLM-5.3-Flash never cracked that band, and Math is a low-confidence mess this week, thin preliminary votes top to bottom and no pick under $3.75, so treat that row as a suggestion and not a promise.&lt;/p&gt;

&lt;p&gt;The one thing I won’t do is pretend both value picks are fast. They’re not. GLM-5.3-Flash and MiMo are both slow. If you’re wiring one into an autonomous agent that chains dozens of calls, the fifty-cent price tag can balloon into a fifty-cent-per-call wall-clock tax. Cheap and slow is a real trade, not a free lunch. Know which one your workload cares about.&lt;/p&gt;

&lt;h2&gt;
  
  
  Google shipped another Flash while the real thing pretrains
&lt;/h2&gt;

&lt;p&gt;Quick check-in on Google, who continue to run the strangest release cadence in the business.&lt;/p&gt;

&lt;p&gt;On September 2 they shipped Gemini 3.8 Flash. That’s the third Flash release in six weeks. It costs exactly what 3.7 Flash cost, 75 cents in and $3.75 out, and beats it on every benchmark they published, plus there’s a locked-down 3.8 Flash Cyber sibling for security work. It’s a genuinely good, cheap coding-and-agent workhorse, and it already shows up at number eight overall on Arena.&lt;/p&gt;

&lt;p&gt;But it’s built on the 3.7 base, not a new one, and it’s shipping instead of the Gemini 3.5 Pro they promised back in the spring, which has been quietly shelved after missing so many dates I lost count. The real model, Gemini 4, just “cleared pretraining” with strong preliminary results and the largest training run in Google’s history. No benchmarks, no price, no date, late 2026 if you believe the tea leaves. So Google’s strategy remains: ship a steady drip of excellent Flash models to stay in the headlines while the actual next-generation model bakes in the background. It’s working, in the sense that they’re still in the conversation. It’s also the fourth time this summer I’ve written that exact paragraph.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is coming
&lt;/h2&gt;

&lt;p&gt;Two things worth watching.&lt;/p&gt;

&lt;p&gt;Grok 4.7 is the loud one. Musk said on September 2 it would be out “in ten days,” which points at roughly September 12. It’s a 2.1 trillion parameter model, up 40 percent from Grok 4.6, and part of its training data comes from SpaceX internal engineering records, which is either the most interesting or the most concerning detail depending on your mood. There’s no model card, no price, no benchmark table, and no API id yet, so treat the date as a tweet and not a commitment.&lt;/p&gt;

&lt;p&gt;Gemini 4, covered above, is the other. Pretraining done, everything else unknown.&lt;/p&gt;

&lt;h2&gt;
  
  
  The honest version
&lt;/h2&gt;

&lt;p&gt;The clean narrative would be that the West is back and the story is over. It’s not that clean.&lt;/p&gt;

&lt;p&gt;Yes, Fable 5.1 and GPT-6 Astra are real, and yes they’re good, and yes the premium ceiling exists again. But both of them shipped with a footnote you have to read before you trust the headline. One blends a weaker model into its own benchmark and doesn’t tell you how often. The other aced an exploit-writing test and comes with a program to loosen its own safety rails. That’s not a reason to dismiss them. It’s a reason to read the harness section before you quote the number.&lt;/p&gt;

&lt;p&gt;And the more durable story is the one nobody put on a stage. The cheapest genuinely good model on the board is a fifteen-cent open-weight model from Z.ai that just took a crown a Xiaomi model held for a month and a half. The frontier got a loud, expensive reload this week. The floor got quietly cheaper. If you’re actually shipping something and paying the bill yourself, guess which one changes your life more.&lt;/p&gt;

&lt;p&gt;I’ll be back next week to see whether Grok 4.7 shows up on the twelfth or whether “ten days” means what it usually means.&lt;/p&gt;

</description>
      <category>llm</category>
      <category>openrouter</category>
    </item>
  </channel>
</rss>
