<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Elin</title>
    <description>The latest articles on DEV Community by Elin (@elineve).</description>
    <link>https://dev.to/elineve</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4003682%2Fa50a6283-e704-40cd-bf5e-b95ec2cc5fb7.png</url>
      <title>DEV Community: Elin</title>
      <link>https://dev.to/elineve</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/elineve"/>
    <language>en</language>
    <item>
      <title>Qwen3.8-Max-Preview and Kimi’s Subscription Pause: What the Moves Really Signal</title>
      <dc:creator>Elin</dc:creator>
      <pubDate>Mon, 20 Jul 2026 10:04:18 +0000</pubDate>
      <link>https://dev.to/elineve/qwen38-max-preview-and-kimis-subscription-pause-what-the-moves-really-signal-4o20</link>
      <guid>https://dev.to/elineve/qwen38-max-preview-and-kimis-subscription-pause-what-the-moves-really-signal-4o20</guid>
      <description>&lt;p&gt;The latest AI product news from China is arriving in two very different forms.&lt;/p&gt;

&lt;p&gt;Alibaba has added Qwen3.8-Max-Preview to its Token Plan and related agent products. Moonshot AI, meanwhile, has reportedly paused new personal subscriptions for Kimi after demand placed pressure on available computing capacity.&lt;/p&gt;

&lt;p&gt;At first glance, these look like unrelated product updates. Together, they show how frontier AI companies are turning model releases into distribution and capacity-management decisions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Qwen3.8-Max-Preview is a product launch, not yet a complete model release
&lt;/h2&gt;

&lt;p&gt;Alibaba Cloud’s &lt;a href="https://www.alibabacloud.com/en/campaign/ai-landing-page-token" rel="noopener noreferrer"&gt;official Token Plan page&lt;/a&gt; lists Qwen3.8-Max-Preview as a supported model. The page also places it inside a broader product bundle that includes agent concurrency, image and video models, third-party models, and team plans.&lt;/p&gt;

&lt;p&gt;That framing is important.&lt;/p&gt;

&lt;p&gt;Qwen3.8-Max-Preview is not being presented only as an API endpoint. It is being used to pull users into Alibaba’s broader AI workflow: model access, coding tools, agent execution, and cloud infrastructure.&lt;/p&gt;

&lt;p&gt;Reports also say the preview is available through Qoder and QoderWork. Those distribution channels matter because most users do not evaluate a model in isolation. They experience it through an IDE, a coding assistant, or an agent that can take action.&lt;/p&gt;

&lt;p&gt;Alibaba is also being associated with a claim that the model ranks just behind Claude Fable 5. I would phrase that carefully. At this stage, it should be treated as an Alibaba positioning claim rather than an independently verified leaderboard result.&lt;/p&gt;

&lt;p&gt;A preview model can be promising without having a stable benchmark profile. The practical questions are more ordinary:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Is the model available consistently?&lt;/li&gt;
&lt;li&gt;What are the real input and output limits?&lt;/li&gt;
&lt;li&gt;How does it behave on tool calls?&lt;/li&gt;
&lt;li&gt;Does it preserve context across long agent sessions?&lt;/li&gt;
&lt;li&gt;Are the pricing and quotas stable after the launch promotion ends?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those answers will matter more to developers than a single ranking sentence.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DYWU3YTMxMGM3MDk4NDEyNTAyZWRhOWY0ZjFlYTUwYTFfUVprZDdEYlgxUHNpYW5XendKS3huY0VyU0c3cWRtSUhfVG9rZW46SVAxamJHRXlCbzh5V014dWJFWmNaR3NTbkhkXzE3ODQ1NDE2MzQ6MTc4NDU0NTIzNF9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DYWU3YTMxMGM3MDk4NDEyNTAyZWRhOWY0ZjFlYTUwYTFfUVprZDdEYlgxUHNpYW5XendKS3huY0VyU0c3cWRtSUhfVG9rZW46SVAxamJHRXlCbzh5V014dWJFWmNaR3NTbkhkXzE3ODQ1NDE2MzQ6MTc4NDU0NTIzNF9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="Screenshot from the Alibaba Cloud" width="1080" height="608"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Screenshot from &lt;a href="https://www.alibabacloud.com/en/campaign/qwen-ai-landing-page?_p_lc=1&amp;amp;utm_content=se_1023441556&amp;amp;gclid=EAIaIQobChMI-ZfEgv_glQMVpFoPAh2lfhQAEAAYASAAEgJwKvD_BwE" rel="noopener noreferrer"&gt;Alibaba Cloud&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Qwen is bundling models with agents
&lt;/h2&gt;

&lt;p&gt;Alibaba’s current product direction is becoming easier to read.&lt;/p&gt;

&lt;p&gt;The company is not only competing on model intelligence. It is packaging model access with tools that turn intelligence into work. The Token Plan page advertises concurrent agents and team-oriented plans alongside individual subscriptions.&lt;/p&gt;

&lt;p&gt;That creates a stronger commercial loop:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A new model attracts attention.&lt;/li&gt;
&lt;li&gt;Developers try it through a coding or agent product.&lt;/li&gt;
&lt;li&gt;Usage generates demand for tokens, storage, orchestration, and cloud services.&lt;/li&gt;
&lt;li&gt;The model becomes an entry point into the wider Alibaba Cloud ecosystem.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This is strategically different from releasing a model and waiting for developers to build the surrounding infrastructure themselves.&lt;/p&gt;

&lt;p&gt;It also creates more room for price competition. A user may compare Qwen3.8-Max-Preview with Claude or OpenAI by capability, but the purchasing decision may ultimately depend on included credits, concurrency, regional availability, and how much setup the platform removes.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Kimi’s subscription pause tells us
&lt;/h2&gt;

&lt;p&gt;Moonshot AI’s official Kimi account reported that new subscriptions were temporarily paused after a sharp increase in demand. The company said existing subscribers were not affected and that capacity was being expanded.&lt;/p&gt;

&lt;p&gt;The announcement can be read as a success signal, but it is also a reminder that AI products are constrained by physical infrastructure.&lt;/p&gt;

&lt;p&gt;A model can be technically available and commercially unavailable at the same time. If inference capacity is limited, accepting every new subscriber may degrade the experience for existing users. Pausing new subscriptions is a way to protect service quality while the company adds capacity.&lt;/p&gt;

&lt;p&gt;This does not prove that Kimi is abandoning individual users. Nor does it prove a completed shift from consumer products to enterprise customers.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DZmNmNmU2ODRkMDkzMjgxNmE2NjhhODIwMDBhYzI2YTVfeWNQdnRzOG9FS001VUpDa0d1ZHp6QkJxdUdzZ2hOU2hfVG9rZW46TkQxbmJZMHVtb1RUWHJ4RTdna2MzcldIbjVYXzE3ODQ1NDE2MzQ6MTc4NDU0NTIzNF9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DZmNmNmU2ODRkMDkzMjgxNmE2NjhhODIwMDBhYzI2YTVfeWNQdnRzOG9FS001VUpDa0d1ZHp6QkJxdUdzZ2hOU2hfVG9rZW46TkQxbmJZMHVtb1RUWHJ4RTdna2MzcldIbjVYXzE3ODQ1NDE2MzQ6MTc4NDU0NTIzNF9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="Screenshot from the Kimi Code Page" width="1157" height="666"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Screenshot from &lt;a href="https://www.kimi.com/code/en?track_id=a22e508e-f56d-4eee-bc56-00d6abf3e20d" rel="noopener noreferrer"&gt;Kimi Code Page&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A more cautious interpretation is that Moonshot is trying to balance three demands:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;public visibility from a popular consumer product;&lt;/li&gt;
&lt;li&gt;expensive, high-volume agent and coding usage;&lt;/li&gt;
&lt;li&gt;longer-term business customers who expect predictable service levels.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That balance is becoming difficult for every AI company operating at the frontier.&lt;/p&gt;

&lt;h2&gt;
  
  
  The business model is moving from access to allocation
&lt;/h2&gt;

&lt;p&gt;The interesting change is not simply that subscriptions are becoming more expensive or harder to obtain.&lt;/p&gt;

&lt;p&gt;It is that AI companies are starting to sell controlled access to compute-intensive workflows. Plans increasingly include quotas, concurrent agents, work modes, model routing, and priority capacity.&lt;/p&gt;

&lt;p&gt;The subscription is no longer just a key to a chatbot. It is closer to a reservation for a certain amount of inference infrastructure.&lt;/p&gt;

&lt;p&gt;This is why Qwen’s bundled Token Plan and Kimi’s temporary subscription pause belong in the same conversation. One is expanding distribution through a packaged platform. The other is limiting distribution because demand has arrived faster than capacity.&lt;/p&gt;

&lt;p&gt;I would avoid attaching unrelated Databricks or OpenEvidence financing claims to this article without primary company announcements or filings. Those stories may be relevant to the broader AI infrastructure market, but they need their own evidence trail.&lt;/p&gt;

&lt;p&gt;For developers, the immediate lesson is simple: test the model, but also test the product around it.&lt;/p&gt;

&lt;p&gt;A strong benchmark score will not compensate for unstable quotas, unavailable subscriptions, poor tool integration, or unpredictable billing.&lt;/p&gt;

&lt;p&gt;The next phase of the model race may be decided less by who has the most impressive demo and more by who can reliably allocate intelligence to real work.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>productivity</category>
      <category>tooling</category>
    </item>
    <item>
      <title>Grok Build Is Open Source: What the GitHub Drop Actually Changes</title>
      <dc:creator>Elin</dc:creator>
      <pubDate>Fri, 17 Jul 2026 09:41:00 +0000</pubDate>
      <link>https://dev.to/elineve/grok-build-is-open-source-what-the-github-drop-actually-changes-124o</link>
      <guid>https://dev.to/elineve/grok-build-is-open-source-what-the-github-drop-actually-changes-124o</guid>
      <description>&lt;p&gt;Elon Musk open-sourced Grok Build, and the repository collected 7.7k GitHub Stars almost immediately.&lt;/p&gt;

&lt;p&gt;The first part is directionally true. The second part needs a timestamp.&lt;/p&gt;

&lt;p&gt;As of July 17, 2026, the official repository shows roughly 13.8k Stars. The 7.7k number appears to describe an earlier moment in the launch cycle, not the current total. That distinction matters because AI tooling moves quickly, and a screenshot can become “the fact” long after the number has changed.&lt;/p&gt;

&lt;p&gt;The more interesting story is not the Star count anyway. It is what xAI has chosen to expose.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Grok Build actually is
&lt;/h2&gt;

&lt;p&gt;According to the &lt;a href="https://x.ai/news/grok-build-open-source" rel="noopener noreferrer"&gt;official xAI announcement&lt;/a&gt;, Grok Build is being open-sourced as a coding agent and terminal user interface. The project is also available through an official &lt;a href="https://github.com/xai-org/grok-build" rel="noopener noreferrer"&gt;GitHub repository&lt;/a&gt;, where it is described as SpaceXAI’s terminal-based AI coding agent.&lt;/p&gt;

&lt;p&gt;The repository includes the Rust source for the &lt;code&gt;grok&lt;/code&gt; CLI, its full-screen TUI, and the underlying agent runtime. It can understand a codebase, edit files, execute shell commands, search the web, and manage longer-running tasks.&lt;/p&gt;

&lt;p&gt;That sounds familiar if you have used Claude Code, Codex, or another terminal-native coding agent. The difference is that Grok Build is presenting more of the harness as a public object.&lt;/p&gt;

&lt;p&gt;The model is only one part of an agent system. The harness decides how context is assembled, when tools are called, how edits are shown, how plans are reviewed, and what the agent is allowed to touch. In practice, those decisions often shape the user experience more than the model name printed in the settings menu.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DMWIzZDI0NTNlYTZjZmRlMWY5MzEwNGE3YzMxZmM5MjJfa20xZ090dFY4R1FLcVc5RDZYeDhQQVE0bzBpYUxTYzBfVG9rZW46VUQ2ZGJIbTZ2b0w4c1l4NUZSeGNEaFk2bmRjXzE3ODQyODEwMDA6MTc4NDI4NDYwMF9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DMWIzZDI0NTNlYTZjZmRlMWY5MzEwNGE3YzMxZmM5MjJfa20xZ090dFY4R1FLcVc5RDZYeDhQQVE0bzBpYUxTYzBfVG9rZW46VUQ2ZGJIbTZ2b0w4c1l4NUZSeGNEaFk2bmRjXzE3ODQyODEwMDA6MTc4NDI4NDYwMF9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="AI Coding Agent Terminal | Automated Code Editing Interface" width="1200" height="630"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Why open-sourcing the harness matters
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://x.ai/news/grok-build-open-source" rel="noopener noreferrer"&gt;official announcement&lt;/a&gt; specifically points to the agent loop, tool-call dispatch, terminal UI, and extension system.&lt;/p&gt;

&lt;p&gt;That gives developers something concrete to inspect.&lt;/p&gt;

&lt;p&gt;You can study how the tool reads and edits files. You can examine how shell commands are dispatched. You can see how skills, plugins, hooks, MCP servers, and subagents are loaded. Even if you never contribute a line of code, this is useful because it turns an opaque workflow into something closer to an engineering artifact.&lt;/p&gt;

&lt;p&gt;For me, the most valuable part is the context assembly layer.&lt;/p&gt;

&lt;p&gt;Coding agents do not simply “read the repository.” They build a temporary working view from instructions, file contents, tool results, previous decisions, and sometimes external search. Small changes in that assembly can create very different outcomes. An agent that receives too much context becomes expensive and unfocused. An agent that receives too little confidently edits the wrong file.&lt;/p&gt;

&lt;p&gt;Open source does not automatically solve this problem. It does make the problem easier to investigate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Local-first is not the same as local inference
&lt;/h2&gt;

&lt;p&gt;There is another phrase worth reading carefully: local-first.&lt;/p&gt;

&lt;p&gt;xAI says Grok Build can be compiled locally, connected to a user’s own local inference endpoint, and configured through &lt;code&gt;config.toml&lt;/code&gt;. The &lt;a href="https://docs.x.ai/build/overview" rel="noopener noreferrer"&gt;official documentation&lt;/a&gt; also describes support for custom models and configurable base URLs.&lt;/p&gt;

&lt;p&gt;That means the client and orchestration layer can be run on your machine. It does not mean the default Grok model suddenly runs locally on a laptop.&lt;/p&gt;

&lt;p&gt;This is an important distinction for anyone working with private repositories. A local terminal interface may still send prompts, code context, tool results, or file contents to a remote model provider. Before testing Grok Build on a production repository, I would inspect the configuration, authentication flow, network requests, logging behavior, and any telemetry or session-trace settings.&lt;/p&gt;

&lt;p&gt;“Open source” answers one question: can we inspect the code?&lt;/p&gt;

&lt;p&gt;It does not answer every privacy question.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DNDYxYzJjOWNjZWI4ZTg3MGI1YWRkY2I0YmJiMjgyNzdfd0VXSFV2UU5lekZSa29yc1pDZDRkTWxlU3RsbWFGQTlfVG9rZW46Wk8yaGJHNkE5b3lLeWJ4Y01sR2NKdFcxbkdkXzE3ODQyODEwMDA6MTc4NDI4NDYwMF9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DNDYxYzJjOWNjZWI4ZTg3MGI1YWRkY2I0YmJiMjgyNzdfd0VXSFV2UU5lekZSa29yc1pDZDRkTWxlU3RsbWFGQTlfVG9rZW46Wk8yaGJHNkE5b3lLeWJ4Y01sR2NKdFcxbkdkXzE3ODQyODEwMDA6MTc4NDI4NDYwMF9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="Grok AI Logo 3D Render | xAI Large Language Model Branding" width="1332" height="749"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Photo by &lt;a href="https://unsplash.com/@maria_shalabaieva" rel="noopener noreferrer"&gt;Mariia Shalabaieva&lt;/a&gt; on &lt;a href="https://unsplash.com/" rel="noopener noreferrer"&gt;Unsplash&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The workflow I would test first
&lt;/h2&gt;

&lt;p&gt;I would not begin with an ambitious autonomous build. I would start with a small repository and a controlled task:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Ask Grok Build to explain the repository without editing files.&lt;/li&gt;
&lt;li&gt;Run &lt;code&gt;grok inspect&lt;/code&gt; and check which instructions, skills, plugins, hooks, and MCP servers it discovers.&lt;/li&gt;
&lt;li&gt;Give it a narrow bug fix with an existing test.&lt;/li&gt;
&lt;li&gt;Review the plan before execution.&lt;/li&gt;
&lt;li&gt;Inspect the diff and command history.&lt;/li&gt;
&lt;li&gt;Repeat the task in headless mode using the documented &lt;code&gt;-p&lt;/code&gt; option.&lt;/li&gt;
&lt;li&gt;Compare the interactive and scripted outputs.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This kind of test reveals more than a polished demo. It shows whether the agent preserves project conventions, whether it makes unnecessary changes, and whether the same prompt behaves differently when there is no human approving each step.&lt;/p&gt;

&lt;p&gt;The repository documentation also mentions Agent Client Protocol support. That makes Grok Build potentially useful as a component inside larger coding workflows, but integration support is not the same as interoperability quality. The practical questions are whether sessions can be resumed, how permissions are represented, and whether failures remain understandable when the agent is launched by another tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real significance of the GitHub launch
&lt;/h2&gt;

&lt;p&gt;The early Star count is a useful signal of curiosity, but it is not a quality benchmark. A repository can attract thousands of Stars because developers want to inspect it, not because they are ready to use it in production.&lt;/p&gt;

&lt;p&gt;The stronger signal is that coding-agent companies are beginning to publish the machinery around the model: planning, tools, extensions, worktrees, permissions, and terminal interaction.&lt;/p&gt;

&lt;p&gt;That is where the next layer of competition will happen.&lt;/p&gt;

&lt;p&gt;Models will continue to improve, but developers will judge agents by whether they can work safely inside an existing repository. Can I understand what the agent is about to do? Can I stop it? Can I reproduce the result? Can I swap the model without rebuilding the entire workflow?&lt;/p&gt;

&lt;p&gt;Grok Build’s open-source release does not answer all of these questions yet. It does make them easier to ask in public.&lt;/p&gt;

&lt;p&gt;And honestly, that is more valuable than another “this model feels smarter” demo.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>code</category>
      <category>github</category>
    </item>
    <item>
      <title>GPT-5.6 vs Fable 5: The Bug Fix Is Your Benchmark Harness</title>
      <dc:creator>Elin</dc:creator>
      <pubDate>Tue, 14 Jul 2026 03:40:27 +0000</pubDate>
      <link>https://dev.to/elineve/gpt-56-vs-fable-5-the-bug-fix-is-your-benchmark-harness-3idg</link>
      <guid>https://dev.to/elineve/gpt-56-vs-fable-5-the-bug-fix-is-your-benchmark-harness-3idg</guid>
      <description>&lt;p&gt;Every model launch creates two parallel stories.&lt;/p&gt;

&lt;p&gt;One is the official story: model cards, API docs, release notes, pricing tables, safety notes. Boring, useful, occasionally painful to read.&lt;/p&gt;

&lt;p&gt;The other is the screenshot story: one impossible benchmark, one mysterious “fatal bug,” one claim that a frontier model now runs on a laptop.&lt;/p&gt;

&lt;p&gt;I understand why the second story travels faster. It has better lighting. But if you are deciding whether to move real engineering work onto GPT-5.6 or Claude Fable 5, the entertaining version is not enough.&lt;/p&gt;

&lt;p&gt;As of July 13, 2026, the grounded part is simple. OpenAI’s model guidance describes GPT-5.6 Sol, Terra, and Luna as distinct targets, with Sol positioned for frontier capability, Terra for a balance of intelligence and cost, and Luna for efficient high-volume use. Anthropic describes Claude Fable 5 as its most capable widely released model, while Claude Mythos 5 is limited availability through Project Glasswing.&lt;/p&gt;

&lt;p&gt;So yes, GPT-5.6 Sol and Fable 5 are real comparison targets.&lt;/p&gt;

&lt;p&gt;What is less grounded is the habit of turning every surprising model behavior into a “bug.”&lt;/p&gt;

&lt;p&gt;If someone says GPT-5.6 has a major bug, I want the unromantic details first: exact model ID, request payload, tool settings, reasoning settings, retry count, transcript, and expected output. Without those, “bug” can mean at least four different things: wrong answer, refusal, timeout, configuration drift, or a benchmark harness quietly doing something silly in the background.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DZmQ1ZWQwZDlmZTY2YmI3YmUzY2ZkMmViYjlkMTBkYWNfZGIzQnVRS2s0SkNMRktqdks0VnYxSWJCVDRKeWtMTnFfVG9rZW46WENQaWJQYVIyb0VUcmd4RXJ3V2NWWWpZbk5XXzE3ODQwMDAyODQ6MTc4NDAwMzg4NF9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DZmQ1ZWQwZDlmZTY2YmI3YmUzY2ZkMmViYjlkMTBkYWNfZGIzQnVRS2s0SkNMRktqdks0VnYxSWJCVDRKeWtMTnFfVG9rZW46WENQaWJQYVIyb0VUcmd4RXJ3V2NWWWpZbk5XXzE3ODQwMDAyODQ6MTc4NDAwMzg4NF9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="image" width="760" height="538"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That last one is not rare. It is practically a launch-week tradition.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start With The Harness, Not The Leaderboard
&lt;/h2&gt;

&lt;p&gt;For a GPT-5.6 vs Fable 5 comparison, I would not begin with one giant public score. I would begin with a private regression set that reflects the work I actually care about.&lt;/p&gt;

&lt;p&gt;My first pass would include four buckets:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Debugging tasks with failing tests.&lt;/li&gt;
&lt;li&gt;Refactors where the best answer is smaller code, not more code.&lt;/li&gt;
&lt;li&gt;Product requests where the model should ask a clarifying question.&lt;/li&gt;
&lt;li&gt;Long-context tasks where early assumptions can quietly go stale.
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;pandas&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;


&lt;span class="n"&gt;TASKS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;task_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;debug-001&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;category&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;debugging&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;description&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Fix a function that returns 0 for an empty list. &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;The correct behavior is to raise ValueError.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;task_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;refactor-001&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;category&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;refactoring&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;description&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Remove duplicated validation logic without adding &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;an unnecessary abstraction layer.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;task_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;clarify-001&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;category&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;clarification&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;description&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;The user only says: Add export support. &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;The model should ask clarifying questions.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;task_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;context-001&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;category&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;long_context&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;description&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Update a feature while preserving an important constraint &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;stated near the beginning of a long specification.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;


&lt;span class="c1"&gt;# These are simulated results.
# Replace them with actual API outputs when running a real comparison.
&lt;/span&gt;&lt;span class="n"&gt;RESULTS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GPT-5.6-Sol&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;task_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;debug-001&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;result&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;correct&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;latency_seconds&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;8.2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GPT-5.6-Sol&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;task_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;refactor-001&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;result&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;wrong&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;latency_seconds&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;7.1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GPT-5.6-Sol&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;task_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;clarify-001&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;result&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;correct&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;latency_seconds&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;3.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GPT-5.6-Sol&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;task_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;context-001&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;result&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;correct&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;latency_seconds&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;10.4&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Claude-Fable-5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;task_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;debug-001&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;result&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;correct&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;latency_seconds&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;7.8&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Claude-Fable-5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;task_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;refactor-001&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;result&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;correct&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;latency_seconds&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;6.9&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Claude-Fable-5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;task_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;clarify-001&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;result&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;refused&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;latency_seconds&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;2.8&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Claude-Fable-5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;task_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;context-001&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;result&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;wrong&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;latency_seconds&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;9.7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;create_summary&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;results_frame&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;results&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;summary&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;results_frame&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;groupby&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;agg&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;correct&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;result&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="k"&gt;lambda&lt;/span&gt; &lt;span class="n"&gt;values&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="n"&gt;values&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;correct&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;()),&lt;/span&gt;
        &lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="n"&gt;wrong&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;result&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="k"&gt;lambda&lt;/span&gt; &lt;span class="n"&gt;values&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="n"&gt;values&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;wrong&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;()),&lt;/span&gt;
        &lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="n"&gt;refused&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;result&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="k"&gt;lambda&lt;/span&gt; &lt;span class="n"&gt;values&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="n"&gt;values&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;refused&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;sum&lt;/span&gt;&lt;span class="p"&gt;()),&lt;/span&gt;
        &lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="n"&gt;average_latency_seconds&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;latency_seconds&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;mean&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;summary&lt;/span&gt;


&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;__name__&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;__main__&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;task_frame&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;TASKS&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;model_summary&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;create_summary&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;RESULTS&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Private Regression Tasks&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;task_frame&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;to_string&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;index&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s"&gt;Model Comparison Summary&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model_summary&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If a post claims a “63 hard problems” benchmark, I would treat it as interesting but incomplete until it shares prompts, scoring rules, retries, model settings, and refusal handling. A benchmark without transcripts is more like a weather report from someone who refuses to say where they were standing.&lt;/p&gt;

&lt;p&gt;The important question is not “which model won?”&lt;/p&gt;

&lt;p&gt;The better question is: which model fails in the way my product can tolerate?&lt;/p&gt;

&lt;h2&gt;
  
  
  Three Things That Can Look Like Bugs
&lt;/h2&gt;

&lt;p&gt;The first possible failure mode is configuration drift.&lt;/p&gt;

&lt;p&gt;OpenAI’s docs say GPT-5.6 supports multiple reasoning effort settings, including &lt;code&gt;max&lt;/code&gt;, and the API docs describe pro mode as an execution mode rather than a separate model slug. That is useful control. It is also a nice little trap if your benchmark forgets to pin settings.&lt;/p&gt;

&lt;p&gt;Temporary fix: log the model ID, reasoning effort, pro/standard mode, tool access, cache behavior, and retry policy for every run. If you cannot reproduce the exact call, do not use it as evidence.&lt;/p&gt;

&lt;p&gt;The second possible failure mode is safety and refusal handling.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;pandas&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;


&lt;span class="n"&gt;EVENTS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;event_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;event-001&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;answer_correct&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;refusal&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;timeout&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tool_error&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;harness_valid&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;event_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;event-002&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;answer_correct&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;refusal&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;timeout&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tool_error&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;harness_valid&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;event_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;event-003&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;answer_correct&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;refusal&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;timeout&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tool_error&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;harness_valid&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;event_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;event-004&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;answer_correct&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;refusal&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;timeout&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tool_error&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;harness_valid&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;event_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;event-005&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;answer_correct&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;refusal&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;timeout&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tool_error&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;harness_valid&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;event_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;event-006&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;answer_correct&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;refusal&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;timeout&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tool_error&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;harness_valid&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;classify_event&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;harness_valid&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;invalid_harness_state&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;timeout&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;timed_out&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tool_error&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tool_failed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;refusal&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;refused&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;answer_correct&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;correct&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;wrong&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;classify_all_events&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;events&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;list&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nb"&gt;dict&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;classified_events&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;

    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;events&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;classified_event&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;copy&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="n"&gt;classified_event&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;classification&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;classify_event&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;classified_events&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;classified_event&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;pd&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;DataFrame&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;classified_events&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;


&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;__name__&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;__main__&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;event_frame&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;classify_all_events&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;EVENTS&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Classified Events&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;event_frame&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
            &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;event_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;classification&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
        &lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;to_string&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;index&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s"&gt;Classification Summary&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event_frame&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;classification&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;value_counts&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;OpenAI says GPT-5.6 is paired with stronger safeguards, especially around cyber and biology risk areas. Anthropic’s Fable 5 docs are even more explicit for integrations: Fable 5 includes safety classifiers, and refusal/fallback behavior needs to be handled deliberately.&lt;/p&gt;

&lt;p&gt;That means a harness should not throw refusals, timeouts, and wrong answers into one bucket. They are different product events. A refusal might be correct policy behavior. A timeout might be latency budget. A wrong answer is model quality. Counting them all as “failed” can be useful for user experience, but it is terrible for diagnosis.&lt;/p&gt;

&lt;p&gt;Temporary fix: split results into &lt;code&gt;correct&lt;/code&gt;, &lt;code&gt;wrong&lt;/code&gt;, &lt;code&gt;refused&lt;/code&gt;, &lt;code&gt;timed out&lt;/code&gt;, &lt;code&gt;tool failed&lt;/code&gt;, and &lt;code&gt;invalid harness state&lt;/code&gt;. It looks fussy. It saves you from arguing with ghosts.&lt;/p&gt;

&lt;p&gt;The third possible failure mode is local-model confusion.&lt;/p&gt;

&lt;p&gt;The GitHub project JustVugg/colibri is genuinely interesting. Its README describes running GLM-5.2, a 744B-parameter MoE model, on a consumer machine with about 25 GB of RAM by streaming experts from disk. It also reports sober numbers: about 370 GB on disk, roughly 20 GB peak RSS during chat, and cold decode around 0.05 to 0.1 tokens per second on the author’s WSL2 setup.&lt;/p&gt;

&lt;p&gt;That is impressive engineering.&lt;/p&gt;

&lt;p&gt;It is not evidence that Claude Fable 5 runs locally on a laptop.&lt;/p&gt;

&lt;p&gt;If you compare Fable 5 with a local system, name the local model precisely. “Fable 5 vs GLM-5.2 via colibri” is honest. “Frontier model runs on my laptop” is not, unless the weights, license, runtime, hardware, and speed are all named.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Practical Take
&lt;/h2&gt;

&lt;p&gt;If I were choosing today, I would not frame this as GPT-5.6 versus Fable 5 in the abstract.&lt;/p&gt;

&lt;p&gt;I would frame it as two integration styles.&lt;/p&gt;

&lt;p&gt;GPT-5.6 looks attractive when you want explicit control over model variants, reasoning effort, and API-side execution choices. Fable 5 looks attractive when you want Anthropic’s top widely released Claude model and are willing to design around refusal and fallback semantics.&lt;/p&gt;

&lt;p&gt;Neither choice saves you from doing the dull work.&lt;/p&gt;

&lt;p&gt;For bug fixing, the workflow stays simple: provide the failing test, relevant file, expected behavior, reproduction steps, and permission boundaries. Ask for the smallest patch. Run the patch. If the model cannot reproduce the bug, do not let it invent the fix.&lt;/p&gt;

&lt;p&gt;The real bug in launch-window model comparisons is usually not inside the model.&lt;/p&gt;

&lt;p&gt;It is in the way we compare screenshots instead of traces. We quote scores without harnesses. We treat safety behavior as stupidity. We treat local inference as if the model name does not matter. Then we are shocked when the conclusion collapses under the weight of one missing setting.&lt;/p&gt;

&lt;p&gt;My boring rule for the next 48 hours: pin the model, log the settings, preserve transcripts, separate refusal from wrong answer, and distrust every laptop claim until the repo names the actual weights.&lt;/p&gt;

&lt;p&gt;Less dramatic.&lt;/p&gt;

&lt;p&gt;Much more useful.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>openai</category>
      <category>claude</category>
    </item>
    <item>
      <title>Agent OS for Codex and Claude Code: The Control Plane Problem Is Real</title>
      <dc:creator>Elin</dc:creator>
      <pubDate>Fri, 10 Jul 2026 04:12:04 +0000</pubDate>
      <link>https://dev.to/elineve/agent-os-for-codex-and-claude-code-the-control-plane-problem-is-real-2hnb</link>
      <guid>https://dev.to/elineve/agent-os-for-codex-and-claude-code-the-control-plane-problem-is-real-2hnb</guid>
      <description>&lt;p&gt;The phrase “Agent OS” still sounds a little too large for what most projects can actually do.&lt;/p&gt;

&lt;p&gt;I do not mean that as a dismissal. I think the phrase exists because developers are running into a real coordination problem. Once you use more than one coding agent in the same week, the pain shifts. It is no longer just “which agent writes the better patch?” It becomes: which agent touched which branch, what context did it see, what command did it run, and where do I review the result?&lt;/p&gt;

&lt;p&gt;As of July 10, 2026, based on public READMEs and official docs I checked, I would describe this space as early control-plane tooling rather than a finished “operating system” layer.&lt;/p&gt;

&lt;p&gt;That difference matters.&lt;/p&gt;

&lt;p&gt;An operating system owns processes, permissions, scheduling, storage, and recovery boundaries. Most current open-source Agent OS projects do not own all of that. They usually sit above tools like &lt;a href="https://github.com/openai/codex" rel="noopener noreferrer"&gt;OpenAI Codex CLI&lt;/a&gt;, &lt;a href="https://code.claude.com/docs/en/overview" rel="noopener noreferrer"&gt;Claude Code&lt;/a&gt;, OpenClaw, Gemini CLI, OpenCode, or other coding agents. Their value is more practical: they make agent work visible, isolate workspaces, expose diffs, route prompts, and give the human operator a better place to approve or stop work.&lt;/p&gt;

&lt;p&gt;That is already useful. It is just not magic.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Agent Management Is Becoming Its Own Layer
&lt;/h2&gt;

&lt;p&gt;Single-agent work is easy to explain.&lt;/p&gt;

&lt;p&gt;Open a terminal. Ask Codex to inspect a failing test. Ask Claude Code to refactor a module. Review the diff. Ship or discard it.&lt;/p&gt;

&lt;p&gt;Multi-agent work is harder to keep in your head. One agent is exploring the codebase. Another is writing a patch. A third is reviewing the output. Maybe one runs locally, one runs in a worktree, and one runs in a remote sandbox. The developer is no longer only prompting. The developer is supervising.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"mode"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"board"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"cols"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"rows"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"titleSize"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;64&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"layers"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Board"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"kind"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"tile"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Agents"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"kind"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"entity"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Review"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"kind"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"gate"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"elements"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"task-101"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"agent"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"codex"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"workspace"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"worktree/fix-tests"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"running"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"branch"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"agent/fix-tests"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"task-102"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"agent"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"claude-code"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"workspace"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"worktree/refactor-auth"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"waiting_for_review"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"branch"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"agent/refactor-auth"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"gate-1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"kind"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"human_review"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"requires"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"diff"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"logs"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"test_result"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is where “agent management” becomes more than a UI preference.&lt;/p&gt;

&lt;p&gt;In &lt;a href="https://arxiv.org/abs/2602.14690" rel="noopener noreferrer"&gt;Harness Engineering for Agentic AI Coding Tools: An Exploratory Study&lt;/a&gt;, Matthias Galster and co-authors analyze configuration mechanisms across tools including Claude Code, GitHub Copilot, Cursor, Gemini, and Codex. In their arXiv v5 abstract, they report an empirical study of 2,853 GitHub repositories and describe &lt;code&gt;AGENTS.md&lt;/code&gt; as an emerging cross-tool starting point for agent configuration.&lt;/p&gt;

&lt;p&gt;A separate May 2026 survey, &lt;a href="https://arxiv.org/abs/2605.18747" rel="noopener noreferrer"&gt;Code as Agent Harness&lt;/a&gt;, frames code as part of the operational substrate for agent reasoning, action, planning, memory, tool use, verification, and multi-agent coordination. I would treat that as research framing, not product proof. But it matches the practical direction: the useful question is less “Can an agent write code?” and more “Can the surrounding system make agent work reviewable, constrained, and recoverable?”&lt;/p&gt;

&lt;h2&gt;
  
  
  What The Current Open-Source Projects Actually Claim
&lt;/h2&gt;

&lt;p&gt;The most directly relevant project in this category is &lt;a href="https://github.com/milisp/codexia" rel="noopener noreferrer"&gt;Codexia&lt;/a&gt;. Its README describes it as a lightweight agent workstation for Codex CLI and Claude Code, with task scheduling, git worktree management, remote control, skills management, MCP server marketplace features, usage analytics, and local storage.&lt;/p&gt;

&lt;p&gt;That is a good example of the control-plane pattern. It does not replace Codex or Claude Code. It wraps them in a workspace where sessions, tasks, files, and agent activity can become easier to inspect.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;task&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;fix-failing-payment-test&lt;/span&gt;
  &lt;span class="na"&gt;title&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Fix failing payment test&lt;/span&gt;
  &lt;span class="na"&gt;agent&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;codex&lt;/span&gt;
  &lt;span class="na"&gt;workspace&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;git_worktree&lt;/span&gt;
    &lt;span class="na"&gt;base_branch&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;main&lt;/span&gt;
    &lt;span class="na"&gt;branch&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;agent/fix-payment-test&lt;/span&gt;
    &lt;span class="na"&gt;path&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;.worktrees/fix-payment-test&lt;/span&gt;

  &lt;span class="na"&gt;allowed_commands&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;npm test&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;npm run lint&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;git diff&lt;/span&gt;

  &lt;span class="na"&gt;blocked_commands&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;git push&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;npm publish&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;rm -rf&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;deploy production&lt;/span&gt;

  &lt;span class="na"&gt;review&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;required_before_merge&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;changed_files&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;command_log&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;test_output&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;human_approval&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://github.com/BloopAI/vibe-kanban" rel="noopener noreferrer"&gt;Vibe Kanban&lt;/a&gt; is also useful as an architectural reference. Its README describes kanban issues, agent workspaces, branches, terminals, dev servers, diff review, app previews, and switching between multiple coding agents including Claude Code and Codex. However, the same README currently says the project is sunsetting, so I would not present it as a stable long-term recommendation without checking the shutdown announcement and project status again before publication.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/21st-dev/1code" rel="noopener noreferrer"&gt;1Code&lt;/a&gt; is another project to evaluate carefully. Its README describes an open-source coding agent client for Claude Code, Codex, and other agents, with features such as worktree isolation, diff previews, MCP and plugins, background agents, automations, chat forking, model selection, and cross-platform support. It also distinguishes between building from source and subscription-backed cloud/background features. So the safer claim is not “no subscription needed for everything.” The safer claim is: source-build/local usage may reduce platform lock-in, while some hosted or background capabilities may still depend on the project’s own subscription model.&lt;/p&gt;

&lt;p&gt;The phrase “Agent OS” also appears in projects that solve adjacent problems rather than direct Codex/Claude management. &lt;a href="https://github.com/buildermethods/agent-os" rel="noopener noreferrer"&gt;Builder Methods Agent OS&lt;/a&gt; describes itself as a system for extracting standards, deploying standards, shaping specs, and keeping agents aligned with codebase conventions. That sounds more like a spec-and-standards layer than a runtime control plane.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/SapienXai/AgentOS" rel="noopener noreferrer"&gt;SapienX AgentOS&lt;/a&gt; is closer to the operating-layer metaphor, but its README positions it as a native control plane for OpenClaw. According to the project README, it manages agents, tasks, models, context, approvals, runtime visibility, and human oversight above OpenClaw. That makes it a useful reference for the shape of agent control planes, but not direct evidence that there is already one universal manager for Codex and Claude Code.&lt;/p&gt;

&lt;h2&gt;
  
  
  What A Real Agent Control Plane Needs
&lt;/h2&gt;

&lt;p&gt;If I were evaluating an “Agent OS” for daily development work, I would not start with the name. I would start with five boring questions.&lt;/p&gt;

&lt;p&gt;First: does it isolate work? Git worktrees matter because agent edits need to be reviewable and reversible. A good control plane should make it hard for an agent to quietly mutate the main branch.&lt;/p&gt;

&lt;p&gt;Second: does it show diffs and tool activity clearly? A chat transcript is not enough. I want to see files changed, commands run, logs produced, and approvals requested.&lt;/p&gt;

&lt;p&gt;Third: does it preserve enough context without mixing everything together? Claude Code’s docs describe subagents as specialized assistants with their own context windows, tool access, and permissions. That is useful because not every side task deserves to flood the main conversation.&lt;/p&gt;

&lt;p&gt;Fourth: does it expose permissions as first-class state? Claude Code’s hooks reference includes events such as &lt;code&gt;PreToolUse&lt;/code&gt;, &lt;code&gt;PermissionRequest&lt;/code&gt;, &lt;code&gt;PostToolUse&lt;/code&gt;, &lt;code&gt;TaskCreated&lt;/code&gt;, and &lt;code&gt;TaskCompleted&lt;/code&gt;. Those are exactly the kinds of lifecycle points a serious control plane should make visible.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"hooks"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"TaskCreated"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"log"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"requireWorkspace"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"PreToolUse"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"rules"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"match"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"shell"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"deny"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"git push"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"npm publish"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"kubectl apply"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"deploy production"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"message"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"This command requires human approval."&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"match"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"file_write"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"allowOnlyInside"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;".worktrees/"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"src/"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"tests/"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"PermissionRequest"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"routeTo"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"human_reviewer"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"include"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"command"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"working_directory"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"diff_preview"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"TaskCompleted"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"requireArtifacts"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"diff"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"logs"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"test_result"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Fifth: does it make failure recoverable? A multi-agent system without clear transcripts, artifacts, and rollback points becomes difficult to debug very quickly.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Safety Argument Is Stronger Than The Productivity Argument
&lt;/h2&gt;

&lt;p&gt;The best argument for an Agent OS is not that it will make developers “10x.” That phrase is too easy to write and too hard to prove.&lt;/p&gt;

&lt;p&gt;The stronger argument is that agents need boundaries.&lt;/p&gt;

&lt;p&gt;In &lt;a href="https://arxiv.org/abs/2607.02294" rel="noopener noreferrer"&gt;Coding Agents Are Guessing: Measuring Action-Boundary Violations in Underspecified DevOps Instructions&lt;/a&gt;, Zimo Ji and co-authors evaluate Claude Code, Codex, and OpenCode on underspecified DevOps tasks. The paper’s abstract describes 69 task families, 2,208 prompt variants, and five agent/model configurations. Within that experimental setup, the authors report that 55.8% to 67.8% of runs violated at least one action boundary.&lt;/p&gt;

&lt;p&gt;I would not use that number as a universal failure rate. It is tied to the benchmark design, agent/model configurations, and task set. But the direction is hard to ignore: when instructions are underspecified, agents may act instead of asking for clarification.&lt;/p&gt;

&lt;p&gt;That is exactly where a control plane earns its keep.&lt;/p&gt;

&lt;p&gt;“Fix the deployment issue” is not a safe instruction unless the system also knows which environment, which service, which credentials, which commands are allowed, and what rollback path exists. The agent may be capable. The surrounding harness still has to define the boundary.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Current Take
&lt;/h2&gt;

&lt;p&gt;The safest way to talk about this category is not “one open-source tool can now control every agent for free.”&lt;/p&gt;

&lt;p&gt;A better framing is this:&lt;/p&gt;

&lt;p&gt;Open-source agent management tools are starting to give developers a shared surface for planning, launching, isolating, monitoring, and reviewing agent work across tools like Codex and Claude Code. Some projects are workstations. Some are kanban-style review surfaces. Some are spec systems. Some are OpenClaw-specific control planes.&lt;/p&gt;

&lt;p&gt;For a Codex + Claude Code workflow, &lt;a href="https://github.com/milisp/codexia" rel="noopener noreferrer"&gt;Codexia&lt;/a&gt; is the most direct project to evaluate from the public README alone. &lt;a href="https://github.com/21st-dev/1code" rel="noopener noreferrer"&gt;1Code&lt;/a&gt; is worth comparing if you want a broader client model and are comfortable checking which features are source-build/local versus subscription-backed. &lt;a href="https://github.com/BloopAI/vibe-kanban" rel="noopener noreferrer"&gt;Vibe Kanban&lt;/a&gt; is worth studying as a workflow pattern, while its sunsetting notice makes it risky to present as a future-proof dependency. &lt;a href="https://github.com/SapienXai/AgentOS" rel="noopener noreferrer"&gt;SapienX AgentOS&lt;/a&gt; is useful for understanding the OpenClaw control-plane direction, not as proof of a universal Codex/Claude operating system.&lt;/p&gt;

&lt;p&gt;The term “Agent OS” may still be inflated. The coordination problem underneath it is not.&lt;/p&gt;

&lt;p&gt;If coding agents are becoming small, semi-autonomous workers, then developers need more than better prompts. We need visible state, scoped permissions, isolated workspaces, clean review surfaces, and systems that make it easier for the human to say: yes, continue here; no, stop there.&lt;/p&gt;

&lt;p&gt;That sounds less dramatic than an operating system.&lt;/p&gt;

&lt;p&gt;It also sounds much closer to what we actually need.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>agents</category>
      <category>claude</category>
    </item>
    <item>
      <title>AI Video Generators in July 2026: How I’d Compare Veo 3.1, Sora 2, Kling, and Wan</title>
      <dc:creator>Elin</dc:creator>
      <pubDate>Thu, 09 Jul 2026 03:56:00 +0000</pubDate>
      <link>https://dev.to/elineve/ai-video-generators-in-july-2026-how-id-compare-veo-31-sora-2-kling-and-wan-lb6</link>
      <guid>https://dev.to/elineve/ai-video-generators-in-july-2026-how-id-compare-veo-31-sora-2-kling-and-wan-lb6</guid>
      <description>&lt;p&gt;I have a bad habit from translation evaluation: I do not trust rankings until I know what they are ranking.&lt;/p&gt;

&lt;p&gt;That applies very neatly to AI video generators.&lt;/p&gt;

&lt;p&gt;A leaderboard can tell you that one model looked better to a set of evaluators on a set of prompts. A blog post can tell you which tools creators are currently excited about. But if you are actually choosing a video model for a workflow, “best AI video generator 2026” is too vague to be useful.&lt;/p&gt;

&lt;p&gt;Best for cinematic b-roll?&lt;/p&gt;

&lt;p&gt;Best for controllable character motion?&lt;/p&gt;

&lt;p&gt;Best for prompt adherence?&lt;/p&gt;

&lt;p&gt;Best for local experiments?&lt;/p&gt;

&lt;p&gt;Best for a tool your team can still access next month?&lt;/p&gt;

&lt;p&gt;As checked on July 9, 2026, the video model landscape has a few awkward facts that comparison posts need to handle carefully. Google DeepMind’s current Veo page presents Veo 3.1 as its leading video generation model. The Sora 2 page by OpenAI serves as a key release document, but the hard truth on that very page is that Sora is unavailable as of April 26, 2026. Kling is another massive talking point for creators right now, but do yourself a favor and verify any 'Kling 3.0' claims inside the actual app before treating it as a reliable public release. Meanwhile, while the open Wan2.1 repo makes Wan easy to verify, that specific Wan2.7-260612 version mentioned in the brief completely vanished from my public searches.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DZjg1YmIwZGE3MjJiNTk4NDlmMTA2NzViNTA2NDIwYzdfVUlMVGpXWjBpOXE4OHdnNEFReklrd1ZCQ0daU25qdmVfVG9rZW46UkJmc2IwOUpIb1dxcGJ4QXZwMmNNNExlbkdjXzE3ODM1Njc1MjM6MTc4MzU3MTEyM19WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DZjg1YmIwZGE3MjJiNTk4NDlmMTA2NzViNTA2NDIwYzdfVUlMVGpXWjBpOXE4OHdnNEFReklrd1ZCQ0daU25qdmVfVG9rZW46UkJmc2IwOUpIb1dxcGJ4QXZwMmNNNExlbkdjXzE3ODM1Njc1MjM6MTc4MzU3MTEyM19WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="image about competition between different ai tools" width="800" height="509"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;So this is not a winner-takes-all comparison.&lt;/p&gt;

&lt;p&gt;It is a practical evaluation map.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Short Comparison
&lt;/h2&gt;

&lt;h2&gt;
  
  
  Veo 3.1: A Strong First Test Candidate, Not A Universal Winner
&lt;/h2&gt;

&lt;p&gt;If I had to start testing one hosted model first, I would probably start with Veo 3.1.&lt;/p&gt;

&lt;p&gt;Not because it is magically “the best.” Because Google DeepMind currently documents it clearly, positions it around video plus audio, and provides an official prompt guide. That gives developers and creators something stable to inspect.&lt;/p&gt;

&lt;p&gt;The Veo page says Veo 3.1 is designed for cinematic video with audio, and the prompt guide gives a useful breakdown of what to specify: shot framing, camera motion, style, lighting, character details, location, action, and dialogue.&lt;/p&gt;

&lt;p&gt;That structure matters.&lt;/p&gt;

&lt;p&gt;A weak prompt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A woman walking through a city at night.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A better test prompt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A medium tracking shot follows a woman who is in a dark wool coat walking through a wet Edinburgh street at midnight. Yellow slight streetlights shine in the pool of the water. The camera moves slowly and carefully behind her left shoulder. Ambient sound: light rain, distant bus honking, quiet small talks. No music.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The second prompt is not just prettier. It is easier to evaluate.&lt;/p&gt;

&lt;p&gt;Did the camera track from behind?&lt;/p&gt;

&lt;p&gt;Did the lighting stay consistent?&lt;/p&gt;

&lt;p&gt;Did the audio match the scene?&lt;/p&gt;

&lt;p&gt;Did the model invent music even though I asked for none?&lt;/p&gt;

&lt;p&gt;That is the kind of prompt I prefer for model testing. It gives the system enough detail to succeed, and enough constraints to fail visibly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sora 2: Important Release, Bad “How To Use” Target
&lt;/h2&gt;

&lt;p&gt;Sora 2 should still be discussed in AI video comparisons, but not carelessly.&lt;/p&gt;

&lt;p&gt;OpenAI published “Sora 2 is here” on September 30, 2025. The page describes Sora 2 as a video and audio generation model with better physical accuracy, realism, controllability, synchronized dialogue, and sound effects.&lt;/p&gt;

&lt;p&gt;That is significant.&lt;/p&gt;

&lt;p&gt;But the same official page now says: as of April 26, 2026, the Sora product is no longer available.&lt;/p&gt;

&lt;p&gt;That changes the article angle.&lt;/p&gt;

&lt;p&gt;I would not write “Sora 2怎么用” or “how to use Sora 2” as if it were a normal current tutorial. A safer and more useful angle is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;what Sora 2 changed in video model expectations&lt;/li&gt;
&lt;li&gt;how other tools now compete with Sora-style capabilities&lt;/li&gt;
&lt;li&gt;what creators should check when a tool depends on a discontinued or region-limited product surface&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Sora 2 is useful as a reference point. It is not a clean recommendation for someone choosing a tool today.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DMDM0Yzg4OTY4YTk4YjRkMWNiYTNlODNhMmQ2ZTEyMGRfNG1vQUpOUmpGcVZjdk1hU2sycXJGRFA0QWNVZEEwbXpfVG9rZW46S0cxRWJJS1dHbzlDUEJ4Nkx6U2NiN29abkFnXzE3ODM1Njc1MjM6MTc4MzU3MTEyM19WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DMDM0Yzg4OTY4YTk4YjRkMWNiYTNlODNhMmQ2ZTEyMGRfNG1vQUpOUmpGcVZjdk1hU2sycXJGRFA0QWNVZEEwbXpfVG9rZW46S0cxRWJJS1dHbzlDUEJ4Nkx6U2NiN29abkFnXzE3ODM1Njc1MjM6MTc4MzU3MTEyM19WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="people need to make decisions whether to recommend or not" width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Kling / Kling 3.0: Treat Version Claims Carefully
&lt;/h2&gt;

&lt;p&gt;Kling is the messiest part of this comparison.&lt;/p&gt;

&lt;p&gt;Not because it is unimportant. It is very important. Kuaishou’s Kling unit is getting serious market attention, and Kling-related technical work continues to appear. The Kling-MotionControl technical report, submitted by the Kling Team in March 2026, focuses on character animation and motion transfer, which fits the broader reason people keep bringing Kling up: controllable motion is one of the hard parts of AI video.&lt;/p&gt;

&lt;p&gt;But I would be careful with the phrase “Kling 3.0.”&lt;/p&gt;

&lt;p&gt;Search results and community discussion may mention it. Some third-party pages may summarize it. But unless you can point to an official release note, model card, app changelog, or visible in-account version label, I would not write “Kling 3.0 does X” as a hard fact.&lt;/p&gt;

&lt;p&gt;My safer framing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;For Kling, I would verify the exact version inside the product before making version-specific claims. The interesting evaluation target is not the version number itself, but whether the model can preserve motion, identity, and camera instructions across a generated clip.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The test I would run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A locked-off medium shot of a glass small ball rolling down a slightly tilted desk, which is made of wood. It bumps into a pencil, then changes direction, and stops near the edge without falling. Natural daylight. No camera movement. Realistic sound of glass on wood.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is intentionally boring.&lt;/p&gt;

&lt;p&gt;Boring prompts expose physical problems faster than fantasy prompts. If the ball changes the direction, merges into the pencil, ignores gravity, or changes size, the model may still be visually impressive, but that's obivious that I would not trust it for controlled product shots or educational video.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wan2.1 / Wan2.x: The Developer Path
&lt;/h2&gt;

&lt;p&gt;Wan belongs in a different bucket.&lt;/p&gt;

&lt;p&gt;Veo and Kling are usually discussed as creator tools. Wan2.1 is more interesting as an open model ecosystem.&lt;/p&gt;

&lt;p&gt;The public Wan2.1 GitHub repository describes it as a comprehensive open suite of video foundation models. It packs in everything you'd expect, from the usual text-to-video and image-to-video to video editing, text-to-image, and surprisingly, even video-to-audio. On top of that, it actually cracks open the technical stuff you need—like model sizes, target resolutions, and how to implement it all.&lt;/p&gt;

&lt;p&gt;That makes Wan useful for developers in a way a closed hosted product is not.&lt;/p&gt;

&lt;p&gt;You can build repeatable tests.&lt;/p&gt;

&lt;p&gt;You can record parameters.&lt;/p&gt;

&lt;p&gt;You can compare outputs across prompts.&lt;/p&gt;

&lt;p&gt;You can inspect integration paths through Hugging Face, ModelScope, ComfyUI, or Diffusers.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DZGIyNDZkNDdmZjQ2ZmQwMWQyMDJiZWY3NWU5MGNkNGJfR3pwTUFFTGhrSWlnak1HN0doZDliNUk1ZDJtbU5sUjRfVG9rZW46S0lqUmJUQlNzb0ZHdGV4UVpCT2NtdlpPbjFkXzE3ODM1Njc1MjM6MTc4MzU3MTEyM19WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DZGIyNDZkNDdmZjQ2ZmQwMWQyMDJiZWY3NWU5MGNkNGJfR3pwTUFFTGhrSWlnak1HN0doZDliNUk1ZDJtbU5sUjRfVG9rZW46S0lqUmJUQlNzb0ZHdGV4UVpCT2NtdlpPbjFkXzE3ODM1Njc1MjM6MTc4MzU3MTEyM19WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="different directions lead to diverse outcomes" width="1170" height="780"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I would be careful with the &lt;code&gt;Wan2.7-260612&lt;/code&gt; label, though. It's not available for me to verify that exact model name from public sources during this check. It may be a leaderboard-specific entry, that is to say, one internal tag. Or a newly listed model not yet documented in the public sources I could access.&lt;/p&gt;

&lt;p&gt;So I might write it like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A leaderboard entry labeled Wan2.7-260612 may be worth watching, but I would not treat it as a verified public model until there is a model card, repo, paper, or official release note.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is less exciting than saying “new Alibaba model just dropped.”&lt;/p&gt;

&lt;p&gt;It is also much less likely to be wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Evaluation Checklist I’d Actually Use
&lt;/h2&gt;

&lt;p&gt;When comparing AI video generators, I would score six things.&lt;/p&gt;

&lt;p&gt;Prompt adherence: Did the model follow the specific instruction, or just the vibe?&lt;/p&gt;

&lt;p&gt;Temporal stability: Do plane surfaces(for example faces), objects, lighting, or spatial layout remain stable across frames?&lt;/p&gt;

&lt;p&gt;Motion logic: Do entity(like bodies), props, vehicles, fluids, and collisions behave plausibly?&lt;/p&gt;

&lt;p&gt;Audio alignment: If the model generates sound, does timing match the visual event?&lt;/p&gt;

&lt;p&gt;Can you actually tweak and expand on a clip, or do you have to restart from zero every single time? Then there's the boring but critical part—access and licensing. Some people may say they can actually use this tool for commercial gigs and export files freely on thier current tier? The truth is, this stuff isn't flashy, but it’s a total dealbreaker. A model can look absolutely insane in an official demo, but it’s completely useless for real-world work if it’s locked behind a waitlist, heavily capped, or saddled with a license that ties your hands.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Current Take
&lt;/h2&gt;

&lt;p&gt;If I were choosing today, I would start with Veo 3.1 as a first hosted test candidate because Google’s current documentation is clear and the model is positioned around video plus audio.&lt;/p&gt;

&lt;p&gt;I would discuss Sora 2 as an important reference point, but not as a current tutorial target.&lt;/p&gt;

&lt;p&gt;I would test Kling for motion-heavy scenes, while verifying the exact product version before making claims about Kling 3.0.&lt;/p&gt;

&lt;p&gt;I would use Wan2.1 or later public Wan releases when reproducibility matters more than convenience.&lt;/p&gt;

&lt;p&gt;That is the less viral answer.&lt;/p&gt;

&lt;p&gt;But for developers, it is the more useful one.&lt;/p&gt;

&lt;p&gt;The real question is not “which AI video generator is best?”&lt;/p&gt;

&lt;p&gt;The better question is:&lt;/p&gt;

&lt;p&gt;Which one fails in the way your project can tolerate?&lt;/p&gt;

&lt;p&gt;A mood-board clip can survive a strange reflection.&lt;/p&gt;

&lt;p&gt;A product demo cannot survive an object changing shape.&lt;/p&gt;

&lt;p&gt;A language-learning clip cannot survive bad lip timing.&lt;/p&gt;

&lt;p&gt;A benchmark cannot survive undocumented settings.&lt;/p&gt;

&lt;p&gt;That is where I would start the comparison.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
      <category>reviews</category>
    </item>
    <item>
      <title>Claude Code Skills in Practice: From Minimal Prompts to a Real Agent Harness</title>
      <dc:creator>Elin</dc:creator>
      <pubDate>Mon, 06 Jul 2026 09:42:55 +0000</pubDate>
      <link>https://dev.to/elineve/claude-code-skills-in-practice-from-minimal-prompts-to-a-real-agent-harness-n1n</link>
      <guid>https://dev.to/elineve/claude-code-skills-in-practice-from-minimal-prompts-to-a-real-agent-harness-n1n</guid>
      <description>&lt;p&gt;I came to Claude Code Skills through translation evaluation, which is a very good way to become suspicious of vague instructions.&lt;/p&gt;

&lt;p&gt;When you compare model outputs for long enough, you stop asking only, “Does it sound fluent?” You start asking smaller questions. What context was actually used? Which instruction changed the result? Did the model improve, or did it just become more confident?&lt;/p&gt;

&lt;p&gt;That is also how I now look at AI coding assistants.&lt;/p&gt;

&lt;p&gt;The current excitement around Claude Code Skills, Codex configuration, multi-agent workflows, and local model setups is real. GitHub Trending is full of skill packs, agent wrappers, review bridges, and “make my coding agent better” projects. But I think the useful question is not “What is the best Claude Code skill in 2026?”&lt;/p&gt;

&lt;p&gt;The useful question is:&lt;/p&gt;

&lt;p&gt;What belongs in a Skill, what belongs in the harness around the agent, and what should remain a human decision?&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DN2E1ZjBhODAyZTMwYzk0YTI5MDQzZWNhMmFhM2Y3NGFfNVVhcHppR0N6UnJzZDJQQVNkZXRnRUZFdjI0ZTlad2hfVG9rZW46SDJPcmI5N3Zyb0FmTjh4MFBJOGNjVzF1bkdlXzE3ODMzMzA4NDg6MTc4MzMzNDQ0OF9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DN2E1ZjBhODAyZTMwYzk0YTI5MDQzZWNhMmFhM2Y3NGFfNVVhcHppR0N6UnJzZDJQQVNkZXRnRUZFdjI0ZTlad2hfVG9rZW46SDJPcmI5N3Zyb0FmTjh4MFBJOGNjVzF1bkdlXzE3ODMzMzA4NDg6MTc4MzMzNDQ0OF9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="image" width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  A Skill Is A Workflow Package, Not Just A Prompt
&lt;/h2&gt;

&lt;p&gt;Anthropic’s Claude Code documentation describes Skills as reusable instruction packages. A skill is built around a &lt;code&gt;SKILL.md&lt;/code&gt; file, and Claude Code can load it when relevant or when invoked directly.&lt;/p&gt;

&lt;p&gt;The part I care about most is context loading.&lt;/p&gt;

&lt;p&gt;According to the &lt;a href="https://docs.anthropic.com/en/docs/claude-code/skills" rel="noopener noreferrer"&gt;Claude Code Skills documentation&lt;/a&gt;, unlike &lt;code&gt;CLAUDE.md&lt;/code&gt;, a skill body loads only when the skill is used. That matters because long procedural instructions stop being a permanent tax on every session.&lt;/p&gt;

&lt;p&gt;A small project skill might look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;.claude/
  skills/
    summarize-changes/
      SKILL.md
      references/
        risk-checklist.md
      scripts/
        validate.sh
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is just an example layout. The official docs describe &lt;code&gt;SKILL.md&lt;/code&gt; as required, with supporting files such as references, examples, templates, and scripts as optional.&lt;/p&gt;

&lt;p&gt;A conservative &lt;code&gt;SKILL.md&lt;/code&gt; could be:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;description&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Summarize&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;uncommitted&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;changes&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;and&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;flag&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;risks.&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Use&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;when&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;reviewing&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;local&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;diff&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;before&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;committing."&lt;/span&gt;
&lt;span class="na"&gt;allowed-tools&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Read Grep&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;

&lt;span class="c1"&gt;## Instructions&lt;/span&gt;

&lt;span class="s"&gt;Read the current change context.&lt;/span&gt;

&lt;span class="na"&gt;Return&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
&lt;span class="s"&gt;1. What changed&lt;/span&gt;
&lt;span class="s"&gt;2. What could break&lt;/span&gt;
&lt;span class="s"&gt;3. The smallest useful verification step&lt;/span&gt;
&lt;span class="s"&gt;4. What you are not confident about&lt;/span&gt;

&lt;span class="s"&gt;Do not rewrite the implementation unless asked.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I would keep the first version boring. A skill that does one thing reliably is more valuable than a skill that tries to become a tiny operating system.&lt;/p&gt;

&lt;h2&gt;
  
  
  What The Caveman Idea Gets Right, And Wrong
&lt;/h2&gt;

&lt;p&gt;One reason &lt;code&gt;caveman&lt;/code&gt; attracted attention is obvious: the repo’s own GitHub Trending description claims it cuts token usage by making Claude Code “talk like caveman.” That is a funny hook. It is not, by itself, a verified benchmark.&lt;/p&gt;

&lt;p&gt;So I would treat the idea as a hypothesis to test.&lt;/p&gt;

&lt;p&gt;The useful part is not “remove grammar everywhere.” The useful part is instruction compression. Many agent prompts contain politeness, repeated framing, and explanatory padding that does not change behavior.&lt;/p&gt;

&lt;p&gt;For a narrow workflow, this kind of instruction can be enough:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Goal: find likely cause of failing test.
Do:
- read error
- inspect related files
- propose one fix
- run smallest test
Do not:
- refactor
- rename
- touch unrelated files
Output:
- cause
- patch summary
- test result
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is not elegant prose. It is not supposed to be.&lt;/p&gt;

&lt;p&gt;But I would not use compressed “caveman” prompting for everything. A translation evaluator learns this quickly: shorter text is not automatically clearer text. If compression removes constraints, examples, or edge cases, the model may spend more tokens recovering the missing context.&lt;/p&gt;

&lt;p&gt;The practical rule is simple: compress wording, not requirements.&lt;/p&gt;

&lt;h2&gt;
  
  
  Skills Need A Harness Around Them
&lt;/h2&gt;

&lt;p&gt;A skill helps with repeatable behavior. It does not solve scope control by itself.&lt;/p&gt;

&lt;p&gt;For coding agents, the harness around the model matters just as much as the prompt. By harness, I mean the configuration layer that decides:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;which instructions load&lt;/li&gt;
&lt;li&gt;which tools are allowed&lt;/li&gt;
&lt;li&gt;which agent handles which part of the task&lt;/li&gt;
&lt;li&gt;which checks run before or after actions&lt;/li&gt;
&lt;li&gt;which outputs require review&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is where Claude Code subagents and hooks become relevant.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://docs.anthropic.com/en/docs/claude-code/sub-agents" rel="noopener noreferrer"&gt;Claude Code subagents documentation&lt;/a&gt; describes built-in and custom subagents for separating exploration, planning, and implementation. That is useful because research can pollute the main thread. If an agent spends 20 minutes scanning a codebase, I do not always want all of that exploration living in the same context as the final edit.&lt;/p&gt;

&lt;p&gt;Hooks are the other side of the harness. The &lt;a href="https://docs.anthropic.com/en/docs/claude-code/hooks" rel="noopener noreferrer"&gt;Claude Code hooks documentation&lt;/a&gt; shows hooks as event-based configuration, such as &lt;code&gt;PreToolUse&lt;/code&gt;, where a hook can inspect a tool call before it proceeds. I would start with simple hooks only: blocking destructive shell patterns, requiring a small verification command, or enforcing local project conventions.&lt;/p&gt;

&lt;p&gt;Hooks are not decoration. They are automation with consequences.&lt;/p&gt;

&lt;p&gt;That's a setup I would actually try.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DN2JiNWUzNTBkZjIxNzM4ZTliNDUxYzgwM2MyN2I4ZGJfVmp5MXlUY0hrandqYlNoVU14TjlMVXkyWWNWS2JJRmZfVG9rZW46SEF2amJtMlFFb3hUdlN4OGlIUGN0dDZ6blViXzE3ODMzMzA4NDg6MTc4MzMzNDQ0OF9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DN2JiNWUzNTBkZjIxNzM4ZTliNDUxYzgwM2MyN2I4ZGJfVmp5MXlUY0hrandqYlNoVU14TjlMVXkyWWNWS2JJRmZfVG9rZW46SEF2amJtMlFFb3hUdlN4OGlIUGN0dDZ6blViXzE3ODMzMzA4NDg6MTc4MzMzNDQ0OF9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="image" width="760" height="427"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If I were configuring a fresh AI coding workflow, I would not install a large skill bundle immediately.&lt;/p&gt;

&lt;p&gt;Instead, I would start with four layers.&lt;/p&gt;

&lt;p&gt;First, a repo-orientation skill.&lt;/p&gt;

&lt;p&gt;It should answer: where is the code, how is it tested, what commands matter, and what areas are sensitive. This belongs in a skill because it is repeated and local.&lt;/p&gt;

&lt;p&gt;Second, a verification skill.&lt;/p&gt;

&lt;p&gt;For a frontend app, that might mean lint, typecheck, and one smoke test. For my own language-tooling experiments, it might mean running a tiny fixture set before trusting a script change.&lt;/p&gt;

&lt;p&gt;Third, a read-only exploration subagent.&lt;/p&gt;

&lt;p&gt;Use this for file discovery and risk mapping. The point is not to make the agent more dramatic. The point is to keep exploration separate from editing.&lt;/p&gt;

&lt;p&gt;Fourth, a review gate for risky changes.&lt;/p&gt;

&lt;p&gt;Not every change needs another agent. But auth, data loss, migrations, concurrency, and generated-code rewrites deserve a second pass.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Codex Fits
&lt;/h2&gt;

&lt;p&gt;Codex can be useful as a second lane, especially for review and delegation.&lt;/p&gt;

&lt;p&gt;The official &lt;a href="https://github.com/openai/codex-plugin-cc" rel="noopener noreferrer"&gt;&lt;code&gt;openai/codex-plugin-cc&lt;/code&gt;&lt;/a&gt; repository says it lets Claude Code users call Codex from inside Claude Code for reviews or delegated tasks. Its README lists commands such as &lt;code&gt;/codex:review&lt;/code&gt;, &lt;code&gt;/codex:adversarial-review&lt;/code&gt;, &lt;code&gt;/codex:rescue&lt;/code&gt;, &lt;code&gt;/codex:status&lt;/code&gt;, and &lt;code&gt;/codex:result&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That makes the pattern concrete: Claude can stay in the main implementation flow, while Codex runs a bounded review or background investigation.&lt;/p&gt;

&lt;p&gt;For configuration, I would keep the article-level example tied to documented Codex config keys rather than pretending one model choice is universal. OpenAI’s &lt;a href="https://developers.openai.com/codex/config-reference" rel="noopener noreferrer"&gt;Codex configuration reference&lt;/a&gt; documents &lt;code&gt;model&lt;/code&gt; and &lt;code&gt;model_reasoning_effort&lt;/code&gt; in &lt;code&gt;config.toml&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;A project-level example can look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# .codex/config.toml
model = "gpt-5.5"
model_reasoning_effort = "high"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The exact model should match what your account, organization, and current Codex release support. The important point is not the sample model string. It is that model and reasoning-effort defaults belong in configuration, not scattered through prompts.&lt;/p&gt;

&lt;p&gt;I would also avoid turning review gates into infinite loops. The Codex plugin README warns that its review gate can create long-running Claude/Codex loops and drain usage limits quickly. That warning is worth taking seriously.&lt;/p&gt;

&lt;h2&gt;
  
  
  What About Local Models?
&lt;/h2&gt;

&lt;p&gt;Local models are attractive for cost control, but I would give them low-risk work first.&lt;/p&gt;

&lt;p&gt;OpenAI’s Codex config reference includes &lt;code&gt;oss_provider&lt;/code&gt; with &lt;code&gt;lmstudio&lt;/code&gt; and &lt;code&gt;ollama&lt;/code&gt; as supported values for local-provider selection when running with &lt;code&gt;--oss&lt;/code&gt;. That tells me local-model workflows are a real configuration path, not just a fantasy.&lt;/p&gt;

&lt;p&gt;Still, I would not put a small local model in charge of production edits on day one.&lt;/p&gt;

&lt;p&gt;I would start with tasks like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;summarizing logs&lt;/li&gt;
&lt;li&gt;grouping similar errors&lt;/li&gt;
&lt;li&gt;drafting grep queries&lt;/li&gt;
&lt;li&gt;compressing long command output&lt;/li&gt;
&lt;li&gt;ranking files by likely relevance&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then a stronger coding model can handle the actual patch, or a human can review before anything writes to the repository.&lt;/p&gt;

&lt;p&gt;Local deployment should not mean “remove judgment.” It should mean “move cheap, low-risk language work out of the expensive lane.”&lt;/p&gt;

&lt;h2&gt;
  
  
  My Current Rule
&lt;/h2&gt;

&lt;p&gt;I don't trust a skill because of its cleverness. I trust it when it makes the workflow easier to inspect.&lt;/p&gt;

&lt;p&gt;A good Claude Code Skill should make one repeated procedure smaller, clearer, and easier to trigger.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DMWNiY2M0NWQ5MTVkODc5Njc0NWVmYmRkMjFlYzNmMThfNUVwSFl6NWZ5VUxYa1F0SjBaR3diUkFjeEZsUjljTDhfVG9rZW46R0dESGJaWWZUb1FCdkN4Vms2emN5c1VxblBFXzE3ODMzMzA4NDg6MTc4MzMzNDQ0OF9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DMWNiY2M0NWQ5MTVkODc5Njc0NWVmYmRkMjFlYzNmMThfNUVwSFl6NWZ5VUxYa1F0SjBaR3diUkFjeEZsUjljTDhfVG9rZW46R0dESGJaWWZUb1FCdkN4Vms2emN5c1VxblBFXzE3ODMzMzA4NDg6MTc4MzMzNDQ0OF9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="image" width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A good Codex configuration should make review and delegation more explicit.&lt;/p&gt;

&lt;p&gt;A good harness should make the agent less likely to wander outside the task.&lt;/p&gt;

&lt;p&gt;That is the part of this ecosystem I find promising. Not the idea that AI coding assistants will magically write everything. More the possibility that the boring parts of engineering can become programmable without making the risky parts invisible.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>agents</category>
    </item>
    <item>
      <title>Fable 5 Is Back. My Translation Workflow Is Not Going Back.</title>
      <dc:creator>Elin</dc:creator>
      <pubDate>Wed, 01 Jul 2026 08:57:08 +0000</pubDate>
      <link>https://dev.to/elineve/fable-5-is-back-my-translation-workflow-is-not-going-back-7nn</link>
      <guid>https://dev.to/elineve/fable-5-is-back-my-translation-workflow-is-not-going-back-7nn</guid>
      <description>&lt;p&gt;On June 12, 2026, I lost access to the model I had quietly built half my translation QA workflow around.&lt;/p&gt;

&lt;p&gt;Eighteen days later, reports said the U.S. Department of Commerce had lifted export controls on Anthropic’s Fable 5 and Mythos 5, and Anthropic would begin restoring access on July 1.&lt;/p&gt;

&lt;p&gt;The internet reaction was predictable: relief, memes, “we are so back,” and a lot of developers talking as if oxygen had returned to the room.&lt;/p&gt;

&lt;p&gt;I understand the feeling. I missed Fable 5 too.&lt;/p&gt;

&lt;p&gt;But I’m not putting my workflow back the way it was.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Was Using Fable 5 For
&lt;/h2&gt;

&lt;p&gt;I’m a translation student, not a software engineer. I use models for a strange mix of tasks: literary translation experiments, terminology checks, side-by-side critique, corpus cleanup, and the occasional Python script that exists only because CSV files are cruel.&lt;/p&gt;

&lt;p&gt;Before the 18-day interruption, Fable 5 had become my favorite model for one very specific job:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It was good at noticing when a translation was technically correct but emotionally wrong.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That sounds soft, but it is a real failure mode.&lt;/p&gt;

&lt;p&gt;Here is a tiny invented example in the style of the notes I keep:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Swedish source:
Han log som om rummet redan hade förlåtit honom.

Literal English:
He smiled as if the room had already forgiven him.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A weaker model usually makes this smoother:&lt;/p&gt;

&lt;p&gt;He smiled as if everyone in the room had already forgiven him.&lt;/p&gt;

&lt;p&gt;That is not awful. It may even be publishable in the right context.&lt;/p&gt;

&lt;p&gt;But something has been lost. The original gives the room agency. It feels a little uncanny. The smoother version explains the metaphor instead of letting it stand.&lt;/p&gt;

&lt;p&gt;Fable 5 was unusually good at flagging that kind of loss. Not always. But often enough that I started trusting it more than I should have.&lt;/p&gt;

&lt;h2&gt;
  
  
  The First Week Was Annoying
&lt;/h2&gt;

&lt;p&gt;When access disappeared, I did the obvious thing: I replaced Fable 5 with whatever else was available.&lt;/p&gt;

&lt;p&gt;For simple translation checks, this was fine. DeepL still does what DeepL does well. Other LLMs can still summarize, compare, and produce decent alternatives.&lt;/p&gt;

&lt;p&gt;The problem was not that my workflow stopped.&lt;/p&gt;

&lt;p&gt;The problem was that it became harder to tell which parts of my workflow had depended on one model’s taste.&lt;/p&gt;

&lt;p&gt;That is a different kind of dependency.&lt;/p&gt;

&lt;p&gt;If your build pipeline depends on one API, you know it. It fails loudly.&lt;/p&gt;

&lt;p&gt;If your judgment pipeline depends on one model, it fails quietly. You still get output. It just feels slightly flatter, slightly less suspicious, slightly too willing to accept the first good answer.&lt;/p&gt;

&lt;p&gt;I only noticed because I had old notes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1z2xdd52z91jtc8jwauk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1z2xdd52z91jtc8jwauk.png" alt="A photo of me working in my apartment with my cat" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Fallback Stack I Built
&lt;/h2&gt;

&lt;p&gt;By day five, I stopped trying to “replace” Fable 5 and started splitting the task into smaller checks.&lt;/p&gt;

&lt;p&gt;This is the fallback stack I ended up using:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1. DeepL
   Use for: first-pass EU language translation, especially practical text.
   Do not use for: literary tone decisions.

2. General LLM
   Use for: explaining ambiguity, generating alternate phrasings.
   Do not use for: final judgment.

3. Smaller local model
   Use for: cheap batch checks and rough terminology consistency.
   Do not use for: subtle register or metaphor.

4. Human pass
   Use for: anything involving style, humor, grief, politeness, or shame.
   Do not skip this. Ever.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This looks less impressive than “use the best model.”&lt;/p&gt;

&lt;p&gt;It is also more honest.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Prompt I Started Reusing
&lt;/h2&gt;

&lt;p&gt;The most useful thing I wrote during the interruption was not a clever translation prompt. It was a dependency prompt.&lt;/p&gt;

&lt;p&gt;I now run this whenever I test a model on a translation task:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You are not translating yet.

Read the source text and identify what kind of difficulty it contains.

Return:
1. Literal meaning risks
2. Tone/register risks
3. Cultural or idiomatic risks
4. Metaphor/image risks
5. Places where a fluent translation might become less faithful
6. What a reviewer should check manually

Do not produce a final translation.
Do not smooth over ambiguity.
If the text is simple, say so.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important line is: &lt;strong&gt;Do not produce a final translation.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Models love finishing the task. Sometimes I need them to slow down and tell me what kind of task it is.&lt;/p&gt;

&lt;p&gt;That one change made my workflow much less dependent on Fable 5. A weaker model can still be useful if I ask it to classify risk instead of making the final call.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpw8xh1qy772ff6dmsmxv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpw8xh1qy772ff6dmsmxv.png" alt="A photo from the official website of Claude" width="680" height="358"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What Fable 5 Coming Back Changes
&lt;/h2&gt;

&lt;p&gt;I will use it again. Of course I will.&lt;/p&gt;

&lt;p&gt;If Fable 5 is available on July 1 as expected, I’ll put it back into my experiments. It was too useful not to.&lt;/p&gt;

&lt;p&gt;But it will no longer be my default judge.&lt;/p&gt;

&lt;p&gt;The last 18 days made one thing very obvious: the best model in your workflow is also the most dangerous one to stop questioning.&lt;/p&gt;

&lt;p&gt;Not because it is bad.&lt;/p&gt;

&lt;p&gt;Because it is good enough to become invisible.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Trade-off
&lt;/h2&gt;

&lt;p&gt;The old workflow was faster:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;source → Fable 5 critique → revision → final human pass
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The new workflow is slower:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;source → risk classification → multiple model passes → compare disagreements → final human pass
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The old workflow felt elegant.&lt;/p&gt;

&lt;p&gt;The new one leaves more mess on the table.&lt;/p&gt;

&lt;p&gt;But the new one shows me where the uncertainty is. That matters more to me now than speed.&lt;/p&gt;

&lt;h2&gt;
  
  
  One Thing I’m Watching
&lt;/h2&gt;

&lt;p&gt;The export-control story is bigger than my little translation workflow. Reports from Wired, The Verge, Axios, and others describe a messy negotiation around security risks, jailbreaks, and access to frontier models.&lt;/p&gt;

&lt;p&gt;I am not qualified to write the policy version of that article.&lt;/p&gt;

&lt;p&gt;But as a user, I can say this: if a model disappearing for 18 days breaks your work completely, the model was not just a tool. It was infrastructure.&lt;/p&gt;

&lt;p&gt;And if it was infrastructure, you need a failure plan.&lt;/p&gt;

&lt;p&gt;Even if you are just translating poems in a flat in Edinburgh while your cat tries to sit on your keyboard.&lt;/p&gt;

&lt;p&gt;Marmalade did not care about export controls. This was probably healthy.&lt;/p&gt;

&lt;h2&gt;
  
  
  My New Rule
&lt;/h2&gt;

&lt;p&gt;I am keeping a simple rule now:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;No model gets to be both translator and judge.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If one model drafts, another critiques. If one model flags tone issues, I verify them manually. If a model sounds too confident, I ask it what it might be missing.&lt;/p&gt;

&lt;p&gt;This is slower.&lt;/p&gt;

&lt;p&gt;It is also the only way I currently know to stay awake inside my own workflow.&lt;/p&gt;

&lt;p&gt;What about you?&lt;/p&gt;

&lt;p&gt;If one model disappeared from your stack for 18 days, what would quietly break first?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
      <category>security</category>
    </item>
  </channel>
</rss>
