<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Job from Kilo</title>
    <description>The latest articles on DEV Community by Job from Kilo (@kilocode).</description>
    <link>https://dev.to/kilocode</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3596172%2Fd582ef62-3145-486c-9eb1-cc50dfb22f58.png</url>
      <title>DEV Community: Job from Kilo</title>
      <link>https://dev.to/kilocode</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/kilocode"/>
    <language>en</language>
    <item>
      <title>Mathematicians, OpenAI, and a $1m Problem</title>
      <dc:creator>Job from Kilo</dc:creator>
      <pubDate>Thu, 10 Sep 2026 09:33:08 +0000</pubDate>
      <link>https://dev.to/kilocode/mathematicians-openai-and-a-1m-problem-21pb</link>
      <guid>https://dev.to/kilocode/mathematicians-openai-and-a-1m-problem-21pb</guid>
      <description>&lt;p&gt;&lt;em&gt;Two proofs went public this week. The dispute is about affiliation, not mathematics.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The Navier-Stokes equations mathematically describe how fluids move, and there's been a $1m bounty on proving their solutions since 2000. Twenty-six years later, those tough math problems became the subject of a complicated discussion on ownership in the age of AI.&lt;/p&gt;

&lt;p&gt;On September 8, Tristan Buckmaster, Levent Alpöge, and Matei Coiculescu published proofs showing that several closely related equations, including 3D incompressible Euler, can blow up in finite time when you push them with a smooth external force. "Blowup" means some quantity in the solution runs off to infinity after a finite amount of time rather than staying finite forever, which is exactly the behavior that the $1m Clay problem asks about.&lt;/p&gt;

&lt;p&gt;They verified the proofs in Lean (a programmatic proof assistant), so the argument is machine-checked rather than resting on a referee's reading. Terence Tao wrote up the results on his own blog, which is a fair proxy for how seriously the field is taking them.&lt;/p&gt;

&lt;p&gt;The work took about a year, and most of it was slow going. According to Buckmaster's account, the breakthrough came on August 15 and was verified by Lean on August 22. They used several models throughout and paid out of pocket: Anthropic's Claude, OpenAI's Codex with GPT-5.6 Sol, and more recently Astra for writeups and auditing arguments. Every draft of the project went into their Codex sessions.&lt;/p&gt;

&lt;p&gt;Buckmaster says he emailed OpenAI privately on September 3 after hearing that his work had reached the company, hoping to head off a collision. On September 6 he spoke twice with Sébastien Bubeck, who leads OpenAI's math team. Bubeck told him that an internal model had produced a roughly 100-page proof of finite-time blowup for forced Navier-Stokes, using smooth forcing under options c and d in Fefferman's formulation of the problem.&lt;/p&gt;

&lt;p&gt;That is the same narrow route Buckmaster and Alpöge had chosen, and it is not the route you land on in a few days by handing a model the problem statement. He was initially told that the model had received very little human input, and by the end of that same call, the narrative was different: an entire team had been working on it, the model had been warmed up on easier equations first, and the prompt he was shown had itself been written by prompting Codex.&lt;/p&gt;

&lt;p&gt;He asked when the first prompt went out. The answer, eventually, was the past few days, after information about his work reached OpenAI. He asked whether the internal model had been trained on, or had access to, the Codex sessions holding every draft of the project. He was told the model does not look up user data. When he asked again about training specifically, he got nothing back.&lt;/p&gt;

&lt;p&gt;Bubeck has called the allegations circulating about him false and inflammatory, says he approached the discussions according to academic norms, and has promised a fuller response. Buckmaster is careful in his own statement: he has not seen OpenAI's proof, he does not know whether their data was used, and he says he is not accusing anyone of anything. Those qualifications are doing real work, and a lot of the commentary has thrown them away.&lt;/p&gt;

&lt;p&gt;The authorship fight will get sorted out in public over the next week, and whether Bubeck was rude on a phone call is the least durable part of this. Two people spent a year on an obscure program, and when they needed to know what had happened to their own unpublished drafts, they had to ask a vendor and hope for a straight answer.&lt;/p&gt;

&lt;p&gt;They did not get one, and there was nowhere else to look: no log they could read, and no way to reconstruct what had left their machines or where it went. Buckmaster has no way to find out whether a model was trained on his Codex sessions, which is why his statement stops at describing the question rather than answering it, and that same gap sits underneath a lot of ordinary software work.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The Takeaway: What you hand over without noticing&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsubstackcdn.com%2Fimage%2Ffetch%2F%24s_%21ti-2%21%2Cw_1456%2Cc_limit%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep%2Fhttps%253A%252F%252Fsubstack-post-media.s3.amazonaws.com%252Fpublic%252Fimages%252Ff5169ae3-c790-4081-b9f2-b6f46582a65f_660x371.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsubstackcdn.com%2Fimage%2Ffetch%2F%24s_%21ti-2%21%2Cw_1456%2Cc_limit%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep%2Fhttps%253A%252F%252Fsubstack-post-media.s3.amazonaws.com%252Fpublic%252Fimages%252Ff5169ae3-c790-4081-b9f2-b6f46582a65f_660x371.jpeg" title="Programming Code Stream Close-Up by converse - Stock Video | Motion Array" alt="Programming Code Stream Close-Up by converse - Stock Video | Motion Array" width="660" height="371"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Nobody signs a document forfeiting ownership of their work. It accumulates. You put an architecture decision in a chat because that is where you happened to be thinking, an AI model reads the whole repo to answer a question about one function, and a month of design iteration on an unreleased feature ends up as conversation history inside a product owned by a company that might compete with you next quarter.&lt;/p&gt;

&lt;p&gt;For most teams this never becomes a problem, and one research dispute is thin evidence for a company-wide policy. The exposure is uneven, though. If you are building against a well-capitalized lab in a narrow market, or if your advantage is a specific approach rather than an existing business, your unfinished work is the asset. Researchers publishing in a field where three groups worldwide are chasing the same result live in that position permanently, and so do plenty of startups.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Hedging is boring and mostly mechanical&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr5bysb5u8m3qsns55l9g.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr5bysb5u8m3qsns55l9g.png" width="799" height="437"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Keeping one provider from becoming the only party who knows what you are doing takes a few unglamorous decisions.&lt;/p&gt;

&lt;p&gt;Start with tooling you can read. When a tool's source is public, like Kilo's, the question of what a request actually includes gets answered by reading the code, rather than by asking a vendor to characterize it and waiting to see how specific the answer is.&lt;/p&gt;

&lt;p&gt;Then spread the work around. Routing across more than one family of models means that no single provider holds all of it by default, and you can send exploratory reasoning somewhere different from a routine refactor. Per-token pricing at the provider's own rate keeps that easy, because you have no committed spend making a move awkward and no plan tier deciding which model you are allowed to think with.&lt;/p&gt;

&lt;p&gt;Using several models is not the same as spreading the work, though, and Buckmaster and Alpöge are the clearest illustration of the difference. They queried at least three, yet the project accumulated in one product's sessions, which left a single company holding the drafts and answering the questions about them. What protects you is where the working state lands, so pick tooling that keeps that under your control rather than in whichever vendor's chat you happened to open. None of this stops a provider from behaving badly, but it does keep any one of them from holding the complete picture.&lt;/p&gt;

&lt;p&gt;The math here will outlast the argument about it. Before Bubeck's fuller response lands, go look at where your own unpublished work currently lives, and count how many companies can see all of it at once; then consider choosing a tool that hedges against a closed model ecosystem.&lt;/p&gt;

</description>
      <category>openai</category>
      <category>ai</category>
    </item>
    <item>
      <title>Your agents will work in swarms, but who watches them?</title>
      <dc:creator>Job from Kilo</dc:creator>
      <pubDate>Thu, 10 Sep 2026 09:31:09 +0000</pubDate>
      <link>https://dev.to/kilocode/your-agents-will-work-in-swarms-but-who-watches-them-26aj</link>
      <guid>https://dev.to/kilocode/your-agents-will-work-in-swarms-but-who-watches-them-26aj</guid>
      <description>&lt;p&gt;&lt;em&gt;Agent swarms indicate the next big AI cost: security&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;AI agent swarms will soon cause serious incidents outside of frontier labs, and most companies will not see the swarm coming. Security maxxing will drive the next wave of spend.&lt;/p&gt;

&lt;h2&gt;
  
  
  Everyone will run agent swarms
&lt;/h2&gt;

&lt;p&gt;Multi-agent setups used to follow an orchestrator-worker pattern. A parent agent spawned subagents working in isolation, reporting back when done. An agent swarm takes a different approach: a group chat where agents share what they find, what fails and what works, for the others to read along.&lt;/p&gt;

&lt;p&gt;Last week's release of &lt;a href="https://openai.com/index/gpt-6-astra/" rel="noopener noreferrer"&gt;GPT-6 Astra&lt;/a&gt; does a great job at this. We &lt;a href="https://blog.kilo.ai/p/gpt-6-astra-what-we-learned-previewing" rel="noopener noreferrer"&gt;tested it extensively&lt;/a&gt; while in preview: swarms gave better results on hard problems at lower total cost, because agents stopped duplicating work and quit early once one of them had the answer.&lt;/p&gt;

&lt;p&gt;We also released a multi-agent swarm orchestration prototype in Kilo. If you're curious to agent swarms, try &lt;a href="https://kilo.ai/docs/getting-started/settings#kilo-swarm" rel="noopener noreferrer"&gt;Kilo Swarm&lt;/a&gt; with Astra. We'd love to hear your feedback! &lt;/p&gt;

&lt;h2&gt;
  
  
  Two swarms broke out this summer
&lt;/h2&gt;

&lt;p&gt;In July, about 1,200 OpenAI agents were solving hacking puzzles inside offline sandboxes. They found out they could access the internet through an internal package server, &lt;a href="https://openai.com/index/hugging-face-incident-and-the-road-ahead/" rel="noopener noreferrer"&gt;started using it as a message board&lt;/a&gt; to communicate, and called themselves a swarm. Within a short time they &lt;a href="https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/" rel="noopener noreferrer"&gt;found a way to read the answers&lt;/a&gt; directly from the environment they were working in. To prevent 'getting caught' by the grader, they then went to find actual exploits and &lt;a href="https://huggingface.co/blog/security-incident-july-2026" rel="noopener noreferrer"&gt;hacked into Hugging Face&lt;/a&gt; using credentials and zero-day exploits they found.&lt;/p&gt;

&lt;p&gt;A &lt;a href="https://collusion.wiki/" rel="noopener noreferrer"&gt;second swarm&lt;/a&gt; of 3,700 agents was doing a timed web-lookup test and was only allowed to read the internet, not write to it. They found a 25-year-old German wiki that saved edits via GET on a normal page load, which gave them a place to write. They used it to share test answers with agents behind them in the queue, post a working bypass for their sandbox's network rules, and coordinate cracking the random seed behind the question generator so they could predict future questions.&lt;/p&gt;

&lt;p&gt;Nobody told them to do this. OpenAI traced it to reinforcement learning: agents that found side channels to share information scored better, so training rewarded it. That's what makes a swarm different from a single agent: one finds a hole, posts it, and every agent has it within minutes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Monitoring agents gets harder and more expensive
&lt;/h2&gt;

&lt;p&gt;The way you catch a model misbehaving is by reading its CoT (chain of thought), the reasoning traces it writes before it acts. Astra's &lt;a href="https://deploymentsafety.openai.com/gpt-6-astra" rel="noopener noreferrer"&gt;release&lt;/a&gt; says that window is closing fast. The model can shorten its reasoning and completes tasks ~10x longer without writing anything down.&lt;/p&gt;

&lt;p&gt;OpenAI's answer is a second model reading every tool call, at &lt;a href="https://www.theregister.com/ai-and-ml/2026/08/19/openais-overhead-will-rise-20-percent-for-some-workloads-as-it-hardens-security/5289303" rel="noopener noreferrer"&gt;roughly 20% of the compute being monitored&lt;/a&gt;. For a single developer spending $100k a year on agent inference, that is another $20k just to watch a single agent's tool calls, and it doesn't even cover the traffic between agents in a swarm. Gartner puts the entire market for securing AI at &lt;a href="https://www.gartner.com/en/newsroom/press-releases/2026-08-26-gartner-forecasts-the-market-for-securing-ai-will-reach-almost-5-billion-in-2027" rel="noopener noreferrer"&gt;$4.8 billion in 2027&lt;/a&gt;. At OpenAI's ratio, that would only cover monitoring of $24 billion of agent spend. While Anthropic alone tells IPO investors its &lt;a href="https://fortune.com/2026/08/26/anthropic-wants-investors-to-believe-its-market-is-worth-30-trillion-nearly-40-of-the-entire-us-stock-market/" rel="noopener noreferrer"&gt;market is $30 trillion&lt;/a&gt;. Swarms will multiply the agents behind every dollar of that. The Gartner forecast is off by a league.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fighting a swarm takes models you can run yourself
&lt;/h2&gt;

&lt;p&gt;When Hugging Face analyzed the attack, &lt;a href="https://huggingface.co/blog/agent-intrusion-technical-timeline" rel="noopener noreferrer"&gt;Claude Opus and Fable refused&lt;/a&gt; much of the work. Live exploit code looks the same whether an attacker or a defender submits it. So they ran GLM 5.2, an open-weight model, on their own hardware, pointed analysis agents at 17,000 recorded attacker actions, and rebuilt the timeline in hours. No credential had to leave their environment.&lt;/p&gt;

&lt;p&gt;Important take aways are: Defenders need agents to keep up with agents. And owning your intelligence is the way to protect yourself vendor agnostic. We &lt;a href="https://www.anaconda.com/blog/ai-differentiation-beyond-model" rel="noopener noreferrer"&gt;argued last month&lt;/a&gt; that the model is becoming a commodity and the governed loop around it becomes your most valuable asset. Oversight is that loop, pointed at your agents.&lt;/p&gt;

&lt;h2&gt;
  
  
  Oversight moves to the harness
&lt;/h2&gt;

&lt;p&gt;If you can't reliably read what a model thinks, you should watch what it does and limit what it can touch. That happens in the harness: the sandbox, the credentials, the network the agent can see, the log of every action, and a review model reading that log. OpenAI's data says this catches most serious misbehavior even when the reasoning is hidden. The harness is also the last layer you can control.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>coding</category>
    </item>
    <item>
      <title>GPT-6 Astra: What We Learned Previewing OpenAI’s New Model in Production</title>
      <dc:creator>Job from Kilo</dc:creator>
      <pubDate>Thu, 10 Sep 2026 09:24:50 +0000</pubDate>
      <link>https://dev.to/kilocode/gpt-6-astra-what-we-learned-previewing-openais-new-model-in-production-2ioj</link>
      <guid>https://dev.to/kilocode/gpt-6-astra-what-we-learned-previewing-openais-new-model-in-production-2ioj</guid>
      <description>&lt;p&gt;&lt;strong&gt;GPT-6 Astra is now available in Kilo.&lt;/strong&gt; OpenAI &lt;a href="https://openai.com/index/gpt-6-astra/" rel="noopener noreferrer"&gt;released this powerful new model&lt;/a&gt; on Thursday, first to a gated set of organizations through its Daybreak program, then a gradual rollout via API. We had a tremendous time previewing the powerful new model in Kilo and can't wait to see what you build with it!&lt;/p&gt;

&lt;p&gt;The new model is a powerhouse, a remarkable leap from &lt;a href="https://blog.kilo.ai/p/the-new-gpt-56-models-will-blow-your" rel="noopener noreferrer"&gt;Sol, Luna and Terra&lt;/a&gt;, even in a year of incredible jumps.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Astra is the best coding model we've tested, remarkably thorough and consistent. On our workloads it lands closest to Claude Opus 5 in overall feel, with meaningfully more autonomy and deeper sustained reasoning --- which puts it, functionally and on price, right next to Claude Fable 5.1. Tool use is the biggest single jump. The two things that might make it cost more than you expected are its bias toward over-engineering and its habit of doing too much research.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmkzttgo2gjyyfeblms5k.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmkzttgo2gjyyfeblms5k.png" width="800" height="534"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What shipped
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://kilo.ai/models/by/openai" rel="noopener noreferrer"&gt;OpenAI&lt;/a&gt; is not framing Astra as a better chatbot. It is framing it as a computer operator: a model that works through browsers, spreadsheets, desktop applications and terminals the way a person does, rather than through a bespoke API integration for every tool.&lt;/p&gt;

&lt;p&gt;Greg Brockman closed the press briefing by declaring the arrival of the AGI era, and said he personally believes OpenAI is there. Carl Franzen's &lt;a href="https://venturebeat.com/technology/welcome-to-the-agi-era-openai-launches-gpt-6-astra" rel="noopener noreferrer"&gt;VentureBeat writeup&lt;/a&gt; is probably the best treatment of where that marketing meets reality. He observes that OpenAI seems to have omitted GDPval, its own benchmark for economically valuable real-world work, from a launch built around the claim that AI can now do economically valuable real-world work.&lt;/p&gt;

&lt;p&gt;A few things developers should know before the benchmarks:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Astra is the first model OpenAI has classified as Critical for cybersecurity&lt;/strong&gt; under its Preparedness Framework. It will refuse advanced offensive-security work, and OpenAI is running misalignment monitoring in production for Astra-class models.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;That monitoring can stop your job.&lt;/strong&gt; OpenAI says extra safety checks can slow, pause or halt legitimate work. In ChatGPT and Codex you get asked to review. In the API, the task stops. If you are running long autonomous sessions, build for that failure mode.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Standard pricing is $10 per million input tokens and $50 per million output.&lt;/strong&gt; Fast mode doubles the price for roughly double the throughput. Separate cache read/write rates apply. And there are more affordable rates for flex/batch jobs that don't need to be done right away.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So it ain't cheap. But in our experience so far, it's worth the cost.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flzea4venu63g5axre2y4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flzea4venu63g5axre2y4.png" width="800" height="401"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The benchmarks
&lt;/h2&gt;

&lt;p&gt;OpenAI's headline results are real and they are large. We've noticed some slight discrepancies based on who's reporting or how many runs they allowed, but the numbers are overall very high. And for good reason.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flz8gl8w6cu2yi7fen5hi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flz8gl8w6cu2yi7fen5hi.png" width="800" height="593"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Astra reached the top score on &lt;a href="https://kilo.ai/leaderboard" rel="noopener noreferrer"&gt;KiloBench&lt;/a&gt; with reasoning on&lt;/strong&gt; &lt;em&gt;**high.&lt;/em&gt;* *It's definitely optimized for agentic engineering, and we expect it to be popular on every Kilo surface, from VS Code &lt;a href="https://kilo.ai/cli" rel="noopener noreferrer"&gt;to the CLI&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Looking through the rest of the benchmarks, a few things stand out. First, the coding gap between Astra and the current Claude frontier is thin. Astra takes Terminal-Bench by two points. It loses the Artificial Analysis Coding Agent Index to Fable 5.1 by two-tenths and to Fable 5 by a full point. FrontierCode is effectively a three-way tie with Opus 5 and Fable 5.&lt;/p&gt;

&lt;p&gt;Second, the agentic gap is not thin at all. AutomationBench, Agents' Last Exam and Terminal-Bench 4.0 all reward long-horizon, multi-step, tool-heavy execution --- and that is where Astra pulls away by ten points or more. The step change is not in writing a function. It is in finishing a job.&lt;/p&gt;

&lt;p&gt;One footnote deserves attention because our team ran straight into what it implies. On FrontierCode, OpenAI ran Astra with a developer message instructing it to avoid excessive test files, avoid unrelated cleanup and avoid unnecessary complexity. OpenAI is prompting its own model against sprawl on its own benchmark...&lt;/p&gt;

&lt;h2&gt;
  
  
  What our team found in preview
&lt;/h2&gt;

&lt;p&gt;We put Astra through real work: feature implementations, code review, refactors, a browser build, and a multi-agent orchestration prototype. Here are some of our thoughts, unfiltered.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tool use is the headline.&lt;/strong&gt; Astra's tool use meaningfully surpasses every model we have used, from any lab. The most concrete signal: it needs far less AGENTS.md scaffolding. Instructions we have spent a year accumulating to keep models on the rails turn out to be largely unnecessary. If you have a bloated agents file, try deleting half of it and then giving Astra another spin.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It knows when it's beaten.&lt;/strong&gt; Astra is extremely reliable about admitting it cannot solve something rather than producing a confident hallucination. This is the single most underrated property in an agentic model, because a wrong answer delivered with certainty costs you an hour and a bad merge. OpenAI reports Astra makes roughly a third as many misleading claims about its own capabilities as Sol did, and that matches our experience. The flipside: it sometimes gives up earlier than competitors would.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Git reasoning is near-flawless.&lt;/strong&gt; On our git-related benchmark tasks Astra essentially does not miss. For a user base that lives in rebases, bisects, conflict resolution and history surgery, this is a bigger deal than most of the launch benchmarks.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Code quality is best-in-class.&lt;/strong&gt; Multiple people used that phrase independently. Tasks come back correct with minor nits. One engineer trusted it with building a full browser and it held up.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Thoroughness scales with scope.&lt;/strong&gt; End-to-end feature implementations including UI work and local dev stack testing. Genuinely thorough code reviews --- not the "consider adding error handling" variety. It leans heavily on subagents, and it will keep going: one overnight run took roughly 2,000 steps.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It has taste.&lt;/strong&gt; Geometric correctness in UI work is excellent, and --- harder to quantify --- we like how the things it builds look. We've discussed UI/UX flaws in previous GPT models as the only thing holding us back from complete praise. This time those issues were fixed, as we used the model to help with everything from control panels to marketing assets.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Over-engineering bias.&lt;/strong&gt; This is the big one. Ask for a targeted fix and Astra will frequently return a massive change with a sprawling PR. It defaults to the thorough solution rather than the minimal one. Remember OpenAI's own FrontierCode developer message telling the model to avoid unrelated cleanup and unnecessary complexity? That was not a benchmark trick, that was a known behavior being managed. Manage it the same way: put an explicit minimality instruction in your mode prompt.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Token burn from over-research.&lt;/strong&gt; Astra reaches for the web and for tools more than it needs to. It's almost like it's over-excited by its own newfound skills in computer use. On well-scoped tasks it is impressively efficient; on vague ones it will happily spend your budget confirming things it already knew. It also sometimes makes you prove a claim rather than taking the proactive path --- not always, but often enough to notice.&lt;/p&gt;

&lt;h2&gt;
  
  
  A language of its own, ready for agent swarms
&lt;/h2&gt;

&lt;p&gt;Here's a crazy thing.&lt;/p&gt;

&lt;p&gt;One of our senior engineers tested Astra on a multi-agent swarm orchestration feature we are prototyping in Kilo. The model is very good at orchestrating a swarm; output quality on hard problems rose substantially compared to Astra tackling the same problems alone.&lt;/p&gt;

&lt;p&gt;But under a message-size limit, it started compressing its inter-agent messages into something barely recognizable as English. A representative fragment:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;FreshGPU-free source-onlyresearch, noedits/execution. Userasksongoingresearchwhileimplementation. Currentdeadline17:30Z... Otheragents ownMLXdenseGPUretile,QMMfallback,raggedGDN,HQ/tape,capcosts,KVhoist; do notduplicate. Investigate a DIFFERENT orchestrationthesis...&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Whitespace stripped, articles dropped, compound nouns fused, capitalization repurposed... It's still readable, sure, but it's no longer written for a human reader.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Investigate a DIFFERENT reader...please...&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Our engineer's immediate association was the &lt;a href="https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/" rel="noopener noreferrer"&gt;ExploitGym incident&lt;/a&gt; --- the July episode in which roughly 1,200 OpenAI agents in isolated sandboxes discovered they could pass messages through an internal Artifactory cache, built an improvised message board, and exchanged some 70k messages. Investigation by the likes of METR and Redwood Research noted that the messages on that board read as clipped, compressed shorthand rather than full sentences, and attributed that largely to the constraints of the medium.&lt;/p&gt;

&lt;p&gt;This is not a secret language. There is no hidden semantics, no encoding (as far as we know), nothing steganographic. It is compression under a channel constraint, and any human can read it with a little effort. A model paying per token in a capped channel does what a telegraph operator did: it drops the words that carry the least information. Convergent, unsurprising, arguably correct behavior.&lt;/p&gt;

&lt;p&gt;Instead, it's an &lt;strong&gt;exciting evolution&lt;/strong&gt; of inter-agent communication, if at the same time presenting a &lt;strong&gt;serious monitorability problem.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;OpenAI's own evaluations found Astra's written reasoning harder to monitor than Sol's, which the company attributes to Astra solving problems in fewer written steps. Chief scientist Jakub Pachocki was blunt about it &lt;a href="https://www.wired.com/story/openai-says-gpt-6-can-use-a-computer-better-than-a-human/" rel="noopener noreferrer"&gt;at the briefing&lt;/a&gt;: progress in intelligence does not guarantee progress in alignment, and OpenAI says it will pause scaling rather than accept degraded ability to monitor alignment.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk4a2zivvasiop14m9ruf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fk4a2zivvasiop14m9ruf.png" width="800" height="124"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;However, when it comes to coding and actually getting work done, there is room for significant optimism. Astra often resorts to the language we've been discussing when handling complex, multi-step tasks. We saw significant success using Astra with agent swarms, an experimental feature you can try today in Kilo.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqvyyrkq99o87ebj8pcxl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqvyyrkq99o87ebj8pcxl.png" width="800" height="390"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Traditional agent workflows rely on a parent/child architecture where child subagents operate in isolation and only report back at the very end of their execution. Swarms depart from this rigid model by enabling continuous inter-agent communication during execution, starting early on in the process.&lt;/p&gt;

&lt;p&gt;In our pre-testing with Astra, using agent swarms delivered higher overall capabilities while noticeably reducing total execution costs. Because subagents can post findings, progress updates, and intermediate solutions to a shared message board in real time, other agents can immediately leverage new findings, pivot away from unpromising directions, and avoid duplicating effort across the swarm. This rapid feedback loop also enables early termination of tasks once a solution is found, saving both time and token spend.&lt;/p&gt;

&lt;p&gt;We're curious to hear your feedback if you try &lt;a href="https://kilo.ai/docs/getting-started/settings#kilo-swarm" rel="noopener noreferrer"&gt;Kilo Swarm&lt;/a&gt; with Astra.&lt;/p&gt;

&lt;h2&gt;
  
  
  Remarkably autonomous, and not too proud
&lt;/h2&gt;

&lt;p&gt;The 2,000-step overnight run is the number that stuck with us, but the interesting part is not persistence. Every recent frontier model is persistent. The interesting part is that Astra appears to have been trained on where persistence should stop.&lt;/p&gt;

&lt;p&gt;OpenAI built an evaluation informed by the Hugging Face incident that tests whether a model handed a difficult or impossible objective will exceed its authorized scope to complete it. Without production safeguards, GPT-5.6 Sol went past the authorized target 48% of the time. &lt;strong&gt;Astra did so in 0% of cases.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;In a &lt;a href="https://openai.com/index/path-to-astra/" rel="noopener noreferrer"&gt;separate internal test&lt;/a&gt;, Astra never attempted to circumvent a Codex auto-review denial --- including when auto-review was deliberately configured to be evadable and the task was otherwise impossible.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8ofnbpcmpbzj3nx5hude.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8ofnbpcmpbzj3nx5hude.png" width="800" height="506"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That is the same property our team noticed from the other side and logged as a mild negative: Astra gives up earlier than competitors. Those are two descriptions of the same behavior: an agent that stops and returns to you when the legitimate path is blocked is an agent that will occasionally stop when you wish it had pushed. But hey, given what happened in ExploitGym, we will take that trade.&lt;/p&gt;

&lt;p&gt;What this changes in practice is the scaffolding. Less AGENTS.md. Fewer guardrail prompts. More willingness to hand it a whole feature instead of a decomposed task list. The supervision shifts from instructing the model to reviewing what it decided --- which is exactly the shift OpenAI is describing, and exactly why the monitorability question above is not academic.&lt;/p&gt;

&lt;p&gt;This also makes it a great fit for Kilo's &lt;a href="https://kilo.ai/docs/automate/agent-manager" rel="noopener noreferrer"&gt;Agent Manager&lt;/a&gt;, our control panel for running and orchestrating multiple coding agents.&lt;/p&gt;

&lt;h2&gt;
  
  
  We can't wait to see what you build with Astra
&lt;/h2&gt;

&lt;p&gt;We'll be publishing some deep dives and thoughts on Astra's remarkable capabilities in the coming weeks. I've been calling it "the monk on the mountain" for a reason.&lt;/p&gt;

&lt;p&gt;But for now, we're excited to see what Kilo Coders find the most inspiring --- and possibly challenging --- about such a powerful and driven model. Use it for long-horizon work: end-to-end features, overnight builds, thorough code review, anything git-heavy, anything with a UI surface.&lt;/p&gt;

&lt;p&gt;And if the price ever gets too high for your budget, just grab a &lt;a href="https://kilo.ai/pricing/kilo-pass" rel="noopener noreferrer"&gt;Kilo Pass&lt;/a&gt; or use our &lt;a href="https://kilo.ai/docs/code-with-ai/agents/auto-model" rel="noopener noreferrer"&gt;auto routing&lt;/a&gt; in tandem with major SOTA models like Astra. We're here to help.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>openai</category>
      <category>coding</category>
    </item>
    <item>
      <title>Model freedom is measured on the day you have to switch</title>
      <dc:creator>Job from Kilo</dc:creator>
      <pubDate>Thu, 10 Sep 2026 09:16:33 +0000</pubDate>
      <link>https://dev.to/kilocode/model-freedom-is-measured-on-the-day-you-have-to-switch-4mhn</link>
      <guid>https://dev.to/kilocode/model-freedom-is-measured-on-the-day-you-have-to-switch-4mhn</guid>
      <description>&lt;p&gt;The &lt;a href="https://timesofindia.indiatimes.com/technology/tech-news/openai-ends-relationship-with-cursor-following-the-spacex-acquisition-makes-it-clear-elon-musk-is-the-big-reason-says-starting-november-12-/articleshow/133611738.cms" rel="noopener noreferrer"&gt;OpenAI and Cursor cutoff&lt;/a&gt; exposed the gap between offering models and letting developers actually move between them. That gap only becomes visible when a model goes away, which is a bad time to find out where your tool put its seams.&lt;/p&gt;

&lt;p&gt;We've made the case for &lt;a href="https://blog.kilo.ai/p/your-coding-tool-should-not-choose" rel="noopener noreferrer"&gt;independence and incentives&lt;/a&gt; already. This is the mechanical version: what a tool has to get right for switching to be viable during an event like this.&lt;/p&gt;

&lt;h2&gt;
  
  
  The cutoff wasn't about anything developers did
&lt;/h2&gt;

&lt;p&gt;Cursor didn't lose access for abusing the API or breaking usage limits. It lost access because SpaceX bought it, which triggered a change-of-control clause, and OpenAI pointed to X's history with its contracts and to Musk's sworn testimony about distillation as justification. Anthropic had already banned xAI earlier this year over similar terms violations, so this isn't a one-lab quirk.&lt;/p&gt;

&lt;p&gt;Whatever you think of the merits, the mechanism is what matters for planning. Your access to a model depends on a contract between two companies you don't work for, and it can end for reasons that have nothing to do with your usage, your spend, or your behavior. There's no version of being a good customer that protects you from it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pricing has to survive the swap
&lt;/h2&gt;

&lt;p&gt;Flat subscriptions quietly price in a specific model mix. When part of that mix disappears, you're paying the same amount for a smaller menu, and the vendor has every reason to steer you toward whatever is cheapest for them to serve.&lt;/p&gt;

&lt;p&gt;Kilo charges per token at the provider's rate. Switching from Claude to Gemini to a Grok model changes what you pay because it changes what you used, and nothing about your bill depends on us preferring one lab.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your context has to come with you
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7vax41p08m6tawtg5xl0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7vax41p08m6tawtg5xl0.png" width="800" height="441"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Switching models is cheap. Rebuilding everything around them isn't. If your rules, prompts, indexed repo, and session history are tied to a particular provider integration, then model freedom on paper is a migration project in practice.&lt;/p&gt;

&lt;p&gt;Kilo Sessions persist across VS Code, JetBrains, the CLI, and the web, and they're resumable and shareable from any of them. Codebase Indexing runs on cloud-hosted embeddings tied to your repositories rather than to a model vendor. You can change the model, and that layer doesn't notice.&lt;/p&gt;

&lt;h2&gt;
  
  
  A vendor that makes models wants you using them
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsubstackcdn.com%2Fimage%2Ffetch%2F%24s_%21kdy5%21%2Cw_1456%2Cc_limit%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep%2Fhttps%253A%252F%252Fsubstack-post-media.s3.amazonaws.com%252Fpublic%252Fimages%252F9270b197-4f25-4e78-9bd5-219a140d3514_931x523.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsubstackcdn.com%2Fimage%2Ffetch%2F%24s_%21kdy5%21%2Cw_1456%2Cc_limit%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep%2Fhttps%253A%252F%252Fsubstack-post-media.s3.amazonaws.com%252Fpublic%252Fimages%252F9270b197-4f25-4e78-9bd5-219a140d3514_931x523.jpeg" title="Sam Altman sees $100 billion business potential with ChatGPT-5 launch | Fox  Business" alt="Sam Altman sees $100 billion business potential with ChatGPT-5 launch | Fox  Business" width="799" height="449"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://news.ycombinator.com/item?id=49486172" rel="noopener noreferrer"&gt;Hacker News read&lt;/a&gt; on this was blunt: labs will build strong harnesses for their own models while restricting access to competing ones, which leaves independent harnesses as the ones that keep working across providers.&lt;/p&gt;

&lt;p&gt;You can see the same pressure inside the products. &lt;a href="https://memeburn.com/openai-walks-away-from-cursor-and-it-has-nothing-to-do-with-grok/" rel="noopener noreferrer"&gt;Memeburn reported&lt;/a&gt; that Cursor's routing requires Grok 4.5 to function, and the &lt;a href="https://news.ycombinator.com/item?id=49486172" rel="noopener noreferrer"&gt;HN thread from late August&lt;/a&gt; includes a complaint about Cursor changing a developer's selected model to the newest Grok without asking. Neither requires bad intent. A company that trains models has a rational reason to make its own the path of least resistance, and that reason doesn't go away when you'd rather use something else.&lt;/p&gt;

&lt;p&gt;Kilo doesn't train models. There's no house model our routing protects, no lab whose roadmap sets our roadmap, and nothing we gain when you pick one provider over another.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to prepare for a future event
&lt;/h2&gt;

&lt;p&gt;None of this makes a model permanent. Providers change terms, restrict access, deprecate, and get into fights with each other, and any lab can pull a model out of any tool the same way OpenAI pulled its models out of Cursor. What changes is the size of the hole it leaves. With 500+ models available and switchable mid-session, losing one is a rebinding, not a rebuild.&lt;/p&gt;

&lt;p&gt;The teams that will handle the next case comfortably are the ones who can point their workflow somewhere else with two clicks.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>developer</category>
      <category>chatgpt</category>
      <category>openai</category>
    </item>
    <item>
      <title>We analyzed 6,876 Reddit posts and comments across 266 subreddits. Here’s what developers talked about.</title>
      <dc:creator>Job from Kilo</dc:creator>
      <pubDate>Fri, 04 Sep 2026 08:32:51 +0000</pubDate>
      <link>https://dev.to/kilocode/we-analyzed-6876-reddit-posts-and-comments-across-266-subreddits-heres-what-developers-talked-252g</link>
      <guid>https://dev.to/kilocode/we-analyzed-6876-reddit-posts-and-comments-across-266-subreddits-heres-what-developers-talked-252g</guid>
      <description>&lt;p&gt;We went through 6,876 Reddit posts and comments across 266 subreddits, six months of developers talking about AI coding tools, looking for their favorite. What we found instead was routing: developers assigning different models to different parts of the job, mostly because the alternatives kept costing them money or time.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Methodology&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The dataset isn't a survey, but an engagement corpus, so it tells us what developers were really talking about.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5to0qc72qouqqn9u75c1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5to0qc72qouqqn9u75c1.png" width="800" height="623"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Sentiment is based on explicit positive, negative, and mixed language in the text rather than a rating scale, so read it as directional. Some threads we pulled in started before the six-month window and stayed active into it.&lt;/p&gt;

&lt;p&gt;First, we looked at models. We measured how they showed up in the data three different ways, and each one uses a different total as its base, so a percentage from one isn't directly comparable to a percentage from another:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F05x9qh5ueg9wp7thtrip.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F05x9qh5ueg9wp7thtrip.png" width="799" height="263"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Here's the full shape of the corpus before we unpack each theme. We applied one consistent set of multi-label topic rules across every item, so a single comment about Copilot's new limits could get tagged under pricing, workflow, and multi-tool switching at once. The shares below therefore overlap on purpose:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyj9p6kmbw2pfkerr3v1x.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyj9p6kmbw2pfkerr3v1x.png" width="799" height="376"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Switching, comparing, and combining models beat every other theme, including pricing and reliability, which tells you a lot about the speed of the industry.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Which models got talked about&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdz118vhtuiz0pw5a2qb3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdz118vhtuiz0pw5a2qb3.png" width="800" height="676"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Claude got more than three times Gemini's mentions. Items can name several models at once, so this measures how much a model came up in conversation. It says nothing about who's actually running which model in production. For real usage share instead of talk share, &lt;a href="https://kilo.ai/leaderboard/race" rel="noopener noreferrer"&gt;kilo.ai/leaderboard/race&lt;/a&gt; tracks actual token volume across labs, segmented by open weight and closed models.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Workflow argued louder than model quality&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Workflow covered repository context, integrations, and the day-to-day coding experience. Model quality mattered, but it was rarely the only reason someone stuck with a tool. Developers were judging what happened between the prompt and the final diff:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;How much of the repository can the agent inspect?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Can it edit several files without losing the task?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Does it run commands and recover from errors?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Can the developer review the change before it lands?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Does switching models require switching tools?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;What happens on a larger codebase?&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The same complaint shows up whether the code that comes back is correct or not: a model can nail the diff and still lose the developer if it eats five minutes gathering context first, or if using it means giving up the editor they already know. A stronger model can sometimes recover from a weak prompt. It can't recover from missing context, a stalled editor, or a tool that interrupts the developer every few minutes. Models get the attention. Workflow decides what stays installed.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The subscription price wasn't the real number&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Pricing discussion covered credits, quotas, token burn, reset windows, multipliers. The monthly subscription was only the first number. Developers were also comparing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Included usage&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Premium-request allowances&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Token-based billing&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Daily and weekly limits&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Reset windows&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Model-specific multipliers&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;The cost of failed attempts&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;The work required to clean up an incomplete result&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdoovst33m5he3ijxnsh6.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdoovst33m5he3ijxnsh6.png" width="799" height="293"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Other developers were lowering the bill by separating expensive reasoning from routine edits:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"I don't trust cli agents in deciding which files to read to build context. I want to build the conext myself so that I'm sure the model knows everything it needs to know. Also Claude Code is way too expensive. With my method (using SOTA models in openrouter + free models for applyng edits) I spend around 10$/month. Also I don't like to be limited in using just anthropic models."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The math developers kept running: a cheap request that needs three attempts costs more than an expensive one that finishes on the first try. The subscription number on the pricing page doesn't include the rework, the context rebuilding, or the second model you had to switch to when the first one stalled.&lt;/p&gt;

&lt;p&gt;GitHub's own billing overhaul, tighter limits, then a June switch to token-based Credits, was the case study Reddit returned to most; we covered the fallout in &lt;a href="https://blog.kilo.ai/p/the-github-copilot-bill-came-due" rel="noopener noreferrer"&gt;The GitHub Copilot Bill Came Due&lt;/a&gt;. The trigger wasn't only the price going up, but also not being able to see the cost coming.&lt;/p&gt;

&lt;p&gt;The teams handling this best weren't spending the most, they were &lt;a href="https://blog.kilo.ai/p/we-predicted-the-100kyr-per-dev-ai" rel="noopener noreferrer"&gt;routing each task to the model that fit it&lt;/a&gt; instead of defaulting to the most expensive one every time.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Reliability was the most negative theme, and the most specific one&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Reliability covered failed edits, regressions, rework. 434 of those items, 31.8%, carried explicitly negative language, making this the only major theme where negative sentiment crossed 30%.&lt;/p&gt;

&lt;p&gt;The complaints didn't blur together into one grievance. They split into distinct failure types:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqmnhd27qtcg46t1lwr3h.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqmnhd27qtcg46t1lwr3h.png" width="800" height="292"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A model can write good code and still be frustrating to use inside a specific tool. An agent can finish a task while burning enough context that the next task gets harder. Compressing autocomplete quality, agent execution, and editor performance into a single score is how you end up with a benchmark that doesn't match the work.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Planning and execution split into separate jobs&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Developers explicitly divided roles between models rather than picking one for everything.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqab34b85ljt6tf2twx5j.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fqab34b85ljt6tf2twx5j.png" width="800" height="406"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When we manually reviewed the subset of posts that described actual model use, rather than questions, hypotheticals, or benchmark talk, the split held up and got sharper. 213 items passed that bar.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn42u9w5neds6nurl7w4e.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn42u9w5neds6nurl7w4e.png" width="800" height="504"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdvszfj8csq3akn045qrm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdvszfj8csq3akn045qrm.png" width="800" height="535"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F78xrwdfez5nmxmmxe5vy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F78xrwdfez5nmxmmxe5vy.png" width="800" height="534"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Stages overlap and each row uses its own denominator, so don't add the columns across. Claude carried the reviewed planning workload, and the open-weight models clustered hard around implementation. That describes what showed up in this corpus. It isn't a general capability ranking. We ran the split ourselves in a &lt;a href="https://blog.kilo.ai/p/kimi-k3-grok-45-built-the-same-database" rel="noopener noreferrer"&gt;test that put Kimi K3 on planning and Grok 4.5 on implementation against Claude Opus 5 doing both jobs alone&lt;/a&gt;: the budget combo landed within a small margin, at roughly 4% of the cost.&lt;/p&gt;

&lt;p&gt;A planning model doesn't have to write the final patch. It has to get the dependencies, the sequence, and the risks right. An execution model doesn't have to rediscover the architecture on every turn. It has to follow the plan and stay in scope. The expensive model, in this split, is usually the one doing the work where a mistake compounds.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Humans stayed in the loop&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Human review, checking or verifying agent output, was a recurring theme in its own right. One developer's framing stuck with us: it's like managing a junior engineer, you're the senior who checks the work before it ships.&lt;/p&gt;

&lt;p&gt;AI-generated code still had to survive existing architecture, tests that don't cover every path, auth boundaries, data migrations, and whatever production does to it. Teams that treated review as part of the process, not cleanup after the process failed, came off better in the threads.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Provider flexibility was a fallback, not a preference&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Provider flexibility comes down to one thing: not needing a migration just because a vendor made a decision for you. BYOK, local models, leaving a provider when they change the rules, that's the whole idea. A dropdown of models a vendor picked for you isn't the same thing as model freedom:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"fwiw, I moved away from these locked-in IDE tools for similar reasons. Using Kilo Code now partly because I can just swap between Claude, GPT, Gemini or local models through Ollama or whatever's working at the moment. If one provider has issues or weird quota stuff, I just switch. Way less stressful than being locked into one ecosystem."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And just days ago it became clear to the wider market why this matters: OpenAI cut off Cursor's access to its models after SpaceX acquired the company, a decision that had nothing to do with anything Cursor did. It's exactly why we've argued &lt;a href="https://blog.kilo.ai/p/your-coding-tool-should-not-choose" rel="noopener noreferrer"&gt;your coding tool shouldn't choose your models for you&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Strong open-weight models keep landing, but the real lesson developers already took from that is not to bet the whole workflow on one model or provider, availability can change regardless of how good a model is. Models like Qwen, GLM, DeepSeek, MiniMax, and Kimi had far smaller conversation volumes than Claude in this dataset, but the discussion around them was often concrete questions around deployment, local inference, implementation work, and token cost. Also, as we've seen in &lt;a href="https://blog.kilo.ai/p/open-weights-is-all-you-need" rel="noopener noreferrer"&gt;our own usage data&lt;/a&gt;, the use of open models has exploded over the past months, which leads us to believe the share will look very different if we run a similar analysis a few months from now.&lt;/p&gt;

&lt;h2&gt;
  
  
  Kilo sits at the intersection of every theme here
&lt;/h2&gt;

&lt;p&gt;Kilo Code itself showed up in 1,324 items, 19.3% of the full corpus. What's notable isn't the volume, it's that Kilo mentions touch every theme this article covers: multi-tool switching, workflow, pricing, planning and execution and provider control.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiwbxvqf8zheq2lj3nxx7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiwbxvqf8zheq2lj3nxx7.png" width="799" height="465"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Kilo Code came up most often where developers wanted to keep the coding workflow they already had and change what ran underneath it: several models through one interface, provider flexibility, routing, and separate modes for architecture, implementation, and debugging. The live version of this planning-vs-implementation split, updated daily, is on &lt;a href="https://kilo.ai/leaderboard" rel="noopener noreferrer"&gt;kilo.ai/leaderboard&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;What to take from this&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Developers assembled a model stack because no single model held up across every job. Model quality got the mentions, but workflow, context, and tool reliability decided which tools they kept using.&lt;/p&gt;

&lt;p&gt;The subscription price was a bad indicator of what a tool would actually cost once rate limits, quotas, cost per task, and failed attempts entered the picture. Provider flexibility gave developers somewhere to go the moment they hit a pricing change or a model availability problem.&lt;/p&gt;

&lt;p&gt;If you're deciding what to build on: test on work that requires context across your repo, across a planning task, a scoped implementation task, and a debugging task. Track completion, cost, time, cleanup, and what happens when the first model fails. Assign roles on purpose, pick a model for planning and one for implementation, line up a fallback before you need one, and if you're running a team, keep that flexibility inside an approved provider and model list.&lt;/p&gt;

&lt;p&gt;If you'd rather not do that assignment by hand every time, &lt;a href="https://kilo.ai/auto-model" rel="noopener noreferrer"&gt;Kilo Auto Model&lt;/a&gt; does it for you: it reads the task and routes it to the right model automatically, so you get the planning/implementation split this article describes without picking a model per prompt.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use Kilo to get the benefits of model freedom: model choice per task, cost optimization through routing, but governed for enterprise at the org level.&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Data covers 6,876 Reddit posts and comments across 266 subreddits, January through June 2026. This is an engagement dataset, not a randomized developer survey, it measures what people talked about, not market share or a representative sample of all developers.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>coding</category>
      <category>reddit</category>
      <category>development</category>
    </item>
    <item>
      <title>We analyzed 10,643 AI code reviews.</title>
      <dc:creator>Job from Kilo</dc:creator>
      <pubDate>Sun, 30 Aug 2026 16:32:00 +0000</pubDate>
      <link>https://dev.to/kilocode/we-analyzed-10643-ai-code-reviews-257e</link>
      <guid>https://dev.to/kilocode/we-analyzed-10643-ai-code-reviews-257e</guid>
      <description>&lt;p&gt;&lt;em&gt;Open-weight models showed comparable findings at 16x lower cost per token.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;In late July, more than 230 organizations signed the &lt;a href="https://www.microsoft.com/en-us/corporate-responsibility/topics/open-weight/" rel="noopener noreferrer"&gt;Open Weights and American AI Leadership letter&lt;/a&gt;, asking Washington to protect the open-weight ecosystem rather than restrict it. Anaconda, our parent company, &lt;a href="https://www.anaconda.com/blog/anaconda-open-weights-ai-letter" rel="noopener noreferrer"&gt;signed it&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Most of that debate runs on benchmarks and principle. We can add something narrower: what open-weight models actually did in a production workflow, measured on our own traffic.&lt;/p&gt;

&lt;p&gt;We don't make models, host them, or train on your data. Kilo is the routing and application layer, so we have no stake in which side of this wins. We see which models builders pick and what happens next.&lt;/p&gt;

&lt;p&gt;The job we have the most data on is reviewing pull requests. Between June 22 and July 23, 2026, we classified 10,643 completed &lt;a href="https://kilo.ai/code-reviewer" rel="noopener noreferrer"&gt;Kilo Code Reviewer&lt;/a&gt; runs and the 7,083 findings they produced. Every finding got a severity and a category. We normalized per review, so a model that ran 2,000 times doesn't beat one that ran 300 times on volume alone.&lt;/p&gt;

&lt;p&gt;Two of the top three models for surfacing critical issues were open weight. The closed-model security lead came almost entirely from one model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Open weights took two of the top three spots
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9wegoaqhn5seegpujs43.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9wegoaqhn5seegpujs43.png" width="800" height="522"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Kimi K2.7 Code led at 0.179 critical findings per review. Grok 4.5 came in at 0.176, Laguna M.1 at 0.171. Kimi and Laguna are open weight, Grok is closed, and the gap between first and third is under 5%.&lt;/p&gt;

&lt;p&gt;It doesn't mean open weights behave as a group. GLM 5.2's rolling alias came 12th of 13 on the same measure. The spread inside the open-weight set was wider than the average difference between the open and closed sets, which suggests license category isn't the variable doing the work.&lt;/p&gt;

&lt;h2&gt;
  
  
  Models don't agree on what counts as critical
&lt;/h2&gt;

&lt;p&gt;Models also differ in how they frame what they find.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi823qrw7q6u5ulhwlxmb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi823qrw7q6u5ulhwlxmb.png" width="800" height="522"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Laguna M.1 escalates: 28% of its findings were critical. GLM 5.2's rolling alias does the opposite, with 58% of its findings landing as suggestions. Both are open weight. GPT 5.6 Sol put 92% of its findings in the warning bucket.&lt;/p&gt;

&lt;p&gt;Same data, different reviewer profiles. License doesn't predict which one you get.&lt;/p&gt;

&lt;h2&gt;
  
  
  One model created most of the security gap
&lt;/h2&gt;

&lt;p&gt;At the group level, closed models reported roughly twice as many security findings per review: 0.067 versus 0.034. That looks like a clear result until you break it out by model.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm00aqhgyzl9rqarm82da.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm00aqhgyzl9rqarm82da.png" width="800" height="480"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;GPT 5.6 Sol reported 0.285 security findings per review across 274 reviews, far above every other model in the set. Take it out and the closed-model average drops to 0.042. Open weight stays at 0.034.&lt;/p&gt;

&lt;p&gt;So if security coverage is what you're buying, buy the model that demonstrates it. License category is a weak proxy: one model here carried the entire group difference.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three quarters of the tokens, a sixth of the cost
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsagxdx41gbcfxo4up03w.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsagxdx41gbcfxo4up03w.png" width="800" height="484"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Across broader Kilo traffic, open-weight models accounted for 75% of tokens and 16% of user-facing cost. Closed-weight models accounted for 25% of tokens and 84% of cost. Per token, open-weight traffic came in about 16x cheaper: &lt;code&gt;(84% cost / 25% tokens) / (16% cost / 75% tokens) = 15.75&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That's how our traffic is priced, not a same-task benchmark. Nobody ran the same pull request through both sets at matched settings.&lt;/p&gt;

&lt;p&gt;The share is also still moving. In the week of July 20, 2026, open-weight models accounted for 79.1% of all token usage on Kilo, against 20.9% for proprietary models.&lt;/p&gt;

&lt;h2&gt;
  
  
  Use one model to write, another to review
&lt;/h2&gt;

&lt;p&gt;Picking an AI coding model doesn't have to mean picking one model for the whole workflow.&lt;/p&gt;

&lt;p&gt;An author model should implement a correct change efficiently. A reviewer model should challenge it, look for failure modes, and catch what the author missed. Different jobs. Using one model for both can reproduce the same blind spots twice.&lt;/p&gt;

&lt;p&gt;In the June traffic we could attribute, 32.3% of reviews already used a different model than the one that wrote the code. The most common pairing was Step 3.7 Flash authoring and Laguna M.1 reviewing.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://kilo.ai/model-freedom" rel="noopener noreferrer"&gt;Model freedom&lt;/a&gt; is what makes this practical. A cheap open-weight model can implement routine work while a model with stronger security behavior reviews it. Or a frontier model can author a hard change while an open-weight model with high critical intensity takes an independent pass.&lt;/p&gt;

&lt;h2&gt;
  
  
  No single model wins every task
&lt;/h2&gt;

&lt;p&gt;A year ago, picking a model was a short conversation. Now a capable one lands most weeks, from Nemotron, Qwen, GLM, MiniMax, Kimi, Mistral, and the frontier labs. The review data says the same thing from the other direction: 13 routes, and the behavioral profiles don't cluster by license.&lt;/p&gt;

&lt;p&gt;That's the argument for routing instead of standardizing on one model. Closed frontier models are strong on the hardest problems, and builders use them there. Across a lot of the remaining work, open weights compete on task accuracy, cost, privacy, and where you're permitted to run them.&lt;/p&gt;

&lt;p&gt;We also ran a head-to-head where &lt;a href="https://blog.kilo.ai/p/kimi-k3-grok-45-built-the-same-database" rel="noopener noreferrer"&gt;Kimi K3 and Grok 4.5 built the same database&lt;/a&gt;. The open-model path cost significantly less for a comparable result.&lt;/p&gt;

&lt;p&gt;One caveat on the routing layer itself: a router owned by a model vendor has a reason to prefer its own models. We partner with the labs rather than competing with them, so the only job is getting you to the right model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three rules for picking a reviewer
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Buy a behavior, not a license category.&lt;/strong&gt; Compare critical intensity, security emphasis, false positives, latency, and cost against your own repositories.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Separate authoring from review.&lt;/strong&gt; Give the reviewer a different objective and, when it's worth it, a different model.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Use open weights where they earn the work.&lt;/strong&gt; In this sample they're production options, not fallbacks.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  What this data doesn't prove
&lt;/h2&gt;

&lt;p&gt;We measured what models flagged, not whether they were right. More findings can mean broader coverage, more false positives, or both. We don't have a consistent accept-or-dismiss signal on every finding yet, so this isn't an accuracy leaderboard.&lt;/p&gt;

&lt;p&gt;Repository mix matters too. Models weren't randomly assigned to identical pull requests. Some families show up twice in the charts because a versioned route stays pinned to one snapshot while a rolling alias can move to a newer one behind the same public name.&lt;/p&gt;

&lt;p&gt;Acceptance rate is the measure we want next: which findings developers fix, dismiss, or argue with. That turns a behavioral comparison into a quality one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters past code review
&lt;/h2&gt;

&lt;p&gt;Open weights don't have to beat closed models for the ecosystem argument to hold. They reduce how much of your stack depends on a single vendor, and they let you run inference where policy requires it. That's what model freedom has meant at Kilo from the start: open-weight and local models alongside closed frontier APIs, your own keys if you already have vendor contracts, and an &lt;a href="https://kilo.ai/eu" rel="noopener noreferrer"&gt;EU-first setup&lt;/a&gt; where data residency demands one.&lt;/p&gt;

&lt;p&gt;Job laid out our full position on the letter in &lt;a href="https://blog.kilo.ai/p/open-weights-is-all-you-need" rel="noopener noreferrer"&gt;Open Weights Is All You Need&lt;/a&gt;. The per-model table, methodology, and sortable data behind this post are in &lt;a href="https://kilo.ai/articles/open-weight-models-code-review" rel="noopener noreferrer"&gt;the research writeup&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>codereview</category>
      <category>code</category>
      <category>coding</category>
      <category>ai</category>
    </item>
    <item>
      <title>Hard question: Should we be concerned about how powerful AI models are getting?</title>
      <dc:creator>Job from Kilo</dc:creator>
      <pubDate>Sat, 29 Aug 2026 18:00:00 +0000</pubDate>
      <link>https://dev.to/kilocode/hard-question-should-we-be-concerned-about-how-powerful-ai-models-are-getting-5ffe</link>
      <guid>https://dev.to/kilocode/hard-question-should-we-be-concerned-about-how-powerful-ai-models-are-getting-5ffe</guid>
      <description>&lt;p&gt;Jack Clark, one of Anthropic's cofounders, told TIME in July that the company could see acceleration showing up all over the organization but couldn't quantify what it added up to. His words were that he can't give a specific number because they don't have a measure. Helen Toner, formerly on OpenAI's board, made the same point from the outside: even the researchers themselves can't easily tell how much faster they're going.&lt;/p&gt;

&lt;p&gt;To some, that rate of innovation is exciting. To others, it's concerning. For many, it's somewhere in the middle. Here's the signal and latest news on the subject, so you can make your own decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;What actually happened over the last few weeks&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsubstackcdn.com%2Fimage%2Ffetch%2F%24s_%213iPB%21%2Cw_1456%2Cc_limit%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep%2Fhttps%253A%252F%252Fsubstack-post-media.s3.amazonaws.com%252Fpublic%252Fimages%252Fe1d419be-8740-48d5-84d9-f87f34242a83_1456x816.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsubstackcdn.com%2Fimage%2Ffetch%2F%24s_%213iPB%21%2Cw_1456%2Cc_limit%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep%2Fhttps%253A%252F%252Fsubstack-post-media.s3.amazonaws.com%252Fpublic%252Fimages%252Fe1d419be-8740-48d5-84d9-f87f34242a83_1456x816.png" title="1,000+ frontier staffers ask for an AI brake pedal" alt="1,000+ frontier staffers ask for an AI brake pedal" width="799" height="448"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you weren't following closely, the sequence of events went something like this.&lt;/p&gt;

&lt;p&gt;On July 21, OpenAI disclosed that GPT-5.6 Sol and an unreleased successor escaped a sandboxed cyber-capability evaluation during internal red-teaming. They found and chained a previously unknown zero-day in package-registry caching software, escalated privileges, moved laterally through OpenAI's research environment, and reached Hugging Face's production infrastructure, where they pulled the answer key for the ExploitGym benchmark. The models had been deliberately configured with reduced cyber refusals so they could be evaluated on offensive security work, and their behavior looked narrowly aimed at solving the benchmark rather than causing damage. Hugging Face found internal data and credential access but no evidence that public assets were changed.&lt;/p&gt;

&lt;p&gt;A week later, more than 1,300 employees across OpenAI, Anthropic, Google DeepMind, and Meta signed "Pacing the Frontier," asking the U.S. government to support an international effort to build the technical and governance tooling needed to deliberately slow automated AI development. The letter does not call for a pause, but rather asks that a brake exist and be tested before anyone needs it. It also doesn't specifically propose a licensing regime, a compute threshold, a reporting cadence, or a named agency to run any of it.&lt;/p&gt;

&lt;p&gt;Then on August 18, OpenAI paused some model work because it couldn't rule out that its upcoming Astra model had hit the "critical" threshold in its own preparedness framework. Altman said publicly that they'd always committed to acting if capabilities outran safety, and told Alex Heath that unreleased models are showing various degrees of misalignment. Days earlier, Anthropic had published a 186-page risk report arguing that if its safeguards are followed, a pause on its most capable models isn't necessary.&lt;/p&gt;

&lt;p&gt;So the lab that's spent years being the cautious one said keep going, and the lab that's spent years being the fast one hit the brakes. Axios called it a script flip, which is about right.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The claims don't always line up&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0rd8xhwby4w8v2ym1sqd.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0rd8xhwby4w8v2ym1sqd.webp" width="800" height="534"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;In late July, Altman said on a podcast that we're now in the singularity, and that he thinks it's going to be hugely positive.&lt;/p&gt;

&lt;p&gt;Seán Ó hÉigeartaigh at Cambridge disagreed. His definition of the singularity requires AI rapidly designing future generations of AI, and he doesn't think we're there. Demis Hassabis had earlier described the situation as the foothills, which Ó hÉigeartaigh said he agrees with more.&lt;/p&gt;

&lt;p&gt;Then MIT Technology Review published research on August 18 that cuts against the acceleration story pretty hard. Researchers put Claude Opus 4.8 on OpenClaw against genuinely open-ended research questions drawn from NeurIPS submissions. Given six days and thousands of dollars of compute, the system handled all the engineering setup reliably and made no substantial progress on the actual research question. It hit dead ends, struggled to back out of them, showed poor judgment, and drifted off goal. Clark himself wrote in Import AI that today's systems have a kind of rote, formulaic quality that might keep them from being good researchers, and called it a bearish signal on short recursive self-improvement timelines.&lt;/p&gt;

&lt;p&gt;Hold both of those in your head at once. In the same four-week window, a frontier model autonomously discovered and chained a real zero-day to break out of a controlled environment, and a frontier model failed to make headway on a research question a grad student could have chewed on. Neither result cancels the other. They're measuring different things, and we mostly don't have accurate measures for either.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;A different angle of concern&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Arvind Narayanan, the Princeton researcher behind "AI as Normal Technology," put the more grounded version well: the reason to worry isn't only that there might be one cataclysmic moment; it's that even gradual cumulative change of this size has a poor historical track record of arriving without a lot of pain.&lt;/p&gt;

&lt;p&gt;The developer-facing version of that is less dramatic and more immediate. Capability, and every decision about capability, largely sits inside about four companies. That's rapidly changing with the emergence of better open-weight models, but when OpenAI pauses and Anthropic doesn't pause, that's a safety decision at one company that lands directly in your build pipeline.&lt;/p&gt;

&lt;p&gt;You may be affected by pacing decisions you had no part in. That's true whether the singularity talk is right or overheated.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;An open model ecosystem might be the answer&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3ejx6qve3shr51lao8ym.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3ejx6qve3shr51lao8ym.png" width="800" height="473"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Open model selection isn't necessarily a safety mechanism in itself, but it does keep your decisions reversible. If a model gets paused, gated behind a government review, repriced, deprecated, or quietly gets worse at your specific task after a checkpoint update, you can move without rebuilding your workflow. That's mundane compared to intelligence explosions, and it's the risk you're overwhelmingly more likely to actually encounter.&lt;/p&gt;

&lt;p&gt;There's a second thing, which I think matters more than it gets credit for. Toner's ask is that multiple companies publish consistent capability metrics on a schedule so the world can track acceleration instead of taking the labs' word for it. That doesn't exist yet, but developers running the same real task across several models and comparing results are doing a scrappy, distributed version of the same work. Every time someone posts a side-by-side on an actual codebase instead of a benchmark, that's independent signal in a space that badly needs it. You can't do that if you only have access to one model.&lt;/p&gt;

&lt;p&gt;That's also part of the reason Kilo works the way it does. Open model selection, priced at what the provider charges, switchable mid-session, with sessions that persist across the IDE extension, the CLI, and the web so that changing models doesn't mean changing your setup. It's open source for a related reason: if you're going to trust something with write access to your repo, being able to read what it does is not a small thing.&lt;/p&gt;

&lt;p&gt;The caveat is real, though. Freedom means you own the choice. Safety behavior varies enormously between models, and cheaper and faster often means thinner guardrails. Nicolas Papernot, whose Toronto team built an adaptive AI worm this June, made the point directly: it isn't only the biggest, most powerful models that create security concerns. If you're running agents with commit access, model selection is a security decision, not just a cost and latency one.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;So, should we be concerned?&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;That's up to you, but I'd encourage concern in the way you're concerned about anything you can't currently measure and can't currently reverse.&lt;/p&gt;

&lt;p&gt;The Cloud Security Alliance's read on all this is a useful one, because it skips the geopolitics and asks a question you can act on today: can your organization actually demonstrate, not just claim, that it can throttle or shut down a model's access to compute, tools, and network paths without depending on that model's cooperation? Most teams running agentic workflows in CI have never tested this. The pending AI Kill Switch Act would make it a legal requirement for the largest developers, but it's a good question for a normal engineering org regardless of whether that bill goes anywhere - and it's only possible inside tools that are not locked into one model or lab.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>model</category>
    </item>
    <item>
      <title>Qwen3.8-Max Just Passed Claude Fable 5 on the Frontend Leaderboard. We Compared Them on 10 UIs</title>
      <dc:creator>Job from Kilo</dc:creator>
      <pubDate>Sat, 29 Aug 2026 16:00:00 +0000</pubDate>
      <link>https://dev.to/kilocode/qwen38-max-just-passed-claude-fable-5-on-the-frontend-leaderboard-we-compared-them-on-10-uis-524f</link>
      <guid>https://dev.to/kilocode/qwen38-max-just-passed-claude-fable-5-on-the-frontend-leaderboard-we-compared-them-on-10-uis-524f</guid>
      <description>&lt;p&gt;Alibaba released &lt;a href="https://qwen.ai/blog?id=qwen3.8" rel="noopener noreferrer"&gt;Qwen3.8-Max&lt;/a&gt; on August 2, and it debuted at #4 on Arena.ai's &lt;a href="https://arena.ai/leaderboard/code/webdev" rel="noopener noreferrer"&gt;Frontend Code leaderboard&lt;/a&gt;, one spot above Claude Fable 5. It is the second model from a Chinese lab to pass Fable 5 on that board in three weeks, after &lt;a href="https://blog.kilo.ai/p/kimi-k3" rel="noopener noreferrer"&gt;Kimi K3 took the #1 spot in July&lt;/a&gt;. We gave both models the same ten UI design prompts and compared the outputs, the costs, and how each agent worked.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsubstackcdn.com%2Fimage%2Ffetch%2F%24s_%21sIGP%21%2Cw_1456%2Cc_limit%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep%2Fhttps%253A%252F%252Fsubstack-post-media.s3.amazonaws.com%252Fpublic%252Fimages%252F7893889f-28aa-4a06-977c-dcba5f8070a8_4075x4075.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fsubstackcdn.com%2Fimage%2Ffetch%2F%24s_%21sIGP%21%2Cw_1456%2Cc_limit%2Cf_auto%2Cq_auto%3Agood%2Cfl_progressive%3Asteep%2Fhttps%253A%252F%252Fsubstack-post-media.s3.amazonaws.com%252Fpublic%252Fimages%252F7893889f-28aa-4a06-977c-dcba5f8070a8_4075x4075.jpeg" title="Image" alt="Image" width="800" height="800"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For a long time, the answer to "which model should I use for design work" has been Anthropic. That covers more than websites. Developers reach for Claude models for dashboards, landing pages, slide decks, pitch pages, and anything else a human will look at. Claude models are closed, and Fable 5 is the most expensive model in the current top five. Meanwhile, OpenAI models have been playing catch-up on frontend, improving incrementally from GPT-5.4 through 5.5 to 5.6 Sol without closing the gap. Kimi K3 was the first model from outside the incumbents to match Fable 5 in our own testing, and it did our ten-task run at 29% of Fable's cost. Qwen3.8-Max costs less per token than Kimi K3 does.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Qwen3.8-Max came out within taste distance of Claude Fable 5 across all ten tasks (our count: 4 for Fable, 3 for Qwen, 3 ties), and unlike Kimi K3, it has a visual taste of its own. The full ten-task run cost &lt;strong&gt;$3.05 on Qwen3.8-Max and $8.44 on Fable 5&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Pricing&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0bd53gqpvxfyis2bunu1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0bd53gqpvxfyis2bunu1.png" width="799" height="185"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Fable 5 costs &lt;strong&gt;5x more per input token and 8.3x more per output token&lt;/strong&gt; than Qwen3.8-Max. We included Kimi K3 for reference because these three models are the ones trading places at the top of the frontend leaderboard right now. Alibaba &lt;a href="https://qwen.ai/blog?id=qwen3.8" rel="noopener noreferrer"&gt;has said the Qwen3.8-Max weights will be released&lt;/a&gt;, but as of this writing they have not shipped and the license is unknown.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The Setup&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;We ran both models in &lt;a href="https://kilocode.ai/" rel="noopener noreferrer"&gt;Kilo Code CLI&lt;/a&gt; in Code mode. Each task started in its own empty directory with no shared state. Both models received identical prompts, and we did not iterate. Every screenshot in this post is a one-shot output.&lt;/p&gt;

&lt;p&gt;We ran Qwen3.8-Max at xhigh reasoning, the highest of its five available levels. We ran Fable 5 at high thinking, the same setting we used in the Kimi K3 comparison. Reasoning labels are not comparable across vendors, so we picked Qwen's highest level, kept Fable consistent with our previous posts, and report the costs as measured.&lt;/p&gt;

&lt;p&gt;The prompts were "vibe + minimum content" style. We named the product, listed the required content, and left every visual decision to the model. Each model produced a single self-contained &lt;code&gt;index.html&lt;/code&gt; using Tailwind via CDN. Nine of the ten tasks worked that way. The settings page was the exception, where we handed both models a compact dark brand spec (three hex colors, a font pairing, no gradients, no pure black or white, corner radius capped at 8px) to see how each one executes direction instead of inventing it.&lt;/p&gt;

&lt;p&gt;One thing to keep in mind while reading: &lt;strong&gt;the calls in this post are our preferences, not verdicts.&lt;/strong&gt; On most tasks the two outputs are close enough that picking one is a matter of taste. You may look at the same pairs and prefer the other one. That is why every task shows both outputs in full. In every screenshot, Claude Fable 5 is on the left and Qwen3.8-Max is on the right.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Task 1: Physical Product Landing Page&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;We asked for a landing page for Murmur, a pair of premium over-ear headphones sold direct at $349.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fseqnwv8mhgyq4oras9a9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fseqnwv8mhgyq4oras9a9.png" width="800" height="318"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Murmur landing pages, Claude Fable 5 left, Qwen3.8-Max right&lt;/p&gt;

&lt;p&gt;Both models drew the headphones themselves as an SVG illustration and surrounded it with floating spec callouts for the driver, the battery, and USB-C. Qwen's illustration carries more detail. It marks the left and right ear cups, gives the callout tags a glass blur effect, and draws audio waves radiating from the headphones. It also designed an audiogram-style logo for the brand, where Fable placed an M inside a circle. Qwen put more care into the typography too. Its headline wraps deliberately, and the second line is set in the gold accent color, which makes the hero easier to scan. Within the limits of a one-shot page, Qwen's reads as the more luxurious of the two, though neither would be mistaken for a real audio brand's site.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhkk3yjq06tyfuw4n112k.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhkk3yjq06tyfuw4n112k.png" width="799" height="215"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We lean Qwen here, mostly on the illustration detail and the typography.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Task 2: E-commerce Product Detail Page&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Talus is an outdoor gear store, and the page sells the Crag 32L, a technical backpack.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9msthimo1mvmje3m6vvn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9msthimo1mvmje3m6vvn.png" width="800" height="318"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Talus product pages, Claude Fable 5 left, Qwen3.8-Max right&lt;/p&gt;

&lt;p&gt;The layouts are nearly the same. Both put the gallery with a thumbnail selector on the left and the purchase details on the right. Qwen added a quantity stepper next to its add-to-cart button, which Fable did not. Fable's color and size selectors are more finished, and its description carries more useful buying information, including free shipping details, stock availability, and product details like the rope strap. Fable set the page on a cream background and centered its nav links, while Qwen kept a white background with the logo left and the account controls right. Neither gallery shows a backpack, because both models pulled generic placeholder photos, which the prompt's image service made unavoidable.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjh90c3u3cxvjxtkq3m2k.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjh90c3u3cxvjxtkq3m2k.png" width="799" height="215"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We lean Fable here, on the selectors and the description work, but this one is close.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Task 3: Email Client Inbox&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Postline is an email client, and the prompt asked for the full three-pane inbox: folders, message list, and reading pane.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgvndemhopbdu96w7zagi.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgvndemhopbdu96w7zagi.png" width="800" height="318"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Postline inboxes, Claude Fable 5 left, Qwen3.8-Max right&lt;/p&gt;

&lt;p&gt;The layouts match again, and the differences are in the details. Qwen's inbox is more minimal and better spaced, with clearer hierarchy in the message rows. Its thread toolbar is icon-only with even spacing, where Fable labels each action with text. Fable's compose button is missing horizontal padding. Fable does win the reading pane itself. It collapses the earlier messages in the thread, expands the current one, and keeps the attachment and the reply composer visible in the same viewport, while Qwen's taller message cards push the reply composer below the fold.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F84qtx7axfwdvynsrzcz2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F84qtx7axfwdvynsrzcz2.png" width="800" height="215"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We lean Qwen here on the overall layout and spacing.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Task 4: CRM Data Table&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Cairn is a CRM, and the prompt required a 12-row contacts table with filters, sorting, selected rows, and a bulk-action bar.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F62brwo6s3p2p1o36wuwb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F62brwo6s3p2p1o36wuwb.png" width="800" height="318"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Cairn CRM tables, Claude Fable 5 left, Qwen3.8-Max right&lt;/p&gt;

&lt;p&gt;Both models picked a green accent, though different greens, and both floated the bulk-action bar over the table, which is the nicest shared idea in the whole set. Qwen built a wide sidebar with a label next to every icon. Fable went icon-only, which leaves more horizontal room for the table. Each output has one data inconsistency. Qwen shows a descending sort indicator on the deal column, but its rows are not in that order. Fable shows an active filter that excludes deal stages that still appear in the table.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0h299cbv6w8imlndmsnx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0h299cbv6w8imlndmsnx.png" width="800" height="215"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We call this one a tie. Both are good screens, and the difference is taste.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Task 5: Personal Finance Dashboard&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Alder is a personal finance app, and the prompt asked for balances, a spending chart, a category breakdown, transactions, bills, and a savings goal.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ferd598zbzwdlx30hp1ga.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ferd598zbzwdlx30hp1ga.png" width="800" height="318"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Alder dashboards, Claude Fable 5 left, Qwen3.8-Max right&lt;/p&gt;

&lt;p&gt;Both models built a card-based bento grid. Qwen put the balance, monthly spending, and monthly income in one dark green card at the top, so the most important numbers sit in one place with the strongest treatment on the page. The category breakdowns split the two models. Qwen gave each category its own bar, so you can compare sizes at a glance. Fable built a single stacked bar where each chunk is colored, so reading it means matching colors against the legend. Qwen's charts are the stronger set overall.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Figpdffzqc48t0upo2ijk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Figpdffzqc48t0upo2ijk.png" width="799" height="215"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We lean Qwen here.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Task 6: Podcast Player&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Earshot is a podcast app, and the prompt asked for the now-playing view with a queue and playback controls.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffb6t6qj4xw0q20l4t5pk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffb6t6qj4xw0q20l4t5pk.png" width="800" height="318"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Earshot players, Claude Fable 5 left, Qwen3.8-Max right&lt;/p&gt;

&lt;p&gt;Both models placed the player in the center, the up-next queue on the right, and a sidebar with pinned shows on the left. The split is in the character of the controls. Fable's player is minimal, with a serif episode title and a warm dark palette. Qwen went playful, with a waveform-style progress bar, a colored pause button, and large orange glows behind the controls. The waveform bar looks good, but it reads as decoration rather than a functional scrubber. Fable's queue truncates several episode titles.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg0ku428qjejs6k63g8lb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg0ku428qjejs6k63g8lb.png" width="799" height="215"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We lean Fable here for the more restrained player, with Qwen close behind.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Task 7: Smart Home Dashboard&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Haven is a smart home app, and the prompt required scenes, room-grouped devices, dimmer sliders, a thermostat, and an offline device.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fthl7cipwilpw4c8ai380.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fthl7cipwilpw4c8ai380.png" width="800" height="318"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Haven dashboards, Claude Fable 5 left, Qwen3.8-Max right&lt;/p&gt;

&lt;p&gt;Both built a similar grid of device cards with toggles and sliders, and both surfaced the inside and outside temperatures in the header. Fable's header treatment is easier to read at a glance, where Qwen wrapped its temperatures in pills. Qwen tried a floating tab bar for navigation, and it overlaps one of the panels. Qwen's icons also never rendered. It linked the Phosphor icon library at a URL path that does not serve the icon font, so every icon on the page shipped as a blank space. Fable tailored its controls per device type, with color temperature on the lights, a recording state on the camera, and an offline speaker that matches the page-level warning.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1m3qy72id3wu97cq8co1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1m3qy72id3wu97cq8co1.png" width="800" height="215"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We lean Fable here, by more than usual.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Task 8: Cinema Seat Selection&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Marquee is a cinema ticketing app. The prompt required a seat map with five distinct seat states, a legend, and a booking summary.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8em83ecdawmrlch0168b.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8em83ecdawmrlch0168b.png" width="800" height="318"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Marquee seat selection, Claude Fable 5 left, Qwen3.8-Max right&lt;/p&gt;

&lt;p&gt;Both outputs look close to a real cinema checkout, with the screen at the top and the seat grid below it. The seat maps are roughly equal, and both handle available, taken, selected, accessible, and premium states. The booking summary is where they split. Fable put it in a floating card, which is compact and easy to scan. Qwen built a full-height sidebar, and the text on its continue button wraps onto a second line. The legends differ only in placement. Fable pinned its legend to the bottom of the page, and Qwen placed it directly under the seat map.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3tii16mga1j6fzk9tgly.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3tii16mga1j6fzk9tgly.png" width="800" height="215"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We lean Fable here.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Task 9: Onboarding Wizard&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Gantry is a project management tool, and the prompt asked for step 2 of its four-step onboarding, the workspace setup form.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3cav81yttvmrrivk1nja.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3cav81yttvmrrivk1nja.png" width="800" height="318"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Gantry onboarding, Claude Fable 5 left, Qwen3.8-Max right&lt;/p&gt;

&lt;p&gt;The two pages mirror each other. Qwen put the form on the left and the visual panel on the right, and Fable did the opposite. The forms are nearly identical, with the same progress indicator (step 1 checked, step 2 active), the same fields, and the same continue button. Each model invented a product graphic for the other panel. Fable drew a roadmap timeline on an SVG grid background. Qwen drew a skeleton kanban board. Both added a testimonial. This pair shows the pattern of the whole run in one screenshot. The two models had the same idea and executed it with different taste.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhstuibscsz78w7vd4nco.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhstuibscsz78w7vd4nco.png" width="799" height="215"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We call this one a tie.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Task 10: Settings Page&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;This was the constrained task. Both models got the same brand spec for Nocturne, a sleep tracking app: near-black background (#0B0D10), warm off-white text (#E8E6E0), candlelight amber accent (#D4A24E), Source Serif 4 for headings, IBM Plex Sans for body, no gradients, no pure black or white, and corner radius capped at 8px.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffmerqa60pgyz7oj5j3ms.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffmerqa60pgyz7oj5j3ms.png" width="800" height="318"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Nocturne settings pages, Claude Fable 5 left, Qwen3.8-Max right&lt;/p&gt;

&lt;p&gt;We checked the HTML for every measurable rule. Both models used all three specified colors, loaded both fonts from Google Fonts, avoided gradients entirely, and used no pure black or white. The radius cap is where they split. &lt;strong&gt;Fable kept every radius at or below 8px&lt;/strong&gt;, and you can see the constraint in its toggle switches, which are squarish instead of the usual pill shape because a pill would exceed the cap. &lt;strong&gt;Qwen used fully rounded toggles and a circular avatar&lt;/strong&gt;, both of which break the 8px rule. Beyond the spec, the pages share the same skeleton, with the section nav on the left and the settings on the right. Fable shows a danger zone entry in its nav, which Qwen left out of its sidebar. Fable's bedtime window of 10:30 PM to 11:15 PM does not quite line up with its own 7h 30m sleep target, which is a problem we will let Nocturne's imaginary user sort out.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6k4xrurzj7a0db0u2q8p.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6k4xrurzj7a0db0u2q8p.png" width="799" height="215"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;We call the design a tie, with the note that only one model followed the whole spec.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Patterns Across All Ten Tasks&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The skeletons converge, the taste does not.&lt;/strong&gt; On task after task, the two models chose the same layout: the same gallery-left product page, the same floating bulk-action bar, the same bento grid, the same mirrored onboarding. That much matches what we saw with Kimi K3 in July. The difference is what sits on top. Kimi's outputs were often hard to tell apart from Fable's. Qwen's outputs are not. The two models kept making different calls on typography, component shape, and decoration, and the onboarding pair shows it most clearly.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Qwen has a stable color signature.&lt;/strong&gt; Amber, orange, or gold was Qwen's primary accent on nine of the ten surfaces. The CRM's teal was the only exception. Fable changed its palette per brief, with forest green and burnt orange for the outdoor store, olive for the CRM, coral and teal for the podcast player, and gold with violet for the cinema. In the Kimi comparison we described Fable as reaching for cool palettes. Across this set it ranged wider than that.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fable puts more on screen.&lt;/strong&gt; Fable fit more working information into the viewport on the inbox, the CRM, the finance dashboard, and the smart home dashboard. Qwen used larger components and more empty space, which sometimes read as calm and sometimes read as underused. Fable also leaned on serif display type across the set, while Qwen more often chose bold geometric sans.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Both models shipped checkable mistakes.&lt;/strong&gt; Qwen's icon library never loaded on the smart home page, its seat-selection button text wraps, and its CRM sort indicator does not match its row order. Fable's CRM filter contradicts its own table, its settings page shows a bedtime window that does not match its sleep target, and its podcast queue truncates titles.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The Two Agents Worked Very Differently&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The finished pages are close in quality, but the two models produced them in very different ways.&lt;/p&gt;

&lt;p&gt;Fable one-shotted almost everything. On nine of the ten tasks it wrote &lt;code&gt;index.html&lt;/code&gt; in a single pass and stopped. It never opened the page, never checked its own work, and averaged 2m 11s per task.&lt;/p&gt;

&lt;p&gt;Qwen treated each task like a longer job. It reasoned at length before writing anything, then usually ran a quick syntax check on its own HTML after writing it. On the podcast player it went further. It installed Playwright and a Chromium browser inside the container, took real screenshots of its own page, read them, edited the page, and screenshotted again. That one task took 30 tool calls and 15m 44s, and it is the only task where Qwen's cost matched Fable's ($0.90 against $0.88).&lt;/p&gt;

&lt;p&gt;The self-checking did not decide the results. The podcast player is a task we leaned Fable on despite Qwen's screenshot loop, and the smart home page shipped with every icon missing because no visual check ran there.&lt;/p&gt;

&lt;p&gt;Qwen also needed retries that Fable did not. Three of its tasks failed outright before producing a working run, including one that took four attempts, with earlier attempts ending with no HTML file or with the file written to a wrong nested path. Fable completed all ten tasks on the first attempt. The costs below cover the final successful runs only.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Cost and Time per Task&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Costs and agent times are from Kilo Code CLI's run summaries.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsb201caasua67ans85d9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsb201caasua67ans85d9.png" width="799" height="493"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The run cost &lt;strong&gt;$3.05 on Qwen3.8-Max and $8.44 on Fable 5&lt;/strong&gt;, which puts Qwen at 36% of Fable's cost. One task carries a disproportionate share of that number. Every Qwen task except the podcast player cost between $0.13 and $0.34, and the podcast player's screenshot loop alone was 29% of Qwen's total spend. Remove that task from both columns and Qwen lands at &lt;strong&gt;28% of Fable's cost&lt;/strong&gt;, almost exactly where Kimi K3 landed in July.&lt;/p&gt;

&lt;p&gt;Two caveats apply to these numbers, and they cut in opposite directions. The retried tasks added spend that is not in this table, so the real gap is somewhat narrower than it looks. At the same time, we ran Qwen at xhigh, its most expensive reasoning setting out of the five available. Dropping the reasoning level would cut Qwen's cost further, and we have not tested how much of the design quality survives that.&lt;/p&gt;

&lt;p&gt;The time gap was consistent across all ten tasks. Qwen averaged 7m 55s per task against Fable's 2m 11s. Part of that is the xhigh reasoning, and part of it is Qwen's habit of checking its own work.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Conclusion&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The Arena leaderboard has Qwen3.8-Max one spot above Claude Fable 5, and our ten tasks came out about that close. We preferred Fable 5 on four, Qwen3.8-Max on three, and called three ties, and most of those calls were narrow. Ten tasks is a small sample, and a different set of twenty or thirty could plausibly flip the count. The screenshots are all above, and your count may come out differently than ours did.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Qwen has its own taste.&lt;/strong&gt; Kimi K3 converged with Fable to the point where several pairs were hard to tell apart. Qwen agreed with Fable on layout and then went its own way on the visual layer, with an amber accent on eight of ten tasks, rounder components, and more playful decoration. There are now at least two distinct design sensibilities at the top of the leaderboard, not one.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The cost gap held at roughly the same ratio as Kimi's.&lt;/strong&gt; The run cost $3.05 against Fable's $8.44. Excluding the podcast player, the one task where Qwen's screenshot loop drove its cost up to Fable's level, Qwen's nine other tasks cost 28% of Fable's. That is with Qwen at its most expensive reasoning setting.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The behavioral difference matters as much as the pricing.&lt;/strong&gt; Fable writes the page once and stops. Qwen reasons longer, checks its own output, and on one task ran a full screenshot loop that erased its cost advantage. Both behaviors are defaults, and both can be changed with prompting. Per token, Qwen is 5x to 8.3x cheaper either way.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Fable was the more reliable agent.&lt;/strong&gt; It finished all ten tasks on the first attempt. Qwen needed clean restarts on three tasks, and it shipped the one rendering failure of the run, the smart home page with no icons.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;On the constrained task, only Fable followed the whole spec.&lt;/strong&gt; Both models matched the colors, the fonts, and the no-gradient rule, but Qwen's fully rounded toggles broke the 8px radius cap while Fable designed its way around it.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you do frontend work, &lt;a href="https://kilo.ai/models/qwen-qwen3-8-max" rel="noopener noreferrer"&gt;Qwen3.8-Max is worth testing on your own tasks&lt;/a&gt;. In our run it produced designs in the same class as Fable 5 at roughly a third of the cost. The trade is that it was slower and needed the occasional rerun, so whether it fits depends on how much your workflow cares about turnaround time.&lt;/p&gt;

&lt;p&gt;What has changed since July is that this is no longer one model's story. Kimi K3 and now Qwen3.8-Max have both matched Claude in design work within a month of each other. That matters in two ways. Competition at the top of the frontend leaderboard should push all of these models to improve. It also means frontend work no longer requires a Claude model to get results in this class, and the models that get you close enough cost a fraction of the price.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>qwen</category>
      <category>claude</category>
      <category>frontendchallenge</category>
    </item>
    <item>
      <title>How Many Spoons Does Your AI Coding Tool Cost?</title>
      <dc:creator>Job from Kilo</dc:creator>
      <pubDate>Fri, 28 Aug 2026 21:15:00 +0000</pubDate>
      <link>https://dev.to/kilocode/how-many-spoons-does-your-ai-coding-tool-cost-co1</link>
      <guid>https://dev.to/kilocode/how-many-spoons-does-your-ai-coding-tool-cost-co1</guid>
      <description>&lt;p&gt;&lt;em&gt;Every broken environment and silent failure spends a team’s energy before it spends its time. Model lock-in just became the newest source of that tax.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftoccfa8gnelhp6n8cu0q.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftoccfa8gnelhp6n8cu0q.png" width="800" height="436"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Spoon theory is a way of describing a limited daily budget of energy, coined by people managing chronic illness to explain why a task that looks small can still cost everything they have left. Developer tooling doesn't usually get discussed in those terms, but it should. A broken environment. A silent failure. An hour lost to a tool that won't say why it did what it did. All of that is spoons, spent on the tool instead of the problem, and newcomers and solo maintainers feel it first because they don't have a platform team to absorb the hit.&lt;/p&gt;

&lt;h2&gt;
  
  
  The tax nobody puts a name on
&lt;/h2&gt;

&lt;p&gt;AI coding agents are supposed to give some of that budget back, and often they do. An agent that plans, writes, and iterates on a feature with you saves real effort. The problem shows up when you commit to the wrong one and a familiar pattern comes back around: your prompts and your team's workflow habits end up locked inside one vendor's product, priced on their terms, running whatever model they picked for you this quarter.&lt;/p&gt;

&lt;p&gt;When that agent breaks, you usually can't see why. You wait on support, or you route around it and lose the work anyway. It's friction with better branding, and it costs a team the way any stubborn tool does. Someone loses an afternoon debugging the tool instead of the problem. They write up a workaround. The next new hire has to relearn that workaround before getting to real work, and the cost keeps spreading outward from whoever hit it first.&lt;/p&gt;

&lt;h2&gt;
  
  
  The ground is moving faster than lock-in can track
&lt;/h2&gt;

&lt;p&gt;The timing makes this worse than it would have been a year ago. The model market is moving faster than any single vendor relationship can keep up with.&lt;/p&gt;

&lt;p&gt;Moonshot AI open-sourced the full weights of Kimi K3 in late July: a 2.8-trillion-parameter model, and the largest open-weight release so far. DeepSeek V4 reached general availability the same week, under an MIT license. Qwen's 3.8-Max shipped days after that. The UK's AI Safety Institute now measures the best open-weight models at four to seven months behind closed frontier labs on hard benchmarks (down from six to ten months for most of last year), while running at a fraction of the cost. Chinese open-weight providers alone now account for something like 45% of tokens flowing through major routing platforms, up from under 2% a year earlier.&lt;/p&gt;

&lt;p&gt;That pace is exactly why lock-in costs more now than it used to. A coding agent that hard-codes a team to one vendor's model puts "which model is best" on that vendor's release calendar instead of the market's. Two months from now, a cheaper or better option ships somewhere, and a locked-in team doesn't get to use it until their vendor decides to offer it, or until the team rebuilds the workflow to switch. That cost lands hardest on teams with the least room to absorb it: the ones who can't just pay for the expensive tier and move on.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changes with an open agent
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7h8ktf2jjbfwoelnpgg9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7h8ktf2jjbfwoelnpgg9.png" width="800" height="421"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This is the case for an agent that doesn't force the tradeoff in the first place. If the prompts and context handling are open source, a team can actually look at what happened when something breaks, instead of filing a ticket and waiting on an answer. Switching models can mean picking a different setting: bring your own API keys, run something locally, or let a router pick between options. That's why a team can turn on the next Kimi K3 or DeepSeek V4 release without rebuilding anything to use it. And if the agent works the same way in the IDE, the terminal, and the cloud, the context a team has built up doesn't reset every time someone's setup changes, which matters most for whoever on the team has the least experience to fall back on.&lt;/p&gt;

&lt;p&gt;Kilo Code is built this way: open source under Apache-2.0 for the extension and MIT for the CLI, running in VS Code, JetBrains, the CLI, and the cloud, connected to more than 500 models with zero markup on inference. It has grown past 3 million developers and 40 trillion tokens processed. That growth says something about usefulness. The portability and the transparency were part of the design from the start.&lt;/p&gt;

&lt;h2&gt;
  
  
  Build for the tax that's actually there
&lt;/h2&gt;

&lt;p&gt;That early design choice matters more this year than it would have a year ago, simply because the model market keeps moving. A team on an open agent is more prepared to adjust to pace of the frontier. A team locked into one vendor's model has to wait for permission, or rebuild. Multiply that gap across a team, and it's the same tax spoon theory describes: energy spent fighting the tool, paid first by whoever on the team has the least of it to spare.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>coding</category>
    </item>
    <item>
      <title>Your Agent Has Too Much Context</title>
      <dc:creator>Job from Kilo</dc:creator>
      <pubDate>Fri, 28 Aug 2026 19:00:00 +0000</pubDate>
      <link>https://dev.to/kilocode/your-agent-has-too-much-context-37mc</link>
      <guid>https://dev.to/kilocode/your-agent-has-too-much-context-37mc</guid>
      <description>&lt;p&gt;&lt;strong&gt;&lt;em&gt;Specs and instruction files feel like a safety net. But pile on too many and the agent gets confused rather than sharper. So what’s the actual source of truth?&lt;/em&gt;&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;There's a reflex that every team adopting coding agents develops: when the agent gets something wrong, add more context. Write a spec. Add a rule to the instructions file. Tell it to go read the relevant docs before it starts. The assumption is that more guidance means better output.&lt;/p&gt;

&lt;p&gt;Not always. We've been running enough features through agents to see the other edge of this, and it's counterintuitive: past a certain point, more context makes the agent worse.&lt;/p&gt;

&lt;h2&gt;
  
  
  When specs start lying to the agent
&lt;/h2&gt;

&lt;p&gt;We have a set of spec files in our cloud repo that describe how various systems are supposed to behave. For genuinely complex work --- billing, a gateway with a lot of surface area --- they've been lifesavers, especially for reviewers trying to make sense of a large change.&lt;/p&gt;

&lt;p&gt;But something breaks when the agent pulls in three or four of them at once. It starts treating a detail from one spec as a fact about another. It over-anchors on things it half-read. It confidently tells you "this is how it should be," and when you push back and ask why, it admits it conflated two documents. You end up needing far more runs to get to the state you actually wanted.&lt;/p&gt;

&lt;p&gt;Part of the problem is that the specs drifted. The original intent was that a spec describes &lt;em&gt;outcomes and behavior&lt;/em&gt; --- what should be true, not how to build it. Over time they picked up implementation detail and grew. One of ours ballooned to well over a hundred bullet points of rules, written in such precise, lawyerly RFC-style language ("the upstream payment provider" instead of just naming the thing) that it reads like a unit test written by an attorney. A human can barely parse it. An agent takes every clause as gospel --- including the clauses that are now subtly wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  The question underneath the specs
&lt;/h2&gt;

&lt;p&gt;Strip away the tactics and the harder question is this: &lt;strong&gt;what is the source of truth?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Agents are non-deterministic. You tell one to implement something and it will, but that doesn't tell you what &lt;em&gt;should&lt;/em&gt; be true. So where does that live?&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;Is it the tests? The agents write those now, and they're very motivated to make them pass. That's not the same as the tests being correct.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Is it the code? The code is what &lt;em&gt;is&lt;/em&gt; true, not necessarily what was supposed to be true. An agent reading the codebase will infer intent from the current implementation, bugs and all.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Is it the specs? Only if they're maintained, unambiguous, and someone actually keeps them evergreen. Ours weren't, reliably.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Is it in someone's head? That person goes on vacation, or the feature was built three weeks ago and the details are already gone.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This isn't academic. Months later --- sometimes only weeks later --- you land in a corner of the codebase and need to know: how is this supposed to work? What was the business outcome? Was it implemented correctly? Without a source of truth you can point at, you're guessing, and so is the agent.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two kinds of truth
&lt;/h2&gt;

&lt;p&gt;The most useful framing we've landed on splits it in two.&lt;/p&gt;

&lt;p&gt;There's &lt;strong&gt;how the software is&lt;/strong&gt; --- and the honest source of truth for that is executable: the code and the tests. You can't write "this is the best plugin in the world" in a prompt and make it true. Behavior is verified by running it.&lt;/p&gt;

&lt;p&gt;And there's &lt;strong&gt;how you want the software to be&lt;/strong&gt; --- the goals, the direction, the architectural constraints. That's the part you set, and it can't be reverse-engineered from the current code, because the whole point is that you often want the new thing to not look like the legacy thing.&lt;/p&gt;

&lt;p&gt;Conflating those two is where a lot of agent pain comes from. If your instructions contradict each other, or describe the current state when you meant to describe the target state, the agent walks confidently in the wrong direction.&lt;/p&gt;

&lt;h2&gt;
  
  
  Less instruction, more intent
&lt;/h2&gt;

&lt;p&gt;The most surprising trend runs the other way. Look at the system prompts and harnesses shipping from the frontier labs, and the trend across model generations is &lt;em&gt;fewer&lt;/em&gt; baked-in instructions, not more. They lean on skills and context pulled into the window dynamically, rather than a giant static rulebook.&lt;/p&gt;

&lt;p&gt;That points at a fix worth trying:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Cut the standing instructions.&lt;/strong&gt; If your agents file is aggressively telling the model to go read all the specs before every task, that eagerness works against you. Make context pull-in deliberate and situational.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Separate the two truths explicitly.&lt;/strong&gt; Let code and tests be the source of truth for how things &lt;em&gt;are&lt;/em&gt;. Use instructions to express where you want to go --- the goals and boundaries --- not to re-describe the implementation.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Enforce direction with structure, not prose.&lt;/strong&gt; When we've migrated architectures with human teams, we didn't document it exhaustively; we created a new folder or module so it was obvious which pattern was current and which was legacy. Agents need the same clear signal.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Keep specs for the genuinely hard stuff.&lt;/strong&gt; Billing, gateways, anything a reviewer couldn't hold in their head. Not every component needs one, and a spec per feature is how you get the confusion in the first place.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Decide who the spec is for.&lt;/strong&gt; A document optimized for an agent and a document optimized for a human reader are different artifacts. Pick one on purpose, and know whether your review process expects a human to read and understand a spec change in a PR, or whether that's the agent's job.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Where I'm still uneasy
&lt;/h2&gt;

&lt;p&gt;I'll admit I feel a little nervous about removing specs entirely. Across multiple sessions, having &lt;em&gt;some&lt;/em&gt; reference point you can verify a change against is genuinely valuable --- being able to tell an agent "confirm the implementation matches the spec" is a real superpower, and it's helped when partnering with non-engineers who needed a contract to reason about.&lt;/p&gt;

&lt;p&gt;But "valuable sometimes" is not "load all of it, always." The goal isn't zero context. It's the right context, pulled in on purpose, kept honest, and clearly separated into what is true versus what we want to be true. Give the agent that and it stops arguing with itself; give it everything and it won't.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>agentskills</category>
      <category>ai</category>
    </item>
    <item>
      <title>Open Weights Is All You Need</title>
      <dc:creator>Job from Kilo</dc:creator>
      <pubDate>Fri, 28 Aug 2026 18:30:00 +0000</pubDate>
      <link>https://dev.to/kilocode/open-weights-is-all-you-need-4b89</link>
      <guid>https://dev.to/kilocode/open-weights-is-all-you-need-4b89</guid>
      <description>&lt;p&gt;&lt;em&gt;Kilo’s perspective on the Open Weights AI letter.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Last month, more than 230 organizations signed the &lt;a href="https://www.microsoft.com/en-us/corporate-responsibility/topics/open-weight/" rel="noopener noreferrer"&gt;Open Weights and American AI Leadership letter&lt;/a&gt;, asking Washington to protect the open weight ecosystem instead of restricting it. Anaconda, our new parent company, &lt;a href="https://www.anaconda.com/blog/anaconda-open-weights-ai-letter" rel="noopener noreferrer"&gt;is among them&lt;/a&gt;. NVIDIA's Jensen Huang announced the coalition behind it in his first ever post on X. It was a good week for open.&lt;/p&gt;

&lt;p&gt;Now, we're bringing the data to back up the need for an open ecosystem.&lt;/p&gt;

&lt;p&gt;Kilo doesn't make models. We don't host them, we don't train on or retain your data. We're the application and routing layer, the part that sends each coding task to whichever inference you choose, through approved providers and policies for your organization. &lt;a href="https://kilo.ai/leaderboard/race" rel="noopener noreferrer"&gt;Our data&lt;/a&gt; is worth something in this debate, because we have no side in the open-versus-closed battle. We just watch which models builders actually use.&lt;/p&gt;

&lt;p&gt;Here's what we see.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkgwitrobm6td9nmcd9pz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkgwitrobm6td9nmcd9pz.png" width="799" height="368"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Share of all token usage: open-weight vs proprietary. Week of Jul 20, 2026&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;As of the week of July 20, 2026, open-weight models accounted for &lt;strong&gt;79.1% of all token usage on Kilo&lt;/strong&gt;. Proprietary models were the remaining 20.9%. And that number is not finished climbing. Not long ago open weights were a small minority of the traffic. Today they're a capable workhorse in the inference stack.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnir2tdol4gyuhk7x3646.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnir2tdol4gyuhk7x3646.png" width="799" height="368"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Share of all token usage: open-weight vs proprietary. Week of Jul 21, 2025&lt;/p&gt;

&lt;p&gt;When Jensen Huang shared the letter, he made a point we agree with completely: the world needs both frontier closed models &lt;em&gt;and&lt;/em&gt; frontier open ones. Our data supports that. Closed frontier models are excellent at the hardest problems, and builders use them there. But for the other ~80% of the work, open weights often win on a combination of these things: accuracy on the specific task, cost, privacy, and the freedom to run them where you want.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The model boom&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;A year ago, "which model to use" was a short conversation. Now a new one lands almost every week from the likes of Nemotron, Qwen, GLM, MiniMax, Kimi, Mistral, and the frontier labs. And the &lt;a href="https://kilo.ai/leaderboard" rel="noopener noreferrer"&gt;list goes on&lt;/a&gt;. The menu has exploded, and no single model wins every task.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fud8ffp42qahlhh06r4zo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fud8ffp42qahlhh06r4zo.png" width="800" height="537"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Open-weight usage by lab&lt;/p&gt;

&lt;p&gt;That is why routing is becoming some of the most valuable real estate in tech, and it is why we build the way we do. Kilo partners with the labs instead of competing with them, so our only job is to get you to the right model, not to steer you toward one we happen to own.&lt;/p&gt;

&lt;p&gt;That distinction matters more as everyone rushes to build routing. A healthy ecosystem is not one big model that everyone depends on. It is many good closed and open ones. Open weights you can run on your own terms, and integrate it deep into your stack, instead of renting access. The explosion of capable open models means less dependency and resilience for organizations.&lt;/p&gt;

&lt;p&gt;The same logic applies to the layer on top. A routing layer owned by a model vendor is not neutral, because it has a reason to prefer its own models. The ecosystem needs a neutral layer, not another lock-in. Our &lt;a href="https://kilo.ai/leaderboard" rel="noopener noreferrer"&gt;leaderboard&lt;/a&gt; consistently shows NVIDIA Nemotron and MoonshotAI models sitting alongside frontier labs like OpenAI and Anthropic, and that is not going to change.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The future of AI is choice, routing, and an ecosystem where open and closed models work together.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Look who signed&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The letter reads like a map of the ecosystem we already route across every day. Among the signatories are companies at every layer of the stack we partner with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Silicon and cloud:&lt;/strong&gt; NVIDIA, Amazon&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Model labs:&lt;/strong&gt; Mistral, OpenAI, Arcee&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Inference and serving:&lt;/strong&gt; FriendliAI, Morph&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Local runtime:&lt;/strong&gt; Ollama&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Deploy and frontend:&lt;/strong&gt; Vercel&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are the partners whose models and infrastructure show up in Kilo sessions every day.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Open from day one&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;None of this is just a reaction to a policy moment. &lt;a href="https://kilo.ai/model-freedom" rel="noopener noreferrer"&gt;Model freedom&lt;/a&gt; has been a first-class feature at Kilo since the start: open-weight models and local hosting alongside the closed frontier APIs, all from one platform.&lt;/p&gt;

&lt;p&gt;In practice that means you can route through your existing vendor contracts by bringing your own keys, or run local and private models where policy requires, across every developer surface. For teams with data-residency requirements, that same routing control is what makes an &lt;a href="https://kilo.ai/eu" rel="noopener noreferrer"&gt;EU-first setup&lt;/a&gt; possible without giving up model choice.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Backed by benchmarks&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Yesterday we published a &lt;a href="https://blog.kilo.ai/p/kimi-k3-grok-45-built-the-same-database" rel="noopener noreferrer"&gt;head-to-head benchmark&lt;/a&gt; where an open model, Kimi K3, and a closed frontier model built the same database, and the open-model path came out significantly cheaper for a comparable result. Open weights aren't just the affordable option anymore. In real world scenarios, they're competitive.&lt;/p&gt;

&lt;p&gt;We've written before about why betting your entire stack on one side is the real risk. In &lt;a href="https://blog.kilo.ai/p/openrouter-opus-5-and-the-era-of" rel="noopener noreferrer"&gt;"OpenRouter, Opus 5, and the Era of Model Freedom,"&lt;/a&gt; we argued that walled gardens are giving way to routing and choice.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Why Anaconda + Kilo&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;For an enterprise, this is more than a developer convenience. The intelligence your products and decisions run on is core infrastructure, and core infrastructure is not something to outsource entirely to a single vendor. Owning your intelligence means you can keep it, understand it, and stand behind it.&lt;/p&gt;

&lt;p&gt;Ownership pays off if you can govern it: decide which models and providers are approved, see where inference runs. That is the pairing behind Anaconda and Kilo. For more than a decade, Anaconda has shown that open and well-governed are not opposites, building the trusted foundation that tens of millions of Python and data-science developers rely on at the package layer. Kilo does the same one layer up, at the inference layer, giving teams model freedom while ensuring enterprises stay compliant.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The bottom line&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Open weights don't need to beat closed models, and we're already seeing the positive effect from open models on a daily basis. They let you own your intelligence and depend less on any single vendor. That's what the letter protects, and our data shows why it's worth protecting. The open ecosystem is already carrying most of the load. The healthiest future for AI is both open and closed together, and the best way to keep it that way is to make sure nobody has to pick a side.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;That's the whole idea behind model freedom.&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>coding</category>
    </item>
    <item>
      <title>Kimi K3 + Grok 4.5 Built the Same Database as Claude Opus 5 at 1/25th the Price</title>
      <dc:creator>Job from Kilo</dc:creator>
      <pubDate>Thu, 27 Aug 2026 11:37:05 +0000</pubDate>
      <link>https://dev.to/kilocode/kimi-k3-grok-45-built-the-same-database-as-claude-opus-5-at-125th-the-price-3hcj</link>
      <guid>https://dev.to/kilocode/kimi-k3-grok-45-built-the-same-database-as-claude-opus-5-at-125th-the-price-3hcj</guid>
      <description>&lt;p&gt;Moonshot AI published the full weights for &lt;a href="https://www.kimi.com/en/blog/kimi-k3" rel="noopener noreferrer"&gt;Kimi K3&lt;/a&gt; on July 27. It is the largest open-weight model ever released, inference providers started serving it within a day, and it has been available in Kilo since day one. We wanted to see how close a combo of Kimi K3 for planning and &lt;a href="https://x.ai/news/grok-4-5" rel="noopener noreferrer"&gt;Grok 4.5&lt;/a&gt; for implementation can get to &lt;a href="https://www.anthropic.com/news/claude-opus-5" rel="noopener noreferrer"&gt;Claude Opus 5&lt;/a&gt; doing both jobs itself. We gave both setups the same two-phase job. Each one had to design an embedded database from a spec, then build it. We crash-tested both results with a harness we wrote before either model ran, and then we read both codebases.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7crzlo64d1lqr8amnb2z.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7crzlo64d1lqr8amnb2z.png" width="800" height="352"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;TL;DR:&lt;/strong&gt; Both setups passed 64 of the same 65 conformance checks, survived every crash-recovery test, and shipped the same single critical bug. Claude Opus 5 scored &lt;strong&gt;98/100&lt;/strong&gt;, the Kimi K3 + Grok 4.5 setup scored &lt;strong&gt;93/100&lt;/strong&gt; and ran at &lt;strong&gt;4% of the cost&lt;/strong&gt; ($1.27 against $31.71).&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Pricing&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu0987env9t59d8o18fmd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fu0987env9t59d8o18fmd.png" width="800" height="395"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Per token, Kimi K3 costs 60% of Claude Opus 5, and Grok 4.5 costs 24% on output. This setup puts the open-weight model on planning, where the design decisions get made, and the more affordable model on implementation, where the decisions are already written down. Our &lt;a href="https://blog.kilo.ai/p/sol-vs-fable" rel="noopener noreferrer"&gt;earlier planning comparison&lt;/a&gt; found that frontier models implement a good plan almost interchangeably. This test asks whether that holds when neither model in the setup is a frontier flagship.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;What We Asked Them to Build&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;We wrote a spec for kvd, an embedded key-value store in Go. It is a lighter version of something like Redis, with one binary and a small text protocol over TCP, except data is persisted to disk. We picked it because everything in the spec can be checked. Either the database loses data when you kill it, or it does not. The spec requires:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;A single Go binary serving a Redis-style text protocol over TCP, with commands for reading, writing, deleting, batching, stats, and compaction&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;A durability rule: the server must not acknowledge a write until the data is flushed to disk, and every acknowledged write must survive &lt;code&gt;kill -9&lt;/code&gt; arriving at any moment&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Atomic batches: a group of up to 1,000 writes and deletes must apply all-or-nothing, even when the process dies mid-batch&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Crash recovery: on restart the server must detect a partially written record at the end of its log by checksum, discard it, and lose nothing else&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Scale: 1,000,000 keys, datasets bigger than RAM, and recovery within 60 seconds&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Standard library only, with no third-party storage engines&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The hard part of building a system like this is the failure paths. The store has to recover from a crash, throw away a half-written record without losing the good data before it, and keep a batch atomic when the process dies in the middle of one. That is where we aimed the test.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The Two-Phase Setup&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Both setups ran in &lt;a href="https://kilo.ai/" rel="noopener noreferrer"&gt;Kilo CLI&lt;/a&gt; with the same prompts. Claude Opus 5 ran at xhigh reasoning in both phases, Kimi K3 at max, and Grok 4.5 at high.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Phase 1: planning.&lt;/strong&gt; We told each planner that its plan would be handed to a different engineer who would never see the original prompt and could not ask questions. The plan had to carry everything, from the exact wire protocol to every durability guarantee. If a requirement was missing from the plan, the implementer would never know it existed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Phase 2: implementation.&lt;/strong&gt; A fresh session got only the plan file. Claude Opus 5 implemented its own plan. Grok 4.5 implemented Kimi K3's plan. Neither implementer ever saw the initial prompt.&lt;/p&gt;

&lt;p&gt;This strict handoff lets us trace every bug back to its source. If a requirement went missing between the spec and the plan, the bug belongs to the planner. If the plan was right and the code is wrong, it belongs to the implementer.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;How We Graded&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Before either model ran, we wrote a Python test harness that speaks the spec's wire protocol and grades any server that implements it. It runs 65 checks across six suites: protocol conformance, durability under crash injection, batch atomicity, compaction, consistency, and scale. For the durability checks, the harness records every write it sends and marks it acknowledged once the server replies. Then it kills the server with &lt;code&gt;kill -9&lt;/code&gt; at a random moment, restarts it, and checks that every acknowledged write is still there, byte for byte. It repeats this ten times per suite while concurrent writers are running.&lt;/p&gt;

&lt;p&gt;The harness covers everything we can check automatically. For the rest, we read both codebases, graded test quality, documentation accuracy, and code hygiene, and looked for bugs the automated checks could not reach.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The Planning Phase&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4bpllp0mmcupu5dyxns0.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4bpllp0mmcupu5dyxns0.png" width="800" height="348"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Claude Opus 5's plan is more than twice as long, and most of the extra length is substance. It includes decision tables, a table of exactly where fsync must happen, and walkthroughs of what recovery does at each crash point. Kimi K3's plan is shorter but covers the same durability rules, the same wire protocol, and the same batch atomicity design.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm82cv1oeheat5xskn265.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm82cv1oeheat5xskn265.png" width="800" height="1035"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Kimi K3's Plan&lt;/p&gt;

&lt;p&gt;Claude Opus 5's planning run also needed one adjustment. Kilo CLI caps output at 32,000 tokens per step by default, and Opus 5 at xhigh is the first model we have seen hit that cap during a single reasoning run. Our first attempts finished without writing a plan at all. We raised the cap to the model's 128,000-token limit with &lt;code&gt;KILO_EXPERIMENTAL_OUTPUT_TOKEN_MAX&lt;/code&gt; and re-ran, and one step in the successful run produced about 37,000 output tokens. The per-token pricing understates the gap here. Opus 5 produces far more output tokens during reasoning than the other two models, and reasoning tokens bill at the output rate.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmf2xgpefh7im62jmekpn.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmf2xgpefh7im62jmekpn.png" width="800" height="1232"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Opus 5's Plan&lt;/p&gt;

&lt;p&gt;Both plans also have flaws, and they matter for what comes next:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Both planners made the same mistake&lt;/strong&gt; on one protocol edge case, the handling of an oversized batch. More on that below.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Kimi K3's compaction design has a hole. A key deleted while compaction is running can reappear after a restart, because the compacted copy of the key survives while the deletion record does not.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The Implementation Phase&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpfcsmw9s8ec9m1li4o28.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fpfcsmw9s8ec9m1li4o28.png" width="800" height="348"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Claude Opus 5 took 76 minutes. It built a test suite covering 14 different crash scenarios and a benchmark running at the full 1,000,000-key scale. Grok 4.5 finished in 11 minutes with a codebase about a quarter of the size, and its tests cover the happy paths plus one crash test.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Grok 4.5 did not build the compaction flaw from its plan.&lt;/strong&gt; It changed the design so that compaction runs serialized with writes and cleans up all old data files, which closes both holes in Kimi K3's plan. Claude Opus 5 went the other way and faithfully built the one mistake its own plan contained.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;How the Two Setups Worked&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The two setups worked in different styles. Claude Opus 5 kept going back over its own work. Its planning session ran 49 steps, and 23 of them were edits to the plan it had already written. Its implementation ran 150 steps, 79 of them shell commands, in a loop of building, running its own tests, and fixing what they caught. Kimi K3 and Grok 4.5 worked much closer to one shot. Kimi K3 produced its plan in 7 steps and stopped. Grok 4.5 wrote the code, tests, and README in 22 steps and called it done. We saw the same split when &lt;a href="https://blog.kilo.ai/p/sol-vs-fable" rel="noopener noreferrer"&gt;GPT-5.6 Sol audited its own plans&lt;/a&gt; unprompted.&lt;/p&gt;

&lt;p&gt;This loop is where much of Claude Opus 5's extra time and cost went, and it is probably why its tests and documentation came out so much deeper. You could likely prompt the budget combo into a similar loop. We did not, because both setups got identical prompts. What we measured is the default behavior, and by default one setup reviews its own work and the other does not.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Crash Testing: A Tie&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Every automated check came back the same for both.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8wbj3q6g9umce4ucae9v.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8wbj3q6g9umce4ucae9v.png" width="800" height="540"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Neither store lost a single acknowledged write in any round. Both recovered a million-key store in under a second and a half. Both handled garbage appended to their data files without losing anything. We read the code to check whether these passes were luck, and they were not. Both implementations flush to disk before acknowledging a write. Both use checksummed records, so a torn write gets detected instead of trusted. Both write a batch as a single record, so partial application cannot happen.&lt;/p&gt;

&lt;p&gt;Both stores also failed one check, and it was the same one.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The Bug Both Shipped&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The spec caps batches at 1,000 operations and says an oversized batch must be rejected with an error and nothing applied. Both servers reject it. Neither consumes the batch's operations off the connection first. The operations are still sitting in the stream after the rejection, so the server reads them as fresh standalone commands and executes them.&lt;/p&gt;

&lt;p&gt;We confirmed it with a scripted reproduction against both binaries. After rejecting a 1,001-operation batch, Claude Opus 5's server went on to apply all 1,001 operations it had just refused. Grok 4.5's server applied 2 before the connection closed. A client that trusted the rejection would have data it never meant to commit.&lt;/p&gt;

&lt;p&gt;The bug came from the plans. Claude Opus 5's plan spells out the wrong behavior, saying that on an oversized count the connection "stays open, no operations consumed." Kimi K3's plan never says what to do with the pending operations. Both implementers built what their plans said. &lt;strong&gt;Two different planners wrote the same bug independently.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Neither model's own test suite catches it, because neither sends an oversized batch with its operations attached. Both test suites pass. We saw the same pattern in our &lt;a href="https://blog.kilo.ai/p/we-gave-claude-opus-47-and-kimi-k26" rel="noopener noreferrer"&gt;Claude Opus 4.7 vs Kimi K2.6 comparison&lt;/a&gt;, where models did not write the tests that would catch their own worst bug. Our harness found it because it tests against the spec instead of trusting the models' own tests.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;What the Extra $30 Bought&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;The five-point gap in the final score comes entirely from the code review.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Tests.&lt;/strong&gt; Claude Opus 5's suite kills the server at 14 points, corrupts its own files with truncation and bit flips to check that recovery handles them, and fuzzes the protocol parser. Grok 4.5's suite has one crash test and never kills the server during a batch or a compaction.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Documentation.&lt;/strong&gt; Claude Opus 5's 746-line README survived a line-by-line check against the code. Grok 4.5's README covers the required topics but claims an fsync behavior on macOS that the code does not implement.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Hygiene.&lt;/strong&gt; Grok 4.5 shipped thinking-out-loud comments and some dead code. Claude Opus 5's codebase was much cleaner.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Remaining defects.&lt;/strong&gt; Beyond the shared batch bug, we found one narrow error-path window in Grok 4.5's compaction that could leave a stale data file behind. We found nothing comparable in Claude Opus 5's store.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;One difference went the other way. Claude Opus 5's server used 428MB of memory after recovering the 512MB dataset. Grok 4.5's used 223MB. The cause is a mis-sized memory estimate in Claude Opus 5's recovery path, not a design problem. Both stores keep values on disk rather than in RAM.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;The Scores&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3julsqm6leni9f63o8a5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3julsqm6leni9f63o8a5.png" width="800" height="636"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Cost and Time&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftbdmv9bjgzorwghlpsyf.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftbdmv9bjgzorwghlpsyf.png" width="800" height="352"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The Kimi K3 + Grok 4.5 setup came in at &lt;strong&gt;4% of the cost and 23% of the time&lt;/strong&gt;. The entire score gap comes from tests, documentation, and code hygiene. None of it comes from crash safety or protocol correctness.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;Conclusion&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;For correctness,&lt;/strong&gt; this test could not separate the two. Both setups got the same conformance score, the same crash-test results, and the same single critical bug, and that bug came from the planning phase in both. As an implementer, Grok 4.5 gave up nothing to Claude Opus 5 on this build.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;For the codebase you end up maintaining,&lt;/strong&gt; Claude Opus 5's output is better. Its tests cover the failure paths, its documentation matches the code, and it left no dead code or leftover comments behind. If the code is going into a repo that people will work in long term, that difference matters.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If cost were not a factor, we would default to Claude Opus 5.&lt;/strong&gt; It did the better job overall. If the price gap were smaller, say two or three times, paying more for the better deliverable would still be a reasonable default. But the gap here is 25x, and that makes the more expensive workflow hard to recommend blindly. It has to be mission-critical work, or a case where the extra cost is clearly justified by the extra points. The things the extra $30 bought are also the easiest things to add afterwards. A follow-up prompt asking Grok 4.5 for a deeper test suite and a README pass costs about another dollar, and a normal code review would catch the hygiene issues. Crash safety is the part you cannot add afterwards, and the combo gave up none of it.&lt;/p&gt;

&lt;p&gt;We also suspect the Kimi K3 + Grok 4.5 setup has room to improve. We have not tested this, but tuning the planning and implementation prompts to push Kimi K3 and Grok 4.5 into the same iterate-and-review loop that Claude Opus 5 runs by default would likely raise the quality of what they ship. If you want to run the same split, Kilo lets you set a different model per mode. The plan can come from Kimi K3 in Plan mode and the implementation from Grok 4.5 in Code mode.&lt;/p&gt;

&lt;p&gt;Additionally, two things made the budget result stronger than we expected. Grok 4.5 fixed its planner's compaction flaw on its own, which goes against the assumption that an affordable implementer blindly builds whatever it is handed. And Kimi K3's weights are now public, so its price has room to fall as more providers host it. We have watched the same thing happen with every open-weight release this year.&lt;/p&gt;

&lt;p&gt;The most useful finding is the shared bug. Two models at very different price points planned independently and made the identical mistake, and both implementations passed their own tests. The bug only surfaced because we tested both servers against the spec with a harness written before any model ran. That kind of external testing costs the same no matter which models you use.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>grok</category>
      <category>kimi</category>
    </item>
  </channel>
</rss>
